Immunocto: a massive immune cell database auto-generated for histopathology
Medical Image Analysis. 2026;109:103905. DOI 10.1016/j.media.2025.103905.
How to cite
AMA
Simard M, Shen Z, Brautigam K, Abu-Eid R, Hawkins MA, Collins Fekete CA. Immunocto: a massive immune cell database auto-generated for histopathology. Med Image Anal. 2026;109:103905. doi:10.1016/j.media.2025.103905
APA
Simard, M., Shen, Z., Brautigam, K., Abu-Eid, R., Hawkins, M. A., & Collins Fekete, C. A. (2026). Immunocto: a massive immune cell database auto-generated for histopathology. Medical Image Analysis, 109, 103905. https://doi.org/10.1016/j.media.2025.103905
BibTeX
@article{simard2026immunocto,
title = {Immunocto: a massive immune cell database auto-generated for histopathology},
author = {Simard, Mika{\"e}l and Shen, Zhuoyan and Br{\"a}utigam, Konstantin and Abu-Eid, Rasha and Hawkins, Maria A. and Collins Fekete, Charles-Antoine},
journal = {Medical Image Analysis},
volume = {109},
pages = {103905},
year = {2026},
doi = {10.1016/j.media.2025.103905}
}
A routine H&E slide shows tissue structure, but it cannot say which cells are immune cells or which subtype they belong to. Teaching a model to read that from H&E needs a large set of accurately labeled cells, and labeling them by hand is slow and expensive.
This group built Immunocto, a database of millions of labeled cells generated with little manual work. They segmented cells on H&E with the Segment Anything Model, then read each cell's immune identity from co-registered multiplexed immunofluorescence on the same tissue, using the publicly available Orion colorectal-cancer dataset as the labeled imaging source.
Models trained on the result reached state-of-the-art lymphocyte detection from H&E alone.
Key findings
- Immunocto is a 6.8-million-object labeled database. It provides an H&E image and a nucleus mask for 6,848,454 cells and objects, including 2,282,818 immune cells across four subtypes: CD4+ and CD8+ T cells, CD20+ B cells, and CD68+/CD163+ macrophages.
- The labels came from the public Orion colorectal-cancer dataset. That dataset supplied 40 whole-slide images pairing H&E with co-registered 18-channel immunofluorescence at 40x magnification and 0.325 µm per pixel.
- Models trained on Immunocto reached state-of-the-art lymphocyte detection from H&E alone. The SAM-plus-classifier ensemble that generated the labels scored above 98% validation accuracy for each immune-cell type.
How this study used Orion data
“The model trained on joint H&E and IF data is applied to the entire cohort of the Orion platform [7], which contains 40 colorectal cancer WSIs.”
— Simard et al., Medical Image Analysis (2026), Methods, “Step #5: Automatic generation of a massive database of immune cells”
Provenance: this study used the publicly available Orion CRC dataset (Lin et al., Nature Cancer 2023); the authors did not run a RareCyte instrument.
Disclosure: RareCyte is listed as an author affiliation on the publication cited above.
Disclosure: RareCyte is named in the competing-interests statement of the publication cited above.
Why it matters for Orion users
If you are weighing Orion for spatial work, look at what its public data was asked to do here. The authors ran no instrument of their own. They built Immunocto, a database of more than two million labeled immune cells, with its labels drawn entirely from the Orion colorectal-cancer dataset. What makes that possible is how Orion data is made. Each colorectal-cancer slide carries an H&E image and an 18-channel immunofluorescence readout registered to the same cells, produced in a single staining and imaging round rather than from serial sections. Because the morphology and the markers come from one piece of tissue, the immunofluorescence channels can label the very nuclei segmented on H&E, and 40 whole-slide images were enough to auto-generate millions of training cells. For your own work, that same-section registration is the payoff: when your H&E and your marker readout share coordinates, one can supply ground-truth labels for the other, and the labeled result can be released for the whole field to reuse.






