MIPHEI-ViT: Multiplex Immunofluorescence Prediction from H&E Images using ViT Foundation Models
arXiv. Posted May 15, 2025. arXiv 2505.10294.
How to cite
AMA
Balezo G, Trullo R, Pla Planas A, Decenciere E, Walter T. MIPHEI-ViT: Multiplex Immunofluorescence Prediction from H&E Images using ViT Foundation Models. arXiv. Posted May 15, 2025. arXiv:2505.10294
APA
Balezo, G., Trullo, R., Pla Planas, A., Decenciere, E., & Walter, T. (2025). MIPHEI-ViT: Multiplex Immunofluorescence Prediction from H&E Images using ViT Foundation Models. arXiv. https://arxiv.org/abs/2505.10294
BibTeX
@misc{balezo2025miphei,
title = {MIPHEI-ViT: Multiplex Immunofluorescence Prediction from H&E Images using ViT Foundation Models},
author = {Balezo, Guillaume and Trullo, Roger and Pla Planas, Albert and Decenciere, Etienne and Walter, Thomas},
year = {2025},
eprint = {2505.10294},
archivePrefix = {arXiv},
primaryClass = {eess.IV}
}
A routine H&E slide shows tissue structure, but it cannot tell you which proteins mark which cells. Multiplex immunofluorescence can read many markers at once on the same section, yet its staining, instruments, and cost keep it off most slides.
This group trained MIPHEI-ViT, a model built on a vision-transformer pathology foundation encoder, to predict marker signals and single-cell type calls directly from an ordinary H&E image. They trained and tested it on the publicly released Orion colorectal-cancer dataset, which pairs H&E with multiplex immunofluorescence from the same tissue, then checked it on five more independent datasets.
The result points toward reading marker-level, cell-type-aware signal from H&E slides that already exist, with one honest limit: the model learns everything it knows from real multiplex data.
Key findings
- MIPHEI-ViT calls single-cell types from H&E alone. On the OrionCRC test set it reached F1 scores of 0.93 for Pan-CK, 0.83 for alpha-SMA and 0.68 for CD3e.
- Training and evaluation ran on the public OrionCRC dataset. It supplied 41 colorectal-cancer whole-slide images pairing H&E with 18-channel immunofluorescence on the same tissue sections, split into 37 training, 2 validation and 2 test slides.
- The H&E-trained model generalized to 5 independent external datasets. Beyond OrionCRC, it was validated on HEMIT, PathoCell, IMMUcan, Lizard and PanNuke, each carrying domain shifts in mIF technology and H&E appearance.
How this study used Orion data
“All models are trained on H&E-mIF images from OrionCRC extracted as 256x256 pixel tiles at 0.5 mpp.”
— Balezo et al., arXiv (2025), Training Setup, “DataConfiguration”
Provenance: this study used the publicly available Orion CRC dataset (Lin et al., Nature Cancer 2023); the authors did not run a RareCyte instrument.
Disclosure: RareCyte is listed as an author affiliation on the publication cited above.
Disclosure: RareCyte is named in the competing-interests statement of the publication cited above.
Why it matters for Orion users
If you are weighing Orion for spatial work, look at what its public data was asked to do here. The authors ran no instrument of their own. They trained a foundation-model pipeline entirely on the OrionCRC dataset, and that data was strong enough to teach a model to call cell types from H&E alone. What makes Orion data usable for a problem like this is how it is made. Each colorectal-cancer slide carries H&E and an 18-channel immunofluorescence readout registered to the same cells, produced in a single staining and imaging round rather than from serial sections. Because the morphology and the markers come from one piece of tissue, a model can learn the mapping between them cell for cell, and 41 whole-slide images gave the study its training and test material. For your own work, that same-section registration is the payoff: when your H&E and your marker readout share the same coordinates, either one can feed tools built for the other.






