Virtual Multiplex Staining for Histological Images using a Marker-wise Conditioned Diffusion Model
arXiv. 2025. arXiv 2508.14681. Accepted to AAAI 2026.
How to cite
AMA
Oh HJ, Kim J, Shi Z, Wu Y, Chen YA, Sorger PK, et al. Virtual Multiplex Staining for Histological Images using a Marker-wise Conditioned Diffusion Model. arXiv. 2025. arXiv:2508.14681
APA
Oh, H.-J., Kim, J., Shi, Z., Wu, Y., Chen, Y.-A., Sorger, P. K., et al. (2025). Virtual Multiplex Staining for Histological Images using a Marker-wise Conditioned Diffusion Model. arXiv. https://arxiv.org/abs/2508.14681
BibTeX
@misc{oh2025virtual,
title = {Virtual Multiplex Staining for Histological Images using a Marker-wise Conditioned Diffusion Model},
author = {Oh, Hyun-Jic and Kim, Junsik and Shi, Zhiyi and Wu, Yichen and Chen, Yu-An and Sorger, Peter K. and Pfister, Hanspeter and Jeong, Won-Ki},
year = {2025},
eprint = {2508.14681},
archivePrefix = {arXiv},
primaryClass = {eess.IV}
}
A routine H&E stain shows tissue structure, but it cannot tell you which proteins mark which cells. Multiplex immunofluorescence can, imaging many markers on one section at once, yet it needs specialized staining, instruments, and cost that most labs cannot reach for on every slide.
This group trained a diffusion model to produce that multiplex view from an ordinary H&E image, generating up to 18 marker channels at once from a single input. They trained and tested it on two public datasets that pair H&E with real multiplex images, one of them a colon-cancer atlas, and it scored ahead of earlier image-translation methods.
The work is a step toward reading multiplex-style signal from slides that already exist, with one honest limit: the model learns everything it knows from genuine multiplex data.
Key findings
- The model generates up to 18 immunofluorescence marker channels from a single H&E image. One marker-wise conditioned diffusion network produces every channel, a jump from the 2-3 markers most earlier virtual-staining methods managed.
- Training and benchmarking used two public paired datasets, including the Orion-CRC atlas. HEMIT supplied 3 mIHC markers, while Orion-CRC contributed 41 colon-cancer whole-slide images carrying 18 registered immunofluorescence channels each.
- On the Orion-CRC benchmark, the method led on average across all 18 markers. It improved on the next-best method by +0.039 SSIM and +1.358 PSNR, and posted the top PSNR on 18 of 18 markers and the top SSIM on 13.
How this study used Orion data
“the Orion-CRC (Lin et al. 2023) dataset”
— Oh et al., arXiv (2025), Setup on Orion-CRC, “Patch filtering”
Provenance: this study used the publicly available Orion CRC dataset (Lin et al., Nature Cancer 2023); the authors did not run a RareCyte instrument.
Disclosure: RareCyte is listed as an author affiliation on the publication cited above.
Disclosure: RareCyte is named in the competing-interests statement of the publication cited above.
Why it matters for Orion users
If you are weighing Orion for spatial work, notice what this paper did not have to do. The authors never operated an instrument. They trained entirely on the public Orion-CRC atlas, and that data was strong enough to anchor a method accepted at AAAI 2026. That is the quiet signal worth reading: when a machine-learning group needed paired H&E and multiplex immunofluorescence on the same tissue to teach a model, Orion-CRC was the reference they reached for. What makes that data usable as a training target is how it was made. Each colon-cancer slide carries H&E and 18 marker channels registered to the same cells, from one staining and imaging pass. You cannot learn an H&E-to-marker mapping from images that do not line up, and single-round, whole-slide acquisition is what keeps them aligned. For your own work the takeaway is narrow but real: Orion output already circulates as benchmark-grade data other groups build on, a different kind of proof than any single study offers.






