Precision Medicine vs Persuasion Medicine: Measuring vs Reading HER2
Why a Yale pathologist argues that HER2 should be measured in attomoles per square millimeter, not read by eye, right at the IHC 0-versus-1+ boundary that now decides who receives HER2 antibody-drug conjugates.
In this Yale Cancer Center Grand Rounds, David Rimm, the Anthony Brady Professor of Pathology and Medicine and director of translational pathology at Yale, argues that HER2 should be measured, not read. His case is direct: pathologists reading HER2 by eye agree well on 3+ tumors but split almost evenly between IHC 0 and 1+, which is exactly the call that now decides whether a patient receives trastuzumab deruxtecan (T-DXd). To close that gap he proposes a high-sensitivity quantitative immunofluorescence assay, HS-HER2, that reports HER2 in attomoles per square millimeter instead of a subjective 0/1+/2+/3+ score, and he ends by warning against what he calls “persuasion medicine,” where a score that gates a drug pressures pathologists to reclassify a read.
In this video:
- The stakes: trastuzumab deruxtecan and a wave of antibody-drug conjugates benefit HER2-low and some HER2-zero tumors, so the IHC 0-vs-1+ call the drug hinges on must be right.
- Reader disagreement: across CAP surveys and an 18-pathologist, 170-case JAMA Oncology study, readers agree on 3+ but ~50/50 on 0 vs 1+: the reader you draw decides the drug.
- Why reading fails: chromogenic IHC is, in a phrase Rimm borrows, a scale built for elephants used to weigh mice, saturating at high HER2, blind to the low range of 0 and 1+.
- Proposed fix: HS-HER2, a quantitative immunofluorescence assay on a cell-line standard curve, HER2 in attomoles/mm², a reflex test for IHC 0 and 1+ like reflexing a 2+ to FISH.
- From bench to chart: run as a CLIA lab-developed test: antibody titration; accuracy, precision, sensitivity/specificity validation; signing a number into the patient record.
- “Persuasion medicine”: once a score gates a drug, Rimm warns, clinicians may push pathologists to upgrade an IHC 0 to 1+; he asks oncologists for measurements, not new reads.
Full transcript
Introduction (Yale Cancer Center): It is my pleasure to introduce David Rimm, the Anthony Brady Professor of Pathology and Medicine here at Yale. David is a Hopkins alumnus, did his pathology residency here, and completed a cytopathology fellowship at the Medical College of Virginia. He has been at Yale for almost 30 years, directs the pathology tissue service, and serves as director of translational pathology. He has been at the forefront of quantitative pathology for many years and is well known throughout the field. He has developed many novel assay techniques for identifying predictive markers that determine which tumors are sensitive to which targeted therapies, and that has become increasingly important as the number of targeted therapies has grown. Today he will focus on the development of companion diagnostics for HER2-directed therapies. This is particularly timely, because the first HER2-targeted therapy for non-FISH-amplified breast cancers was approved just six months ago, and exactly how we identify which patients will benefit is a huge question the field is struggling with.
David Rimm (Yale Cancer Center): Thanks, and thanks especially to Ian for introducing me; he is a world leader in the HER2 antibody-drug-conjugate space I am going to talk about. My title is “precision medicine versus persuasion medicine,” and I will get to what persuasion medicine is toward the end. The other half of the title is reading versus measuring. Measuring is what you do quantitatively; reading is what pathologists do when they look at slides. The difference is between subjective and objective assessment of tissue.
Let me start with my disclosures. I do a fair bit of consulting, and much of the research in my lab, including the work that led to this, was sponsored by companies including Cepheid and Konica Minolta. Over the next 55 minutes I will give a quick introduction to the new drugs, propose a new assay for them that I will call high-sensitivity HER2, or HS-HER2, cover what it takes to move an assay from a research lab into a CLIA lab where results go into the patient chart, and finish with precision medicine versus persuasion medicine, where I will try to talk the oncologists in the room into focusing on precision rather than persuasion.
This is the drug that reportedly earned the first standing ovation in 25 years at ASCO. Underneath it is the same old trastuzumab, but eight topoisomerase-inhibitor payloads have been conjugated to it, and that gives you some special tricks. It delivers highly toxic payloads right to the cell, so you avoid the toxicity you would get from giving the drug systemically at those doses; and when the payload uncouples inside the cell, it can spill out and kill neighboring cells, the bystander or proximity effect. It worked really well. Very few patients were resistant, most had some response, and there were 11 complete responses in the early trials. It worked for essentially all patients, but especially for patients who were not amplified for HER2. The initial trials were in HER2-amplified patients, but then trials opened in HER2-low patients, IHC 2+ and IHC 1+, and the curves look pretty similar. In those low patients, in the DESTINY-Breast04 trial, the survival curve improved median survival in advanced breast cancer from five to nine months, and that is what ultimately led to the popularization and the success of the drug.
But what about HER2-zero? What if a tumor does not express any HER2 at all, and can we even tell the difference between HER2-zero and HER2-low? There is a trial underway for HER2 greater than zero but less than 1+, the DESTINY-Breast06 trial, which has not reported yet. There is also the DAISY trial in France, a small trial where the waterfall plots clearly showed patients who benefited even with HER2 equal to zero. Why is it important to understand this and to have good diagnostics for it? Because this drug is the tip of the iceberg. There is a long list of other targets for antibody-drug conjugates in clinical trials, so ADCs may become very important for oncology over the next few years, and equally important will be companion diagnostics that pick the right patients, because it is important to pick patients who express the right amount of target.
So what do we do now? The ASCO/CAP guidelines from 2018 are how we practice as pathologists in assessing HER2 expression. The algorithm looks at circumferential staining that is complete and intense in greater than 10 percent of cells for a 3+, a 2+ and a 1+ below that, and a zero for no membrane staining, with weak-to-moderate partial staining being the subjective middle. It used to be important mainly to tell the 3+ tumors from the rest. But now the new category that matters is HER2-low, and there are a lot of them, as many as 65 to 70 percent of patients thought to fall into the HER2-low category. That means many patients could get the drug, but it also means we need to be as accurate as we can, because we do not want HER2-zero patients getting the drug if they will not benefit.
How well do we actually do this? I am fortunate to be on the immunohistochemistry committee of the College of American Pathologists, so I get access to the proficiency surveys that CLIA labs must pass twice a year to be allowed to return results to the chart. On the HER2 tissue-microarray survey from 2020, three of ten cases, cases four, six, and seven, did not reach consensus. Of the roughly 1,400 labs that did this, they could not reach the 90 percent agreement required. One failed case showed a big discordance between called zero and called one, almost 50/50, with some labs even calling it two or three. That is troubling. If we are assuring twice a year that labs give the right answer for patients, how can there be that much difference between zero and one? Because I am on the committee, I could ask for the data from the past few years, and of 80 cases from 2019 and 2020, 15 showed a discordance greater than 25 percent between zero and one.
We thought, this is tissue microarray, not the real world. So we studied real-world core biopsies. We enrolled 18 pathologists from 15 institutions around the United States and asked them to read actual core biopsies that had been read at Yale, 170 cases, scored by the ASCO/CAP guidelines before the popularization of the 1+-versus-zero distinction, so they were simply scoring zero, one, two, and three as they always had. Of those cases, 92 were discordant, and 69 of the 92 were discordant between zero and one; only 20 were discordant between two and three. That work, led by Eileen Fernandez in my lab, ultimately got us published in JAMA Oncology, although we were not allowed to say what we wanted to say, which is that there is great discordance between zero and one and not so much between two and three. For a two we have a solution, because we can reflex to FISH, an orthogonal assay. Between zero and one we do not yet have a solution.
You can also look at the per-pathologist calls, work done by Jack Robbins with Eileen Fernandez, showing the percentage of readers who called a case zero versus one. These are all currently signing-out pathologists, most with more than five years of experience, not residents. If you are pathologist number 18 you scored only 15 percent of the patients as zero, but if you are pathologist number one you scored 44 percent. So whether or not you get trastuzumab deruxtecan can depend on who your pathologist is, and that does not sound like a great idea to me.
So we asked, with Gang Han, how many readers you need to be sure an assay agrees with itself. There are about 21,000 pathologists just in the United States and 100,000 in the world. There is actually no statistical method for this, so we simply plotted the overall percent agreement against the number of observers, or readers. What you see is that the more observers you have, the less agreement you get, which makes sense mathematically. Does this actually work to assess assays? For estrogen receptor it turns out we are really good; if a quartet of pathologists reads estrogen receptor, all four agree somewhere between 85 and 95 percent of the time. For HER2 we do not do so well. For 3+ versus not-3+ we are really good, but for a quartet deciding zero versus not-zero, agreement is between 40 and 85 percent. How many readers do we need to build a good assay? It is where the curve plateaus. In one case you probably need nine or ten; in the case of telling ones from not-ones, no number is sufficient, because it goes all the way down to baseline.
So I hope I have convinced you that we need a new assay, one that is measured, not read. We started from the beginning with cell lines, some that are gene-amplified and some that express HER2 but are not gene-amplified. With the current FDA-approved assay you can separate the highs from the lows or negatives, but you cannot stratify the negatives. With a new assay that uses about ten times more antibody, a pretty simple change, you can then stratify the low cell lines and tell the zeros from the ones. The current assay is the wrong tool for the job; as a group in France put it, the current FDA-approved assay is like weighing mice on a scale built for elephants. Everybody gets that. If you have a scale for elephants, it does not work for weighing mice, and it is all about dynamic range.
So here is the assay we invented. We take a series of cell lines and do something like a Bradford assay from college chemistry, making a standard curve. We used our tissue microarray of cell lines and, with the help of Array Science, made a standard curve; with a mass-spectrometry lab we figured out how many attomoles per microgram were in each cell line; and then, using QuPath, we converted that to attomoles per square millimeter. So now we have an assay that reports attomoles per square millimeter. Like all assays it saturates when it gets too high, so the amplified cases saturate and we cannot use them, but since we do not really care about 2+ and 3+, which pathologists tell apart just fine, we discard those two and are left with a very nice linear standard curve that we can use to assign each case a value in attomoles per square millimeter.
A little assay terminology, taken straight from the FDA handbook. The limit of detection is the lowest concentration of the analyte that can be reliably distinguished from zero but not necessarily quantified. What we really want is the limit of quantification, because then we can do it right every time. And there is a limit of linearity, above which the response is no longer linear. Because we do not yet know how much HER2 is required to benefit from trastuzumab deruxtecan, we measure all the way down to the limit of detection and below to see what we get. On the tissue microarray, most of the pathologist-read 3+ cases sit above our limit of linearity, but look how many ones and twos there are in the middle range, which is further evidence that we need a measured assay to pick the right patients. Surprisingly, some cases called zero, and some called one or two, actually fall below our limit of quantification, or even below our limit of detection.
Then we did what you have to do in a CLIA lab, running 40 cases (normally 20 positives and 20 negatives, per Fitzgibbons and colleagues, although we have a continuous scale rather than positives and negatives). These are actual core biopsies, not tissue microarrays, and you see the same thing: a broad range of attomoles per square millimeter for the zeros and ones, while the amplified 3+ cases are pretty tight. In our first 40, about 20 percent of cases appeared to be below the limit of quantification for HER2 protein but potentially still present as a target for the drug. To summarize to this point: about 70 percent of cases have HER2-low, defined as above the limit of quantification and below the levels associated with gene amplification; about 8 to 10 percent are below our limit of quantification, and probably about 6 percent below our limit of detection; and many cases called HER2-zero, as many as 60 percent, and in our studies maybe 75 percent, have detectable HER2 between about three and twenty attomoles. So the quantitative HER2 assay could be envisioned as a reflex test, so that if a pathologist reads IHC zero, it reflexes to the quantitative test, in the same way we reflex a 2+ IHC to FISH today.
Now let us take it to the clinic. A colleague of mine from Brigham and Women’s once said that when the assay works in your research lab, you are five percent of the way there, and I think that is really true. Bringing this assay to the clinical setting, with help from Trish Gaul, Nay Chan, and Reva Kamova, required antibody titration and maximization of signal to noise, analytic validation, and characterization of accuracy, precision, sensitivity, specificity, and serial-core reproducibility, and then working out how to report it to our oncology colleagues. Peak signal to noise was at one microgram per milliliter for a new, higher-sensitivity antibody. Our accuracy is only 87 percent, because we are more sensitive than the status-quo IHC 0/1/2/3 assay we have to compare against, but overall we have quite good concordance and more resolution in the low range. Inter-assay precision at 10 percent sounds like it might not be great, and the assay we just bridged to is now under 10 percent, but it is acceptable; intra-assay precision, calculated from three slides run on separate trays of the same machine, is about five percent. Our sensitivity compared to the historical assay is 100 percent, and our specificity is 84 percent, low because we are more sensitive and call some cases positive that IHC called negative.
Here is the proposed clinical workflow, which we are doing now. Labs come to what I have called the QDAP lab, for quantitative diagnostics and anatomic pathology, a new lab that is now open for business, running QDAP assay number one, the high-sensitivity HER2 assay. We batch the stains and run them on a Leica Bond autostainer, then read them; we originally read on some old legacy hardware, but we recently completed a bridging study to a much higher-throughput device that scans a slide in about four minutes instead of an hour, and we sign the results out in CoPath as a procedure so they ultimately reach Epic and clinicians can see them. The pathologist picks a representative region rather than measuring the entire core biopsy, sees a pseudo-IHC image of what the region looked like, the number of fields of view (23 in this case), and the score (15.4 attomoles per square millimeter in this case). That number goes into the report, along with an interpretation: positive for expression high, meaning above our limit of linearity; positive for expression intermediate, meaning like a one or a two; positive for expression low, meaning present but possibly not reproducible, that is, above our limit of detection but not necessarily above our limit of quantification; or negative, below the limit of detection.
Our vision is that we currently offer HS-HER2 in the QDAP lab; tests must be requested by an oncologist, patients are billed, and there are diagnostic codes for the work. We began a prospective study on all breast biopsies to build a year of prospective data, and we are about seven months in. So far we have received a grand total of two clinical specimens, but we hope the assay gains traction, especially in the more distant future once we know how much target is necessary for patients to respond. It is worth noting that only about 15 percent of lab tests in the United States are provided by academic labs; the other 85 percent are provided by private labs, so if we want this to affect patients broadly it needs to reach the large private lab companies, and those discussions are beginning.
The last thing I want to talk about is precision versus persuasion medicine. Our original vision for this assay was to adjudicate the IHC-equals-zero cases, measuring them and telling you whether they are above the limit of detection or above the limit of response, which we do not know yet. But something happened in the last three or four months that I have not been able to document yet, probably because it is not mature enough: suddenly IHC zero is rare. Pathologists are people too, and a pathologist might be a little more lenient about what they call IHC one, call it a sympathy vote, because then the patient can get this new drug.
Here are some real quotes I have heard, and I will not name the people. One: “Hi, Dr. Pathologist, I see you called Mrs. X’s biopsy IHC zero; that means I am going to have to offer brain radiation. Are you sure it is not a one, so I could give her trastuzumab deruxtecan?” Should the pathologist go look at that slide again? Does that mean the first view was not accurate, or was it accurate, and should the diagnosis change because it is persuaded to be better for the patient? I am not sure that is a great idea. From a West Coast director of a pathology service: “We do not have any IHC zeros anymore.” And from a Midwestern oncologist: “I am not seeing the response rates in HER2 patients that they saw in the clinical trial; they are getting a lot of IHC zeros, and maybe IHC zeros really do not respond.” We know that 8 to 10 percent of cases really do not express any target, and this is a targeted therapy.
So what has happened is that now we may need to adjudicate the 1+ cases, because the zeros have minimized. I do not say they have gone away, and if you ask pathologists they will sternly tell you they of course still call IHC zero, and data in a year or so will tell us how our IHC-zero calls actually changed. But IHC one is now more common, so if it is more common, maybe that is the one we should be measuring, and in fact that is the plan. We are going to study IHC-equals-one several ways. The first is the QDAP lab prospective study, directed by Nate Chan, which began August 1 and runs to July 2023; we have 226 cases today and I anticipate around 400, and the primary objective is to determine how many IHC-zero cases have detectable HER2, with a secondary objective of how many IHC ones fall below the limit of detection. We have already started some quantitative work on prospective tissue, and you can see many cases called zero that are above our limit of detection, with, so far, not very many below it, though time will tell as the study matures.
There are two other studies in progress. One is a proposal to the Translational Breast Cancer Research Consortium, a group of 16 or 17 institutions that do translational studies together, which we joined with the arrival of colleagues at Yale; its goal is to evaluate HER2 measurement in 1+ metastatic cases, so that if we collect two or three hundred 1+ cases from 17 institutions, we can tell how often patients called 1+ actually have no target, and vice versa, and see response, since those patients will have received trastuzumab deruxtecan. The second, led by Miriam Lustberg here, is a proposed study of patients who are IHC zero and then prospectively given trastuzumab deruxtecan, much the way the DAISY trial worked; it is not yet fully designed or approved. These are the kinds of studies we need, with real-world or clinical-trial patient response, to figure out the attomoles per square millimeter above which patients benefit. It probably will not be a single cut point, because there are other mechanisms of resistance beyond simply not enough HER2, which many labs, including my own, are working on. A HER2-TROP2 assay is also well along, so we can help clinicians decide between sacituzumab govitecan, a TROP2-targeting therapy, and trastuzumab deruxtecan.
For my last slide: overall, HS-HER2 is a laboratory-developed test, not FDA-approved, so if you only do FDA-approved tests you probably do not do many here, because most of our assays are LDTs. Many people do not realize that if you change even one step of the protocol of an FDA-approved assay, it becomes an LDT that you must then validate, so most assays we do, including many molecular and gene-mutation assays, are actually LDTs rather than FDA-approved tests. HS-HER2 is in the correct dynamic range; we are not weighing mice on a scale built for elephants. The level of target required for trastuzumab deruxtecan is still unknown, and I do not want to hide that from you; it is very clear we do not know the answer to that question yet. But if we waited until we knew before we started the assay, we would be years behind. So now that we have the tools, I ask the oncologists in the audience: ask for measurements, not for readings, and please do not ask the pathologist to change their mind. That is persuasion medicine, not precision medicine. We all respect our pathology colleagues; when they give you a reading, they really believe it is right, and just as you should not go back and change your answer on a test, that first impression is probably the true impression and the best reading. With that, I want to thank the people in the lab who do all the work I get to talk about, and I have left about 20 minutes for questions. Thank you very much.
Audience question: What about discordance when pathologists read the same slides again after a washout period?
David Rimm (Yale Cancer Center): That is a pivotal question. When you do any kind of pathologist study and you read a case once and then read it again, you should have a washout period so you do not remember the case, because pathologists have a really good memory for the morphology of cases and can even remember the patient name on the label. We did not need a washout period in this study, because the readers only saw the slides once. If we were going to show them the slides again, or do any kind of intra-observer reproducibility, which we did not do, we would need a washout period; but in this study it was not required.
Audience question: Is heterogeneity within a tumor an important issue? Is it more important to have a small percentage of cancer cells that express a high amount of HER2, or to know that a high number of cells express at least the minimum amount of HER2?
David Rimm (Yale Cancer Center): A phenomenal question, and it is essentially Jack’s thesis project. HER2 is very heterogeneous, not only within a slide but between cuts, and all the pathologists in the audience know that when we sample one core biopsy, that is less than one percent of the tumor, so there is no way for us to answer the question about the true heterogeneity of the whole tumor. What we can ask about is heterogeneity on the slide, that is, how important high expression in a single cell is versus high expression in the average cell. We started with the average, because you have to start somewhere, and I do not know that the average is the correct answer; you could argue that, because of the bystander effect of the drug, it is actually the highest cells that make the most difference, but we do not know that, and that is just speculation at this point.
Audience question: How much heterogeneity do you see in the attomole expression, given that you are averaging across so many fields of view?
David Rimm (Yale Cancer Center): The heterogeneity within a core biopsy is quite substantial; as you know, when you read them you see bright areas and less-bright areas. How do we handle that? Someday we will know whether it is the highest cell, the average cell, or the lowest cell that matters most for response, but we do not know that yet. So in the same way we take a core biopsy and say it represents the whole tumor, we take the average and say it represents the expression of HER2. I should add that when I applied for tissue from the DESTINY-Breast04 trial to test this, I was quickly told I would never see it, and I do not fault them for that; they have their own people who can do quantitative work, they have FDA approval for IHC 0/1/2/3, and it would not be in their interest to give me access to that tissue.
Audience question: We see situations with heterogeneity where a clearly 3+ tumor has a complete response to trastuzumab, while another tumor in the same patient that was HER2-negative or 1+ did not respond. Would such patients benefit from a second round of the drug when they have two distinct HER2 profiles?
Ian (HER2 ADC discussant): It depends on the clinical situation. We know pretty clearly that with the previous generation of HER2 therapies you do not see any benefit in non-3+ or non-amplified cancer, so the HER2-lows did not respond to any of the previous generation of HER2 therapies. With trastuzumab deruxtecan you could make a case that you might see benefit in both the clearly amplified and the HER2-low components. Prior to that, we would look at a case like that on a case-by-case basis and use the HER2 therapy to address the usually more aggressive HER2-positive cancer, then worry about the HER2-negative or HER2-low cancer later. The field is evolving now that we have drugs that work across different levels of HER2. On the heterogeneity point, with the first-generation HER2 therapies it was very clear: we ran a large prospective trial with the first antibody-drug conjugate, which does not have a bystander effect, and a heterogeneous cancer responded much less effectively than a non-heterogeneous cancer; quantitatively, what mattered was the percent of HER2-negative cells, not the intensity of HER2 on the positive cells. With a drug that has a bystander effect, as David was alluding to, that might switch: you may only need a certain number of strongly HER2-positive cells to get the drug in, and then the bystander effect takes care of the HER2-negative cells. We would like to test that prospectively but have not yet had the funding.
Audience question: Do you have any information on whether the conjugate drug can be activated in the extracellular microenvironment of the tumor cell?
David Rimm (Yale Cancer Center): I would defer to Ian, who is much more of an expert on this than I am, but my understanding is that the payload has to be cleaved inside the cell, and once it comes off it survives in the extracellular environment, and that is how the bystander effect can kill neighboring cells. The drug is a topoisomerase inhibitor, so it has to get to the nucleus to have its effect. The reason you conjugate it to the antibody is to increase the local dose at the tumor. One consequence worth flagging is toxicity: this drug has pulmonary toxicity, and patients get interstitial lung disease about 10 percent of the time, which is another reason you need a companion diagnostic. One wonders whether the interstitial lung disease is due to extracellular cleaving of the drug even in the absence of HER2, although we have also found that HER2 is present in normal airways at about a “1+” level, roughly four to six attomoles per square millimeter.
Transcript reproduced from the recorded Yale Cancer Center Grand Rounds and cleaned from an automated caption source for readability: speaker labels were added, verbal disfluencies and a brief audio-visual interruption were removed, and the presenter’s substantive words are otherwise as delivered. Obvious caption garbles of established drug, trial, and assay names were normalized (for example trastuzumab deruxtecan, DESTINY-Breast04, DESTINY-Breast06, the DAISY trial, sacituzumab govitecan, HS-HER2, and attomoles per square millimeter). One ambiguous, unverifiable proper-noun token in the caption was omitted rather than guessed. Statements of affiliation, quantities, and study details are reproduced as spoken by the presenter and other participants and may differ from formally published values; the quantitative HER2 argument presented here is the speaker’s own academic work.







