A 2026 head-to-head comparison of 34 pulse oximeters found that device performance across skin pigmentation varies by manufacturer, and that 18 of them could pass or fail depending on which participants you tested. A separate counterfactual study fed the same model SpO2 or SaO2 and watched its mortality prediction degrade. Every fairness audit in healthcare AI examines the algorithm. Nobody is auditing the number the algorithm reads.
I want to start with a number that I think is the most important one in clinical AI right now, and it has nothing to do with a model.
Researchers compared 34 pulse oximeters using current and emerging regulatory frameworks. Eleven of them showed more positive bias in participants with dark versus light dorsal finger pigmentation across the 70 to 100 percent saturation range, with a median difference of 1.42 and an interquartile range of 1.35 to 1.90. More devices passed when the testing cohort was expanded beyond 24 participants. And eighteen of the thirty-four could pass or fail depending on which cohort you happened to select for the analysis.
Eighteen devices whose regulatory verdict depends on who volunteered.
That study was published in Anesthesia and Analgesia in April of this year, and it landed while both the International Organization for Standardization and the FDA are actively revising pulse oximeter requirements to address exactly this problem.
Now hold that next to something else.
A separate group ran a counterfactual experiment on the BOLD dataset, which links 163,396 nearly simultaneous SpO2 and SaO2 measurements to clinical features and outcomes. They trained two machine learning models with identical data, identical features, and identical settings. The only difference was the oxygen input. One model got the pulse oximeter number. The other got the blood gas number.
In patients whose pulse oximeter overestimated saturation by 3 percent or more, mortality prediction recall fell from 0.63 to 0.59, with a P value below 0.001. More false negatives. The model was reassured for the same reason we are reassured, and it was wrong for the same reason we are wrong.
Every fairness framework I have read for healthcare AI tells you to audit the algorithm. Check the training data for representation. Check performance stratified by subgroup. Check for disparate impact in the output.
Not one of them tells you to check whether the sensor feeding the algorithm reads accurately on all of your patients.
In the ICU, that is where the problem starts.
ICCN Update
The new ICCN website is live at iccn.io. Every article published in the past two weeks is now archived in one place, and our new Research section pulls recent published data from 26 major critical care and medical journals into a single curated feed for subscribers. Bookmark iccn.io.
Why This Matters
Let me lay out the clinical foundation carefully, because this literature is strong in direction and weak in design, and I do not want to overstate it.
The signal that reopened this question was a research letter in the New England Journal of Medicine in 2020, which reported that occult hypoxemia, meaning an arterial saturation below 88 percent when the pulse oximeter read 92 to 96 percent, occurred roughly three times more often in Black than in White patients. A larger analysis in JAMA Network Open the following year examined discrepancies between SpO2 and SaO2 by race and ethnicity and reported associations with organ dysfunction and mortality. A JAMA Internal Medicine analysis found that discrepancy was associated with delayed identification of eligibility for COVID-19 treatment. A multicenter Veterans Health Administration cohort published in The BMJ found the same pattern outside the ICU, among general medical and surgical inpatients, and examined reproducibility as well as bias. A 2024 systematic review in the British Journal of Anaesthesia pulled the skin tone literature together.
Every one of those is retrospective and observational, and I will come back to that.
Now, why this matters more in an ICU than anywhere else, in four steps.
SpO2 is the most heavily reused number in critical care. It titrates oxygen. It gates weaning protocols. It triggers rapid response criteria. It feeds SOFA scoring. It appears in ARDS phenotyping work, in deterioration models, in sepsis alerts, and in the closed-loop ventilation systems I wrote about three weeks ago. A single measurement error does not stay in one place.
Low perfusion makes it worse, and the ICU is where low perfusion lives. A prospective study in Anesthesia and Analgesia examined low perfusion and missed diagnosis of hypoxemia in darkly pigmented skin. Shock, vasopressors, hypothermia, and edema are the daily conditions of critical care, and they degrade the signal on top of any pigmentation effect.
The error is directional, not random. Random noise widens a confidence interval. Directional error moves a decision threshold. A pulse oximeter that overestimates does not make you uncertain. It makes you falsely reassured, and it does that more often in one group of patients than another.
Models inherit it silently. A model trained on SpO2 has learned a version of oxygenation that is systematically shifted for some patients. Nothing in the model’s performance metrics will tell you that, because the model is being evaluated against outcomes that were themselves shaped by the same biased input.
“You cannot audit your way to a fair model when the number going into it is not the same number for every patient. The fairness problem in critical care is a metrology problem wearing an algorithm’s clothes.”
What Stood Out
1. The regulatory verdict depends on who showed up. Eighteen of thirty-four devices could pass or fail depending on cohort selection. That is not a finding about pigmentation. That is a finding about the fragility of the entire clearance standard, and it applies to every patient regardless of skin tone.
2. The problem is not uniform across manufacturers. Eleven of thirty-four showed the pattern. That means this is not an inherent limitation of the physics of two-wavelength oximetry that we all have to live with. Some devices do better. That reframes it from a scientific limit to a procurement decision.
3. The counterfactual design is the right experiment. Holding everything constant except the oxygen input is the cleanest way anyone has shown that measurement error propagates into model behavior. It is one dataset and it is retrospective, but the design is exactly right.
4. The direction of model failure is the direction of clinical failure. More false negatives. The model missed deteriorating patients for the same reason a bedside clinician would have. The algorithm did not introduce a new bias. It automated an existing one at scale and at speed.
5. There is a serious statistical counterargument and it deserves airtime. A published critique in Annals of Intensive Care has argued that some of what is reported as device bias in these retrospective paired-measurement studies is at least partly statistical bias arising from the study design, and has separately called for manufacturers to disclose their calibration algorithms. I do not think that critique overturns the clinical concern, but any honest account has to put it on the table, and I have never seen it cited in an AI ethics paper on this topic.
Ethical Interpretation
Here is what I think is the structurally interesting part, and why this belongs on an ethics channel rather than only a devices channel.
Algorithmic fairness has been built around a particular story about how bias enters healthcare AI. The story is that historical data encodes historical inequity, the model learns that inequity, and the model reproduces it. That story is correct and important, and the classic examples, including cost-based risk scores that under-identified Black patients, fit it well.
But that story locates the problem in the training data and the objective function. It assumes the measurements themselves are neutral.
In critical care, the measurements are not neutral. They are physical readings taken through a patient’s skin by an optical device calibrated on a population, and the calibration does not hold equally across that population.
Three consequences follow.
The first is that fairness auditing as currently practiced would not catch this. If you stratify a deterioration model’s performance by race and find no difference in AUROC, you have not established that the model treats patients equally, because the ground truth labels and the input feature were generated by the same shifted instrument for the same patients. The audit and the object being audited share a contaminant.
The second is that this is a procurement and biomedical engineering problem that has been filed under clinical ethics, and it has therefore landed on nobody’s desk. The people who read AI ethics papers do not buy pulse oximeters. The people who buy pulse oximeters are not in the AI governance meeting. In most hospitals I know of, there is no single person who could tell you both which oximeter models are on the floors and which deployed algorithms consume their output.
The third, and this is the one I feel most strongly about, is that respiratory therapy sits at the exact point where this could be caught. We place the probe. We choose the site. We assess the perfusion. We decide whether the reading is believable and whether to draw a gas. Every one of those is a data-quality decision, and none of them has ever been described that way in a job description or a competency checklist.
I am not going to claim that RTs can solve a device calibration problem at the bedside. We cannot. But the discipline that generates the number has a legitimate and currently unclaimed voice in the conversation about what is done with it, and I would like us to claim it.
Bedside and Workplace Takeaways
1. Respiratory therapists: treat the probe as a data-quality decision, and confirm when the picture disagrees.
When the clinical picture and the pulse oximeter disagree, the gas settles it. Work of breathing, mental status, and the number all pointing different directions is the exact scenario the occult hypoxemia literature describes. Check the site, check perfusion, reposition, and if the disagreement persists, advocate for a blood gas. This is not new clinical behavior. What is new is understanding that you are also generating the ground truth that everything downstream depends on.
2. ICU nurses: cross-check on the patient, not on the number.
Trend, waveform quality, perfusion, and work of breathing are your independent read. Escalate on the patient. When you escalate and the SpO2 looks acceptable, say so out loud, because that disagreement is clinically informative and it is also the local signal that something upstream may be off.
3. Critical care pharmacists: find the order sets that use SpO2 as a hard gate.
Oxygen titration protocols, weaning criteria, and some sedation and mobility pathways use a saturation threshold as a binary trigger. Ask which of your institution’s order sets treat SpO2 as a gate rather than as one input among several. A hard threshold is where a systematic 1 to 2 percent shift changes an actual decision.
4. Advanced practice providers: you are the ones who generate ground truth.
Every arterial blood gas you order produces a measurement that the pulse oximeter cannot. When the reading and the examination do not match, order the gas. Beyond the individual patient, a unit with a reasonable rate of paired measurements is a unit that could actually study this locally, and a unit that never draws gases has no way to know.
5. Intensivists and medical directors: build the SpO2 dependency list, then ask the vendors.
Produce a written list of every deployed model and protocol that takes SpO2 as an input. Then ask each monitoring and algorithm vendor for performance stratified by skin pigmentation. Expect not to receive it. The absence of an answer is itself the finding, and it belongs in your governance minutes.
6. Perfusionists: extracorporeal support compounds this for separate reasons.
Circuit flow, recirculation, peripheral perfusion, and cannulation site all degrade peripheral oximetry independently of any pigmentation effect. On support, confirm oxygenation with co-oximetry rather than relying on the peripheral probe, and name that limitation when the team is making decisions off a saturation number.
Teaching Pearl
There are three places bias can enter a clinical AI system, and they need three different fixes.
The label. The outcome the model was trained to predict encodes a historical inequity. The fix is choosing a better outcome.
The algorithm. The model learns a shortcut that tracks a protected characteristic. The fix is auditing and constraining the model.
The sensor. The input measurement is systematically shifted for some patients. The fix is metrology: better devices, better calibration cohorts, and confirmatory measurement at the bedside.
Almost all of healthcare AI ethics addresses the first two. In critical care, the third is where our biggest exposure sits, and it is the one no fairness toolkit currently touches.
What We Should Not Over-Assume
We should not assume any deployed ICU model is currently producing inequitable output. The mechanism is documented. The propagation has been shown in one retrospective linked dataset. Whether any specific model in any specific unit is producing worse decisions for darker-skinned patients has never been measured, and I am not going to claim it has.
We should not assume the occult hypoxemia literature is causal. It is retrospective and observational, it depends on paired measurements that were clinically indicated rather than protocolized, and a published statistical critique argues that part of the reported effect reflects design rather than device.
We should not assume all devices behave the same way. Eleven of thirty-four showed the pattern in the head-to-head comparison. Treating this as universal is inaccurate and would let the better-performing manufacturers off the hook for no reason.
We should not correct SpO2 by race. I want this stated flatly. Applying an offset based on a patient’s race or skin tone would embed a racial variable into a physiologic measurement, which is the same category of error the field has spent a decade removing from other clinical algorithms. The answer is a better instrument and a confirmatory measurement, not a correction factor.
We should not assume the proposed fixes work. Larger pigment-stratified premarket cohorts, objective skin tone measurement, multi-wavelength hardware, and algorithmic signal correction have all been proposed. None has been shown to change a patient outcome.
Limitations
There is no Tier 1 evidence in this entire article. Not one randomized trial has tested any intervention against pulse oximeter bias or its downstream propagation, and no accumulation of observational studies substitutes for that.
The clinical literature is retrospective, uses clinically indicated rather than protocolized paired measurements, and has a live statistical critique against it.
The 34-device comparison is a laboratory desaturation study in volunteers, not a study of critically ill patients, and volunteer desaturation studies do not reproduce shock, vasopressors, or edema.
The counterfactual machine learning study uses a single linked dataset, examines three prediction tasks, and was published in a workshop proceedings volume rather than a clinical journal.
And the central gap: nobody has measured what any of this does to a deployed ICU algorithm in real use. The argument in this article connects a documented measurement problem to a documented propagation mechanism and asks a question about deployed systems that no one has answered.
Bottom Line
Pulse oximeters read differently across skin pigmentation, the effect varies by manufacturer, the current clearance standard is fragile enough that eighteen of thirty-four devices could pass or fail depending on who was tested, and a model fed that biased number predicts worse in exactly the patients where the number is most wrong.
None of that is proven to be harming anyone through an algorithm today. All of it is a mechanism sitting in plain sight in the most measurement-dense unit in the hospital.
The fairness toolkits will not find it, because they are pointed at the model.
So the work is unglamorous and it is available immediately. Build the list of what consumes SpO2. Ask the vendors the question they cannot answer. Find out which oximeters are actually on your floors. Draw the gas when the patient and the number disagree.
And if you are a respiratory therapist reading this, understand that the discipline that places the probe is the discipline that determines the quality of the data everything downstream is built on. That is a larger role than the profession has claimed, and it is exactly the right moment to claim it.
References
Hughes C, Chen D, Law T, et al. Pulse oximeter performance and skin pigment: comparison of 34 oximeters using current and emerging regulatory frameworks. Anesth Analg. Published online April 20, 2026. doi:10.1213/ANE.0000000000008048
Martins I, Matos J, Gonçalves T, Celi LA, Wong AKI, Cardoso JS. Evaluating the impact of pulse oximetry bias in machine learning under counterfactual thinking. In: Wu S, Shabestari B, Xing L, eds. Applications of Medical Artificial Intelligence. AMAI 2024. Lecture Notes in Computer Science, vol 15384. Springer; 2025. doi:10.1007/978-3-031-82007-6_21
Sjoding MW, Dickson RP, Iwashyna TJ, Gay SE, Valley TS. Racial bias in pulse oximetry measurement. N Engl J Med. 2020;383(25):2477-2478. doi:10.1056/NEJMc2029240
Wong AKI, Charpignon M, Kim H, et al. Analysis of discrepancies between pulse oximetry and arterial oxygen saturation measurements by race and ethnicity and association with organ dysfunction and mortality. JAMA Netw Open. 2021;4(11):e2131674. doi:10.1001/jamanetworkopen.2021.31674
Valbuena VSM, Seelye S, Sjoding MW, et al. Racial bias and reproducibility in pulse oximetry among medical and surgical inpatients in general care in the Veterans Health Administration 2013-19: multicenter, retrospective cohort study. BMJ. 2022;378:e069775. doi:10.1136/bmj-2021-069775
Fawzy A, Wu TD, Wang K, et al. Racial and ethnic discrepancy in pulse oximetry and delayed identification of treatment eligibility among patients with COVID-19. JAMA Intern Med. 2022;182(7):730-738. doi:10.1001/jamainternmed.2022.1906
Martin D, Johns C, Sorrell L, et al. Effect of skin tone on the accuracy of the estimation of arterial oxygen saturation by pulse oximetry: a systematic review. Br J Anaesth. 2024;132(5):945-956. doi:10.1016/j.bja.2024.01.023
Shachar C, Drabo EF, Iwashyna TJ, Ferryman K. Addressing racial and ethnic bias in pulse oximeters, a wicked problem. JAMA. 2025;333(7):563-564. doi:10.1001/jama.2024.25443
Matos J, Struja T, Gallifant J, et al. BOLD: blood-gas and oximetry linked dataset. Sci Data. 2024;11(1):535. doi:10.1038/s41597-024-03225-z
Gudelunas MK, Lipnick M, Hendrickson C, et al. Low perfusion and missed diagnosis of hypoxemia by pulse oximetry in darkly pigmented skin: a prospective study. Anesth Analg. 2024;138(3):552-561. doi:10.1213/ANE.0000000000006755
Cabanas AM, Sáez N, Collao-Caiconte PO, Martín-Escudero P, Pagán J, Jiménez-Herranz E, Ayala JL. Evaluating AI methods for pulse oximetry: performance, clinical accuracy, and comprehensive bias analysis. Bioengineering. 2024;11(11):1061. doi:10.3390/bioengineering11111061
Tobin MJ, Jubran A. Unreliable pulse oximetry in dark-skin patients: a plea for algorithm disclosure. Ann Intensive Care. 2022. doi:10.1186/s13613-022-00993-y
Saidy S, Iqbal A, Baig SH. Pulse oximetry discrepancies and occult hypoxemia in ICU patients: predictors and clinical outcomes. J Intensive Care Med. 2025. doi:10.1177/08850666251351594
Clinical Disclaimer
This content is provided for educational and professional development purposes only. It does not constitute medical advice, a clinical protocol, a practice guideline, or a substitute for institutional policy, local governance requirements, or independent clinical judgment. Nothing in this article should be interpreted as a recommendation to apply any correction factor to a physiologic measurement. Clinicians remain responsible for all decisions regarding patient care and for compliance with the standards of their licensing bodies and employers.
Javier Amador-Castaneda, BHS, RRT, FCCM
Founder and CEO, Interprofessional Critical Care Network




