Publication:

Epistemic Limits of Trustworthy Machine Learning

dash.author.emaillucasmonteiropaes@gmail.com
dash.depositing.authorMonteiro Paes, Lucas W.
dash.licenseLAA
dc.contributor.advisordu Pin Calmon, Flavio
dc.contributor.authorMonteiro Paes, Lucas W.
dc.contributor.committeeMemberLakkaraju, Himabindu
dc.contributor.committeeMemberLu, Yue M.
dc.contributor.committeeMemberSudan, Madhu
dc.date.accessioned2026-01-13T23:13:39Z
dc.date.available2026-01-13T23:13:37Z
dc.date.created2025
dc.date.issued2025-08-05
dc.date.submitted2025
dc.description.abstractTheoretical understanding of a system’s limits has long driven technological breakthroughs. Carnot delineated the fundamental limits of heat engine efficiency, paving the way for the design of modern state-of-the-art engines. More than a century later, Claude Shannon unraveled the fundamental limit of communication, known as channel capacity. This insight revolutionized communication systems, enabling continual improvements that ultimately led to wireless communication as we know it today. This thesis discusses the epistemic limits of machine learning (ML) and leverages them to improve the trustworthiness of ML systems. ML models have an epistemic limit when proving one of their properties is impossible. Epistemic refers to the impossibility of providing theoretical guarantees (knowledge) about a model's property. Epistemic limits are information-theoretic converse results on the hypothesis test that checks a model's property. First, we prove a limit on how much information personalized models can use while ensuring reliable test for performance gains across all users -- epistemic limits of personalization. We leverage this limit to develop a tool to help with feature selection. Second, we show a limit for reliably testing if model performance is equitable across multiple demographic groups --epistemic limit of fairness testing. We exploit this limit to design a metric for efficient algorithmic bias detection. Third, we prove a limit for testing if one model outperforms another on average -- epistemic limit of model selection. We use this result to delineate the set of indistinguishably good models --Rashomon set. Finally, we argue that the epistemic limits in model selection imply that explaining the predictions of ML models is necessary. Then, we develop efficient methods for explaining the content produced by large language models.
dc.description.sponsorshipEngineering and Applied Sciences - Applied Math
dc.format.mimetypeapplication/pdf
dc.identifier.citationMonteiro Paes, Lucas W.. 2025. Epistemic Limits of Trustworthy Machine Learning. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.
dc.identifier.orcid0000-0003-0129-1420
dc.identifier.other32122281
dc.identifier.urihttps://p2p8-sa-zuvru-a9vusux.re-cotta.com/handle/1/42725112
dc.language.isoen
dc.subjectArtificial Inteligence
dc.subjectExplainability
dc.subjectFairness
dc.subjectHypothesis Testing
dc.subjectInformation Theory
dc.subjectPredictive Multiplicity
dc.subjectApplied mathematics
dc.subjectStatistics
dc.subjectArtificial intelligence
dc.titleEpistemic Limits of Trustworthy Machine Learning
dc.typeThesis or Dissertation
dc.type.materialtext
dspace.entity.typePublication
oaire.licenseConditionLAA
thesis.degree.date2025
thesis.degree.departmentEngineering and Applied Sciences - Applied Math
thesis.degree.grantorHarvard University Graduate School of Arts and Sciences
thesis.degree.levelDoctoral
thesis.degree.namePh.D.

Open/View Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Epistemic_Limits_of_Trustworthy_Machine_Learning.pdf
Size:
4.8 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license_Monteiro Paes.pdf
Size:
72.05 KB
Format:
Adobe Portable Document Format
Description: