The Construction of Algorithmic Predictions in Society

DSpace Repositorium (Manakin basiert)


Dateien:

Zitierfähiger Link (URI): http://hdl.handle.net/10900/183611
http://nbn-resolving.org/urn:nbn:de:bsz:21-dspace-1836116
Dokumentart: Dissertation
Erscheinungsdatum: 2026-09-22
Sprache: Englisch
Fakultät: 7 Mathematisch-Naturwissenschaftliche Fakultät
Fachbereich: Informatik
Gutachter: Williamson, Robert (Prof. Dr.)
Tag der mündl. Prüfung: 2026-09-07
DDC-Klassifikation: 004 - Informatik
Schlagworte: Maschinelles Lernen , Wahrscheinlichkeit
Lizenz: http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=de http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=en
Zur Langanzeige

Inhaltszusammenfassung:

Immer mehr Entscheidungen in gesellschaftlichen Kontexten werden automatisiert. Algorithmische Vorhersagen werden dabei häufig so interpretiert, dass sie Zusammenhänge zwischen realen Aspekten der Welt identifizieren (oder zumindest approximieren), die unabhängig von den gewählten technischen Methoden sind. Diese Annahme prägt sowohl die Fachliteratur als auch die öffentliche Debatte und wird oft als Rechtfertigung von Entscheidungen herangezogen, die auf solchen Vorhersagen basieren. Wie ich argumentiere, beruht dieses Verständnis statistischer Regelmäßigkeiten auf (i) einem von Naturgesetzen geprägten Wissenschaftsverständnis, (ii) einem „Casino-Verständnis“ von Wahrscheinlichkeiten als objektiven Größen sowie (iii) der zentralen Rolle, die Gesetzmäßigkeiten in der Statistik zugesprochen wird. Im ersten Teil dieser Arbeit zeige ich, dass Wahrscheinlichkeiten generell konstruiert sind. Dies wird oft übersehen, da sie nicht direkt evaluierbar sind. Ich systematisiere die Evaluierung probabilistischer Vorhersagen mittels Kalibrierung und vereinheitliche unterschiedliche Auffassungen von Wahrscheinlichkeiten über Gemeinsamkeiten in der Abhängigkeit von Modellen. Im zweiten Teil entwickle ich ein Verständnis von maschinellem Lernen für gesellschaftliche Kontexte, welches ohne die Annahme von datengenerierenden Wahrscheinlichkeitsverteilungen auskommt, und zeige, dass diese Annahme in der Praxis irreführend ist. Ich widerlege die daraus abgeleitete Vorstellung, dass eine Berücksichtigung von Fairness notwendigerweise mit schlechteren Vorhersagen einhergehen muss, sowohl theoretisch als auch empirisch. Zudem zeige ich, dass---entgegen der gängigen Auffassung---die Einbeziehung demografischer Merkmale Ungleichheiten und Stereotype verstärken kann. Im dritten Teil erweitere ich den Anwendungsbereich der vorangegangenen Überlegungen. Ich entwickle ein mathematisches Modell für kausale Inferenz, dessen Annahmen und Vorhersagen leichter überprüfbar sind, und schließe mit einer Kritik verbreiteter Auffassungen über künstliche Intelligenz. Zusammengefasst lege ich Probleme des vorherrschenden Verständnisses algorithmischer Vorhersagen dar und leiste sowohl konzeptionelle als auch technische Beiträge, die den konstruierten Charakter von Vorhersagen aufzeigen und die theoretische Analyse näher an die soziale Realität heranführen.

Abstract:

As data-driven decision-making is increasingly pervasive in society, algorithmic predictions are often understood as identifying, or at least approximating, prop- erties of the world that are independent of the technical tools; this applies to the technical literature as much as to the public. Individual decisions based on such predictions are then justified by the claim that the models track real relationships between measured attributes. I argue that this view of statistical regularities is based on (i) the common understanding of science as identifying natural laws, (ii) the casino understanding of chances or risks as something that can be identified, and (iii) the fundamental role of statistical regularities in the standard mathematical framework of statistics and machine learning. In the first part of this thesis, I demonstrate that probabilities are generally constructed and that contrary intuitions rest at least partly on the fact that we cannot directly evaluate them. In particular, I systematise the evaluation of prob- abilistic predictions through notions of calibration and unify different perspect- ives on probability by emphasising the ubiquity of model-dependence. In the second part, I present a constructive perspective on machine learn- ing for societal contexts in particular. I show that the assumption of a data- generating distribution is not necessary for modelling machine learning and that it can be misleading in practice. Theoretically and empirically, I disprove the idea derived from this assumption that fairness considerations necessarily require a deviation from better predictions. I also show that (against received wisdom) the inclusion of demographic attributes can exacerbate inequality and that their current use can be harmful even in auditing for racial discrimination. In the third and last part, I extend the scope of the foregoing considerations beyond the supervised learning setting. For causal inference tasks, I present a framework dispensing with data-generating distributions, which not only cap- tures established estimators but also paves the way for greater reliability and accountability by making the required assumptions testable. I conclude with a more general critique of common understandings of artificial intelligence. Taken together, I demonstrate problems with the prevalent understanding of algorithmic predictions and offer conceptual and technical contributions that, in various ways, emphasise their constructed nature and bring theoretical analysis closer to social reality.

Das Dokument erscheint in: