While platforms such as Hugging Face, PyTorch, and scikit-learn are often viewed as democratic alternatives to concentrated corporate AI development, it remains unclear whether they meaningfully distribute participation and influence or instead reproduce existing hierarchies in new forms. This project develops the first integrated empirical framework for measuring participation, influence, and value capture across the open-source AI ecosystem. We construct a large-scale, multi-layer dataset linking scientific publications, authors, institutions, software artifacts, contributors, and downstream usage. By connecting software components to their scientific origins and tracing their adoption across GitHub repositories and AI development pipelines, we quantify how ideas move from research into practice and identify which actors benefit from that process. Using network analysis, statistical modeling, and inequality measures grounded in constrained baseline comparisons, we evaluate how participation and influence are distributed across firms, universities, individuals, and countries. The resulting analyses provide a data-driven assessment of whether open-source AI functions as a counterweight to platform power or remains dependent on concentrated sources of control. Beyond its substantive findings, the project establishes a reusable research infrastructure and openly available dataset that support future work on AI governance, digital sovereignty, innovation systems, and the evolving relationship between scientific knowledge and technological deployment.
Research & Programming
Open-Source AI as Democratic Infrastructure: Mapping Power, Participation, and Control
Overview
Artificial intelligence is increasingly shaped by open-source ecosystems that mediate the translation of scientific discoveries into widely deployed technologies.