How to cite this paper
Retter, Adam. “What does AI mean for Open Source?” Presented at Balisage: The Markup Conference 2026, Washington, DC, August 3 - 7, 2026. In Proceedings of Balisage: The Markup Conference 2026. Balisage Series on Markup Technologies, vol. 31 (2026). https://doi.org/10.4242/BalisageVol31.Retter01.
Balisage: The Markup Conference 2026
August 3 - 7, 2026
Balisage Paper: What does AI mean for Open Source?
Adam Retter
Adam has been the Director of Evolved Binary Ltd since 2014, where
they specialise in software, consultancy, and training in information
storage and retrieval. Their customers include large governmental
organisations, publishers, universities, and one of the worlds largest
social media companies. Adam also co-founded eXist Solutions GmbH, a
software consultancy company in Germany. Adam has been instrumental in
the development of the eXist-db Open Source Native XML Database since
2005, and led its development from 2017. In 2024 Adam forked
eXist-db, to create Elemental, an advanced Open Source replacement for
eXist-db. eXist-db was previously used as the base for the cityEHR
Open Source Electronic Health Records software which is deployed in
the UK, Ukraine, Kenya, and Nigeria; from there in 2024, Adam was
appointed as a professor on the Applied Health Informatics program at
Fordham University.
He is passionate about Open Source software, standards, and open
technical communities, and is an invited expert on several
international standards groups. Additionally Adam serves on the board
of several international conferences specialising in information
markup. He is a recognised expert in several computer programming
languages, and estimates that he has made contributions to over 50
different Open Source projects. He has also published the reference
book on eXist-db with O’Reilly and several papers that advanced the
state of the art in information retrieval in the publishing and
digital humanities domains.
When not travelling, Adam can be found snowboarding or hiking around the peaks on
the Italian French border where he resides.
Copyright Adam Retter 2026
Abstract
The recent meteoric rise of LLMs (Large Language Models) and associated tools was largely unexpected and surprising to most. The rapid ascent
of this technology has caught many software developers unawares, leaving them suddenly
somewhat ignorant, and arguably under-skilled.
LLMs, whilst still advancing, have recently demonstrated impressive capabilities in
their ability to assist software developers in their day-to-day tasks (e.g., coding
new features, and locating and fixing issues). However, the use and adoption of LLMs
presents many larger challenges for society as a whole; many of which are not in themselves
technical concerns.
This paper examines the current and perceived impact of this technology in the context
of Open Source. We identify several social, economic, environmental, political, legal,
and technical concerns regarding the use of LLMs in Open Source projects.
We contribute guidance around defining an AI policy for Open Source projects. We further
offer an AI Policy Score Card to assist projects in clearly defining and declaring
how they wish to work with AI or not.
Table of Contents
- Introduction
- Concerns for AI in Open Source
-
- Social Concerns
- Economic Concerns
- Environmental Concerns
- Political Concerns
- Legal Concerns
- Technical Concerns
- Defining AI Policy for an Open Source Project
-
- Components of an AI Policy
- An AI Policy Score Card for Open Source Projects
-
- 1 Star Rating
- 2 Star Rating
- 3 Star Rating
- 4 Star Rating
- 5 Star Rating
- Conclusion
Introduction
Research into LMs (Language Models) and NLP (Natural Language Processing) has been a subset of AI (Artificial Intelligence) research since the 1960’s. Developments such as ELIZA demonstrated a natural language
chat-like interface between human and machine [Wizenbaum66].
Whilst research continued and incremental progress was made, it wasn’t until more
than 50 years later, in 2017 that Google published their seminal paper introducing
the Transformer Architecture [Vaswani17] that created the foundation for modern LLMs. In the last 9 years we have seen exponential
growth in the development of new and more powerful LLMs:
-
2018: GPT-1, and BERT
-
2019: GPT-2
-
2020: GPT-3
-
2022: GPT-3.5, and ChatGPT
-
2023: GPT-4, Claude 1, Claude 2, Claude 2.1, LLaMA 1, LLaMA 2, Gemini 1, and DeepSeek
-
2024: Gemini 1.5, GPT-4 Turbo / GPT-4o, Claude 3, Claude 3.5, Claude 4.7, LLaMA 3, DeepSeek
v2, DeepSeek v2.5, Mistral, Cohere, and Qwen
-
2025: GPT-4.5, Claude 4, LLaMA 4, Gemini 2, Gemini 3, DeepSeek v3, DeepSeek v3.1, and
DeepSeek v3.2
-
2026: GPT-5.2, GPT-5.3, GPT-5.4, GPT-5.5, GPT-5.6, Gemini 3.x, Claude 4.x, Claude Opus
5, Grok 4.x, and Qwen 3.x
These new generations of LLMs not only improved the human-machine chat interface,
they also added audio voice chat capabilities and autonomous capabilities.
Significantly for Software Developers an Agentic
approach was developed. This provided: (a) APIs (Application Programming Interfaces) to access huge LLMs running in the Cloud, and (b) Software Agents that whilst operating locally within a Software Developer’s development
environment can utilise these APIs to service tasks requested by developers. Such
tasks include automated running of various tools and tests, development of new software
code, and providing fixes to issues in existing code.
Furthermore, the latest trend is towards orchestration of Agents, whereby you may
have a fleet or cluster of LLM powered Software Agents running across multiple systems
and utilising different LLMs collaborate on solving a particular task. Remarkably,
using Open Source tools such as LangChain [Chase22] or n8n [Oberhauser19], it is quite possible today to have a team of Software Agents that act in a hierarchy,
whereby one or more agents may be supervising or providing quality assurance on the
work of one of more other agents.
Whilst the modern capabilities of LLMs were barely imaginable just a few years ago,
and may appear extraordinarily useful, their use and operation is not without concern.
We will examine such concerns in the context of Open Source in section “Concerns for AI in Open Source”.
Open Source, not to be confused with Free Software [Stallman98], as a term is rather broad and arguably overloaded, itself an umbrella that covers
several different aspects of both technical and legal domains.
Officially, The OSD (Open Source Definition) [Perens98] sets out 10 criteria that a piece of computer software and its associated license
must meet in order to qualify as Open Source
. The OSI (Open Source Initiative) recognises over 80 licenses that comply with the OSD; however, in practice there
are far more Open Source licenses available, each offering different grants and restrictions
on use. One may choose to publish their software under such a license if they wish
it to be classed as Open Source.
More organically, there are what are known as Open Source Projects
, these are most usually software projects of a specific purpose that are developed
by one or more persons and/or organisations. These persons and/or organisations collaborate,
often online in public, to produce software that is licensed and published under one
of the Open Source licenses.
These Open Source Projects vary wildly in structure and governance. A project may
be as simple as an individual creating software for their own pleasure, or perhaps
a group of like minded individuals have a shared cause, and have coalesced around
an idea to create software to assist anyone in improving some environmental concern.
Conversely, at the other end of the spectrum, a project, may involve one or more large
corporate/government organisation(s) producing a complex product through the collaboration
of large teams of developers, upon which their business model depends.
Ultimately, Open Source can be thought of as an ideology, part of which is encoded
into a license for use of the product (e.g., software). An open Source license requires
you to allow anyone else to both use your software and modify it, furthermore you
cannot restrict the use of your software by any party that you might consider undesirable
for any reason. This shared ideology infers not only the free exchange and development
of ideas and technology, but from preventing any other entity from restricting this
[GiovanH21].
Whilst there is nothing in the OSD that requires development of Open Source software
to be carried out in the open and/or online, or even for the source code to be published
online, the expectation of users since around the mid-2000’s is that Open Source projects
will host their source code publicly online, and be developed in the public eye in
a real-time fashion [Peterson13, AlMarzouq22]. Therefore, modern code projects are produced in public via online collaboration
through a code repository such as GitHub, Sir Hat, or Code Berg.
It is typical for an Open Source project to have multiple contributors who are geographically
separated and perhaps not known to each other outside of their online presence. Differences
in expectations, technical ability, language barriers, and differing human cultural
and social norms of contributors, can be a strength, but can also contribute friction,
making such projects challenging to govern.
One could argue that any project involving multiple contributors is susceptible to
such conflicts; however, for private companies with bespoke (closed) software projects
and strong governance structures, such issues are easily hidden from public view within
the confines of their organisation. As there is now an expectation that all aspects
of an Open Source project are carried out and documented in public, interactions that
lead to such friction and/or conflict, can often feel amplified or more significant
when under the scrutiny of their peers. In an attempt to avoid such issues, many
projects have adopted Code’s of Conduct such as Contributor Covenant [Ehmke25], The Citizen Code of Conduct [Koehler18], or Anticode [Webb20]; however, such policies are only of use if they are correctly enforced by the administrators
and/or moderators of the project itself.
Historically, all of these problems have been very human centric, and do not directly
relate to the use of AI. However, it is our position that the use of AI, either by
developers, or autonomously via agents, to submit changes to Open Source projects
is likely to not only increase friction between humans, but also potentially introduce
friction between humans and machines.
It is therefore important that both the human and artificial contributors, current
and future, to an Open Source project, all understand the project’s policy with regards
to both human and AI/LLM produced contributions. Many Open Source projects already
include documentation for humans about how to contribute, but very few have yet decided
upon and/or published their AI policy towards human-machine contributions, and few
have documented these in a manner suitable for both humans and machines. To try and
solve this issue, our contribution is an AI Policy Score Card
that may be used by Open Source projects to gauge their AI policy; see section “An AI Policy Score Card for Open Source Projects”.
Concerns for AI in Open Source
The tragic impacts that AI/LLM can affect through sycophantic traits, including Chatbot
Psychosis [WIKI01], continue to be documented by psychiatrists [Østergaard25, Morrin26, Girgis26], academics [Cheng25], The Human Line Project [THLP25], and various news outlets [Moore26, Proven26/2].
Open Source projects and code have historically been solely built by humans, with
the rise of LLMs assisting Software Developers to generate code; Open Source founders
should consider the impact of this upon their projects and contributors, and therefore
must define their policy around AI carefully. First and foremost, they must remember
that Open Source contributors, i.e., software developers, however remote, are also
human beings (to date), and as such are susceptible to the same AI biases and sycophantic
reinforcement as any other sector of society; perhaps more so, as their use of AI
is not just limited to their personal life, but also pervades their chosen career.
The governance and purpose of Open Source projects span a wide spectrum, whilst some
projects adhere to a common foundational model of governance [ASF26, Burcher14, TLF26], many more are completely ad-hoc and structured according to the founders and/or
contributors preferences. Therefore, the concerns affecting the use of AI in Open
Source projects whilst many, will undoubtedly vary from project to project.
We present concerns that our research indicates are likely applicable to most Open
Source projects. Whilst we recognise that some of these concerns are multi-faceted
and cross-cutting, we have loosely grouped them into the following categories: Social, Economic, Environmental, Political, Legal, and Technical.
Social Concerns
For those human founders and/or contributors whom are not financially compensated
for the time and resources they invest in Open Source projects, their contributions
can largely be considered a social activity. Indeed, contributing to software that
is published as Open Source, through the very nature of the OSD means that you are
making something available to society as a whole.
The large majority of people contribute to FOSS because of fun (91%), altruism (85%),
and kinship (80%). Moreover, when analysing differences in motivations to join and
continue, the study found that ideology, own-use, or education-related programs can
be an impetus to join FOSS, but individuals continue for intrinsic reasons (fun, altruism,
reputation, and kinship).
— opensource.com - 2021-04-05 [OS21]
Allowing AI generated code into an Open Source project, regardless of a net-positive
or negative outcome, will have both, an immediate impact on the existing human contributors
to the project, and also the ability of the project to attract contributors in future.
One overarching concern in this area, is that of trust. In the past, potential new
and unknown contributors to an Open Source project were rarely immediately trusted,
instead they gained trust and entry into the project by demonstrating their utility
and building up a reputation over time. Contributors utilising AI/LLMs change that
dynamic; it is now difficult to ascertain their level of experience vs. that of the
AI/LLMs. Previous mechanisms for establishing trust between humans are becoming strained,
and it is yet not clear what, if any, new mechanisms will replace them [Blates26].
On the one hand, human contributors may be keen to embrace AI, with the potential
to help non/less-technical contributors implement code that that would otherwise be
beyond their reach [Karim26], or perhaps allow experienced developers to pair with AI to increase their velocity
[Zheyuan25, Edwards26].
On the other hand, experienced engineers may be disappointed with the code produced
by AI [LinearB26], or lament having to review large change sets that have been partially or entirely
generated by AI (that was not well instructed) and exhibit many issues [Song23].
Whether for or against AI, software developers experiencing fear and/or depression
triggered by AI is well documented [Lițan25, Spirlet26]; they may experience the feeling of loss that they will no longer be able to enjoy
socially crafting code [Orosz26, Lawson26], or the fear that they will be left behind by competitors that have embraced AI
where they have not [Kwon26].
Additionally, as human beings we subscribe to and/or create philosophies, ideologies,
and belief systems. As such, individuals may, understandably and freely, exhibit and
justify ethical and/or non-technical reasoning around the extent to which they embrace
or reject AI. Around these new ideologies a number of colloquial terms are forming
to describe them, such as AI Vegan
[Mahdawi25], and the less strict AI Vegetarian
[Proven26/1].
I have a religious exemption from using all generative “A.I.”
— Catherine Sawers - 2025-09-12 [Sawers25]
If human contributors within an Open Source project are unhappy with its policy towards
AI use, the project may experience a collapse of its social fabric, resulting in a
loss of contributions and/or risk being forked. To avoid social conflict and worse,
alienation, it would follow that some form of consensus will be required between
the founders and/or contributors to the project on, if and how AI may be used in that
project. We argue that that consensus, forms the basis of an acceptable use of AI
policy for the project, and whilst it may of course evolve over time, it should be
clearly communicated to any potential contributors.
Economic Concerns
It remains unclear whether Open Source in and of itself has had a nett-positive or
negative affect on reducing the social-economic divide with regards to developing
software.
On the one hand, the cost of access to software tools has been reduced through eliminating
licensing fees, and self taught developers can study existing code [Pscheidt09, LOS24].
On the other hand, producing and contributing to Open Source software requires hardware,
a reliable internet connection, education (specialised technical knowledge), and potentially
the wealth to allow individuals to contribute without financial compensation; this
can make the cost of entry for those from lower socioeconomic backgrounds impossible
[Blind24].
We argue that the use of AI in Open Source projects may further increase the social-economic
and digital divide between those that can afford to contribute and those that cannot.
For SaaS (Software as a Service) based AI services, which appear to be the most commonly used, these are typically
presented to the end user as either free, or with a subscription and/or token pricing
model. The current startup companies behind these services have huge infrastructure
and staff costs that are fuelled by large VC (Venture Capital) investments that have
yet to be realised. Whilst these companies are striving for market dominance and determining
the product-market fit for their offerings, they are operating their services at a
loss and spending vast sums of VC money [Sachs24]. During this time, the cost to some of using such services may seem reasonable,
whilst to others it may already be prohibitive. We cannot yet determine what the true
future price to the user of such services will be when these companies have to start
making a profit to repay their investors [Aggarwal26]. The financial markets expectation is that the end-user that wishes to continue
using such services will have to pay more than they currently do so.
Sadly, we have not been able to find much evidence of discounted SaaS AI services
to help those in developing countries or those from disadvantaged socially-economic
backgrounds. What little evidence we have points to no, or meagre, discounts that
are out of touch with the economies of developing nations.
The true financial cost of using SaaS AI services cannot yet be determined.
For those who wish to avoid subscription services by operating their own AI/LLM models,
there are also considerable costs, some of which are non-obvious [Ivchenko26].
It would seem to follow that those Open Source projects that are well funded and/or
operated by large private/government organisations may be able to afford to pay for
AI use in their projects, whereas smaller or unfunded projects likely will not.
We are not yet aware of any Open Source projects that strictly mandate the use of
AI by contributors, but there are certainly already many projects that encourage its
use. Unfortunately while the use of AI/LLM comes with a financial cost, which itself
seems likely to increase, it seems inevitable that this will further increase both
the economic-social and technical divides in contributing to Open Source projects.
We believe that those championing the use of AI in Open Source projects should consider
if they are excluding potential contributors due to social-economic status, and how
they might remove that barrier.
Environmental Concerns
A great amount of Open Source software is started under the banner of OSS4SG (Open Source Software for Social Good) [Fang26]; however, serious secondary environmental impacts have been raised in relation to
the operation of AI/LLMs which would seem to be at conflict with such a purpose.
The operation of AI/LLMs at scale require large compute resources that pose a number
of environmental concerns, the main of which can roughly be divided into three categories:
-
Manufacturing of Equipment – the required hardware: compute nodes, computer network
infrastructure, power and cooling systems, and housing. This requires raw materials
(incl. rare earth minerals) that are mined, and then electricity and water during
their construction [AIEC25].
-
Construction – dedicated physical data centres need to be constructed to house all
of the manufactured equipment. This requires the acquisition of suitable land, and
development of it. This has been documented to lead to water pollution of the surrounding
environment, and reducing nearby property prices [Luscombe26].
-
Operation – the day-by-day operation of the data centres require large quantities
of electricity for power, and water for cooling. The by-product is pollution of the
local environment from one or more of waste water, heat, noise, and light [Robinson26].
Early research into the environmental impacts of LLM’s focused on the environmental
costs of training the LLM’s. One such study, found that in 2019 training an LLM required
nearly the same carbon footprint as an average human being creates during 60 years
of their life [Strubell19]. Since that time whilst LLM’s have become much more complex and training them requires
ever more data, it is the scale at which they have been deployed, and subsequently
made available as a service that has changed this landscape.
Whereas previously training the LLM created a larger carbon footprint than its operation,
that position has now been reversed. The latest research focuses on the total environmental
cost, or which day-to-day use of the models forms a huge part. The UNU-INWEH (United Nations University Institute for Water, Environment, and Health) recently published a report suggesting that by 2030, global AI data centres will
consume 945 terawatt-hours of electricity annually. This will require a massive water
footprint for cooling and power generation equivalent to the basic annual needs of
1.3 billion people, alongside a land footprint exceeding 14,500 KM2 [Aczel26].
Such environmental concerns affect us all, and it would seem likely that the impact
of AI on our physical environment may be of particular interest to those operating
under the banner of OSS4G, and in all likelihood, many more.
We believe that such concerns may dissuade people from using AI, and by extension
individuals and organisations may be put-off contributing to, or using, Open Source
projects that make use of AI.
At present it can be hard to determine if an Open Source project uses AI or not, we
believe that such use should be clearly communicated and form part of the projects
AI policy, thus allowing users and contributors to make informed decisions.
Political Concerns
The rise and rapid adoption of AI/LLMs has not gone unnoticed by national and international
governments and organisations. These institutions which have political influence or
govern, have the ability to control if and whom can use AI, when they can use it,
and under what constraints if any. AI/LLMs, like any sufficiently advanced technology,
are a powerful tool that can be wielded to change the dynamic of power at a local
or international level for the gain of those that can assert control over it.
We must not forget that AI/LLMs whilst appearing to be available internationally,
are operated by corporations, and that these corporations are incorporated in, and
operating within, the bounds of a particular nation or state. By extension, these
corporations are subjects of that nation and its politics, they are required to comply
with the laws of their nation, changes to governance of their nation, and their nation’s
foreign policy.
Subsequent assertion of control over such advanced technologies after the fact of
their creation and adoption has plenty of past precedent.
As an example, the USA has in the past tightly controlled the export of cryptographic
algorithms and equipment. During the Cold War, previous legislation was updated in
1976 and as a result the ITAR (International Traffic in Arms Regulations) [WIKI02] officially classified cryptographic algorithms and equipment as munitions. Thus
greatly restricting cryptographic inventions born in the USA from being used outside
of the USA. In 1996, this was reformed and legislation moved from the ITAR to the
EAR (Export Administration Regulations) [EAR96]. Most of these restrictions on cryptography were eventually rescinded in 2000, some
24 years later.
More recently, in October 2022, under the same EAR that imposed cryptographic controls,
and the ECRA (Export Control Reform Act) legislation, the USA restricted the export of hardware related to accelerating AI/LLM
processing [BIS22]. Further export restrictions were added in October 2023.
Recently it has been reported that the US Government is considering passing legislation
to prohibit the use of AI/LLM models produced in China [Curi26]. The FY2026 NDAA (Fiscal Year 2026 National Defense Authorization Act) already prohibits some non-USA developed models from being used in services or products
procured by the US Department of Defense [NDAA26].
For an Open Source project that is adopting, promoting, or delivering the use of AI/LLMs,
consideration must be given to whether users and/or contributors are subject to such
export restrictions. For example, are the models you use also usable by your users
and/or contributors. For those dependent on the use of AI/LLMs consideration should
be given to whether such services will remain available to them. The European Union
AI Act [EU23], China’s Interim Measures for the Management of Generative AI Services [WIKI03], and the proposed AI Kill Switch Act in the USA [Lieu26], offer their respective governments the authority to restrict or shutdown AI/LLM
providers. Whilst politicians are often keen to assure us that such measures are an
action of last resort, we have seen many governments in recent years restrict or disable
access to the Internet and/or mobile telephony services when it suits them, and even
for reasons that may at times be perceived as against the interests of their own constituents
[Howard11, Stremlau24].
Considering political travel in the opposite direction, the companies operating the
largest AI/LLM models are often valued at billions of dollars and have a huge impact
on their user base. These companies are often themselves founded or backed by individuals
with their own fortunes valued in the billions. These organisations and individuals
are able to assert far greater political influence on their nation’s government than
the average citizen [Chesterman26]. This influence may be gained through political lobbying [Tolomia26], forming alliances to further shared interests, or often, direct financial contributions
to a political party or campaign [Williams26, Wilkins26, Dorn26].
From the individual hobbyist Open Source project, to the largest collaborative project,
it is likely impossible that everyone within holds the same set of ethics or political
views. Organisations and individuals behind large AI/LLM services are able to influence
governments in a left or right direction on a particular matter. For many individuals
that concern themselves with local or international politics, it would seem that either
as users of, or contributors to, an Open Source project, they may wish to know its
AI policy. Such information, allows them to check that the political views of their
AI providers (and its backers), align sufficiently with their own.
Legal Concerns
The plethora of legal concerns regarding the use of content produced by AI/LLMs ultimately
boil down to two key areas:
Licensing is built upon copyright. As a human creator of a piece of work, that work
is your Intellectual Property, and you are automatically assigned the copyright of
that work. An exception is that, if you are being compensated by a 3rd-party to create
that work (e.g., your employer), depending on the contract between yourself and that
3rd-party, the copyright may instead be vested with that 3rd-party.
Only the copyright owner(s) of a work may license that work for others to use. A license
is simply a set of terms granted by the copyright owner to one or more parties about
how they may use the copyrighted work. The copyright owner may license their work
as many times as they wish, under whatever terms they wish. There is no restriction
on them with regards to the terms of the licenses they may offer unless some other
contractual agreement to their benefit prevents that. The exception being the legal
doctrine of Fair Use, which allows constrained use of a copyrighted work without permission/licensing
from the copyright owner for the purposes of criticism, commentary, and teaching.
AI/LLMs are created by training them on vast amounts of information. This information
is often scraped from the Internet and includes web pages, artworks, music, video,
blog/forum posts, digitised books, journal articles, government documents, news articles,
and source code repositories amongst others. Each and every such piece of this information
that was used for training an AI/LLM has a pre-existing copyright, and a copyright
owner [Milmo25].
Use of this material to train an AI/LLM is subject to the same copyright law, and
potentially any licenses offered by the copyright holder, as any other copyrighted
work. Until very recently it was argued by many AI/LLM providers that their models
were entitled to use such copyrighted works under the legal doctrine of Fair Use,
whilst copyright owners argued that such training infringed their rights. Neither
argument, had been legally proven until in July 2026 in a US class action lawsuit,
Bartz, et al. v. Anthropic PBC [TCA26], a federal court reached a judgement that:
-
Books purchased by Anthropic to train their models constituted an “exceedingly transformative” approach under US copyright law, and therefore qualified under the legal doctrine
of Fair Use.
-
Books that were ‘pirated’, i.e., downloaded illegally, from the shadow libraries LibGen and PiLiMi to train
their models constituted copyright infringement; Anthropic reached a $1.5 billion
settlement with the class action members [JND26, Milmo26].
Following this judgement, for a provider to legally use copyrighted works for training
an AI/LLM it must at least ensure that firstly, the works have been legally obtained,
and secondly: (a) that the works have been licensed from the copyright owner for this
purpose, or (b) that the desired use of these works fall under the legal doctrine
of Fair Use. Whilst the initial driver for that legal case was focused on books, the
outcome (the legal judgement), was relevant to all copyrighted works in general.
This now makes the law around the use of copyrighted works to train an AI/LLM tangible,
and relatively clear.
Unfortunately, with regards to training an AI/LLM the concerns around licensing of
copyrighted works which builds atop copyright, remain unresolved at this time.
Many AI/LLMs that are aimed at assisting developers have been trained on vast amounts
of Open Source code that was acquired from public repositories such as GitHub [Chen21]. The code from these Open Source projects are copyrighted works, but further, many
of them declare an Open Source license of some variety. These Open Source licenses
impose restrictions on how these works may be used. Whilst there are many different
Open Source licenses, a common theme is that the original authors name, copyright
notice, and license declaration must be preserved. Furthermore, for some Copy-Left
classes of Open Source license, e.g., GNU (GNU’s Not Unix) GPL (General Public License) variants, it can mean that the same license terms have to be adopted Ad infinitum.
The open legal questions around the use of Open Source code to train an AI/LLM revolve
around, whether after training, any substantive code produced by an AI/LLM is in violation
of the DMCA (Digital Millennium Copyright Act). The code produced by an AI/LLM is not accompanied by the original CMI (Copyright Management Information), e.g., the original authors name, copyright notice, and license declaration. It
could be a violation of the DMCA to remove such CMI when reusing Open Source code.
The AI/LLM providers argue that the output of their services rarely, if ever, reproduce
code that they were trained upon, whereas the representatives for Open Source authors,
claim otherwise. This legal uncertainty has been working its way through the US courts
since 2022 in the form of the Doe vs. GitHub class action lawsuit [BH2026, Eslinger26].
An appellate ruling from the ninth court on the case is expected within the next six
months. However, that is likely not the conclusion, and the case could yet take longer
to resolve. The outcome of this case is significant for users of AI/LLMs, as it could
prove or reject, that code produced with the use of AI/LLMs is subject to one or more
original copyrights and set of license terms.
Should it be determined that the code produced by an AI/LLM is legally void of any
original copyrights or licenses used in its training data, then an important question
comes into play – Who owns the copyright on code produced by an AI/LLM?
This question arises because under current copyright law, a copyright owner can only
be a person, or by extension an organisation, it cannot be by an animal [Guadamuz18], or an AI/LLM [Lanquist26]. Whether copyright on code produced by an AI/LLM at the instruction of a human,
is vested with that human revolves around the concept of “meaningful human authorship”, which remains unquantifiable by the US Copyright Office, and is still currently
being tested within the US legal system [Evren2026].
For Open Source projects that wish to make use of AI/LLMs, the legal uncertainties
around the works (e.g., code) produced by AI/LLMs presents a number of significant
issues.
At this time, it is unclear whether such generated works, (a) carry existing copyright
and licenses, use of which may be a legal violation in itself, and/or in conflict
with the project’s own choice of Open Source license, or if not, (b) whether these
works are copyrightable by the human producer, and if not, by extension, they cannot
be licensed by the project to others under the project’s choice of Open Source license.
We argue that it would be eminently sensible for Open Source projects to establish
an AI policy, whereby contributions involving entirely or partially AI/LLM generated
works are clearly labelled as such. By this means, should the courts (a) decide against
AI/LLMs reproducing partial or complete copyright and/or licensed works from their
training models without CMI, or (b) continue to assert that human assisted AI generated
work cannot be copyrighted, then in future, projects can easily identify the contributions
in their projects that may need to be undone, or modified, for the project to become
compliant with current/new legislation. Should the courts instead decide in favour
of AI/LLMs, little has been lost by enacting such a policy as a safe guard.
Technical Concerns
AI/LLMs have demonstrated incredible feats of code generation when successfully directed,
and this can be very appealing to Open Source projects which may have a grand vision,
but lack the human resources to implement it. Early stage code generation and routine
tasks can often be performed faster by AI/LLMs than by humans; one such study found
that AI/LLMs were 31.4% faster than humans [Sankhe25]. Other industry articles on the subject have described developers using AI/LLMs
as a force multiplier, allowing them to be 10x more productive [Klenk25, Utley26].
The main technical concerns for Open Source projects employing AI/LLM generated code
fall into the following two main categories:
-
Code Quality
-
Project Maintainability
Neither of these categories are exclusive to code generated by AI/LLMs, indeed they
were equally applicable in the past when only humans were present. The difference
now, is the volume and speed at which code is produced.
For example, if a project previously had 10 contributors, each producing 1 contribution
per week. If 6 of these contributors decide to adopt an AI assisted workflow, then
the number of contributions the project receives per week could increase from 10,
to up to between 13 and 100. The use of AI/LLMs becomes a force multiplier across
the number of autonomous contributors operating in parallel.
Regarding code quality, code that is produced by AI/LLMs is often verbose, less than
optimal, and includes many repeated patterns rather than abstracting commonality into
constants, functions, or modules [Harding25, Marinho26]. Worse yet, it is not unusual for AI/LLMs to ‘hallucinate’ and produce nonsensical results [Suprmind26], this can lead to code that doesn’t produce the expected result or perhaps isn’t
even syntactically valid.
AI/LLMs as coding assistants are often described as having capabilities commensurate
with a human software developer at the ‘Junior Developer’ career level. AI/LLMs just like their human junior level developer counterparts,
naturally require very careful direction, mentoring, monitoring, review, and feedback,
of their work by a more experienced developer [Iyer26].
Long term project maintainability can be a time consuming task for an Open Source
project. Attention to this is essential to ensure the long term health of the project.
This can be helped by limiting technical debt before it occurs, and/or by addressing
and refactoring it out of the project at regular intervals. Arguably, the most important
mechanism to prevent the build-up of technical debt, is for a project to review, and
potentially request changes to, each submission before it is accepted into the project
[Alami19, Pathirage26].
In many places, AI/LLMs have allowed users without an education in Computer Science
and little or no formal Software Engineering experience to create and contribute to
software projects [Gama25]. This is an incredible feat of achievement for the producers of AI/LLMs, and should
be celebrated. When this works well, it can be very advantageous for Open Source projects
to receive contributions from such users. Unfortunately, there are too often downsides
to this situation for Open Source projects. Additionally some classes of project require
a great deal of rigour to ensure correctness and may not be suited to AI/LLM contributions,
for example:
-
systems that arrange, store, or validate essential data (e.g., schemas, filesystems,
and databases, etc.)
-
systems that manage or transmit highly sensitive information (e.g., electronic health
systems, government voting systems, secure communications, etc.)
-
critical real-time systems (e.g., autonomous vehicle control, military weapons, industrial
or medical control systems, etc.).
Users newly enabled by AI/LLM coding assistants, who are inexperienced developers,
may willingly or ignorantly, submit contributions to projects without understanding
exactly how their code works or what it does. This creates a burden on the human reviewers
for a project to understand each contribution, review it, and guide the submitter.
As the submitter will likely not understand the feedback from the reviewer, this can
often result in a cyclic discussion between the AI/LLM and the reviewer, with the
submitter simply acting as a copy/paste buffer between the two parties. This can introduce
frustrations on both sides, whereby the reviewer feels that they are having a disconnected
conversation with the AI/LLM, and the submitter who cares about their improvements
struggles to make progress. Terms such as “AI Proxy”, “AI Slopper” [Mourzenko26], and the rather more unpleasant “Meat Proxy” [Gruhn26], have been used, presumably as a result of frustration, to describe this variety
of interaction with code submitters. The authors of this paper do not condone these
terms; however, it is important to realise that such frustrations are real and have
an impact on the human reviewers of open source projects.
Whilst the barrier to contribution may be lowered, there may also be a lowering in
quality of submissions [CodeRabbit26]. Driven by the use AI/LLMs in one manner or the other, the volume of submissions
to projects is increasing [Iyer26]. Sometimes these contributions even come from entirely automated AI systems that
are tasked with scanning thousands of Open Source projects, and then identifying and
submitting fixes for various classes of bugs (e.g., security) [Dunn26]. Regardless, more submissions, means that more time is required by humans to review
submissions and either reject, request changes, or accept them. There are finite human
reviewers in any project, which can lead to the project and the reviewers becoming
overwhelmed, and thus spending less time on other key aspects of the project [LiveWyer26].
Some of the technical issues encountered in code generated by AI/LLM contributions
could potentially be lessened by better instructing contributors, including the AI/LLMs
themselves. Detailed documentation as to, the foundational technical architecture
and direction of the project, which technology choices have been made and why, if
those choices are open to re-evaluation, and a description of anything that is technically
unacceptable. Whilst AI/LLMs are quite capable of digesting documentation written
for humans, it can often be beneficial to provide documentation better tailored to
an AI/LLM (e.g., an AGENTS.md file) [Huet25]. In the authors’ experience, the larger technical landscape of an Open Source project
is rarely documented clearly.
Regarding the increased workload for human reviewers that can be caused by AI/LLMs,
again it may be helpful to provide clear documentation on the review criteria for
submissions. This is again equally relevant for both human and AI/LLM contributors.
Additionally, if there are computable quality factors, e.g.: test suites, code quality/analysis
tools that should be run, etc., these and their acceptance criteria should be documented.
Such documentation should again likely be tailored once for the human, and then again,
if applicable, for any AI/LLM audiences.
We argue that it is valuable for the AI policy of an Open Source project to include
documentation about: (a) the technical architecture of the project, (b) its code quality
standards, and (b) its code review policy and acceptance criteria for submissions.
Defining AI Policy for an Open Source Project
Due to the recent exponential advancements made in AI/LLM research, the current period
is one of great change in Software Engineering, this imparts a considerable amount
of uncertainty upon those involved, and as a by-product, stress. Therefore, to reduce
such stresses on both developers, we argue that clear communication of an organisation
or project’s use of AI policy is paramount.
For those employed directly by an organisation, due to the potential legal implications,
it is likely that a policy around the acceptable use of AI already exists (or could
be rapidly developed upon request).
However, for contributors to Open Source projects, individuals or organisations, whether
for fun or profit, the reality is much more complex. Currently, few Open Source projects
publish any information about their policy towards acceptable use of AI [Mendonça26,Holterhoff26]. Of those that do publish such a policy, notable projects include:
-
The Linux Kernel
AI/LLM assisted contributions are allowed but must labelled as such, and must be signed-off
by a human who ensures copyright and licensing compliance (with GPL-2.0) [Levin26].
-
The Fedora Linux Distribution
AI/LLM assisted contributions are allowed but must labelled as such, and must be signed-off
by a human who ensures copyright and licensing compliance. Furthermore use of AI/LLM
without human input in any code review process is forbidden [Brooks25].
-
The GCC (GNU Compiler Collection)
AI/LLM generated/assisted contributions are prohibited if they are legally significant
(more than 15 lines of code), unless they are for test cases. In any case any such
contributions must be labelled as such, and must be signed-off by a human whom ensures
copyright and licensing compliance [Wakely26].
-
LLVM (Low Level Virtual Machine)
Only AI/LLM assisted contributions are allowed; there must always be a human in the
loop. Pull Request descriptions must be written by humans, and code-review feedback
should not be fed back to AI/LLMs but instead addressed by humans. AI/LLM assisted
contributions should be labelled as such. The human contributor must ensure the copyright
and licensing compliance of the entire contribution [Kleckner25].
-
OpenJDK
AI/LLM contributions, either standalone or human assisted, are prohibited [ORACLE26].
-
NetBSD
AI/LLM contributions, either standalone or human assisted, are prohibited [Campbell24].
-
Codeberg
Unlike the projects listed above, Codeberg is a project that enables hosting of other
Open Source projects. Codeberg’s current informal guidelines discourage its use for
projects that are created, developed, or maintained by AI/LLMs, or those projects
which make heavy use of AI/LLMs [Tzovaras26].
An interesting theme in some of these policies, is where AI/LLM creations or assisted
contributions are allowed, but a human contributor is made responsible for, “Sign
Off”, i.e., ensuring the copyright and licensing compliance of the entire contribution.
These projects on the one-hand signal that they are open to AI/LLM contributions;
however, on the other hand they are clearly aware of the copyright and licensing legal
implications that these may bring, and so they attempt to push that burden onto the
human contributor. However, as discussed in section “Legal Concerns”, as the case-law is still being established with regards to whether the copyrighted
and/or licensed data used to train AI/LLMs is also produced in its output, it seems
implausible at this time that any human or organisation can sign-off on AI/LLM contributions
with any legal certainty as to their position. At this point in time, solely from
a legal perspective, the position of the OpenJDK and NetBSD projects would seem most
prudent. From a social community perspective, it yet remains to be seen whether this
will encourage or discourage contributions to these projects.
Before potential contributors or users invest their time into an Open Source project,
due to the plethora or concerns around the use of AI/LLMs, it would then seem important
for them to comprehend how AI may or may not be used in a particular project. Therefrom,
whether you are pro or anti AI/LLM, we argue that it is key for Open Source projects
to clearly document and communicate their policy around the acceptable use of AI.
As a result, users and/or contributors can then understand if the project aligns with
their own set of goals and ethics. They can then make an informed choice as to whether
they wish to interact with the project.
Components of an AI Policy
Our research, specifically in section “Concerns for AI in Open Source”, has highlighted a number of nuanced areas where humans may have motivations for
or against the use of AI in Open Source projects.
As humans, clear communication is paramount, our knowledge and reasoning is built
upon context, and as such we often seek justification to reach an understanding. It
would therefore seem important not to just have a position on the use of AI/LLMs,
but as a project hoping to attract others, it would seem reasonable, that as part
of its outreach, the project can justify its position.
Therefore, we argue that an Open Source project whose AI policy might simply state
that: ‘the use of AI/LLMs is allowed’, is lacking. Such a policy should be augmented with additional components, that likely
include:
-
Guidance around acceptable use:
-
What AI/LLM tools and models may be used?
-
Where may AI/LLMs be used in the project? E.g., Writing code and/or documentation,
language translations, reviewing contributions, etc.
-
How AI/LLMs may be used in the project? E.g., Do AI/LLM contributions need to be identified
as such and attributed, and if so, how?
-
Details of any measures that have been put in-place to ensure inclusion for all who
wish to contribute.
-
If the project is more than a one-person undertaking, then how consensus was reached
on the AI policy.
-
Is there any financial funding or sponsorship available for those that may not be
able to afford to use AI/LLMs?
-
Are there any policies on discrimination between humans and AI/LLMs?
-
Is there any activity on identifying or offsetting the environmental cost of the use
of AI/LLMs by the project?
-
Whom to contact in the case of concerns around the AI policy, and/or the use of AI/LLMs.
Conversely, we also argue that, an Open Source project whose AI policy simply states:
‘any use of AI/LLMs is forbidden’, may be perceived as having not considered the advantages of such technology, and
would likely benefit from a clear and succinct statement as to their concerns.
An AI Policy Score Card for Open Source Projects
Inspired by the ‘5-star deployment scheme for Open Data’ [BernersLee10], we have developed a Score Card for AI Policy in Open Source projects. We believe
that this Score Card can help Open Source projects gauge the quality of their policy
towards their terms for acceptable use of AI.
Our Score Card has 5 levels, and each level builds upon the previous level. That is
to say, for example, that, to achieve a ‘2 Star’ rating, you must have addressed the concerns of both the ‘1 Star’ and ‘2 Star’ levels.
1 Star Rating
The Open Source project has an AI policy.
This policy considers human factors and was developed though consensus of the project
stakeholders.
Guidance: Think about what constitutes acceptable use of AI in your project. Consider,
amongst others, any of the applicable concerns raised in section “Concerns for AI in Open Source”. Consider the components that might be included in your policy, as discussed in section “Components of an AI Policy”
2 Star Rating
Fully or partially machine generated contributions to the Open Source project are
clearly labelled, and remain identifiable as such in future.
Guidance: To safeguard against potential copyright/licensing issues, or human vs.
machine assumptions, consider adopting a DCO (Developer Certificate of Origin) and documenting the AI/LLM tools and models used within each contribution (e.g.,
code commit). If you reject AI/LLM contributions to your project, then simply ensure
that all human contributions are attributable; this may likely already be the case.
3 Star Rating
The Open Source project’s AI policy is documented for humans.
It should be published publicly in an open fashion on the Web, and be easily locatable
by humans.
Guidance: Write this in natural language aimed at a semi-technical human audience.
Include a copy in your project’s source code repository (e.g., an AI_POLICY.md file),
and consider publishing it on your project’s website (e.g., https://example.org/ai-policy, or https://example.org/ai-policy.html).
5 Star Rating
The Open Source project has technical contributor documentation.
It should be linked from the AI policy document(s), and be written to target humans
and/or machines as applicable.
Guidance: Write this twice (if applicable), once aimed at a semi-technical human audience,
and secondly in a form that is best interpreted by AI/LLMs. It should compliment your
AI policy but standalone from it. It should include technical aspects of contributing
to the project, e.g., code and quality standards, applicable tools, code/contribution
review policy, and any contribution acceptance criteria.
Conclusion
Success of an Open Source project may be judged on a number of factors. For some projects,
a measure of success may be as simple as the enjoyment gained from those participating
in it, or perhaps the everyday personal utility that it may yield. For projects with
a larger reach, measures of success often centre around (a) their community, where
the number of participants, and how welcoming they are, can be important factors,
and (b) the impact and/or adoption of the produced software itself, how many users
it has and how valuable the software is to them.
Regardless of the purpose of an Open Source project, participation has historically
been a solely human endeavour. As we move into the future, the recent meteoric rise
of AI/LLMs is changing that narrative.
In section “Concerns for AI in Open Source”, from our research we have highlighted a number of important concerns that may arise
for Open Source projects where humans and AI/LLMs are interacting (or not). Rather
than focusing on only the technology itself, we have taken a holistic view, and also
examined, environmental, economic, legal, and political concerns. Furthermore, we
argue that for Open Source to be a continued success, i.e., attract users and contributors,
projects should clearly document their AI policy. This enables both producers and
consumers of Open Source to make informed decisions about their involvement (if any).
At this juncture, it is perhaps worth stating that whilst the concerns that we have
raised are undoubtedly significant, the authors of this paper have attempted to remain
objective and adopt a position of pro-choice. Our concern is that Open Source projects
should clearly document their AI policy for the purpose of empowering both humans
and machines. It is not our position to prescribe whether any project should, or should
not, use AI/LLMs.
In section “Defining AI Policy for an Open Source Project”, we have both highlighted some existing policies of Open Source projects and contributed
guidance as to what should be addressed in an AI policy of an Open Source project.
Finally, in section “An AI Policy Score Card for Open Source Projects”, for those projects that do wish to develop and publish an AI policy, we contribute
a simple system to allow an Open Source project to gauge the quality and completeness
of their AI policy.
Ultimately, it remains clear that the use of AI/LLMs within Open Source projects comes
with both many advantages and disadvantages. We strive only to make the picture clearer,
and we look forward to seeing how both Open Source and AI/LLMs progress.
References
[Wizenbaum66] Joseph Wizenbaum, ELIZA—a computer program for the study of natural language communication between man
and machine
, 1966. ACM Communications of the ACM. doi:https://doi.org/10.1145/365153.365168, https://dl.acm.org/doi/10.1145/365153.365168
[Vaswani17] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,
Łukasz Kaiser, and Illia Polosukhin, Attention is All you Need
, 2017. Advances in Neural Information Processing Systems. arXiv:1706.03762, doi:https://doi.org/10.48550/arXiv.1706.03762, https://arxiv.org/abs/1706.03762
[Chase22] Harrison Chase, LangChain - The agent engineering platform, 2022. https://github.com/langchain-ai/langchain
[Oberhauser19] Jan Oberhauser, n8n - Secure Workflow Automation for Technical Teams, 2019. https://github.com/n8n-io/n8n
[Stallman98] Richard Stallman, Why Open Source Misses the Point of Free Software
, 1998. https://www.gnu.org/philosophy/open-source-misses-the-point.html
[Perens98] Bruce Perens, The Open Source Definition, 1998. https://opensource.org/osd
[GiovanH21] GiovanH, Ethical Source is Hot Garbage, 2021. https://blog.giovanh.com/blog/2021/10/29/ethical-source-is-hot-garbage/
[Peterson13] Kevin Peterson, The GitHub Open Source Development Process
, 2013. Mayo Clinic.
[AlMarzouq22] Mohammad AlMarzouq, Abdullatif AlZaidan, and Jehad Al Dallal, The Relevance of SourceForge Data in the Age of GitHub
, 2022. ACM SIGMIS Database. doi:https://doi.org/10.1145/3571823.3571830
[Ehmke25] Coraline Ada Ehmke, Contributor Covenant 3.0 Code of Conduct, 2025. https://www.contributor-covenant.org/version/3/0/code_of_conduct/
[Koehler18] Christie Koehler, Citizen Code of Conduct, 2018. https://github.com/stumpsyn/policies/blob/master/citizen_code_of_conduct.md
[Webb20] Webb, The Anticode of Conduct. https://git.sr.ht/~webb/anticode/tree/master/item/ANTICODE.md
[WIKI01] Chatbot Psychosis
, 2026. Wikipedia. https://en.wikipedia.org/wiki/Chatbot_psychosis
[Østergaard25] Søren Dinesen Østergaard, Generative Artificial Intelligence Chatbots and Delusions: From Guesswork to Emerging
Cases
, 2025. Acta Psychiatrica Scandinavica. doi:https://doi.org/10.1111/acps.70022, https://onlinelibrary.wiley.com/doi/10.1111/acps.70022
[Morrin26] Hamilton Morrin, Luke Nicholls, Michael Levin et al., Artificial intelligence-associated delusions and large language models: risks, mechanisms
of delusion co-creation, and safeguarding strategies
, 2026. The Lancet Psychiatry. PII S2215-0366(25)00396-7, doi:https://doi.org/10.1016/S2215-0366(25)00396-7, https://www.thelancet.com/journals/lanpsy/article/PIIS2215-0366(25)00396-7/abstract
[Girgis26] Dr Ragy Girgis, What is AI Psychosis? A Conversation on Chatbots and Mental Health, 2026. National Academy of Medicine. https://nam.edu/news-and-insights/what-is-ai-psychosis/
[Cheng25] Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, and Dan Jurafsky, Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence
, 2025. arXiv:2510.01395. doi:https://doi.org/10.48550/arXiv.2510.01395, https://arxiv.org/abs/2510.01395
[THLP25] The Human Line Project, 2025. https://www.thehumanlineproject.org/
[Moore26] Anna Moore, Marriage over, €100,000 down the drain: the AI users whose lives were wrecked by delusion
, 2026. The Guardian. https://www.theguardian.com/lifeandstyle/2026/mar/26/ai-chatbot-users-lives-wrecked-by-delusion
[Proven26/2] Liam Proven, Bcachefs creator insists his custom LLM is female and ‘fully conscious’
, 2026. The Register. https://www.theregister.com/software/2026/02/25/bcachefs-creator-claims-his-custom-llm-is-fully-conscious/4671792
[ASF26] Apache Software Foundation, How the ASF works, 2026. https://www.apache.org/foundation/how-it-works/
[Burcher14] Richard Burcher, Starting a Project at the Eclipse Foundation, 2014. https://www.eclipse.org/community/eclipse_newsletter/2014/july/article2.php
[TLF26] The Linux Foundation, Host a project, 2026. https://www.linuxfoundation.org/projects/hosting
[OS21] Igor Steinmacher, Georg Link, Anita Sarma, Gregorio Robles, Bianca Trinkenreich, Christoph
Treude, Marco Gerosa, and Igor Wiese, What motivates open source software contributors?, 2021. https://opensource.com/article/21/4/motivates-open-source-contributors
[Blates26] Sebastian Baltes, Marc Cheong, and Christoph Treude, ‘An Endless Stream of AI Slop’: How Developers Discuss the Burden of AI-Assisted Software
Development
, 2026. arXiv:2603.27249. doi:https://doi.org/10.48550/arXiv.2603.27249, https://arxiv.org/abs/2603.27249
[Karim26] S M Rakib UI Karim, Wenyi Lu, and Sean Goggins, Artificial Intelligence in Open Source Software Engineering: A Foundation for Sustainability
, 2026. arXiv:2602.07071. doi:https://doi.org/10.48550/arXiv.2602.07071, https://arxiv.org/pdf/2602.07071
[Zheyuan25] Kevin Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, and Tobias
Salz, The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments
with Software Developers
, 2025. MIT Economics. https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf. See also published work in Management Science, 2026. doi:https://doi.org/10.1287/mnsc.2025.00535
[Edwards26] Rob Edwards, and Richard Appiah, Developer Productivity in the Age of Generative AI: A Psychological Perspective, 2026. Google Research. https://research.google/pubs/developer-productivity-in-the-age-of-generative-ai-a-psychological-perspective/
[LinearB26] Software Engineering Benchmarks Report ’26 - The AI Productivity Edition, 2026. https://assets.linearb.io/image/upload/v1777392920/resources/LinearB_2026_Software_Engineering_Benchmarks_Report.pdf
[Song23] Fangchen Song, Ashish Agarwal, and Wen Wen, The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence
from GitHub Copilot
, 2023. Social Science Research Network. doi:https://doi.org/10.2139/ssrn.4856935, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4856935
[Lițan25] Daniela-Elena Litan, Mental health in the ‘era’ of artificial intelligence: technostress and the perceived
impact on anxiety and depressive disorders—an SEM analysis
, 2025. Front. Psychol. doi:https://doi.org/10.3389/fpsyg.2025.1600013. See also https://pmc.ncbi.nlm.nih.gov/articles/PMC12169247/
[Spirlet26] Thibault Spirlet, Software engineers are facing an ‘identity crisis bordering on depression,’ Menlo
Ventures partner says
, 2026. The Business Insider. https://www.businessinsider.com/software-engineers-face-an-ai-identity-crisis-vc-partner-says-2026-6
[Orosz26] Gergely Orosz, The grief when AI writes most of the code, 2026. The Pragmatic Engineer. https://blog.pragmaticengineer.com/the-grief-when-ai-writes-most-of-the-code/
[Lawson26] Nolan Lawson, We mourn our craft, 2026. https://nolanlawson.com/2026/02/07/we-mourn-our-craft/
[Kwon26] Heesung Kwon, Jeesun Oh, Suyoun Lee, Sunok Lee, and Sangsu Lee, Investigating AI-induced Technostress and Coping Strategies of Professionals
, 2026. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI
’26). doi:https://doi.org/10.1145/3772318.3791671, https://dl.acm.org/doi/10.1145/3772318.3791671
[Mahdawi25] Arwa Mahdawi, Meet the AI vegans
, 2025. The Guardian. https://www.theguardian.com/commentisfree/2025/aug/06/meet-the-ai-vegans
[Proven26/1] Liam Proven, Struggling to put your AI aversion into words? Here’s a handy glossary
, 2026. The Register. https://www.theregister.com/2026/03/19/ai_skeptic_labels/
[Sawers25] Catherine Sawers, i included this in my syllabi this year, 2025. Bluesky. https://bsky.app/profile/catebridget.bsky.social/post/3mcxosc7c322t
[Pscheidt09] Markus Pscheid and Theo P. van der Weide, Bridging the Digital Divide by Open Source - A theoretical model of best practice
, 2009. International Journal of Innovation in the Digital Economy. doi:https://doi.org/10.4018/jide.2010040103
[LOS24] Living Open Source Foundation, Impact of Open Source in Developing Countries, 2024. https://livingopensource.org/impact-of-open-source-in-countries/
[Blind24] Knut Blind and Torben Schubert, Estimating the GDP effect of Open Source Software and its complementarities with R&D
and patents: evidence and policy implications
, 2024. The Journal of Technology Transfer. doi:https://doi.org/10.1007/s10961-023-09993-x
[Sachs24] Goldman Sachs, Gen AI: Too Much Spend, Too Little Benefit?, 2024. https://www.goldmansachs.com/insights/top-of-mind/gen-ai-too-much-spend-too-little-benefit
[Aggarwal26] Gaurav Aggarwal, We will enjoy cheap AI coding assistants while they last, 2026. Linkedin. https://www.linkedin.com/posts/gauagg_we-will-enjoy-cheap-ai-coding-assistants-activity-7463460994980306944-LYRR/
[Ivchenko26] Oleh Ivchenko, Cost-Effective AI: The Hidden Costs of “Free” Open Source AI - What Nobody Tells You, 2026. Stabilarity Hub. https://hub.stabilarity.com/cost-effective-ai-the-hidden-costs-of-free-open-source-ai-what-nobody-tells-you/
[Fang26] Zihan Fang, Yueke Zhang, Thomas Zimmermann, Denae Ford, and Yu Huang, Contribution Patterns in Open Source Software for Social Good: Dynamics, Individuals,
and Impact
, 2026. Proceedings of the ACM on Human-Computer Interaction. doi:https://doi.org/10.1145/3788046
[AIEC25] AI Hardware Team, AI Hardware Sustainability: The Environmental Cost of GPUs and TPUs, 2025. AI Energy Calculator. https://aienergycalculator.com/ai-hardware-environmental-impact-sustainability/
[Luscombe26] Richard Luscombe, Wyoming tightens wastewater rules after Meta datacenter contractor flushed contaminated
water
, 2026. The Guardian. https://www.theguardian.com/us-news/2026/jul/08/meta-datacenter-ai-wyoming-water
[Robinson26] Dan Robinson, Google meets the neighbors and gets both barrels over its new UK datacenter
, 2026. The Register. https://www.theregister.com/on-prem/2026/07/24/google-meets-the-neighbors-and-gets-both-barrels-over-its-new-uk-datacenter/5277188
[Strubell19] Emma Strubell, Ananya Ganesh, and Andrew McCallum, Energy and Policy Considerations for Deep Learning in NLP
, 2019. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. doi:https://doi.org/10.18653/v1/P19-1355, https://aclanthology.org/P19-1355/
[Aczel26] Mariam Aczel, Sanaz Chamanara, Mir Matin, Aria Farsi, Tshilidzi Marwala, and Kaveh
Madani, Environmental Cost of AI’s Energy Use, 2026. United Nations University - Institute for Water, Environment, and Health. doi:https://doi.org/10.53328/INR26RMA002, https://collections.unu.edu/eserv/UNU:10647/UNU-INWEH-Report-The_Env_Cost_of_AI-2026.pdf
[WIKI02] International Traffic in Arms Regulations
. Wikipedia. https://en.wikipedia.org/wiki/International_Traffic_in_Arms_Regulations
[EAR96] Bureau of Export Administration, Export Administration Regulation - Simplification of Export Administration Regulation
, 1996. Federal Register. https://www.govinfo.gov/content/pkg/FR-1996-03-25/pdf/96-4173.pdf
[BIS22] Bureau of Industry and Security, U.S. Department of Commerce, Implementation of Additional Export Controls: Certain Advanced Computing and Semiconductor
Manufacturing Items; Supercomputer and Semiconductor End Use; Entity List Modification
, 2022. Federal Register. https://www.federalregister.gov/documents/2022/10/13/2022-21658/implementation-of-additional-export-controls-certain-advanced-computing-and-semiconductor
[Curi26] Maria Curi, The secret Trump administration battle to fight Chinese AI
, 2026. Axios. https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi
[NDAA26] US Congress, National Defense Authorization Act for Fiscal Year 2026, Public Law No. 119-60, 2025. U.S. Government Publishing Office. https://www.congress.gov/bill/119th-congress/senate-bill/1071/text
[EU23] European Parliament, EU AI Act: first regulation on artificial intelligence, 2023. European Parliament. https://www.europarl.europa.eu/topics/en/article/20230601STO93804/eu-ai-act-first-regulation-on-artificial-intelligence
[WIKI03] Cyberspace Administration of China, Interim Measures for the Management of Generative AI Services
, 2023. Wikipedia. https://en.wikipedia.org/wiki/Interim_Measures_for_the_Management_of_Generative_AI_Services
[Lieu26] AI Kill Switch Act, H.R. 9917 (proposed bill), 2026. US Congress (introduced by Reps. Ted W. Lieu and
Nathaniel Moran). https://lieu.house.gov/sites/evo-subsites/lieu-evo.house.gov/files/evo-media-document/ai-kill-switch-act.pdf. See also https://www.congress.gov/bill/119th-congress/house-bill/9917
[Howard11] Philip N. Howard, Sheetal D. Agarwal, and Muzammil M. Hussain, The Dictators’ Digital Dilemma: When Do States Disconnect Their Digital Networks?
, 2011. Issues in Technology Innovation. The Center for Technology Innovation at Brookings. doi:https://doi.org/10.2139/ssrn.2568619
[Stremlau24] Nicole Stremlau, Internet Shutdowns, Sovereignty, and the Postcolonial State in Africa
, 2024. Global Policy Journal. doi:https://doi.org/10.1111/1758-5899.13483, https://onlinelibrary.wiley.com/doi/full/10.1111/1758-5899.13483?campaign=wolearlyview
[Chesterman26] Simon Chesterman, AI is Giving Tech Companies Power That Once Belonged to Governments, 2026. https://restofworld.org/2026/ai-government-regulation-tech-giants/
[Tolomia26] Chris Tolomia, OpenAI and Anthropic are breaking their own lobbying records as IPOs loom
, 2026. Quartz. https://qz.com/openai-anthropic-lobbying-records-q2-2026-072126
[Williams26] Kylie Williams, The data center and AI donors powering Donalds’ bid for Florida governor
, 2026. Politico. https://www.politico.com/news/2026/07/20/byron-donalds-florida-ai-data-center-fundraising-01005141
[Wilkins26] Emily Wilkins, What AI companies want for the millions they’re spending on elections, 2026. CNBC. https://www.cnbc.com/2026/07/09/ai-companies-election-spending.html
[Dorn26] Sara Dorn, Musk Reportedly Spending Up To $120 Million Helping GOP In Midterms—After Saying He
‘Got A Little Too Involved In Politics’
, 2026. Forbes. https://www.forbes.com/sites/saradorn/2026/07/30/musk-reportedly-spending-up-to-120-million-helping-gop-in-midterms-after-saying-he-got-a-little-too-involved-in-politics/
[Milmo25] Dan Milmo, ‘Impossible’ to create AI tools like ChatGPT without copyrighted material, OpenAI
says
, 2025. The Guardian. https://www.theguardian.com/technology/2024/jan/08/ai-tools-chatgpt-copyrighted-material-openai
[TCA26] Top Class Actions, $1.5B Anthropic settlement resolves AI training lawsuit
, 2026. https://topclassactions.com/lawsuit-settlements/lawsuit-news/1-5b-anthropic-settlement-resolves-ai-training-lawsuit/
[JND26] JND Legal Administration, Welcome to the Anthropic Copyright Settlement Website. https://www.anthropiccopyrightsettlement.com/
[Milmo26] Dan Milmo and Mark Sweney, Harry Potter publisher to receive millions in Anthropic copyright settlement
, 2026. The Guardian. https://www.theguardian.com/technology/2026/jul/22/bloomsbury-book-publisher-anthropic-copyright-settlement
[Chen21] Mark Chen et al., Evaluating Large Language Models Trained on Code
, 2021. arXiv:2107.03374. doi:https://doi.org/10.48550/arXiv.2107.03374, https://arxiv.org/abs/2107.03374
[BH2026] Baker & Hostetler, LLP, Doe v. GitHub, Inc., 2026. www.bakerlaw.com/the-copilot-litigation/
[Eslinger26] Bonnie Eslinger, 9th Circ. Mulls DMCA Claim Against Microsoft And OpenAI, 2026. Law360. https://www.law360.com/articles/2440761/9th-circ-mulls-dmca-claim-against-microsoft-and-openai
[Guadamuz18] Andres Guadamuz, Can the monkey selfie case teach us anything about copyright law?
, 2018. WIPO Magazine. https://www.wipo.int/en/web/wipo-magazine/articles/can-the-monkey-selfie-case-teach-us-anything-about-copyright-law-40287
[Lanquist26] Edward D. Lanquist, Benjamin West Janke, Dominic Rota, and Lesli Harris, Supreme Court Denies Certiorari in Thaler v. Perlmutter: AI Cannot Be an Author Under
the Copyright Act, 2026. Baker, Donelson, Bearman, Caldwell & Berkowitz, PC. https://www.bakerdonelson.com/supreme-court-denies-certiorari-in-thaler-v-perlmutter-ai-cannot-be-an-author-under-the-copyright-act
[Evren2026] Sena Evren, Who Owns the Code Claude Wrote?, 2026. Legal Layer. https://legallayer.substack.com/p/who-owns-the-claude-code-wrote
[Sankhe25] Purvi Sankhe, Neeta Patil, Minakshi Ghorpade, Pratibha Prasad, and Monisha Linkesh,
Empirical Analysis of AI-Assisted Code Generation Tools on Code Quality, Security,
and Developer Productivity
, 2025. International Journal for Multidisiplinary Research. doi:https://doi.org/10.36948/ijfmr.2025.v07i06.61350
[Klenk25] Mathias Klenk, Rethinking the 10x Engineer: From Developer to Force Multiplier in the Age of AI, 2025. Tech Founder Stack. https://www.techfounderstack.com/p/rethinking-the-10x-engineer-from
[Utley26] Jeremy Utley, The AI Multiplier: Why Your Organic Capabilities
Matter More Than Ever, 2026. https://www.jeremyutley.com/blog/the-ai-multiplier
[Harding25] William Harding, AI Copilot Code Quality - Evaluating 2024’s Increased Defect Ratevia Code Quality
Metrics, 2025. GitClear AI Code Quality Research. https://gitclear-public.s3.us-west-2.amazonaws.com/GitClear-AI-Copilot-Code-Quality-2025.pdf
[Marinho26] Renato Marinho, The Problem with AI-Generated 'Shadow' Duplication, 2026. DEV Community. https://dev.to/renato_marinho/the-problem-with-ai-generated-shadow-duplication-2gi3
[Pharaoh26] Pharaoh, How to Prevent Duplicate Functions in AI Coding Workflows
, 2026. Medium.
https://medium.com/@usepharaoh/how-to-prevent-duplicate-functions-in-ai-coding-workflows-797bba3c87c0
[Suprmind26] Suprmind, AI Hallucination Rates,Statistics & Benchmarks in 2026, 2026. https://suprmind.ai/hub/ai-hallucination-rates-and-benchmarks/
[Iyer26] Arjun Iyer, Open source maintainers are drowning in AI-generated pull requests
, 2026. The New Stack. https://thenewstack.io/ai-generated-code-crisis/
[Alami19] Adam Alami, Marisa Leavitt Cohn, and Andrzej Wąsowski, Why Does Code Review Work for Open Source Software Communities?
, 2019. IEEE 41st International Conference on Software Engineering. doi:https://doi.org/10.1109/ICSE.2019.00111, https://ieeexplore.ieee.org/document/8812037
[Pathirage26] Anupama Pathirage, Why Code Reviews Matter in Open Source - And How to Make Them a Daily Habit
, 2026. Medium. https://medium.com/building-tech-teams/why-code-reviews-matter-in-open-source-and-how-to-make-them-a-daily-habit-4c74961494b4
[Gama25] Kiev Gama, Filipe Calegario, Victoria Jackson, Alexander Nolte, Luiz Augusto Morais,
and Vinicius Garcia, ‘Can you feel the vibes?’: An exploration of novice programmer engagement with vibe
coding
, 2025. arXiv:2512.02750v1. doi:https://doi.org/10.48550/arXiv.2512.02750, https://arxiv.org/html/2512.02750v1
[Mourzenko26] Arseni Mourzenko, How to deal with a programmer who acts as a proxy for AI?, 2026. Software Engineering. https://softwareengineering.stackexchange.com/questions/460875/how-to-deal-with-a-programmer-who-acts-as-a-proxy-for-ai
[Gruhn26] Niklas Gruhn, Don’t be a meat proxy, 2026. https://gruhn.me/blog/2026-08-03/
[CodeRabbit26] CodeRabbit, State of AI vs.Human CodeGeneration Report, 2026. CodeRabbit. https://www.coderabbit.ai/content/assets/code-rabbit-state-of-ai-vs-human-code-generation-report-lite.pdf
[Dunn26] John E. Dunn, Open source maintainers are being targeted by AI agent as part of ‘reputation farming’
, 2026. InfoWorld. https://www.infoworld.com/article/4132851/open-source-maintainers-are-being-targeted-by-ai-agent-as-part-of-reputation-farming.html
[LiveWyer26] LiveWyer, AI Disruption to Open Source Software (OSS)
, 2026. Medium. https://medium.com/@livewyer/ai-disruption-to-open-source-software-oss-377f10be2d8a
[Huet25] Romain Huet, AGENTS.md - a simple, open format for guiding coding agents, 2025. https://github.com/agentsmd/agents.md
[Mendonça26] Melissa Weber Mendonça, Open Source AI Contribution Policies, 2026. https://github.com/melissawm/open-source-ai-contribution-policies
[Holterhoff26] Kate Holterhoff, The Generative AI Policy Landscape in Open Source, 2026. RedMonk. https://redmonk.com/kholterhoff/2026/02/26/generative-ai-policy-landscape-in-open-source/
[Levin26] Sasha Levin and Willy Tarreau, AI Coding Assistants, 2026. https://docs.kernel.org/process/coding-assistants.html
[Brooks25] Jason Brooks, AI-Assisted Contributions Policy, 2025. Fedora Council. https://docs.fedoraproject.org/en-US/council/policy/ai-contribution-policy/
[Wakely26] Jonathan Wakely, GNU Compiler Collection - AI Policy, 2026. https://gcc.gnu.org/ai-policy.html
[Kleckner25] Reid Kleckner, Hubert Tong, and Marco Falke, LLVM AI Tool Use Policy, 2025. https://llvm.org/docs/AIToolPolicy.html
[ORACLE26] Oracle, OpenJDK Interim Policy on Generative AI, 2026. https://openjdk.org/legal/ai
[Campbell24] Taylor R. Campbell, NetBSD Commit Guidelines, 2024. https://www.netbsd.org/developers/commit-guidelines.html
[Tzovaras26] Bastian Greshake Tzovaras, Otto Richter, and William Zijl, Protecting our FLOSS commons from LLMs, 2026. Codeberg e.V.. https://blog.codeberg.org/protecting-our-floss-commons-from-llms.html
[BernersLee10] Tim Berners-Lee, 5-Star Deployment Scheme for Linked Data, 2010. https://www.w3.org/2011/gld/wiki/5_Star_Linked_Data
[Cardillo26] Kayla Cardillo, AI.TXT: A Declaration File for AI Usage Preferences, Licensing, and Policy, 2026. IETF Internal Draft. https://datatracker.ietf.org/doc/draft-car-ai-txt-wellknown/
[Howard26] Jeremy Howard, The /llms.txt file, v2, 2026. llms-txt. https://llmstxt.org
[Nottingham19] M. Nottingham, RFC 8615: Well-Known Uniform Resource Identifiers (URIs), 2019. doi:https://doi.org/10.17487/RFC8615, https://www.rfc-editor.org/info/rfc8615/
×Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,
Łukasz Kaiser, and Illia Polosukhin, Attention is All you Need
, 2017. Advances in Neural Information Processing Systems. arXiv:1706.03762, doi:https://doi.org/10.48550/arXiv.1706.03762, https://arxiv.org/abs/1706.03762
×Kevin Peterson, The GitHub Open Source Development Process
, 2013. Mayo Clinic.
×Markus Pscheid and Theo P. van der Weide, Bridging the Digital Divide by Open Source - A theoretical model of best practice
, 2009. International Journal of Innovation in the Digital Economy. doi:https://doi.org/10.4018/jide.2010040103
×Knut Blind and Torben Schubert, Estimating the GDP effect of Open Source Software and its complementarities with R&D
and patents: evidence and policy implications
, 2024. The Journal of Technology Transfer. doi:https://doi.org/10.1007/s10961-023-09993-x
×Zihan Fang, Yueke Zhang, Thomas Zimmermann, Denae Ford, and Yu Huang, Contribution Patterns in Open Source Software for Social Good: Dynamics, Individuals,
and Impact
, 2026. Proceedings of the ACM on Human-Computer Interaction. doi:https://doi.org/10.1145/3788046
×Mariam Aczel, Sanaz Chamanara, Mir Matin, Aria Farsi, Tshilidzi Marwala, and Kaveh
Madani, Environmental Cost of AI’s Energy Use, 2026. United Nations University - Institute for Water, Environment, and Health. doi:https://doi.org/10.53328/INR26RMA002, https://collections.unu.edu/eserv/UNU:10647/UNU-INWEH-Report-The_Env_Cost_of_AI-2026.pdf
×Philip N. Howard, Sheetal D. Agarwal, and Muzammil M. Hussain, The Dictators’ Digital Dilemma: When Do States Disconnect Their Digital Networks?
, 2011. Issues in Technology Innovation. The Center for Technology Innovation at Brookings. doi:https://doi.org/10.2139/ssrn.2568619
×Purvi Sankhe, Neeta Patil, Minakshi Ghorpade, Pratibha Prasad, and Monisha Linkesh,
Empirical Analysis of AI-Assisted Code Generation Tools on Code Quality, Security,
and Developer Productivity
, 2025. International Journal for Multidisiplinary Research. doi:https://doi.org/10.36948/ijfmr.2025.v07i06.61350