Half of the SKILL.md files on GitHub are verbatim copies of another file, according to a paper examining how instructions for AI coding agents spread between repositories. The study finds no registry, versioning system or provenance mechanism to track those copies and deliver updates.

Fahd Seddik at the University of British Columbia, Okanagan, submitted the paper on October 8, 2026. Its GitSkills dataset includes every SKILL.md committed to GitHub from October 2025 through July 2026.

During that ten-month period, the format reached 259,596 repositories and included 1,612,846 distinct skills. The researchers reconstructed 2,193,119 skill adoptions from Git history to map relationships between original files and their copies.

Updates rarely reach downstream copies

The findings point to a supply-chain problem for AI coding agents. SKILL.md files contain instructions executed by these agents, but the copy-based format does not provide a built-in way to push a correction from an original file to repositories that adopted it.

Only 11.1% of changes made to a source skill reached all copies during each observation window. Across the copy genealogies examined, only 4.0% ever changed consistently.

That pattern makes an unchanged copy the expected outcome when an upstream skill is corrected. A repository may retain an earlier version without a visible connection to the source or a dependable update path.

The paper describes this as an unversioned supply chain. Without a registry or provenance system, users and platforms lack a consistent record of where a skill originated, which repositories copied it or whether downstream files reflect later changes.

Star counts miss influential repositories

The study also compared conventional popularity signals with an influence-based ranking. Reviewing the 100 most-starred repositories would have intercepted 0.5% of later high-risk skill adoptions.

The influence-based ranking performed better, intercepting 14.9% of those later high-risk adoptions. The difference suggests that repository stars do not identify the repositories that play the most important role in skill propagation.

The paper therefore recommends that platforms distribute versioned references instead of copies. That approach would separate adoption from duplication and could provide a clearer route for tracking revisions, although the supplied findings do not report an implemented registry or fix system.

The research offers the first empirical map described in the material of how SKILL.md files propagate on GitHub. Its central finding is not simply that duplication is widespread, but that copied instructions can remain disconnected from later changes.

Conclusion

The study finds that GitHub's AI agent skill ecosystem has spread across hundreds of thousands of repositories without a registry, versioning or provenance system. With consistent propagation occurring in only 4.0% of copy genealogies, downstream users should not assume that copied skills receive upstream fixes.

Frequently Asked Questions

Q. What are SKILL.md files?

They are files containing instructions executed by AI coding agents.

Q. How many repositories were covered?

The GitSkills dataset covered 259,596 repositories.

Q. How many distinct skills were identified?

The dataset contained 1,612,846 distinct skills.

Q. How often did source changes reach every copy?

Only 11.1% of source-skill changes reached all copies at each observation window.

Q. How often did copy genealogies change consistently?

Only 4.0% of copy genealogies ever changed consistently.

Q. Did repository stars identify most later high-risk adoptions?

No. The 100 most-starred repositories would have intercepted 0.5% of later high-risk skill adoptions, compared with 14.9% for the influence-based ranking.

Q. What does the paper recommend?

It recommends that platforms distribute versioned references rather than copies.