Researchers at Truffle Security tested 543,699 API keys, database passwords, and access tokens found in public GitHub repositories last July. Every single one authenticated. The median credential had been sitting in publicly readable code for 784 days.
The findings come from a scan of The Stack v3, a 224-million-repository snapshot of public GitHub code assembled to train large language models. The crawl closed on August 7, 2025. Eleven months later, when Truffle Security ran live verification against each issuing provider, more than half a million credentials still worked. That number is more than double the 221,303 live credentials the company found when it ran a similar scan against 7.6 petabytes of Hugging Face training data earlier this year.
The oldest credential in the dataset was last touched on June 13, 2009. It lives inside an Erlang web server configuration file, and it was still valid 16.1 years after it was committed. Behind it: an FTP login inside a GPS logger's C source code from September 2009, replicated across 62 repositories, and an AWS key tucked inside a Rails S3 config from November of that same year. Truffle Security declined to name the repositories because the credentials in them still work.
A Protection That Only Faces Forward
GitHub has progressively tightened its defenses around exposed credentials. The platform made secret scanning alerts free for all public repositories in February 2023. Push protection, which blocks a commit before it reaches the remote branch if it carries a recognised secret, became generally available in May 2023 and was switched on by default for all public repositories on February 29, 2024.
GitHub's secret scanning covers more than 200 token types and patterns from over 180 service providers. The rollout had a measurable effect on new leaks. Among credential shapes the system recognises and blocks, Truffle Security found a 53 percent drop in the rate of fresh exposures across the twelve months following the default rollout, compared to the twelve months before it. Slack tokens fell 64 percent, GitHub's own tokens and AWS access keys each fell 59 percent.
But push protection has no mechanism to reach the credentials already there. Of the 543,699 live credentials, 199,843 landed after push protection became the default in February 2024. Developers either bypassed the block or committed credential types the system does not recognise.
That second category is the larger problem. Truffle Security found that 51.8 percent of every live credential in the dataset is a shape that a default-configured public repository will accept without objection. Database connection strings, private keys, and Google API keys all fall outside the default block list. Push protection focuses on specific, highly identifiable secrets and misses generic ones. Connection strings and private keys are classified as generic patterns, and blocking them requires an organisation to go into settings and explicitly opt in.
The Gemini Problem
The Google API key situation illustrates the limits of pattern-based blocking in particularly sharp terms. The 33,343 live Google API keys in Truffle Security's dataset include 31,374 that authenticate specifically to Gemini, Google's AI model platform. Their median leak date is February 2025, meaning the entire population is younger than the push protection rollout.
Google API keys carry the prefix `AIzaSy` whether they were created for Google Maps, Firebase, or Gemini. GitHub's pattern list recognises the prefix but marks it as not push-protected, because a Maps key sitting in client-side JavaScript is not a secret by design. Google's own approach to API keys was historically built around the assumption that these keys would live in client-side code, exposed to anyone who opened a browser's developer tools. The problem is that Gemini runs on the same key format, turning what developers were trained to treat as a non-sensitive identifier into a billable AI credential. One pattern cannot distinguish between the two uses, so nothing gets blocked, and the keys that matter arrive alongside the keys that do not.
Revocation is the Deciding Variable
The most instructive comparison in the data is between providers that automatically revoke leaked tokens and those that do not.
npm committed 101,886 tokens to public code. One remains live. GitHub committed 73,048 tokens; 260 survived. Hugging Face committed 30,437; 15 are still valid. Each of these platforms runs an automated pipeline that kills a token the moment it is detected in public code.
The contrast with database credentials is stark. Of 12,985 Postgres connection strings in the dataset, 11,465 are still live, an 88 percent survival rate. MySQL connection strings survive at 75 percent. MongoDB, where the detector only reports a URI it successfully connected to, returned all 51,067 live.
Push protection blocks secrets at the door. Automated revocation kills them wherever they are. The Truffle Security data shows that the second mechanism is the one that changes the outcome, and for the majority of credential types sitting in public repositories right now, no provider is running it.
The practical guidance from the researchers: treat any committed credential as compromised regardless of whether anything flagged it, scan your own repository history rather than assuming the push-time block was sufficient, and favour credentials that expire automatically. Most of what Truffle Security found would have been harmless long ago if it had ever been given a finite lifetime.