AI Decoded: 10-K filings - reading between the risks
Explore how AI and NLP analyse 10-K risk disclosures to identify changing corporate risks and support systematic equity investing.
10-K Filings: The language of corporate risk and disclosure
The United States is the leader in making data collected by government institutions available to the public, with the example of the Data.gov Catalog1 listing several hundred thousand data sets. One particularly interesting source is the U.S. Securities and Exchange Commission’s EDGAR system, which contains all 10-K filings since 2001 offered in a machine readable format.
A 10-K filing is an annual report that publicly listed companies in the United States are required to submit to the SEC. It provides a comprehensive and legally binding overview of a company’s business model, financial condition, strategic priorities, and, crucially, its risks.
It is one of the few corporate communications where companies must speak with precision and completeness, not persuasion. That obligation to disclose, in full and without embellishment, is precisely what gives the 10-K filing its unique ability to reveal how a company thinks, not just what it reports.
Among their various sections, the “Risk Factors” section is particularly informative, as it outlines the potential challenges that could materially affect the company’s future performance. Importantly, these disclosures are not purely descriptive; they are shaped by legal and regulatory incentives, which encourage firms to provide a thorough and often cautious account of potential risks to clearly flag possible downside scenarios.
In addition to this annual filing, many more corporate disclosures are required by the SEC. For example, the quarterly 10-Q filing also contains a risk section with related information. In what follows, we focus exclusively on the 10-K as incorporating the 10-Qs does not materially affect our analysis.
Source: Allianz Global Investors, June 2026. For illustrative purposes only.
How to read a 10K
A 10-K filing consists of 16 mandated items grouped into four parts, covering both quantitative financial information and qualitative narrative disclosures. While certain sections, most notably Item 8, contain structured numerical data that is well covered by standard financial databases, other economically relevant sections are predominantly textual in nature. Item 1A (Risk Factor) and Item 7 (Management’s Discussion and Analysis) contain management’s forward-looking assessment of uncertainties, challenges, and structural risks facing the business.
This paper focuses exclusively on Item 1A of the 10-K filings. The Risk Factors section is designed to provide a comprehensive and annually updated overview of the most material risks a company faces. A key limitation of the traditional analysis of Risk Factors is that it often treats disclosures as static checklists, counting words, themes, or references, without adequately capturing how corporate risk profiles evolve over time. In practice, economically meaningful information often lies not in the absolute level of risk disclosure, but in changes: which risks are newly introduced, which are emphasized more strongly, and which are downplayed or removed altogether.
Our approach therefore centers on a year-over-year comparison of Risk Factor disclosures within the 10-K filings. Using NLP, we treat each company’s Item 1A section as a structured body of text that can be decomposed into paragraphs, sentences, and semantic units. This allows us to systematically compare the language used in the successive filings and distinguish between broadly stable, recurring disclosures and genuinely new or evolving risk narratives.
Data collection
Managing filing complexity: the case for NLP
Our database comprises approximately 216,000 10-K filings, covering US listed companies from 1998 onward. The year 2007 represents a natural starting point for our 10-K filing signal. Following a December 2005 amendment to the Securities Exchange Act of 19342, companies were required to include a dedicated Risk Factors section (Item 1A) in their annual 10-K filings. As a result, for most companies, the first filings containing this section appear in 2006. On average, around 7,500 10-K filings are published each year, far more than any human reader could reasonably analyze. Over the period of two decades until today, the Risk Factors section has grown both longer and more complex. As illustrated in the figure below, the average complexity of 10-K risk disclosures has increased steadily over time and is now requiring a significantly higher level of effort and attention than in the past.
The increasingly dense, nuanced risk narratives underscore the growing challenge faced by investors in manually assessing corporate risk disclosures. At the same time, it highlights the potential value of systematic NLP-based approaches that are designed to scale across large textual corpora and to quantify changes in risk communication in a consistent and economically interpretable manner.
Exhibit 1: When risk disclosures outgrow the average reader
Source: Allianz Global Investors, Systematic Equity team. April 2026. For illustrative purposes only. Years of formal education estimated via Flesch-Kincaid grade level.
Comparing risk narratives: from text to signal
With the focus placed on changes of risk narratives, the key question to us becomes how to systematically compare two risk disclosures that often span many pages and are written in dense legal language.
We start by simplifying the problem. Instead of treating each file as one long document, we break it down into individual sentences. This allows us to compare disclosures in a more structured and granular way.
The idea is straightforward: for each sentence in one year, we look for the closest match in the last year’s filing. In other words, we ask ourselves: are there major reformulations or additions potentially representing new risks?
To do this, we compare one sentence with all sentences in the last document and select the most similar one – the “best match”. This similarity is measured using a simple text metric that captures how many small edits (such as adding, deleting, or changing words) are needed to turn one sentence into the other3. The fewer changes required, the more similar the sentences are.
By repeating this process across all sentences, a clear picture begins to emerge. If a sentence has a close match, the underlying risk is broadly unchanged. If the match is weaker, the description has evolved. And if no meaningful match can be found, this points to a genuinely new risk being introduced. The similarity score ranges from 0% to 100%, where higher values indicate nearly identical wording and lower values reflect more substantial differences.
Exhibit 2: How we compare risk sections in 10-K filings over time
Source: Allianz Global Investors, Systematic Equity team. June 2026. For illustrative purposes only.
Exhibit 3: An example from Adobe Inc.
Source:Allianz Global Investors, Systematic Equity team. For illustrative purposes only.
Do risk disclosures keep pace with market stress?
As a plausibility check of our approach, we examine how the signal behaves during periods of major macroeconomic stress - events that are expected to surface new risks and therefore lead to observable changes in 10-K risk disclosures.
To better understand how the signal evolves over time, we compute the average change in risk disclosures across all companies for each year. This aggregate view, as depicted below, reveals clear patterns around major market events over the past decades. In particular, the periods following the great financial crisis (in 2008) and the coronavirus pandemic (in 2020) stand out. During these phases, companies tend to update their risk disclosures more extensively, leading to a noticeable increase in overall change across filings.
These shifts are consistent with intuition. When the macroeconomic environment changes abruptly, firms reassess and communicate new sources of uncertainty, which is reflected in more substantial revisions to their risk sections. The presence of such coordinated, market‑wide adjustments provides an important validation of the algorithm, as it captures exactly those periods where meaningful changes in risk perceptions are expected to occur.
Exhibit 4: Testing the signal in times of market disruption
Source: Allianz Global Investors, Systematic Equity team. Data as of February 2026. For illustrativepurposes only.
Model interpretability – LLM plausibility check
A key advantage of our approach is that it is built on a relatively simple and transparent framework. Unlike more complex “black-box” models, our methodology is easy to monitor, explain, and understand. Each step, from identifying matching sentences to measuring their similarity, can be directly traced and interpreted, which is critical in an investment context where model outputs must be interpretable.
At the same time, recent advances in LLMs have demonstrated a remarkable ability to understand and summarize complex text. These models can provide natural-language explanations of risk disclosures, offering an additional perspective that is intuitive for human decision-makers.
We use this technique not as a replacement for our model, but as a plausibility check. If a portfolio manager has questions about a particular output, an LLM can be used to independently assess the same document and highlight the most relevant risks. This provides a simple, intuitive “second opinion” that helps validate whether the changes identified by our algorithm are economically meaningful.
Going back to our earlier example, we prompt an LLM to summarize the top challenges Adobe Inc. might face according to the 2024 10-K risk section. The model highlights several key challenges, including those in the illustration below.
These outputs broadly align with the areas identified by our similarity based approach. In particular, the emergence of risks related to GenAI is captured both qualitatively by the LLM and quantitatively through lower similarity scores in our methodology. This consistency provides additional confidence that the signal is capturing economically relevant changes in corporate risk narratives.
A natural question is why not rely on LLMs directly. While powerful, such models come with important limitations: their outputs are not fully controlled, can evolve over time, and may introduce forward-looking biases that are difficult to detect or manage. In contrast, our approach is fully transparent, stable, and reproducible, which are key requirements for oursystematic equity investment process.
That said, LLMs can play a valuable complementary role. By translating model outputs into intuitive summaries, they help make the signal more accessible to portfolio managers and support the interpretation of model-driven insights. In this way, we combine the strengths of both approaches: a robust, controlled model, complemented by flexible, human-readable explanations.
- Keeping pace with rapid AI-driven technological change: uncertain customer adoption, high AI compute costs that could compress margins, and risk of failed or delayed innovation across Creative, Document and Experienc Clouds.
- Managing expanding, fragmented regulation: evolving AI rules, global privacy and datatransfer restrictions (e.g., GDPR, PIPL, US state laws) that may force product or practice changes, add compliance cost, and increase liability.
- Intensifying competition: AI-native entrants and established rivals integrating AI may erode differentiation, pressure pricing and renewals, and slow growth in both core and standalone AI offerings.
- Security and reliability risks: cyberattacks, fraud/abuse, and dependence on third-party cloud and delivery platforms could lead to outages, data loss, fines, customer churn, and reputational damage.
- Talent and execution under macro uncertainty: difficulty hiring/retaining scarce AI and cybersecurity talent, hybrid work challenges, long/complex enterprise sales cycles, and global economic/geopolitical volatility impacting demand and operations.
From risk changes to investment outcomes
To assess the applicability of the signal in a long-only investment framework, we conducted the following simulation: We rank stocks based on their 10-K similarity score, construct an equal-weighting portfolio based on these rankings and compare their performance against the Russell 1000 equal-weighted index. The results show a clear pattern whereby stocks exhibiting the most significant changes in risk disclosures consistently underperform the broader market.
This finding suggests that large shifts in risk narratives are associated with subsequent underperformance. Intuitively, significant changes in the Risk Factor section may reflect emerging challenges, rising uncertainty, or deteriorating business conditions that are not yet fully reflected in financial metrics. From an investment perspective, this signal can therefore provide valuable incremental insight. In particular, it may help identify Value traps4 and “Growth torpedoes”5.
As such, incorporating this signal into a long-only strategy can enhance risk awareness and support more informed security selection, complementing traditional financial analysis with a forward-looking view on evolving corporate risks.
Exhibit 5: Underperformance of the Quintile 5 simulated portfolio relative to the Russell 1000 index
Source: Allianz Global Investors, Systematic Equity team. Data as of May 2026. For illustrative purposes only. The hypothetical performance and simulationsshown are for illustrative purposes only and do not represent actual performance; they do not predict future returns. Please see important informationregarding back-testings and hypothetical or simulated performance data on the last page. Allianz Global Investors, Systematic Equity team. Data as ofFebruary 2026. For illustrative purposes only.
Conclusion
2 Please refer to https://www.sec.gov/news/press/2005-99.htm for additional information.
3 We are using the normalized Levenshtein distance to measure the similarity. The Levenshtein distance is a string metric for measuring the difference between two sequences. The Levenshtein distance between two words is the minimum number of single-character edits (insertions, deletions or substitutions) required to change one word into the other. It is named after Soviet mathematician Vladimir Levenshtein, who defined the metric in 1965. Please find more information via https://en.wikipedia.org/wiki/Levenshtein_distance.
4 Companies that appear attractive based on traditional metrics but face worsening underlying risks.
5 Companies with strong growth profiles that may be vulnerable to newly emerging challenges.