An enterprise risk assessment method and system based on a large language model and text quantification analysis
By combining text parsing, semantic recognition, and differential comparison analysis with large language models and expert logic, quantitative indicators for enterprise risk assessment are generated. This solves the problems of insufficient target focus, reliability, and quantitative systematicity of large language models in enterprise risk assessment, and achieves efficient and accurate risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- E FUND MANAGEMENT CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-26
AI Technical Summary
Existing large language models suffer from insufficient focus on objectives, poor reliability of results, lack of quantitative systematicity, and insufficient integration with expert experience in enterprise risk assessment, which limits the application of unstructured text information in quantitative analysis.
By combining text parsing with semantic recognition and classification, and utilizing a large language model for semantic chapter division and difference comparison analysis, quantitative analysis indicators are generated. Combined with expert logical guidance, a complete chain from unstructured text to structured quantitative indicators is constructed to achieve risk assessment.
It significantly improves the systematicness, accuracy, and practicality of extracting quantitative information from unstructured text. The generated risk assessment results combine the efficiency of automated processing with the depth of expert analysis logic, thereby improving the effectiveness and accuracy of enterprise risk assessment.
Smart Images

Figure CN122088482A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text semantic analysis technology, and in particular to a method and system for enterprise risk assessment based on large language models and text quantitative analysis. Background Technology
[0002] In the field of enterprise management information analysis, information is a core decision-making element. Traditional quantitative analysis mainly relies on structured information such as transaction data and financial indicators. However, with the deepening of information disclosure, unstructured data in text form, such as listed company annual reports, industry research reports, and policy documents, is increasingly becoming a key information source, containing a large number of signals with predictive value for enterprise operations. Due to the inherent difficulties in storing, retrieving, and processing unstructured data, its quantitative application has long been limited. To overcome this bottleneck, related technologies have undergone continuous evolution: early methods used shallow feature extraction methods such as keyword matching and word frequency statistics, which achieved preliminary text quantification, but could not capture deep semantics; subsequently, semantic vector technology based on deep learning and pre-trained models (such as BERT) was applied, which improved semantic representation capabilities by converting text into high-dimensional vectors, but its "black box" nature led to poor interpretability and a lack of standardized quantification rules across scenarios.
[0003] In recent years, large language models (LLMs) have opened up new avenues for text analysis due to their superior semantic understanding, logical reasoning, and generative capabilities. LLMs can identify complex logical relationships in text, such as causal, conditional, and degree modifiers, overcoming the limitations of traditional models in deep semantic understanding and providing new technical possibilities for directly extracting structured signals from unstructured text. However, directly applying LLMs to text-based quantitative analysis for risk assessment still has significant drawbacks: First, when lacking clear task guidance, the model's output is often broad or deviates from the analytical focus, making it difficult to directly transform into targeted structured signals; second, requiring LLMs to directly provide analytical conclusions can easily lead to factual "illusions" or logical jumps, resulting in insufficient stability and interpretability of the results; third, existing methods lack a systematic quantitative transformation mechanism, making it difficult to reliably extract numerical indicators from text that can be used for quantitative modeling; finally, current technical approaches are often disconnected from the text analysis logic and experience accumulated by domain experts, making the output results difficult to adopt in actual business scenarios.
[0004] In summary, although large language models have made breakthroughs in semantic understanding, they still have shortcomings in terms of target focus, result reliability, quantitative systematicity, and integration with expert experience. These deficiencies restrict the in-depth, reliable, and large-scale application of unstructured text information in quantitative analysis, becoming a pressing technical problem to be solved. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides a method and system for enterprise risk assessment based on large language models and text quantitative analysis, thereby improving the effectiveness and accuracy of risk assessment for target enterprises.
[0006] In a first aspect, embodiments of this application provide a method for enterprise risk assessment based on large language models and text quantitative analysis, including: Obtain the target company's current first annual report text; The first annual report text is parsed to obtain the corresponding first parsed text, wherein the first parsed text includes several subheadings and several corresponding subtexts; The subheadings are semantically identified and categorized using a pre-defined large language model. Based on the semantic identification and categorization results, the subtexts corresponding to each subheading are divided into corresponding semantic chapters, thereby generating the corresponding first reconstructed text. The first target text is extracted from the first reconstructed text according to the preset target semantic chapter; According to the preset analysis logic, the first target text and the second target text are compared and analyzed by the big language model to obtain several different texts and several corresponding change trends. The second target text is obtained by sequentially parsing, semantic recognition and classification and text extraction of the second annual report text of the target company in the previous year. Based on the number of each of the aforementioned differential texts, the aforementioned trends of change, and the average number of words per chapter, quantitative analysis indicators for the first annual report text are calculated, wherein the average number of words per chapter is calculated based on the total number of words in the first and second annual report texts and the number of semantic chapters. A risk assessment is conducted on the target company based on the quantitative analysis indicators, and a risk assessment result is generated.
[0007] This application provides a corporate risk assessment method based on a large language model and text quantitative analysis, significantly improving the systematicness, accuracy, and practicality of extracting quantitative information from unstructured text. First, this embodiment combines text parsing with semantic recognition and classification to transform loosely structured annual report text into reconstructed text with clearly defined semantic chapters, effectively solving the inherent problem of directly quantifying unstructured data. Utilizing the deep semantic understanding capabilities of the large language model, the semantic connotations of subheadings can be accurately identified and classified according to preset semantic definitions. This avoids semantic biases caused by traditional keyword matching or shallow statistical methods, ensuring the accuracy and consistency of information segmentation. Second, through cross-year difference comparison analysis, this embodiment can automatically identify changes in key text content such as business descriptions and risk disclosures, and extract corresponding trends. This semantic-based difference analysis goes beyond simple text addition and subtraction statistics; it captures the evolution of deeper information such as tone, emphasis, and risk exposure levels, providing richer and more insightful observations for judging corporate business dynamics. Finally, by combining qualitative information such as the number of differentially expressed texts and trends with the average number of words per chapter, a series of quantitative analysis indicators were innovatively calculated and used for risk assessment. This process successfully constructed a complete technical chain from unstructured text to structured quantitative indicators, and then to risk assessment conclusions, greatly enhancing the measurability, comparability, and direct application value of the analysis results. This embodiment effectively suppresses the defects of "illusion" and unstable output by guiding the large language model to perform standardized and logical analysis tasks, making the generated risk assessment results combine the efficiency of automated processing with the depth of expert analysis logic, thereby improving the effectiveness and accuracy of risk assessment for target enterprises.
[0008] Furthermore, the first annual report text is parsed to obtain a corresponding first parsed text, which includes several subheadings and corresponding subtexts, including: Convert the first annual report text into markup language format to obtain the corresponding first format converted text; The first formatted text is matched with headings and divided into heading levels according to a preset regular expression, and several subheadings, corresponding heading levels and corresponding subtexts are determined. Based on each of the subheadings, the corresponding heading levels, and the corresponding subtexts, the first format-converted text is subjected to structured processing to generate the first parsed text with a heading hierarchy.
[0009] This application provides a text parsing method that significantly improves the automation, structural standardization, and convenience of subsequent text processing. Specifically, by first converting the annual report text into a markup language format, a unified, machine-readable foundation is provided for programmatic processing, effectively addressing the parsing challenges caused by the complex formats and inconsistent styles of original PDF or Word documents. More importantly, by using preset regular expressions for title matching and title level classification, it can intelligently and accurately identify subheadings and their hierarchical relationships from the formatted text. This rule-and-pattern matching method, compared to relying entirely on model recognition, has higher accuracy and reliability in extracting elements with obvious format features, such as titles, reducing the risk of misidentification and omission. After determining the titles and their levels, the text is further structured based on this information to generate a first parsed text with a clear title hierarchy, providing high-quality input data for subsequent large language model analysis, laying a solid foundation for subsequent semantic chapter division and content extraction, and improving the effectiveness and accuracy of subsequent risk assessments of target enterprises.
[0010] In one possible implementation, the step of semantically recognizing and classifying each subheading using a preset large language model, and dividing the subtext corresponding to each subheading into corresponding semantic chapters based on the semantic recognition and classification results, thereby generating the corresponding first reconstructed text, includes: The preset first prompt word is input into the large language model so that the large language model can perform semantic recognition and classification of each subheading in the first parsed text according to the semantic definition of each semantic chapter in the first prompt word, and determine the semantic chapter to which each subheading belongs; Based on the semantic chapter to which each subheading belongs, the subtext corresponding to each subheading is divided into the corresponding semantic chapter, thereby determining the text content of each semantic chapter; The first reconstructed text is generated based on each of the semantic chapters, the corresponding text content, and the order and output format settings of each of the semantic chapters in the first prompt word.
[0011] This application provides a method for generating reconstructed text. Through a pre-designed prompting process, a large language model is guided to complete standardized and interpretable semantic analysis and text reconstruction tasks, thereby ensuring the focus of the analysis and the consistency of the results. This embodiment inputs a preset first prompt word into the large language model, which clearly defines the semantic categories and boundaries of each semantic section. This prevents the large language model from performing aimless general text understanding, but rather, within a specific task framework, it performs semantic judgments similar to expert rules, accurately classifying the previously parsed subheadings into preset semantic sections. This prompt-based guidance effectively overcomes the problem of broad or off-focused outputs from the large language model, ensuring that the semantic recognition process is highly aligned with business analysis needs. After classifying each subheading, its corresponding subtext can be classified accordingly. This setting of classifying subheadings first and then subtexts simplifies the length of context the model needs to process and improves its processing efficiency. This embodiment essentially "solidifies" the classification logic and experience of domain experts into machine-executable instructions through prompt words, achieving a deep integration of expert experience and the capabilities of large language models. This provides a reliable guarantee for generating high-quality, trustworthy intermediate analytical products, and improves the effectiveness and accuracy of subsequent risk assessments of target companies.
[0012] In one possible implementation, the step of performing a difference comparison analysis on the first target text and the second target text using the large language model according to a preset analysis logic, to obtain several difference texts and corresponding several change trends, includes: The preset second prompt word is input into the large language model, so that the large language model performs a difference comparison analysis on the first target text and the second target text according to the analysis logic set in the second prompt word, and determines several difference comparison groups in which there are differences in the first target text and the second target text. Each difference comparison group includes a first difference text segment in the first target text and a corresponding second difference text segment in the second target text. Based on each of the first and second difference text segments, semantic summaries are performed on each of the difference comparison groups to obtain the corresponding difference texts. The preset third prompt word is input into the large language model so that the large language model can perform a qualitative analysis of the target company's business status based on the third prompt word and each of the differential texts, and then generate the change trend corresponding to each of the differential texts.
[0013] This application's embodiments define the specific process of difference comparison analysis, greatly improving the depth and quality of the model's extraction of effective signals from changes in annual reports. Instead of simply performing a rough full-text comparison, this embodiment guides the large language model to perform intelligent difference detection by inputting a second prompt word with preset analysis logic. This analysis logic can be designed to focus on specific types of statement changes, such as identifying additions or deletions to arguments on the same topic, shifts in word strength, data updates, and refinements in risk descriptions. Based on this, the large language model identifies pairs of differing text fragments (difference comparison groups), which captures substantial semantic changes more effectively than simply statistically analyzing word frequency changes or using edit distance. Subsequently, these difference comparison groups are semantically summarized to extract the core differences, forming highly condensed "difference texts." This summarization process filters out redundant information and sentence structure differences, directly pointing to changes in the core content. Finally, guided by the third cue word, the large language model conducts a comprehensive qualitative analysis of the company's operating status based on all the extracted differential texts, and infers the changing trends of each differential text. This ensures that the final trends are not an overreaction to a single word or phrase, but rather a comprehensive inference result based on a series of semantic evidence. This significantly improves the stability, rationality, and interpretability of trend judgments, and enhances the effectiveness and accuracy of subsequent risk assessments of target companies.
[0014] In one possible implementation, the step of calculating quantitative analysis indicators for the first annual report text based on the number of each of the differing texts, the respective trends of change, and the average number of words per chapter includes: The first analytical index of the first annual report text is calculated by dividing the number of each of the aforementioned differential texts by the average number of words in each chapter. The difference between the positive and negative trends in each of the aforementioned trends is divided by the average number of words in each chapter to obtain the second analytical indicator of the first annual report text. By combining the first and second analytical indicators, quantitative analytical indicators for obtaining the first annual report text are constructed.
[0015] This application provides a method for calculating quantitative analysis indicators, creatively transforming the semantically rich qualitative analysis results produced by a large language model into concise, objective, and mathematically computable standardized quantitative indicators that can be compared over time. Specifically, regarding the first analysis indicator, a "density of information change per unit text length" indicator is obtained by dividing the number of differing texts by the average number of words per chapter. This indicator effectively eliminates the impact of length differences between annual reports, standardizing the measurement of the activity level or information update intensity of annual report content changes. Regarding the second analysis indicator, a "net sentiment or net trend direction per unit text length" indicator is obtained by calculating the net value (difference) of positive and negative trends and dividing it by the average number of words per chapter. This indicator not only reflects the quantity of change but also captures the "net effect" of the nature of the change, such as whether the enhancement of positive business descriptions outweighs the increase in risk warnings. The set of quantitative analysis indicators constructed by combining these two indicators has profound business implications: the first indicator can be regarded as a proxy variable for "information uncertainty" or "change in management communication"; the second indicator can be regarded as a proxy variable for "net sentiment about business outlook" or "net change in risk exposure". This quantitative approach allows previously vague and qualitative text analysis conclusions to be integrated into financial econometric models, risk scoring cards, or monitoring dashboards in numerical form. This facilitates trend tracking, threshold warnings, cross-sectional ranking, and regression analysis, addressing the pain point of "emphasizing qualitative analysis and neglecting quantitative analysis" in the application of large language models, and improving the effectiveness and accuracy of subsequent risk assessments of target companies.
[0016] Furthermore, the risk assessment of the target enterprise based on the quantitative analysis indicators includes: Obtain historical quantitative analysis indicators from the annual report texts of the target company in different years. Based on the historical first analysis indicators and the numerical values of the first analysis indicators in each of the aforementioned historical quantitative analysis indicators, a first change curve is constructed. The second change curve is determined based on the values of each historical second analysis indicator and the second analysis indicator in each of the aforementioned historical quantitative analysis indicators; Based on the slope value of the first change curve within a preset sliding time window, generate the first evaluation text; Based on the value of the second change curve within a preset sliding time window, a second evaluation text is generated; Based on the numerical values of the first and second analytical indicators, a third evaluation text is generated; The first assessment text, the second assessment text, and the third assessment text are combined to generate the risk assessment result of the target company.
[0017] This application provides a risk assessment method based on quantitative analysis indicators, which greatly enhances the systematicness, objectivity, and forward-looking nature of risk assessment. Specifically, this embodiment does not rely on isolated indicators from a single year, but first obtains historical quantitative analysis indicators of the target company over many years, thereby expanding the analytical perspective from "points" to "lines" and "surfaces." By constructing historical change curves for the first and second analysis indicators, the long-term evolution trajectory of the company's information disclosure behavior, operational communication tone, and risk situation can be intuitively displayed. Based on this, dynamic analysis is further introduced: a first assessment text is generated based on the slope of the first change curve within a sliding time window. Through fluctuations in indicator values, the stability of the company's business expectations and strategic planning is objectively reflected. For example, if the first change curve continues to rise, it indicates that the company's business planning adjustments are frequent and operational uncertainty is increased. A second assessment text is generated based on the values of the second change curve within the sliding window, which can assess the recent level range of the company's business sentiment or risk situation. For example, if the second change curve turns from positive to negative or continues to decline, it indicates that the company's business expectations are shifting towards a conservative or pressured direction. Finally, a third assessment text is generated based on the values of the first and second analysis indicators for the current year, accurately assessing the target company's current operating status. Ultimately, combining the assessment texts from these three dimensions results in a risk assessment that is no longer simply labeled as "high risk" or "low risk," but rather a comprehensive analysis report that includes dynamic interpretation, historical comparison, and current status positioning. This provides more accurate clues and evidence for risk warning and in-depth due diligence, improving the effectiveness and accuracy of risk assessment of target companies.
[0018] Furthermore, the risk assessment results also include a fourth assessment text, which is generated by comparing the target company's various peer companies' current second analysis indicators. Specifically: Obtain the current second analysis indicators of each peer company of the target company; The number of positive and negative values in each of the current second analysis indicators is statistically obtained. The fourth evaluation text is generated based on the number of positive values, the number of negative values, and a preset quantity threshold.
[0019] Furthermore, this embodiment adds a fourth assessment text to the risk assessment results, introducing a dimension of horizontal industry comparison. This allows the risk assessment conclusions to move beyond focusing solely on the company's own historical data, providing a broader market coordinate system and relative evaluation perspective, significantly improving the comprehensiveness and fairness of the risk assessment. This embodiment generates the fourth assessment text based on the number of positive and negative values of the second analytical indicator for each of the target company's peers. This is because the second analytical indicator reflects the net sentiment or net trend direction conveyed in the company's annual report. By statistically analyzing the positive and negative values of this indicator for multiple comparable companies within an industry, it is possible to determine whether the overall sentiment of the industry is optimistic or pessimistic, or the magnitude of the common risk pressure faced by the industry. For example, if the target company's second indicator is negative, but the vast majority of its peers also have negative indicators, then this negative sentiment may stem more from industry-wide shocks than from company-specific problems. Conversely, if the industry as a whole has positive values, but the target company alone has negative values, it strongly suggests that the company may be facing individual operational difficulties or risks, and its risk level requires special attention. Therefore, the fourth assessment text provides crucial contextual information, helping users distinguish between "systemic risk" and "specific risk." It calibrates the company's textual signals within an industry context, avoiding misjudgments of individual companies due to overall industry fluctuations, and further improving the effectiveness and accuracy of risk assessments of target companies.
[0020] Secondly, embodiments of this application provide an enterprise risk assessment system based on a large language model and text quantitative analysis, including an acquisition module, a text parsing module, a semantic recognition module, a text extraction module, an analysis module, an indicator calculation module, and a risk assessment module; The acquisition module is used to acquire the text of the target company's current first annual report. The text parsing module is used to parse the first annual report text to obtain the corresponding first parsed text, wherein the first parsed text includes several subheadings and several corresponding subtexts; The semantic recognition module is used to perform semantic recognition and classification on each of the subheadings using a preset large language model, and to divide the subtexts corresponding to each of the subheadings into corresponding semantic chapters based on the semantic recognition and classification results, thereby generating the corresponding first reconstructed text; The text extraction module is used to extract the corresponding first target text from the first reconstructed text according to the preset target semantic chapter; The analysis module is used to perform a difference comparison analysis on the first target text and the second target text according to the preset analysis logic and the large language model to obtain several difference texts and several corresponding change trends. The second target text is obtained by sequentially parsing, semantic recognition and classification and text extraction of the second annual report text of the target company in the previous year. The indicator calculation module is used to calculate the quantitative analysis indicators of the first annual report text based on the number of each of the different texts, the various change trends, and the average number of words per chapter. The average number of words per chapter is calculated based on the total number of words in the first annual report text and the second annual report text and the number of semantic chapters. The risk assessment module is used to conduct risk assessment on the target enterprise based on the quantitative analysis indicators and generate risk assessment results.
[0021] Furthermore, the analysis module, based on preset analysis logic, performs a difference comparison analysis on the first target text and the second target text using the large language model to obtain several difference texts and corresponding change trends, including: The preset second prompt word is input into the large language model, so that the large language model performs a difference comparison analysis on the first target text and the second target text according to the analysis logic set in the second prompt word, and determines several difference comparison groups in which there are differences in the first target text and the second target text. Each difference comparison group includes a first difference text segment in the first target text and a corresponding second difference text segment in the second target text. Based on each of the first and second difference text segments, semantic summaries are performed on each of the difference comparison groups to obtain the corresponding difference texts. The preset third prompt word is input into the large language model so that the large language model can perform a qualitative analysis of the target company's business status based on the third prompt word and each of the differential texts, and then generate the change trend corresponding to each of the differential texts.
[0022] Furthermore, the indicator calculation module calculates quantitative analysis indicators for the first annual report text based on the number of each differing text, each changing trend, and the average number of words per chapter, including: The first analytical index of the first annual report text is calculated by dividing the number of each of the aforementioned differential texts by the average number of words in each chapter. The difference between the positive and negative trends in each of the aforementioned trends is divided by the average number of words in each chapter to obtain the second analytical indicator of the first annual report text. By combining the first and second analytical indicators, quantitative analytical indicators for obtaining the first annual report text are constructed. Attached Figure Description
[0023] Figure 1 A flowchart illustrating an enterprise risk assessment method based on a large language model and text quantitative analysis, provided for embodiments of this application; Figure 2 A schematic diagram illustrating the specific execution flow of an enterprise risk assessment method based on a large language model and text quantitative analysis at each stage, provided for an embodiment of this application. Figure 3 This is a schematic diagram of the structure of an enterprise risk assessment system based on a large language model and text quantitative analysis, provided as an embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0025] It should be noted that the step numbers in this document are only for the convenience of explaining the specific embodiments and are not intended to limit the order in which the steps are performed. In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0026] Example 1: like Figure 1 As shown, Embodiment 1 provides a method for enterprise risk assessment based on large language models and text quantitative analysis, including steps S1-S7: Step S1: Obtain the target company's current first annual report text; Step S2: Perform text parsing on the first annual report text to obtain the corresponding first parsed text, wherein the first parsed text includes several subheadings and several corresponding subtexts; Step S3: Semantic recognition and classification of each subheading is performed using a preset large language model, and the subtexts corresponding to each subheading are divided into corresponding semantic chapters based on the semantic recognition and classification results, thereby generating the corresponding first reconstructed text; Step S4: Extract the corresponding first target text from the first reconstructed text according to the preset target semantic chapter; Step S5: According to the preset analysis logic, the first target text and the second target text are compared and analyzed by the big language model to obtain several different texts and several corresponding change trends. The second target text is obtained by sequentially parsing, semantic recognition and classification and text extraction of the second annual report text of the target company in the previous year. Step S6: Calculate the quantitative analysis indicators of the first annual report text based on the number of each of the different texts, the change trends, and the average number of words per chapter. The average number of words per chapter is calculated based on the total number of words in the first annual report text and the second annual report text and the number of semantic chapters. Step S7: Conduct a risk assessment of the target enterprise based on the quantitative analysis indicators and generate risk assessment results.
[0027] This application provides a corporate risk assessment method based on a large language model and text quantitative analysis, significantly improving the systematicness, accuracy, and practicality of extracting quantitative information from unstructured text. First, this embodiment combines text parsing with semantic recognition and classification to transform loosely structured annual report text into reconstructed text with clearly defined semantic chapters, effectively solving the inherent problem of directly quantifying unstructured data. Utilizing the deep semantic understanding capabilities of the large language model, the semantic connotations of subheadings can be accurately identified and classified according to preset semantic definitions. This avoids semantic biases caused by traditional keyword matching or shallow statistical methods, ensuring the accuracy and consistency of information segmentation. Second, through cross-year difference comparison analysis, this embodiment can automatically identify changes in key text content such as business descriptions and risk disclosures, and extract corresponding trends. This semantic-based difference analysis goes beyond simple text addition and subtraction statistics; it captures the evolution of deeper information such as tone, emphasis, and risk exposure levels, providing richer and more insightful observations for judging corporate business dynamics. Finally, by combining qualitative information such as the number of differentially expressed texts and trends with the average number of words per chapter, a series of quantitative analysis indicators were innovatively calculated and used for risk assessment. This process successfully constructed a complete technical chain from unstructured text to structured quantitative indicators, and then to risk assessment conclusions, greatly enhancing the measurability, comparability, and direct application value of the analysis results. This embodiment effectively suppresses the defects of "illusion" and unstable output by guiding the large language model to perform standardized and logical analysis tasks, making the generated risk assessment results combine the efficiency of automated processing with the depth of expert analysis logic, thereby improving the effectiveness and accuracy of risk assessment for target enterprises.
[0028] In a preferred embodiment, such as Figure 2As shown, steps S1-S6 can be summarized into four stages: text parsing, semantic analysis, logical guidance, and quantitative conversion. This embodiment follows a human-computer collaborative approach of "text parsing—semantic analysis—logical guidance—quantitative conversion," and is a general methodology for the quantitative conversion of unstructured financial text using a "large language model + expert logic guidance" approach, not limited to specific text types or single expert logic. Its core logic lies in clearly defining the collaborative division of labor between the large language model and expert experience—the large language model focuses on general technical tasks such as text parsing and semantic understanding, while expert experience logic serves as scenario-based guidance rules, which can be flexibly customized according to the information attributes and analytical needs of the target text. Ultimately, through a unified quantitative conversion mechanism, the unstructured text is stably mapped into signals that can be used for quantitative analysis of enterprise business information. The methodology for each stage is summarized as follows: 1. Text Parsing: With "format normalization → core information extraction → hierarchical structuring" as the core logic, various text formats are first converted into easily processed intermediate formats (such as Markdown). Through "rule initial screening + large language model to fill gaps", core information is ensured to be complete. Finally, JSON data with a unified structure is output, providing a standardized input basis.
[0029] 2. Semantic Analysis: The goal of constructing a general semantic dimension and providing standardized prompts is to normalize semantic information. Implementation Logic: The design of general basic dimensions follows the principles of "clear boundaries and scalability." The prompting system employs three elements: "general task definition + dimension judgment criteria + standardized output format," ensuring unbiased structured output, achieving semantic alignment across different texts, and adapting to various scenario requirements.
[0030] 3. Logical Guidance: Based on the "analysis objective + core information relevance points of the text", an expert logical framework is constructed, which includes "comparison benchmark (cross-period / cross-industry / cross-entity) + direction determination (positive / negative / neutral) + priority screening (excluding irrelevant details)". This methodology can flexibly customize guidance rules according to the text type to ensure that the output information is targeted and interpretable, and adapts to the analysis needs of various financial texts.
[0031] 4. Quantitative Transformation Methodology: Core variable selection + universal benchmark calibration, aiming to construct comparable quantitative indicators across scenarios. Implementation Logic: Adhering to the principles of "objectivity, comparability, and interpretability," a dual-variable system of "core variables + benchmark variables" is adopted. Scenario differences are eliminated through relative value calculation; standardized mapping rules are established to transform relative values into quantifiable indicators. Through "relative value calculation + standardized mapping," the indicators are ensured to be comparable across scenarios, stably outputting usable signals.
[0032] Furthermore, in step S2, the first annual report text is parsed to obtain the corresponding first parsed text. The first parsed text includes several subheadings and corresponding several subtexts, including: Convert the first annual report text into markup language format to obtain the corresponding first format converted text; The first formatted text is matched with headings and divided into heading levels according to a preset regular expression, and several subheadings, corresponding heading levels and corresponding subtexts are determined. Based on each of the subheadings, the corresponding heading levels, and the corresponding subtexts, the first format-converted text is subjected to structured processing to generate the first parsed text with a heading hierarchy.
[0033] This application provides a text parsing method that significantly improves the automation, structural standardization, and convenience of subsequent text processing. Specifically, by first converting the annual report text into a markup language format, a unified, machine-readable foundation is provided for programmatic processing, effectively addressing the parsing challenges caused by the complex formats and inconsistent styles of original PDF or Word documents. More importantly, by using preset regular expressions for title matching and title level classification, it can intelligently and accurately identify subheadings and their hierarchical relationships from the formatted text. This rule-and-pattern matching method, compared to relying entirely on model recognition, has higher accuracy and reliability in extracting elements with obvious format features, such as titles, reducing the risk of misidentification and omission. After determining the titles and their levels, the text is further structured based on this information to generate a first parsed text with a clear title hierarchy, providing high-quality input data for subsequent large language model analysis, laying a solid foundation for subsequent semantic chapter division and content extraction, and improving the effectiveness and accuracy of subsequent risk assessments of target enterprises.
[0034] In a preferred embodiment, annual report documents of listed companies are obtained from public channels, and the raw text in PDF or HTML format is converted into hierarchical JSON text. Using parsing tools, core sections such as "Management Discussion and Analysis" and "Future Outlook" can be effectively extracted. Since the formats of annual reports vary significantly between different companies and years, this step requires a combination of rules and a large language model to ensure the stability of the parsing results. The specific implementation process is as follows: 1. For a specific listed company's stock in a specific year, first obtain the name of its annual report for that year, read the corresponding PDF file of the annual report, and convert it into MD format.
[0035] 2. Read the markup file, match the first-level headings according to the regular expression of the headings, and classify all possible headings in the content that match the regular expression using a large language model. Process headings at different levels as necessary to reduce errors caused by excessively long text content or incomplete regularization matching.
[0036] 3. Store the text as a hierarchical JSON text according to the division of subheadings, and then obtain the first parsed text with a heading hierarchy.
[0037] The regularization expression mentioned above is a general regularization expression formed based on the patterns summarized from the annual reports of listed companies, that is, patterns summarized from expert experience. The phrase "processing headings at different levels in a layered manner when necessary" mainly refers to the approach of segmenting the extracted subheadings into levels using a large language model, and then summarizing them all at once when there are too many subheadings, thereby reducing the task difficulty and improving accuracy.
[0038] In one possible implementation, in step S3, the semantic recognition and classification of each subheading is performed using a preset large language model, and the subtext corresponding to each subheading is divided into corresponding semantic chapters based on the semantic recognition and classification results, thereby generating the corresponding first reconstructed text, including: The preset first prompt word is input into the large language model so that the large language model can perform semantic recognition and classification of each subheading in the first parsed text according to the semantic definition of each semantic chapter in the first prompt word, and determine the semantic chapter to which each subheading belongs; Based on the semantic chapter to which each subheading belongs, the subtext corresponding to each subheading is divided into the corresponding semantic chapter, thereby determining the text content of each semantic chapter; The first reconstructed text is generated based on each of the semantic chapters, the corresponding text content, and the order and output format settings of each of the semantic chapters in the first prompt word.
[0039] This application provides a method for generating reconstructed text. Through a pre-designed prompting process, a large language model is guided to complete standardized and interpretable semantic analysis and text reconstruction tasks, thereby ensuring the focus of the analysis and the consistency of the results. This embodiment inputs a preset first prompt word into the large language model, which clearly defines the semantic categories and boundaries of each semantic section. This prevents the large language model from performing aimless general text understanding, but rather, within a specific task framework, it performs semantic judgments similar to expert rules, accurately classifying the previously parsed subheadings into preset semantic sections. This prompt-based guidance effectively overcomes the problem of broad or off-focused outputs from the large language model, ensuring that the semantic recognition process is highly aligned with business analysis needs. After classifying each subheading, its corresponding subtext can be classified accordingly. This setting of classifying subheadings first and then subtexts simplifies the length of context the model needs to process and improves its processing efficiency. This embodiment essentially "solidifies" the classification logic and experience of domain experts into machine-executable instructions through prompt words, achieving a deep integration of expert experience and the capabilities of large language models. This provides a reliable guarantee for generating high-quality, trustworthy intermediate analytical products, and improves the effectiveness and accuracy of subsequent risk assessments of target companies.
[0040] In a preferred embodiment, the hierarchical JSON text obtained after text parsing (i.e., the first parsed text) is input into the large language model. The model classifies and divides the various text fragments in the first parsed text based on the classification dimensions set by the prompting project, which are not limited to a specific domain. Through this classification operation, the analysis benchmark of different subject texts is unified, and the comparable effect of each text in the analysis dimensions is finally achieved, resulting in the first reconstructed text after reclassification and adjustment.
[0041] For example, the specific content of a first prompt word is as follows: "You are a headline classification expert. Given a list of subheadings, please divide them into the following five categories:" Category 1: Industry Landscape and Trends Category 2: Company Business Plan Category 3: Potential Risks the Company May Face and Countermeasures Category 4: Future Outlook Category 5: Other To complete this task, the following requirements must also be met: 1. For categories 1-4, at most one subheading can match this category, or none may match. 2. Subheadings that do not fall under categories 1-4 should all be assigned to category 5. When all subheadings in the list belong to categories 1-4, category 5 will be empty. 3. If the subheading contains empty strings '', they must also be retained. 4. The output format is JSON, as shown in the example below: {{ "Industry Landscape and Trends": ["xxx"], "Company Business Plan": ["xxx"], "Potential Risks and Countermeasures the Company May Face": ["xxx"], "Future Outlook": ["xxx"], Other: ["xxx","yyy"] }} 5. Only output JSON that can be read by the program; do not add any other content.
[0042] 6. The output will be used for subsequent program parsing. Please ensure that the output subheadings do not contain any modified characters.
[0043] Given: A list of subtitles is given by: {subtitle}; Please answer: """ It's important to note that the input to the model consists of subheadings. Once these subheadings are categorized, their corresponding text can also be categorized accordingly. This setup simplifies the amount of context the model needs to process. Then, in step S4, simply extracting the text content of the "Future Outlook" section from the first reconstructed text yields the first target text corresponding to the target section.
[0044] In one possible implementation, in step S5, the step of performing a difference comparison analysis on the first target text and the second target text using the large language model according to a preset analysis logic, to obtain several difference texts and corresponding several change trends, includes: The preset second prompt word is input into the large language model, so that the large language model performs a difference comparison analysis on the first target text and the second target text according to the analysis logic set in the second prompt word, and determines several difference comparison groups in which there are differences in the first target text and the second target text. Each difference comparison group includes a first difference text segment in the first target text and a corresponding second difference text segment in the second target text. Based on each of the first and second difference text segments, semantic summaries are performed on each of the difference comparison groups to obtain the corresponding difference texts. The preset third prompt word is input into the large language model so that the large language model can perform a qualitative analysis of the target company's business status based on the third prompt word and each of the differential texts, and then generate the change trend corresponding to each of the differential texts.
[0045] This application's embodiments define the specific process of difference comparison analysis, greatly improving the depth and quality of the model's extraction of effective signals from changes in annual reports. Instead of simply performing a rough full-text comparison, this embodiment guides the large language model to perform intelligent difference detection by inputting a second prompt word with preset analysis logic. This analysis logic can be designed to focus on specific types of statement changes, such as identifying additions or deletions to arguments on the same topic, shifts in word strength, data updates, and refinements in risk descriptions. Based on this, the large language model identifies pairs of differing text fragments (difference comparison groups), which captures substantial semantic changes more effectively than simply statistically analyzing word frequency changes or using edit distance. Subsequently, these difference comparison groups are semantically summarized to extract the core differences, forming highly condensed "difference texts." This summarization process filters out redundant information and sentence structure differences, directly pointing to changes in the core content. Finally, guided by the third cue word, the large language model conducts a comprehensive qualitative analysis of the company's operating status based on all the extracted differential texts, and infers the changing trends of each differential text. This ensures that the final trends are not an overreaction to a single word or phrase, but rather a comprehensive inference result based on a series of semantic evidence. This significantly improves the stability, rationality, and interpretability of trend judgments, and enhances the effectiveness and accuracy of subsequent risk assessments of target companies.
[0046] In a preferred embodiment, the expert experience logic used for the "Future Outlook" section is intertemporal difference identification. This involves using a large language model to compare similar sections from two adjacent years of the same company, identifying new or missing statements, shifts in viewpoint, and changes in emphasis. For example, if one year emphasizes "expanding the number of stores," while the following year emphasizes "improving the profitability of individual stores," this is identified as a semantic difference. To make the differences interpretable, this embodiment introduces a positive / negative labeling mechanism after difference identification. Guided by the researcher's investment research logic, the large language model judges each difference, clarifying whether the change in business conditions it represents is "better" or "worse." For example, "adding digital construction" is labeled as better, while "no longer emphasizing industry consolidation trends" might be labeled as worse. The output is presented in a structured form, avoiding the problem of traditional NLP models only outputting fuzzy probabilities or non-directional information. It should be noted that the "intertemporal difference identification + positive and negative labeling" expert experience logic used in this embodiment is only a scenario-based selection adapted to the target chapter "future outlook" (focusing on changes in the company's expected business operations). For scenarios with other target chapters, the complete process framework of this methodology can be directly reused by replacing the corresponding expert experience logic to achieve targeted quantitative analysis (for example, for the scenario where the target chapter is "risks that the company may face and countermeasures", the "risk dimension classification + severity assessment" logic can be used; for the scenario where the target chapter is "industry pattern and trends", the "intertemporal data matching verification + abnormal fluctuation attribution" logic can be used, etc.).
[0047] For example, the large language model provides expert logic guidance through the following second cue words: "You are a senior expert in corporate management information analysis, mainly looking for investment opportunities in the A-share market, and hoping to identify changes in corporate business planning and strategy by analyzing the differences in the descriptions of future outlooks in the annual reports of listed companies in different years."
[0048] Now please carefully read the section on the company's future outlook in the annual reports of {stock_name} in {year} and {year1}, and analyze the semantic or expressive differences between the two texts.
[0049] The specific requirements for the task are as follows: 1. If the narrative emotions or perspectives on the same event differ between {year} and {year1}, it is considered a difference. 2. Content mentioned in only one text but not in another text is considered a difference, including content mentioned in {year} but not in {year1}, and content mentioned in {year1} but not in {year}.
[0050] 3. Each occurrence of the above situations is considered a new difference and must be listed one by one. 4. Your overall goal is to identify core changes in the company's business plan by analyzing the differences. Differences that are too detailed or do not affect the business plan analysis do not need to be mentioned. 5. Your overall goal is to comprehensively present the core changes in the company's business plan through the analysis of differences. For sections with significant differences, you can list them in multiple points according to their semantics, using quantity to demonstrate the substantial nature of the differences. 6. Areas where there are no significant differences do not need to be mentioned. 7. The output format is JSON, as shown in the example below: {{"Difference 1":"xxx","Difference 2":"xxx"}} 8. If there is no difference, please output an empty JSON, as shown in the example below: {{}} 9. The output will be used for subsequent program parsing. Please ensure that the output is JSON that can be read by Python and do not add any other content.
[0051] The outlook for {year} is: {content} +'\n' The outlook for year {year1} is: {content1} +'\n Please answer: "" The structured output of the large language model is obtained, namely the various differing texts, as shown in the following example: {"Difference 1": "The 2024 annual report proposed the annual operating policy of 'seizing the window of opportunity and seeking breakthroughs,' while the 2023 annual report proposed the annual operating policy of 'maintaining focus and resilient growth.'" "Difference 2": "The 2024 report mentioned adjusting the store structure and increasing revenue per store, while the 2023 report emphasized maintaining the pace of store expansion and increasing market share." "Difference 3": "The 2024 report emphasized focusing on core business and improving profitability, while the 2023 report, although mentioning increasing market share, did not specifically mention improving profitability." "Difference 4":} "The 2024 report mentioned expanding the membership system and strengthening digital construction, while the 2023 report mentioned diversifying and upgrading marketing methods and continuing to focus on brand rejuvenation. The two reports have different emphases in digitalization and brand marketing." "Difference 5": "The 2024 report mentioned emphasizing risk control and expanding into overseas markets, while the 2023 report also mentioned accelerating brand globalization and expanding into the global market, but the 2024 report places greater emphasis on risk control." "Difference 6": "The 2024 report did not mention continuously optimizing the management structure and strengthening the group's organizational capabilities, while the 2023 report did mention this." Furthermore, based on expert experience in prompting engineering using the aforementioned large language model output results, a pre-set third prompt word is input into the large language model: "You are a business information analysis expert, proficient in interpreting listed company annual reports. It is known that {stock_name} has the following differences in its statements in the annual reports of {year} and {year1}:" {diff} Please analyze whether this difference indicates that the company's operating performance in year {year} improved or worsened compared to year {year1}. Please note that company reports will not include irrelevant information. If you believe the difference does not clearly indicate the company's operating performance, please conduct a more in-depth fundamental analysis and reasonable inference based on your financial knowledge. Your output can only be one of the following two options: improved or worsened.
[0052] Please note that only a two-word conclusion is required; no further analysis or explanation is needed.
[0053] Please answer: """ The changing trends corresponding to each of the aforementioned differing texts are shown in the following examples: {"Difference 1":"Better","Difference 2":"Better","Difference 3":"Better","Difference 4":"Better","Difference 5":"Better","Difference 6":"Better"} In one possible implementation, in step S6, calculating the quantitative analysis indicators of the first annual report text based on the number of each of the differing texts, the respective trends of change, and the average number of words per chapter includes: The first analytical index of the first annual report text is calculated by dividing the number of each of the aforementioned differential texts by the average number of words in each chapter. The difference between the positive and negative trends in each of the aforementioned trends is divided by the average number of words in each chapter to obtain the second analytical indicator of the first annual report text. By combining the first and second analytical indicators, quantitative analytical indicators for obtaining the first annual report text are constructed.
[0054] This application provides a method for calculating quantitative analysis indicators, creatively transforming the semantically rich qualitative analysis results produced by a large language model into concise, objective, and mathematically computable standardized quantitative indicators that can be compared over time. Specifically, regarding the first analysis indicator, a "density of information change per unit text length" indicator is obtained by dividing the number of differing texts by the average number of words per chapter. This indicator effectively eliminates the impact of length differences between annual reports, standardizing the measurement of the activity level or information update intensity of annual report content changes. Regarding the second analysis indicator, a "net sentiment or net trend direction per unit text length" indicator is obtained by calculating the net value (difference) of positive and negative trends and dividing it by the average number of words per chapter. This indicator not only reflects the quantity of change but also captures the "net effect" of the nature of the change, such as whether the enhancement of positive business descriptions outweighs the increase in risk warnings. The set of quantitative analysis indicators constructed by combining these two indicators has profound business implications: the first indicator can be regarded as a proxy variable for "information uncertainty" or "change in management communication"; the second indicator can be regarded as a proxy variable for "net sentiment about business outlook" or "net change in risk exposure". This quantitative approach allows previously vague and qualitative text analysis conclusions to be integrated into financial econometric models, risk scoring cards, or monitoring dashboards in numerical form. This facilitates trend tracking, threshold warnings, cross-sectional ranking, and regression analysis, addressing the pain point of "emphasizing qualitative analysis and neglecting quantitative analysis" in the application of large language models, and improving the effectiveness and accuracy of subsequent risk assessments of target companies.
[0055] In a preferred embodiment, the text and semantic analysis results are converted into numerical indicators and quantitative indicators are constructed by following the principles: 1. The indicators are derived from the objective transformation of text semantic analysis results, eliminating interference from human subjective judgment, avoiding the hollowness of quantitative indicators, and improving the accuracy and reliability of quantitative results.
[0056] 2. Standardized processing logic ensures the comparability of indicators across subjects and time dimensions, breaks down the inherent differences between different texts, and expands the applicability of technical solutions; 3. The hierarchical indicator design balances universality and specificity, making it suitable for overall text quantification as well as for extracting signals from specific dimensions of content, thereby enhancing the explanatory power of the indicators; Specific examples are as follows: Indicator 1: Number of Differences / Average Chapter Word Count. Reason for design: The length of annual reports varies greatly among different companies. By comparing this to the "average chapter word count," the inherent interference of text length can be eliminated, allowing the indicator to truly reflect the degree of information fluctuation within the corresponding text dimension and ensuring the comparability of indicators for companies of different sizes.
[0057] Indicator 2: (Number of positive differences - Number of negative differences) / Average number of chapter words. Reason for design: Introducing the "positive-negative" calculation logic converts positive / negative expressions in the text into numerical signals; then, comparing it with the average number of chapter words further eliminates interference from text length, ultimately allowing the indicator to accurately reflect the company's development trend and operating conditions.
[0058] Furthermore, in step S7, the risk assessment of the target enterprise based on the quantitative analysis indicators includes: Obtain historical quantitative analysis indicators from the annual report texts of the target company in different years. Based on the historical first analysis indicators and the numerical values of the first analysis indicators in each of the aforementioned historical quantitative analysis indicators, a first change curve is constructed. The second change curve is determined based on the values of each historical second analysis indicator and the second analysis indicator in each of the aforementioned historical quantitative analysis indicators; Based on the slope value of the first change curve within a preset sliding time window, generate the first evaluation text; Based on the value of the second change curve within a preset sliding time window, a second evaluation text is generated; Based on the numerical values of the first and second analytical indicators, a third evaluation text is generated; The first assessment text, the second assessment text, and the third assessment text are combined to generate the risk assessment result of the target company.
[0059] This application provides a risk assessment method based on quantitative analysis indicators, which greatly enhances the systematicness, objectivity, and forward-looking nature of risk assessment. Specifically, this embodiment does not rely on isolated indicators from a single year, but first obtains historical quantitative analysis indicators of the target company over many years, thereby expanding the analytical perspective from "points" to "lines" and "surfaces." By constructing historical change curves for the first and second analysis indicators, the long-term evolution trajectory of the company's information disclosure behavior, operational communication tone, and risk situation can be intuitively displayed. Based on this, dynamic analysis is further introduced: a first assessment text is generated based on the slope of the first change curve within a sliding time window. Through fluctuations in indicator values, the stability of the company's business expectations and strategic planning is objectively reflected. For example, if the first change curve continues to rise, it indicates that the company's business planning adjustments are frequent and operational uncertainty is increased. A second assessment text is generated based on the values of the second change curve within the sliding window, which can assess the recent level range of the company's business sentiment or risk situation. For example, if the second change curve turns from positive to negative or continues to decline, it indicates that the company's business expectations are shifting towards a conservative or pressured direction. Finally, a third assessment text is generated based on the values of the first and second analysis indicators for the current year, accurately assessing the target company's current operating status. Ultimately, combining the assessment texts from these three dimensions results in a risk assessment that is no longer simply labeled as "high risk" or "low risk," but rather a comprehensive analysis report that includes dynamic interpretation, historical comparison, and current status positioning. This provides more accurate clues and evidence for risk warning and in-depth due diligence, improving the effectiveness and accuracy of risk assessment of target companies.
[0060] Furthermore, the risk assessment results also include a fourth assessment text, which is generated by comparing the target company's various peer companies' current second analysis indicators. Specifically: Obtain the current second analysis indicators of each peer company of the target company; The number of positive and negative values in each of the current second analysis indicators is statistically obtained. The fourth evaluation text is generated based on the number of positive values, the number of negative values, and a preset quantity threshold.
[0061] Furthermore, this embodiment adds a fourth assessment text to the risk assessment results, introducing a dimension of horizontal industry comparison. This allows the risk assessment conclusions to move beyond focusing solely on the company's own historical data, providing a broader market coordinate system and relative evaluation perspective, significantly improving the comprehensiveness and fairness of the risk assessment. This embodiment generates the fourth assessment text based on the number of positive and negative values of the second analytical indicator for each of the target company's peers. This is because the second analytical indicator reflects the net sentiment or net trend direction conveyed in the company's annual report. By statistically analyzing the positive and negative values of this indicator for multiple comparable companies within an industry, it is possible to determine whether the overall sentiment of the industry is optimistic or pessimistic, or the magnitude of the common risk pressure faced by the industry. For example, if the target company's second indicator is negative, but the vast majority of its peers also have negative indicators, then this negative sentiment may stem more from industry-wide shocks than from company-specific problems. Conversely, if the industry as a whole has positive values, but the target company alone has negative values, it strongly suggests that the company may be facing individual operational difficulties or risks, and its risk level requires special attention. Therefore, the fourth assessment text provides crucial contextual information, helping users distinguish between "systemic risk" and "specific risk." It calibrates the company's textual signals within an industry context, avoiding misjudgments of individual companies due to overall industry fluctuations, and further improving the effectiveness and accuracy of risk assessments of target companies.
[0062] In a preferred embodiment, there are three risk assessment methods: 1. Single-Entity Vertical Tracking: The indicators "Number of Differences / Average Chapter Word Count" and "(Number of Positive Differences - Number of Negative Differences) / Average Chapter Word Count" are applied to the cross-period monitoring of the same entity. By constructing corresponding change curves, the stability of the company's business expectations and strategic planning is objectively assessed based on the fluctuations of these curves. For example, if "Number of Differences / Average Chapter Word Count" continuously increases, the first change curve will show a slope greater than 0 over a continuous period, indicating that the company's business plan is adjusted frequently and business uncertainty is increased. If "(Number of Positive Differences - Number of Negative Differences) / Average Chapter Word Count" changes from positive to negative or continues to decline, the second change curve will show a vertical axis value changing from positive to negative or both falling below a preset threshold over a continuous period, indicating that the company's business expectations are shifting towards a more conservative or pressured direction.
[0063] 2. Risk warning threshold setting: Combining historical industry data and business operation patterns, set reasonable fluctuation thresholds for two indicators (for example, if the indicator is cross-sectionally normalized to a standard normal distribution in a specific stock pool, the threshold can be set to greater than 3 or less than -3). When the first or second analysis indicator exceeds the corresponding threshold range, an operational risk warning is triggered, providing signal guidance for subsequent targeted in-depth analysis of the business situation.
[0064] 3. Horizontal Benchmarking within the Same Industry: By comparing two indicators across different entities within the same industry, we can identify the divergence in operational stability and the degree of positive outlook within the industry, providing objective data support for the assessment of industry operational characteristics. For example, if most companies in the industry have a positive "(number of positive differences - number of negative differences) / average number of chapters" indicator, it reflects a positive overall operational outlook for the industry. If some companies have this indicator significantly lower than the industry average, we can further focus on the rationality of their operational planning adjustments and potential risks.
[0065] Those skilled in the art can freely choose the above-mentioned risk assessment methods or combine them based on actual needs, and then use the text generation capabilities of the large language model to generate corresponding assessment texts based on the numerical analysis results, and then combine them to obtain the final risk assessment result.
[0066] In summary, the embodiments of this application, by focusing the large model on text parsing and semantic analysis, output targeted difference information under the logical guidance of researchers, and transform these differences into calculable quantitative indicators, achieving a stable transformation from unstructured text to quantitative signals. This technical solution leverages the semantic understanding advantages of large language models while avoiding the uncontrollable risks associated with direct investment decisions, thus possessing significant application value.
[0067] Example 2: like Figure 3 As shown, Embodiment 2 provides an enterprise risk assessment system based on a large language model and text quantitative analysis, including an acquisition module 10, a text parsing module 20, a semantic recognition module 30, a text extraction module 40, an analysis module 50, an indicator calculation module 60, and a risk assessment module 70. The acquisition module 10 is used to acquire the current first annual report text of the target enterprise; The text parsing module 20 is used to parse the first annual report text to obtain the corresponding first parsed text, wherein the first parsed text includes several subheadings and several corresponding subtexts; The semantic recognition module 30 is used to perform semantic recognition and classification on each of the subheadings using a preset large language model, and to divide the subtexts corresponding to each of the subheadings into corresponding semantic chapters according to the semantic recognition and classification results, thereby generating the corresponding first reconstructed text; The text extraction module 40 is used to extract the corresponding first target text from the first reconstructed text according to the preset target semantic chapter; The analysis module 50 is used to perform a difference comparison analysis on the first target text and the second target text through the large language model according to the preset analysis logic, and obtain several difference texts and corresponding several change trends. The second target text is obtained by sequentially parsing, semantic recognition and classification and text extraction of the second annual report text of the target company in the previous year. The indicator calculation module 60 is used to calculate the quantitative analysis indicators of the first annual report text based on the number of each of the different texts, the various change trends and the average number of words in each chapter, wherein the average number of words in each chapter is calculated based on the total number of words in the first annual report text and the second annual report text and the number of semantic chapters; The risk assessment module 70 is used to conduct a risk assessment of the target enterprise based on the quantitative analysis indicators and generate risk assessment results.
[0068] Furthermore, the text parsing module 20 performs text parsing on the first annual report text to obtain the corresponding first parsed text. The first parsed text includes several subheadings and corresponding several subtexts, including: Convert the first annual report text into markup language format to obtain the corresponding first format converted text; The first formatted text is matched with headings and divided into heading levels according to a preset regular expression, and several subheadings, corresponding heading levels and corresponding subtexts are determined. Based on each of the subheadings, the corresponding heading levels, and the corresponding subtexts, the first format-converted text is subjected to structured processing to generate the first parsed text with a heading hierarchy.
[0069] Furthermore, the semantic recognition module 30 performs semantic recognition and classification on each subheading using a preset large language model, and divides the subtext corresponding to each subheading into the corresponding semantic chapter according to the semantic recognition and classification results, thereby generating the corresponding first reconstructed text, including: The preset first prompt word is input into the large language model so that the large language model can perform semantic recognition and classification of each subheading in the first parsed text according to the semantic definition of each semantic chapter in the first prompt word, and determine the semantic chapter to which each subheading belongs; Based on the semantic chapter to which each subheading belongs, the subtext corresponding to each subheading is divided into the corresponding semantic chapter, thereby determining the text content of each semantic chapter; The first reconstructed text is generated based on each of the semantic chapters, the corresponding text content, and the order and output format settings of each of the semantic chapters in the first prompt word.
[0070] In one possible implementation, the analysis module 50, based on preset analysis logic, performs a difference comparison analysis on the first target text and the second target text using the large language model to obtain several difference texts and corresponding several change trends, including: The preset second prompt word is input into the large language model, so that the large language model performs a difference comparison analysis on the first target text and the second target text according to the analysis logic set in the second prompt word, and determines several difference comparison groups in which there are differences in the first target text and the second target text. Each difference comparison group includes a first difference text segment in the first target text and a corresponding second difference text segment in the second target text. Based on each of the first and second difference text segments, semantic summaries are performed on each of the difference comparison groups to obtain the corresponding difference texts. The preset third prompt word is input into the large language model so that the large language model can perform a qualitative analysis of the target company's business status based on the third prompt word and each of the differential texts, and then generate the change trend corresponding to each of the differential texts.
[0071] In one possible implementation, the indicator calculation module 60 calculates quantitative analysis indicators for the first annual report text based on the number of each of the differing texts, the respective trends of change, and the average number of words per chapter, including: The first analytical index of the first annual report text is calculated by dividing the number of each of the aforementioned differential texts by the average number of words in each chapter. The difference between the positive and negative trends in each of the aforementioned trends is divided by the average number of words in each chapter to obtain the second analytical indicator of the first annual report text. By combining the first and second analytical indicators, quantitative analytical indicators for obtaining the first annual report text are constructed.
[0072] Furthermore, the risk assessment module 70 performs a risk assessment on the target enterprise based on the quantitative analysis indicators, including: Obtain historical quantitative analysis indicators from the annual report texts of the target company in different years. Based on the historical first analysis indicators and the numerical values of the first analysis indicators in each of the aforementioned historical quantitative analysis indicators, a first change curve is constructed. The second change curve is determined based on the values of each historical second analysis indicator and the second analysis indicator in each of the aforementioned historical quantitative analysis indicators; Based on the slope value of the first change curve within a preset sliding time window, generate the first evaluation text; Based on the value of the second change curve within a preset sliding time window, a second evaluation text is generated; Based on the numerical values of the first and second analytical indicators, a third evaluation text is generated; The first assessment text, the second assessment text, and the third assessment text are combined to generate the risk assessment result of the target company.
[0073] Furthermore, the risk assessment results also include a fourth assessment text, which is generated by comparing the target company's various peer companies' current second analysis indicators. Specifically: Obtain the current second analysis indicators of each peer company of the target company; The number of positive and negative values in each of the current second analysis indicators is statistically obtained. The fourth evaluation text is generated based on the number of positive values, the number of negative values, and a preset quantity threshold.
[0074] This application provides an enterprise risk assessment system based on a large language model and text quantitative analysis, significantly improving the systematicness, accuracy, and practicality of extracting quantitative information from unstructured text. First, this embodiment combines text parsing with semantic recognition and classification to transform loosely structured annual report text into reconstructed text with clearly defined semantic sections, effectively solving the inherent problem of directly quantifying unstructured data. Utilizing the deep semantic understanding capabilities of the large language model, it can accurately identify the semantic connotations of subheadings and classify them according to preset semantic definitions. This avoids semantic biases caused by traditional keyword matching or shallow statistical methods, ensuring the accuracy and consistency of information segmentation. Second, through cross-year difference comparison analysis, this embodiment can automatically identify changes in key text content such as business descriptions and risk disclosures, and extract corresponding trends. This semantic-based difference analysis goes beyond simple text addition and subtraction statistics; it captures the evolution of deeper information such as tone, emphasis, and risk exposure levels, providing richer and more insightful observations for judging enterprise operational dynamics. Finally, by combining qualitative information such as the number of differentially expressed texts and trends with the average number of words per chapter, a series of quantitative analysis indicators were innovatively calculated and used for risk assessment. This process successfully constructed a complete technical chain from unstructured text to structured quantitative indicators, and then to risk assessment conclusions, greatly enhancing the measurability, comparability, and direct application value of the analysis results. This embodiment effectively suppresses the defects of "illusion" and unstable output by guiding the large language model to perform standardized and logical analysis tasks, making the generated risk assessment results combine the efficiency of automated processing with the depth of expert analysis logic, thereby improving the effectiveness and accuracy of risk assessment for target enterprises.
[0075] For a more detailed explanation of the working principle and procedures of this embodiment, please refer to the relevant description in Embodiment 1.
[0076] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.
Claims
1. A method for enterprise risk assessment based on large language models and text quantitative analysis, characterized in that, include: Obtain the target company's current first annual report text; The first annual report text is parsed to obtain the corresponding first parsed text, wherein the first parsed text includes several subheadings and several corresponding subtexts; The subheadings are semantically identified and categorized using a pre-defined large language model. Based on the semantic identification and categorization results, the subtexts corresponding to each subheading are divided into corresponding semantic chapters, thereby generating the corresponding first reconstructed text. The first target text is extracted from the first reconstructed text according to the preset target semantic chapter; According to the preset analysis logic, the first target text and the second target text are compared and analyzed by the big language model to obtain several different texts and several corresponding change trends. The second target text is obtained by sequentially parsing, semantic recognition and classification and text extraction of the second annual report text of the target company in the previous year. Based on the number of each of the aforementioned differential texts, the aforementioned trends of change, and the average number of words per chapter, quantitative analysis indicators for the first annual report text are calculated, wherein the average number of words per chapter is calculated based on the total number of words in the first and second annual report texts and the number of semantic chapters. A risk assessment is conducted on the target company based on the quantitative analysis indicators, and a risk assessment result is generated.
2. The enterprise risk assessment method based on large language models and text quantitative analysis as described in claim 1, characterized in that, The first annual report text is parsed to obtain the corresponding first parsed text, which includes several subheadings and corresponding subtexts, including: Convert the first annual report text into markup language format to obtain the corresponding first format converted text; The first formatted text is matched with headings and divided into heading levels according to a preset regular expression, and several subheadings, corresponding heading levels and corresponding subtexts are determined. Based on each of the subheadings, the corresponding heading levels, and the corresponding subtexts, the first format-converted text is subjected to structured processing to generate the first parsed text with a heading hierarchy.
3. The enterprise risk assessment method based on large language models and text quantitative analysis as described in claim 1, characterized in that, The step of performing semantic recognition and classification on each subheading using a preset large language model, and dividing the subtext corresponding to each subheading into corresponding semantic chapters based on the semantic recognition and classification results, thereby generating the corresponding first reconstructed text, includes: The preset first prompt word is input into the large language model so that the large language model can perform semantic recognition and classification of each subheading in the first parsed text according to the semantic definition of each semantic chapter in the first prompt word, and determine the semantic chapter to which each subheading belongs; Based on the semantic chapter to which each subheading belongs, the subtext corresponding to each subheading is divided into the corresponding semantic chapter, thereby determining the text content of each semantic chapter; The first reconstructed text is generated based on each of the semantic chapters, the corresponding text content, and the order and output format settings of each of the semantic chapters in the first prompt word.
4. The enterprise risk assessment method based on large language models and text quantitative analysis as described in claim 1, characterized in that, According to the preset analysis logic, the first target text and the second target text are compared and analyzed using the large language model to obtain several differing texts and corresponding trends, including: The preset second prompt word is input into the large language model, so that the large language model performs a difference comparison analysis on the first target text and the second target text according to the analysis logic set in the second prompt word, and determines several difference comparison groups in which there are differences in the first target text and the second target text. Each difference comparison group includes a first difference text segment in the first target text and a corresponding second difference text segment in the second target text. Based on each of the first and second difference text segments, semantic summaries are performed on each of the difference comparison groups to obtain the corresponding difference texts. The preset third prompt word is input into the large language model so that the large language model can perform a qualitative analysis of the target company's business status based on the third prompt word and each of the differential texts, and then generate the change trend corresponding to each of the differential texts.
5. The enterprise risk assessment method based on large language models and text quantitative analysis as described in claim 1, characterized in that, The quantitative analysis indicators for the first annual report text are calculated based on the number of each of the aforementioned differential texts, the aforementioned trends in change, and the average number of words per chapter, including: The first analytical index of the first annual report text is calculated by dividing the number of each of the aforementioned differential texts by the average number of words in each chapter. The difference between the positive and negative trends in each of the aforementioned trends is divided by the average number of words in each chapter to obtain the second analytical indicator of the first annual report text. By combining the first and second analytical indicators, quantitative analytical indicators for obtaining the first annual report text are constructed.
6. The enterprise risk assessment method based on large language models and text quantitative analysis as described in claim 5, characterized in that, The risk assessment of the target enterprise based on the quantitative analysis indicators includes: Obtain historical quantitative analysis indicators from the annual report texts of the target company in different years. Based on the historical first analysis indicators and the numerical values of the first analysis indicators in each of the aforementioned historical quantitative analysis indicators, a first change curve is constructed. The second change curve is determined based on the values of each historical second analysis indicator and the second analysis indicator in each of the aforementioned historical quantitative analysis indicators; Based on the slope value of the first change curve within a preset sliding time window, generate the first evaluation text; Based on the value of the second change curve within a preset sliding time window, a second evaluation text is generated; Based on the numerical values of the first and second analytical indicators, a third evaluation text is generated; The first assessment text, the second assessment text, and the third assessment text are combined to generate the risk assessment result of the target company.
7. The enterprise risk assessment method based on large language models and text quantitative analysis as described in claim 6, characterized in that, The risk assessment results also include a fourth assessment text, which is generated by comparing the target company's various peer companies with their current second analytical indicators. Specifically: Obtain the current second analysis indicators of each peer company of the target company; The number of positive and negative values in each of the current second analysis indicators is statistically obtained. The fourth evaluation text is generated based on the number of positive values, the number of negative values, and a preset quantity threshold.
8. A business risk assessment system based on large language models and text quantitative analysis, characterized in that, It includes an acquisition module, a text parsing module, a semantic recognition module, a text extraction module, an analysis module, an indicator calculation module, and a risk assessment module; The acquisition module is used to acquire the text of the target company's current first annual report. The text parsing module is used to parse the first annual report text to obtain the corresponding first parsed text, wherein the first parsed text includes several subheadings and several corresponding subtexts; The semantic recognition module is used to perform semantic recognition and classification on each of the subheadings using a preset large language model, and to divide the subtexts corresponding to each of the subheadings into corresponding semantic chapters based on the semantic recognition and classification results, thereby generating the corresponding first reconstructed text; The text extraction module is used to extract the corresponding first target text from the first reconstructed text according to the preset target semantic chapter; The analysis module is used to perform a difference comparison analysis on the first target text and the second target text according to the preset analysis logic and the large language model to obtain several difference texts and several corresponding change trends. The second target text is obtained by sequentially parsing, semantic recognition and classification and text extraction of the second annual report text of the target company in the previous year. The indicator calculation module is used to calculate the quantitative analysis indicators of the first annual report text based on the number of each of the different texts, the various change trends, and the average number of words per chapter. The average number of words per chapter is calculated based on the total number of words in the first annual report text and the second annual report text and the number of semantic chapters. The risk assessment module is used to conduct risk assessment on the target enterprise based on the quantitative analysis indicators and generate risk assessment results.
9. The enterprise risk assessment system based on large language models and text quantitative analysis as described in claim 8, characterized in that, The analysis module, based on preset analysis logic, performs a difference comparison analysis on the first target text and the second target text using the large language model, obtaining several difference texts and corresponding change trends, including: The preset second prompt word is input into the large language model, so that the large language model performs a difference comparison analysis on the first target text and the second target text according to the analysis logic set in the second prompt word, and determines several difference comparison groups in which there are differences in the first target text and the second target text. Each difference comparison group includes a first difference text segment in the first target text and a corresponding second difference text segment in the second target text. Based on each of the first and second difference text segments, semantic summaries are performed on each of the difference comparison groups to obtain the corresponding difference texts. The preset third prompt word is input into the large language model so that the large language model can perform a qualitative analysis of the target company's business status based on the third prompt word and each of the differential texts, and then generate the change trend corresponding to each of the differential texts.
10. The enterprise risk assessment system based on large language models and text quantitative analysis as described in claim 8, characterized in that, The indicator calculation module calculates quantitative analysis indicators for the first annual report text based on the number of each differentiated text, each changing trend, and the average number of words per chapter, including: The first analytical index of the first annual report text is calculated by dividing the number of each of the aforementioned differential texts by the average number of words in each chapter. The difference between the positive and negative trends in each of the aforementioned trends is divided by the average number of words in each chapter to obtain the second analytical indicator of the first annual report text. By combining the first and second analytical indicators, quantitative analytical indicators for obtaining the first annual report text are constructed.