Financial risk intelligent early warning system and method based on large language model
Through the intelligent early warning system for financial risks based on large language models, problems such as inefficiency and limited coverage of traditional risk analysis methods are solved, efficient and real-time risk identification and management are achieved, and the efficiency and accuracy of risk management are improved.
Patent Information
- Application Number
- CN202510668040.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
AI Technical Summary
Traditional risk analysis methods are inefficient, limited coverage, poor timeliness, insufficient flexibility and over-reliance on subjective judgments, making it difficult to cope with the complexity of modern financial markets and big data challenges.
The intelligent early warning system for financial risks based on large language models is adopted, multi-source data is collected through the data integration module, real-time task scheduling engine is used for data integration, data preprocessing unit cleans and deduplication, the large language model analysis module conducts in-depth analysis of public opinion and financial indicators, the comprehensive evaluation module generates a comprehensive risk score, and triggers the early warning mechanism through the early warning module.
It has achieved efficient processing of massive data and multi-dimensional risk identification, with real-time and dynamic response mechanisms, which can accurately identify potential risks and improve the efficiency and accuracy of risk management.
Smart Images

Figure CN120198231A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to financial risk early warning technology, and in particular to a financial risk intelligent early warning system and method based on a large language model. Background Art
[0002] With the rapid development of financial markets, information related to investment entities is growing exponentially, and traditional risk analysis methods have obvious limitations when facing modern financial markets. First, relying on manual data processing makes analysis inefficient, especially when the amount of data is huge, manual screening of information is prone to omissions or errors, and it is difficult to identify potential risks in a timely manner. Secondly, traditional methods have limited coverage and usually rely on financial data and market conditions, but cannot effectively incorporate unstructured data such as public opinion and social media content, which makes it impossible for analysis to fully capture the multi-dimensional risks in the market. Furthermore, traditional analysis methods have too high requirements for timeliness, rely on periodically updated reports and data sources, and cannot quickly respond to market fluctuations or emergencies, resulting in delayed response. Traditional risk analysis is often based on static models, lacks flexibility, and cannot adapt to the rapid changes in the market environment.
[0003] In addition, traditional methods have weak processing capabilities for unstructured data and are unable to fully tap into important risk signals from news and social media sources. They also lack the ability to process big data. Traditional methods perform poorly when the amount of data is large, and it is difficult to quickly process and analyze data in real time. They usually lack flexibility and scalability, and cannot quickly adapt to market changes or adjustments in business needs. Finally, traditional methods rely too much on the experience and subjective judgment of analysts, are easily affected by human bias and cognitive errors, and lack objectivity and systematicity.
[0004] In short, traditional risk analysis methods are unable to cope with the complexity and big data challenges of modern financial markets, and need to rely on emerging technologies such as artificial intelligence and big data analysis to achieve more efficient and accurate risk assessment and early warning. Summary of the invention
[0005] In view of this, the present invention aims to address the deficiencies in the prior art, and its main purpose is to provide a financial risk intelligent early warning system and method based on a large language model to solve the problems of low efficiency, limited coverage, poor timeliness, lack of flexibility and over-reliance on subjective judgment in traditional risk analysis methods in the prior art.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A financial risk intelligent early warning system based on a large language model, comprising: Data integration module, used to connect to multiple financial data sources and collect multi-dimensional data on listed companies’ financial indicators and authoritative industry white papers; A real-time task scheduling engine, which is used to integrate the collected data, supports dual-mode data collection of timed polling and event triggering, and constructs a 7×24-hour intelligent monitoring network; A data preprocessing unit, which is used to clean, de-duplicate, and convert the format of the collected data, and extract effective information closely related to a specific main company; A large language model analysis module, which is used to analyze the preprocessed public opinion data and financial indicators, including steps of preliminary screening, polarity judgment, and classification rating; A comprehensive evaluation module, which is used to summarize the classification rating results, conduct a comprehensive rating in combination with financial indicators, and generate a comprehensive risk score; An early warning module, which is used to trigger an early warning mechanism according to the comprehensive risk score and send an alarm notification to relevant decision-makers; The data collected by the data integration module is integrated by the real-time task scheduling engine and then transmitted to the data preprocessing unit. The data processed by the data preprocessing unit is transmitted to the large language model analysis module. The analysis results of the large language model analysis module are transmitted to the comprehensive evaluation module. The comprehensive risk score of the comprehensive evaluation module is transmitted to the early warning module.
[0007] As a preferred solution: The data integration module adopts a standardized interface technology and supports docking with financial data platforms such as Caihui, DM, Wind, and YY.
[0008] As a preferred solution: The real-time task scheduling engine is developed based on the SpringBoot framework, realizes task registration through the @EnableScheduling annotation, dynamically loads task policies using Spring Cloud Config, adopts a master-slave architecture design, monitors the node survival status in real time based on the Spring Boot Actuator health check endpoint, and realizes dynamic switching of the master and standby roles through the @ConditionalOnProperty conditional annotation. The task allocation strategy relies on SpringBoot Metrics to collect CPU load and memory usage rate indicators of each node in real time, and dynamically calculates the node weight value through the weighted round-robin algorithm.
[0009] As a preferred solution: The processing steps of the data preprocessing unit include: data cleaning, removing irrelevant noise data and invalid information; de-duplication processing, using efficient algorithms and technologies to identify and eliminate duplicate records; format conversion, unifying data from different sources into a standard format; data association and integration, using the unified social credit code of the enterprise as the core association identifier to achieve entity alignment and information integration across data sources.
[0010] As a preferred solution: when the large language model analysis module conducts classification and rating, it includes detailed ratings of strategic risk, compliance risk, financial risk, market risk, operational risk, and legal risk. The ratings under each risk category generate specific results based on specific preset instructions or question templates.
[0011] As a preferred solution: during the deduplication process of the data preprocessing unit, the SimHash algorithm is used to calculate the text similarity, and the redundancy of the data is evaluated based on the text similarity. When the similarity threshold satisfies the following formula, it is determined as redundant content and corresponding processing is performed: Similarity = 1 - Hamming distance / fingerprint length ≥ 95%; where the fingerprint length is the length of the binary fingerprint generated by SimHash.
[0012] As a preferred solution: during the task allocation process of the real-time task scheduling engine, a specific task allocation strategy is adopted to optimize task allocation, and this strategy is expressed as: Node weight = α・CPU load + β・memory usage rate + γ・IO throughput; where α, β, and γ are preset coefficients, and α + β + γ = 1.
[0013] As a preferred solution: when the comprehensive evaluation module conducts comprehensive rating, it considers the impact of market fluctuations on risks and dynamically adjusts the threshold according to the industry β coefficient. The adjustment formula is: Comprehensive risk score threshold = benchmark threshold × (1 + industry β coefficient × market volatility); where the industry β coefficient is the covariance / variance of the industry index and the market index in the most recent 12 months.
[0014] As a preferred solution: when the comprehensive evaluation module conducts comprehensive rating, it adjusts the market risk weight by combining liquidity factors to determine the final risk weight. This weight adjustment follows the following formula: Final risk weight = market risk weight × [1 + (1 - average trading volume in the past 5 days / benchmark trading volume)]; where the benchmark trading volume is the average daily trading volume in the past year; 1 + (1 - average trading volume in the past 5 days / benchmark trading volume) is the liquidity adjustment factor.
[0015] As a preferred solution: when the comprehensive evaluation module generates a comprehensive risk score, the following formula is used for calculation: ; where i is the risk dimension serial number, n is the total number of risk dimensions, and the normalized weight is adjusted by the dynamic calibration rule engine.
[0016] A method applied to the financial risk intelligent early warning system based on the large language model includes the following steps: S1. Connect to multiple financial data sources through the data integration module, and collect the financial indicators of listed companies and multi-dimensional data of industry white papers; S2. Use the real-time task scheduling engine to integrate the collected data, support dual-mode data collection of regular polling and event triggering, and build a 7×24-hour intelligent monitoring network; S3. The data collected is cleaned, de-duplicated, and format-converted by the data preprocessing unit to extract effective information closely related to the specific subject company; S4. Analyze the preprocessed public opinion data and financial indicators through the large language model analysis module, including steps of preliminary screening, polarity judgment, and classification rating; S5. The comprehensive evaluation module summarizes the classification rating results, combines the financial indicators for comprehensive rating, and generates a comprehensive risk score; S6. If major risks or abnormal signals are detected, trigger the warning module and send an alarm notification to the decision maker.
[0017] Compared with the prior art, the present invention has obvious advantages and beneficial effects. Specifically, as can be seen from the above technical solutions, firstly, by efficiently processing massive data and multi-dimensional risk identification, it can quickly analyze and process massive unstructured data from news reports, social media content, and policy announcement sources, getting rid of the constraints of the low efficiency and limitations of traditional data processing means, enabling financial institutions to capture dynamic changes in real time in the ever-changing market environment and optimize investment strategies. Secondly, it realizes multi-dimensional risk identification and in-depth understanding, extracts risk information from various types of unstructured data, provides a comprehensive and detailed risk assessment framework, helps to reveal subtle trends and signals hidden in a large amount of data, and provides more accurate risk predictions for investors. Thirdly, it has real-time and dynamic response mechanisms, can quickly capture sudden market changes or public opinion fluctuations, provide risk warnings immediately, help investors and decision makers respond quickly, and reduce investment risks caused by information lag. In addition, through the application of deep semantic understanding and context association, it processes complex factors in the text, such as context, emotion, and tone, and extracts potential risk signals more accurately, greatly improving the accuracy of risk identification. It can also automate and personalize risk control, automatically generate personalized risk tips according to real-time data and specific investment needs, improve the efficiency of risk management, and support investors to flexibly adjust strategies according to the latest market dynamics. Finally, relying on the scalability and flexibility of the large language model, it adapts to the changing market environment and new types of risks, integrates data from multiple channels across platforms for comprehensive analysis, and further enhances the comprehensiveness and accuracy of risk identification and decision support.
[0018] To more clearly illustrate the structural features and functions of the present invention, the following will be described in detail in combination with the accompanying drawings and specific embodiments. Brief Description of the Drawings
[0019] Figure 1 It is a schematic diagram of the data flow of the application of the large language model of the present invention in the field of financial risks; Figure 2 It is a system architecture diagram of the application of the large language model of the present invention in the field of financial risks; Figure 3 It is a schematic diagram of the steps of the system warning method of the present invention. Detailed Embodiments
[0020] As shown in the present invention Figures 1 to 3 A financial risk intelligent warning system and method based on a large language model. The system includes a data integration module, a real-time task scheduling engine, a data preprocessing unit, a large language model analysis module, a comprehensive evaluation module, and a warning module, wherein: The data integration module adopts standardized interface technology, supports docking with mainstream financial data platforms such as Caihui, DM, Wind, and YY, and can flexibly expand the data ecosystem according to requirements. The specific docking capabilities are as follows: Caihui: Supports HTTP API and SFTP file transfer protocols, and is suitable for real-time retrieval and batch download of financial data.
[0021] DM: Compatible with FIX protocol and WebSocket real-time data stream, meeting the needs of fixed income market quotation and transaction monitoring.
[0022] Wind: Provides WindPy API and ODBC database connection solutions, covering high-frequency query of structured data.
[0023] YY Rating: Based on REST API and Webhook event subscription, realizes dynamic tracking of credit ratings.
[0024] For example, when it is necessary to obtain the latest financial statement data of a listed company, the data integration module can quickly obtain relevant data by docking with the Wind data platform and using the WindPy API provided by it, ensuring the timeliness and accuracy of the data.
[0025] The real-time task scheduling engine is developed based on the Spring Boot framework. Task registration is achieved through the @EnableScheduling annotation, and Spring Cloud Config is used to dynamically load task policies. It adopts a master-slave architecture design. Based on the Spring Boot Actuator health check endpoint, the survival status of nodes is monitored in real time. When the health metrics of the master node are abnormal, the backup node takeover process is automatically triggered. The dynamic switching of the master and backup roles is achieved through the @ConditionalOnProperty conditional annotation. During the switching process, a distributed lock is used to ensure the consistency of configuration synchronization. The task allocation strategy relies on Spring Boot Metrics to collect CPU load and memory usage rate metrics of each node in real time. The node weight value is dynamically calculated through the weighted round-robin algorithm. When a task execution exception occurs, the retry mechanism is automatically triggered. The exponential backoff strategy is adopted for hierarchical retries within a preset time window. If the continuous failure reaches the threshold, the task is marked as an abnormal state and transferred to a healthy node for execution. At the same time, the system starts a minute-level monitoring cycle after the data source is updated, generates real-time warnings in combination with the anomaly detection model, and realizes breakpoint continuation execution through the persistent task status and checkpoint mechanism in the case of task interruption or node downtime. Its task allocation strategy formula is: Node weight = α・CPU load + β・memory usage rate + γ・IO throughput; where α, β, and γ are preset coefficients, and α + β + γ = 1.
[0026] Taking the monitoring of a company's market public opinion as an example, the real-time task scheduling engine can set up a timed polling task to automatically collect public opinion data related to the company from major news websites and social media platforms at regular intervals. At the same time, according to the event trigger mechanism, when there is a major event report related to the company, the data collection task is immediately started to ensure the real-time and integrity of the data. Compared with the passive scheduling mode of traditional ETL tools that rely on timed polling, the engine forms an active push mechanism through the dynamic policy loading of @EnableScheduling and Spring Cloud Config, compressing the data update response delay from the hour level to the minute level, effectively improving the timeliness of data processing.
[0027] The data preprocessing unit performs cleaning, deduplication, and format conversion on the collected data, and extracts effective information closely related to the specific subject company. The specific processing steps are as follows: Data cleaning: Remove irrelevant noise data and invalid information, including identifying and correcting or deleting error data (such as format errors, logical errors), filling in missing values, and removing redundant parts in duplicate records. For example, for a table containing the financial data of listed companies, data cleaning can remove duplicate rows and correct incorrect date formats.
[0028] Deduplication: Use efficient algorithms and technologies to identify and eliminate duplicate records. Data deduplication uses ID primary key verification to retain the latest timestamp record in the duplicate ID; text data uses the SimHash algorithm to detect redundant content with a similarity of >95% and remove it. The similarity calculation formula is: Similarity = 1-Hamming distance / number of fingerprint bits. In experimental verification, when the similarity threshold is set to 95%, it can effectively balance the deduplication efficiency and semantic integrity, while avoiding misjudgment problems caused by punctuation and stop word replacement.
[0029] Format conversion: Considering that data from different sources may have different formats, all input data will be unified into a standard format for subsequent analysis. Data deduplication uses ID primary key verification (such as user_id / order_id) to retain the latest timestamp record in the duplicate ID; text data (such as news headlines) uses the SimHash algorithm to detect redundant content with a similarity of >95% and remove it. After trial and verification, the similarity threshold is set to 95%, which can effectively balance deduplication efficiency and semantic integrity, while avoiding misjudgment problems caused by non-critical differences such as punctuation and stop word replacement.
[0030] The missing value processing follows the rules: Numerical fields (such as transaction amounts) are filled with the median of similar data. The specific mathematical logic is that the median calculation is based on industry classification grouping and falls back to the global median when the industry is missing. In the time series scenario, the dynamic median of the rolling time window is used, and the category field (such as industry classification) is marked as "Unknown". Time series data (such as stock price) is filled forward. The format conversion logic includes: the date and time are unified as YYYY-MM-DD HH:MM:SS (regular matching correction is triggered when parsing fails), the numerical unit is forced to be converted (such as currency converted to RMB 10,000 according to the real-time exchange rate, and percentages are converted to decimals by removing "%"), special characters are cleaned from text data (such as → space), invisible Unicode characters are removed, and the encoding is unified as UTF-8.
[0031] Compared with traditional rule-based preprocessing methods, this solution has significant advantages in data deduplication, format conversion efficiency, and missing value handling accuracy. Traditional methods rely on hard-coded rules to clean data, which are prone to parsing failures or accidental deletions due to format diversity (such as insufficient date format compatibility requiring manual intervention). In contrast, this solution achieves automated format correction through dynamic regular expression matching and forced conversion mechanisms, reducing the cost of manual verification. In terms of data deduplication, traditional rule-based cleaning only supports the identification of exact duplicates and cannot handle semantically similar texts (such as scenarios where news headlines are replaced with synonyms). This solution uses the SimHash algorithm combined with a semantic sensitivity threshold to improve the recognition rate of redundant content while preserving key information. For missing value imputation, traditional global mean / mode filling is likely to introduce industry-specific biases. This solution is based on a grouped dynamic median calculation and a time window rolling strategy to ensure that the numerical imputation better fits the data distribution pattern. The "Unknown" marking strategy for categorical fields can effectively avoid the problem of category contamination compared to traditional random filling or high-frequency value coverage, improving the reliability of subsequent analysis.
[0032] The data processing flow includes format standardization, deduplication optimization, and missing value imputation: In the format conversion stage, the date is unified into the YYYY-MM-DD HH:MM:SS format (regular expression correction is triggered when parsing fails), numerical units are forcibly converted (such as converting currencies to ten thousand yuan in RMB at the real-time exchange rate), text data is cleaned by automatically detecting the original encoding (such as using the chardet library to identify the encoding type) and converting it to UTF-8, while special characters are cleaned (such as replacing → with a space) and invisible Unicode characters are removed; Data deduplication is based on ID primary key verification and the retention of the latest timestamp record, and the SimHash algorithm is used to filter redundant texts (such as news headlines) with a similarity > 95%; In the handling of missing values, numerical fields are filled with the median of grouped data by industry classification (falling back to the global median when the industry is missing), a rolling window dynamic median is used in time series scenarios, categorical fields are marked as "Unknown", and forward filling logic is used for time series data.
[0033] Data deduplication is based on ID primary key verification and the SimHash algorithm: When the ID primary key (such as user_id / order_id) is repeated, the record with the latest timestamp is retained; For text deduplication, the SimHash algorithm is used. Its mathematical principle is to generate a weighted hash value after text tokenization and compress it into a binary fingerprint, and the similarity is determined by calculating the Hamming distance (similarity = 1 - Hamming distance / fingerprint length). The threshold is set at 95% based on experimental verification: The test set shows that the misjudgment rate is < 5% at this threshold (such as accidental deletions caused by non-semantic differences like punctuation replacement and stop word adjustment), while more than 95% of the valid semantic samples are retained to ensure that semantic integrity is not damaged during redundant text filtering.
[0034] Data Association and Integration: For the independent entity identification systems of different data sources, with the unified social credit code of enterprises as the core association identifier, entity alignment and information integration across data sources are realized. The specific process is as follows: Identification Mapping - Extract the corresponding relationship between the entity ID and the unified credit code in each data source to construct a global credit code mapping table; Code Cleaning - Standardize the credit code format (such as unifying it to 18 - digit uppercase without separators); Conflict Resolution - When merging multi - source data, structured fields (such as registered capital, industry classification) are overwritten according to the authority priority of the data source, and unstructured fields (such as business dynamics) are concatenated and de - duplicated. Through the above - mentioned carefully designed data cleaning, de - duplication, and format conversion steps, the high purity and standardization of the output data set are ensured, providing a solid foundation for subsequent in - depth analysis.
[0035] After being processed by the data pre - processing unit, data from different data sources with various formats can be converted into a unified standard format, providing high - quality data input for subsequent large language model analysis.
[0036] The large language model analysis module is used to analyze the pre - processed public opinion data and financial indicators, including steps of preliminary screening, polarity judgment, and classification and rating. Specifically as follows: Preliminary Screening: Use the large language model to preliminarily filter the input public opinion information to judge whether it is related to the specified entity. For example, when monitoring a news report, by comparing the keywords in the news content with the name and business scope information of the subject company, it is determined whether the news is related to the company, ensuring that subsequent analysis focuses on directly related data.
[0037] Polarity Judgment: For the public opinion information determined to be relevant, further use the large language model to conduct polarity analysis on it to distinguish positive and negative public opinions. For example, by analyzing the sentiment - inclined words in the news report, it is judged whether the report is optimistic or pessimistic about the company's development prospects.
[0038] Classification and Rating: Conduct separate detailed ratings for major risk types such as strategic risk, compliance risk, financial risk, market risk, operational risk, and legal risk. The ratings under each risk category are generated based on specific Prompts to obtain specific results. For example, when evaluating financial risk, the large language model will analyze the company's financial condition according to the company's financial statement data, such as asset - liability ratio, cash - flow indicators, combined with the preset Prompt, and generate corresponding risk rating results. Among them, Prompt refers to the preset instructions or question templates used to guide the large language model to analyze and rate specific risk categories, and its content is designed according to different risk categories to ensure that the model can accurately output effective rating results for the corresponding risk types.
[0039] The large model first identifies the sentiment tendency of public opinion. Negative public opinion (including risk signals) - triggers a risk rating (range: -5 to 0), and neutral / positive public opinion - triggers an opportunity rating (range: 0 to +5). For positive public opinion, it is directly scored according to the clarity of the favorable event and the potential market impact. For negative public opinion, the large language model is used to conduct a separate and detailed rating on the following main types of risks: Strategic risk: Refers to the possibility that the enterprise's strategic goals cannot be achieved due to external environmental changes, decision-making mistakes, or resource misallocation, and evaluates the impact of negative public opinion on the enterprise's long-term development strategy.
[0040] Compliance risk: Refers to the possibility of being punished, having operations restricted, or reputation damaged due to failure to comply with laws and regulations, industry standards, or internal systems, and analyzes the possibility of violating laws and regulations and their potential consequences.
[0041] Financial risk: Refers to the potential threat of financial crisis that may be triggered during the enterprise's fund-raising, capital structure, or cash flow management processes, and examines the company's financial statement data, debt levels, and cash flow status to determine whether there is potential financial instability.
[0042] Market risk: Refers to the business uncertainties caused by changes in market elements such as supply and demand relationships, price fluctuations, and consumer preferences, monitors factors such as market demand fluctuations and changes in the competitive landscape, and evaluates the impact on the enterprise's market share and competitiveness.
[0043] Operational risk: Refers to the risk of efficiency loss or business interruption caused by process defects, human errors, or system failures during the enterprise's daily operations, and evaluates the challenges faced by the enterprise's internal operations through data on supply chain management, production efficiency, and human resources.
[0044] Legal risk: Refers to the risk of mandatory legal consequences arising from changes in legal provisions, contract disputes, or infringement events, and pays attention to changes in laws and regulations, the development of litigation cases, and their specific impacts on the company's business.
[0045] The dynamic adjustment logic of the risk dimension weights of the system establishes a dynamic functional relationship through preset quantitative indicators and real-time data streams: the market risk weight forms a quadratic function with the industry stock index volatility (beta coefficient) and the slope of the market share change. The strategic risk weight constructs a power function with the policy environment change index as the base and the enterprise strategic deviation degree as the exponent. The compliance risk weight forms a linear weighted function with the regulatory penalty frequency and the regulation update rate. The financial risk weight constructs a piecewise function based on the month-on-month increase rate of the leverage ratio and the proportion of the cash flow gap. The operational risk weight adopts a product function of the supply chain interruption probability and the production anomaly event frequency. The legal risk weight is calculated by superimposing the logarithmic function of the lawsuit amount involved in litigation cases and the exponential function of the legal provision update frequency. All function parameters are dynamically calibrated by the rule engine according to the industry benchmark values and the entity's historical data to ensure that the weight allocation conforms to the quantitative laws of risk conduction.
[0046] Under each risk category, the large language model generates specific rating results based on specific Prompts to ensure independent and in-depth analysis of each risk type.
[0047] On this basis, the system uses the large language model to conduct a comprehensive rating of the financial indicators and the results of the previous classification and rating. The specific steps are as follows: The system first aggregates the risk rating results of each dimension of strategic risk, compliance risk, financial risk, market risk, operational risk, and legal risk obtained from the classification rating.
[0048] After the system aggregates the six types of basic rating results such as strategic risk and compliance risk, when the large language model integrates the financial indicators, it synchronously accesses the real-time market data and automatically calibrates the weights of each risk dimension according to the preset business rules: when a bear market signal is detected (such as the continuous decline of the main stock index), the system increases the market risk weight coefficient; if it is detected that the entity's financial indicators break through the industry benchmark values, the proportion of the financial risk weight is correspondingly increased.
[0049] Then, the large language model comprehensively integrates and analyzes these risk rating results in combination with the financial indicators provided by the entity (such as revenue growth rate, profit margin, asset-liability ratio, etc.).
[0050] Finally, based on the comprehensive rating results and the entity's current position status, the system uses the large language model for analysis and gives suggestions of no operation, in-depth investigation, on-site investigation, and prompt exit. These suggestions aim to help decision-makers timely understand potential risks and take corresponding measures, such as adjusting investment strategies, optimizing resource allocation, or strengthening internal management, so as to effectively respond to the challenges brought by various uncertainties.
[0051] Traditional risk models (mainly including methods such as logistic regression and decision trees, which rely on structured information such as financial data and credit scores for prediction) have an accuracy of about 82% in conventional risk assessment. However, by integrating unstructured public opinion data analysis, this system has increased the accuracy to 90% in key scenarios of negative public opinion identification, demonstrating stronger feature extraction capabilities and risk warning effects for multi-source heterogeneous data.
[0052] Each module in the system architecture interacts through a data flow pipeline: The preliminary screening module pushes relevant public opinion data to the polarity judgment module through a message queue. The negative public opinion after polarity judgment triggers a six-dimensional risk analysis in the classification and rating module. The rating results independently generated by each risk dimension are synchronously written into the risk database through the API interface; The comprehensive evaluation module periodically polls the risk database to obtain the latest classification and rating data, and the comprehensive risk score processed by the dynamic weight calculation engine is pushed to the warning module in real time; The warning module distributes the formatted risk signals and recommended instructions to the corresponding business systems through the enterprise communication middleware according to the preset threshold strategy (such as the comprehensive score ≤ -3 triggering a red warning), and at the same time writes the disposal feedback data back to the analysis log for model optimization.
[0053] According to the benchmark test data, the accuracy of this system reaches 90.2% (95% confidence interval: 89.5% - 90.8%) in key scenarios of negative public opinion identification, which is 7.9 percentage points higher than the benchmark accuracy of 82.3% of the traditional logistic regression model (using the same test set). The experiment uses the historical data of 1,200 listed companies disclosed by the China Securities Regulatory Commission as the test set. The fusion analysis of 36 structured financial indicators (such as asset-liability ratio, cash flow, etc.) relied on by the traditional model and public opinion data has increased the F1-score from 0.784 to 0.886, and the AUC value has been optimized from 0.812 to 0.901. The improvement is most significant in evaluation dimensions dominated by unstructured data such as strategic risk (+11.2%) and legal risk (+9.8%). The comparative experiment shows that the large language model extracts unstructured features (such as policy relevance, implicit risks in litigation texts, etc.) through the chain of thought method, shortening the early risk warning response time to 1 / 3 of the traditional model.
[0054] In practical applications, for a news report involving a company's financial fraud scandal, the large language model analysis module can accurately identify it as negative public opinion, and further analyze the impact of this event on the company's financial risk and compliance risk, generating detailed risk rating results to provide a basis for subsequent comprehensive evaluation.
[0055] The comprehensive evaluation module is used to summarize the classification and rating results, conduct a comprehensive rating in combination with financial indicators, and generate a comprehensive risk score. When conducting the comprehensive rating, consider the impact of market fluctuations on risks and make dynamic threshold adjustments based on the industry β coefficient. The adjustment formula is: Comprehensive risk score threshold = Benchmark threshold × (1 + Industry β coefficient × Market volatility); where, the Industry β coefficient is the covariance / variance of the industry index and the market index in the most recent 12 months.
[0056] At the same time, adjust the market risk weight by combining liquidity factors to determine the final risk weight, and this weight adjustment follows the following formula: Final risk weight = Market risk weight × [1 + (1 - Average trading volume in the past 5 days / Benchmark trading volume)]; where, the benchmark trading volume is the average daily trading volume in the past year; 1 + (1 - Average trading volume in the past 5 days / Benchmark trading volume) is the liquidity adjustment factor.
[0057] When the comprehensive assessment module generates the comprehensive risk score, the following formula is used for calculation: ; where, i is the risk dimension serial number, n is the total number of risk dimensions, and the normalized weight is adjusted by the dynamic calibration rule engine.
[0058] The normalized weight refers to converting the original weight values of each risk dimension through mathematical processing into a standardized weight with a sum of 1, ensuring that weights of different dimensions or magnitudes can fairly participate in the comprehensive score calculation. The specific formula is: ; n is the total number of risk dimensions (such as six categories including strategic risk, compliance risk, etc.); The original weight i is the dynamically adjusted weight value of the i-th type of risk; The normalized weight i is the standardized weight finally participating in the comprehensive score.
[0059] The normalized weight can eliminate the scoring deviation caused by the difference in weight ranges of different risk dimensions, ensuring the comparability and rationality of the comprehensive risk score.
[0060] Weight calculation and normalization: The rule engine generates the original weights of each risk dimension according to static rules and dynamic rules; perform normalization processing on the original weights to ensure the sum is 1; write the normalized weights into the configuration database for the comprehensive assessment module to call.
[0061] Suppose the original weights calculated by the rule engine on a certain day are: Market risk 0.4; Strategic risk 0.3; Compliance risk 0.2; Other risks 0.1; then the normalized weights are: ; Strategic risk weight = 0.3, and so on.
[0062] For example, for a listed company during a period of large market fluctuations, based on summarizing various risk rating results, the comprehensive evaluation module adjusts the comprehensive risk score threshold according to the industry beta coefficient and market volatility. At the same time, in combination with the comparison of the average trading volume of the company in the past 5 days with the benchmark trading volume, it adjusts the market risk weight, and finally generates a comprehensive risk score that can comprehensively reflect the company's current risk status, providing an intuitive risk assessment result for decision-makers.
[0063] Take the default warning of "Bond A" as an example: Trigger condition: It is detected that the trading volume has decreased by 30% in the past 5 days (liquidity risk has increased), and the industry beta coefficient has jumped from 1.0 to 1.5; Rule engine response: According to the market risk weight formula, the increase in the beta coefficient causes the original market risk weight to increase from 0.3 to 0.45;
[0064] After normalization, the proportion of the market risk weight increases from 25% to 35%; Result: The comprehensive risk score decreases significantly due to the increase in the market risk weight, triggering a red warning.
[0065] The warning module is used to trigger the warning mechanism according to the comprehensive risk score and send an alarm notification to relevant decision-makers. The notification management module supports configuring the hierarchical warning threshold and notification strategy as needed: The trigger rule adopts a customizable comprehensive risk score threshold (such as starting "in-depth investigation" when > 3; forcing "exit as soon as possible" when > 5), and the threshold setting can be dynamically adjusted in combination with the business cycle; The alarm notification supports checking and freely combining email, SMS, and system message channels, and reaches in real time through a preset interface. At the same time, it automatically generates a task to-do item related to the disposal process, realizing the closed-loop management of risk response.
[0066] When the comprehensive risk score reaches or exceeds the preset threshold, the warning module immediately sends an alarm notification to relevant decision-makers via email and SMS, reminding them to pay timely attention to and handle potential risk issues. For example, when the comprehensive risk score of a certain company exceeds the threshold, the warning module will send a red warning notification to the company's management, suggesting that they take measures as soon as possible, such as adjusting the investment strategy and strengthening internal management, to cope with the possible risks.
[0067] The invention combines the natural language processing capabilities of large language models with the advantages of the Chain of Thought (COT) method to achieve in-depth analysis of public opinion data and financial indicators. The system decomposes the analysis logic into reusable reasoning steps through the COT method: First, the large language model decomposes complex tasks and generates an intermediate reasoning chain. Then, the rule engine verifies the logical rationality. Subsequently, the reasoning path is dynamically adjusted based on the verification results. Finally, the verified reasoning mode is precipitated into a standardized analysis template. This modular process design ensures that key analysis capabilities do not depend on a single model parameter. Users can choose different technology stacks for deployment according to their needs. At the same time, the preset causal verification mechanism guarantees the interpretability of the analysis process.
[0068] The system is fully compatible with mainstream large model architectures such as GPT-4, Llama, Qwen, and DeepSeek. Based on modular design, it realizes seamless switching between different AI engines. Users can freely choose the appropriate adaptation plan according to business scenarios, cost budgets, and performance requirements. This feature significantly reduces the technical migration threshold, ensures the scalability and operation and maintenance efficiency of the system, and provides elastic support for diversified AI applications.
[0069] The innovation of cross-data-source entity alignment in this system lies in constructing a global mapping table with the unified credit code as the core to solve the problem of multi-source ID heterogeneity. By standardizing the code format, the basic differences between data sources are eliminated. For structured and unstructured data, authoritative priority coverage and splicing de-duplication conflict resolution strategies are adopted respectively to form a fusion mechanism that takes into account data authority and information integrity, and systematically overcomes key problems such as identity heterogeneity, format chaos, and information conflicts in multi-source entity alignment.
[0070] Compared with the coverage blind spots caused by relying on manual rule matching in traditional data preprocessing methods (such as being unable to handle complex logical errors or special coding problems) and the limitation of simple de-duplication based only on field complete consistency, this method combines algorithmic deep cleaning (automatically identifying format / logical errors and converting special characters) with the ID primary key verification and timestamp optimization mechanism, avoiding the efficiency bottleneck of high rule maintenance costs and frequent manual intervention. At the same time, the SimHash algorithm is introduced in the de-duplication link to quantify the semantic similarity of texts (threshold 95%). Compared with traditional string matching methods, it can accurately identify the substantial duplicate content caused by non-critical differences (such as punctuation replacement), and avoid the problem of misdeletion caused by minor changes in fields in simple de-duplication.
[0071] A method applied to the above intelligent early warning system includes the following steps: S1. Connect to multiple financial data sources through the data integration module, and collect the financial indicators of listed companies and multi-dimensional data of industry white papers: Use the standardized interface technology of the data integration module to establish connections with major financial data platforms and obtain various types of financial data required. For example, regularly download the financial statement data of listed companies from the Caihui data platform every day, and obtain industry white papers from industry research institutions; S2. Use the real-time task scheduling engine to integrate the collected data, support dual-mode data collection of regular polling and event triggering, and build a 7×24-hour intelligent monitoring network: Set up regular polling tasks to collect data at preset time intervals. At the same time, configure an event triggering mechanism. When a specific event is detected, immediately start the data collection task to ensure the real-time and integrity of the data. For example, for some important financial news websites, set up real-time monitoring tasks. Once major news reports related to the target company appear, immediately collect the news data; S3. The data preprocessing unit cleans, de-duplicates, and converts the format of the collected data, and extracts effective information closely related to a specific subject company: Preprocess the large amount of collected data, remove the noisy data and duplicate data, convert data in different formats into a unified standard format, and extract key information related to a specific subject company. For example, clean the collected public opinion data, remove irrelevant advertising information and duplicate comment content, and retain the effective public opinion information related to the target company; S4. Analyze the preprocessed public opinion data and financial indicators through the large language model analysis module, including steps of preliminary screening, polarity judgment, and classification rating: Use the powerful analysis ability of the large language model to deeply analyze the preprocessed data. First, conduct a preliminary screening of the public opinion data to judge whether it is related to the target subject; then judge the polarity of the relevant public opinion data to determine whether it is positive or negative public opinion; finally, classify and rate various risk types to generate detailed rating results. For example, for a news report about the company's new product release, after the large language model analysis module determines its relevance to the company through preliminary screening, further analyze the sentiment of the report, judge it as positive public opinion, and rate the possible market risks and strategic risks brought by this event according to the preset Prompt; S5. The comprehensive evaluation module summarizes the classified rating results and conducts a comprehensive rating in combination with financial indicators to generate a comprehensive risk score: The comprehensive evaluation module aggregates various risk rating results generated by the large language model analysis module, and in combination with the company's financial indicator information, conducts a comprehensive rating and calculates a comprehensive risk score to comprehensively reflect the overall risk situation of the company. For example, based on aggregating various rating results such as strategic risk, compliance risk, and financial risk, the comprehensive evaluation module combines the company's financial statement data, such as revenue growth rate and profit margin, to calculate the company's comprehensive risk score, providing an intuitive risk assessment result for decision-makers; S6. If a major risk or abnormal signal is detected, trigger the warning module and send an alarm notice to relevant decision-makers: When the comprehensive risk score reaches the preset warning threshold or other major risk signals are detected, the warning module is immediately activated, and an alarm notice is sent to relevant decision-makers via email or text message to remind them to take measures to address the risk in a timely manner. For example, when the company's comprehensive risk score exceeds the preset red warning threshold, the warning module sends an urgent alarm notice to the company's management, suggesting that they immediately hold a meeting to discuss countermeasures, such as adjusting the investment plan and strengthening risk control.
[0072] Through the above steps, the intelligent financial risk early warning system based on the large language model can achieve real-time monitoring, analysis, and early warning of financial risks, providing strong support for financial decision-making, helping financial institutions and investors better cope with market risks, and improving the efficiency and effectiveness of risk management.
[0073] Taking the default risk early warning of "Bond A" as an example, the application effect of the system of the present invention in actual financial risk early warning is demonstrated. Company A defaulted on August 13, 2024. As of that day, the remaining amount of "Bond A" was 489.535 million yuan, and the company's existing monetary funds were unable to pay off "Bond A", and "Bond A" could not make principal and interest payments on schedule. The system predicted a huge change in the default probability as early as August 2, 2024, with the default probability as high as 97.86%, triggering the warning mechanism. The system collected the financial data and public opinion data of Company A through the data integration module, and after being integrated by the real-time task scheduling engine, it was transmitted to the data preprocessing unit. The data preprocessing unit cleaned, deduplicated, and converted the format of the data, and extracted the effective information closely related to the company. The large language model analysis module analyzed the preprocessed public opinion data and financial indicators, initially screened out the public opinion information related to the company, and conducted polarity judgment and classification rating. The comprehensive evaluation module summarized the classification rating results and conducted a comprehensive rating in combination with financial indicators to generate a comprehensive risk score. When the comprehensive risk score reaches the preset threshold, the warning module is immediately triggered, and an alarm notice is sent to relevant decision-makers, timely warning of the default risk of "Bond A", providing sufficient time for investors and decision-makers to take countermeasures, and effectively reducing potential losses.
[0074] The specific working process of this system is as follows: Using the data integration module, it seamlessly connects to diversified data sources such as Caihui, DM, Wind, and YY, and constructs a multi-dimensional data ecosystem. This system can integrate various types of data, including financial indicators of listed companies, authoritative industry white papers, government supervision dynamics, and real-time market sentiment, achieving the intelligent integration of structured and unstructured data. With the support of the distributed task scheduling engine, the system adopts two mechanisms, namely regular collection and event-driven, to build an intelligent monitoring network that operates 7×24 hours, ensuring that the information required for decision-making can be obtained immediately.
[0075] Next, the collected data will be initially cleaned and filtered to remove irrelevant noise data and extract effective information closely related to the specific target company. During the data preprocessing process, data cleaning will be performed to identify and correct or delete incorrect data (such as format errors, logical errors), fill in missing values, and remove redundant parts in duplicate records. For possible special characters or encoding problems in text data, corresponding conversions and cleanings will also be carried out. In addition, efficient algorithms and technologies are used to identify and eliminate duplicate records, ensuring the uniqueness and accuracy of the data set. Considering that the data from different sources may have format differences, all input data will be unified into a standard format for subsequent analysis. In particular, since different data sources use different unique ID systems, it is necessary to associate and integrate the information about the same entity from these data sources to build a comprehensive database, where each entity contains the most complete and up-to-date information integrated from multiple data sources.
[0076] Subsequently, based on the powerful natural language processing capabilities of the large language model and the unique advantages of the Chain of Thought (COT) method, the system conducts in-depth understanding and analysis of the sentiment data and financial indicators. The system first uses the large language model to preliminarily filter the input sentiment information to determine whether it is relevant to the specified entity, ensuring that subsequent analysis focuses on directly related data. For the determined relevant sentiment information, the system further uses the large language model to perform polarity analysis to distinguish positive sentiment and negative sentiment. According to the polarity of the sentiment, the system uses the large language model to conduct separate detailed ratings for the main risk types such as strategic risk, compliance risk, financial risk, market risk, operational risk, and legal risk. The ratings under each risk category are specific results generated based on specific Prompts, ensuring independent and in-depth analysis.
[0077] Finally, in the comprehensive evaluation stage, the system uses a large language model to conduct a comprehensive rating on the financial indicators and the results of the classification ratings. The system first aggregates the risk rating results of each dimension obtained from the classification ratings, and then conducts a comprehensive integration and analysis in combination with the financial indicators provided by the entity. The large language model generates a comprehensive risk score based on different risk types and their interrelationships, reflecting the overall risk situation of the enterprise in multiple dimensions. This comprehensive rating not only considers the direct impact of a single risk type, but also analyzes the interactions between different risks, providing more accurate risk prediction and early warning. Based on the comprehensive rating results and the user's current investment portfolio status, the system also uses the large language model to generate professional risk management suggestions to help decision-makers timely understand potential risks and take corresponding measures, such as adjusting investment strategies, optimizing resource allocation, or strengthening internal management, so as to effectively cope with the challenges brought by various uncertainties.
[0078] The system calculates structured data, including an industry classification list, a risk dimension list, and a three-dimensional data matrix composed of (score × normalized weight), where the score range is locked from -5 to +5, and the weights are normalized between 0 - 100%. After receiving this data, the large language model converts each record in the three-dimensional matrix into the ECharts standard data format, that is, reconstructs the data points according to the "industry name - risk type - calculated value" triple (such as ["Manufacturing", "Strategy", -3.2]). At the same time, it automatically configures the red - white - green gradient color scale mapping rule according to the positive and negative intervals of the calculated value, and binds the industry list and risk type list in the original data to the x-axis and y-axis coordinates of the heat map respectively. For high-risk data points with an absolute value of the calculated value ≥ 3, the model automatically attaches a highlighted identification attribute, and finally generates an ECharts configuration object containing complete coordinate definitions, a color scale scheme, and data sequences. Its JSON structure can be directly used for front-end visual rendering, realizing an interpretable conversion from structured data to a heat map.
[0079] If a major risk or abnormal signal is detected, the system will automatically trigger the early warning mechanism and send an alarm notification to relevant decision-makers or management teams according to the business situation, so that they can take corresponding measures in a timely manner to prevent potential risk events from occurring.
[0080] The real-time early warning mechanism preprocesses and standardizes the original data through a preprocessing unit and then transmits it to the risk identification module for multi-dimensional analysis in real time. When the risk identification module detects abnormal indicators or risk signals based on preset rules or models, it automatically synchronizes the judgment results to the early warning system, triggers a dynamic response process, and sends graded risk alarm notifications to designated personnel according to the preset business scenario thresholds and alarm rules, realizing a closed-loop linkage from data preprocessing to risk identification and then to early warning triggering.
[0081] The data flow diagram of the system is as attachedFigure 1 As shown below: The data flow diagram shows the data flow and processing process in the intelligent financial risk early warning system based on large language models. First, the data integration module collects multi-dimensional data such as listed company financial indicators and industry white papers from multiple financial data sources, and realizes the initial aggregation of data by docking mainstream financial data platforms such as Caihui, DM, Wind, and YY. After being integrated by the real-time task scheduling engine, these data are transmitted to the data preprocessing unit for cleaning, deduplication, and format conversion operations, extracting effective information closely related to specific subject companies to form subject information, laying a foundation for subsequent analysis.
[0082] The preprocessed data is sent to the large language model analysis module. Using the powerful capabilities of the large language model, a series of risk identification operations such as preliminary screening, polarity judgment, and classification rating are performed on public opinion data and financial indicators, covering multiple dimensions such as strategic risk, compliance risk, and financial risk. The rating of each dimension uses specific Prompts to guide the model to generate specific results, realizing in-depth mining and analysis of financial risk data.
[0083] The results and suggestions obtained from the analysis are summarized and sorted through the collection module of analysis results and suggestions to form structured data for subsequent comprehensive evaluation and decision support. At the same time, the event correlation module integrates and analyzes the information in different data sources to identify the relevance between financial events and the potential connections with financial entities, further improving the comprehensiveness and accuracy of risk assessment.
[0084] The comprehensive evaluation module summarizes the classification rating results, comprehensively considers and calculates in combination with multiple factors such as financial indicators to generate a comprehensive risk score, reflecting the overall risk status of financial entities. If the comprehensive risk score reaches the preset threshold or a major risk signal is detected, the early warning module is triggered immediately to send an alarm notice to relevant decision-makers, realizing timely early warning and response to risks, assisting financial institutions or investors to take preventive measures in advance and optimize risk management strategies.
[0085] The entire data flow diagram reflects the complete process of the system from multi-source data collection, processing, analysis to risk assessment and early warning. With the help of the advanced technology of large language models, it realizes the intelligent early warning and accurate assessment of financial risks, providing strong support for financial decision-making.
[0086] The following are the specific explanations of "subject information", "large language model risk identification", "collection of analysis results and suggestions", and "event correlation" in the appendix Figure 1 as follows: Subject Information: Subject information is the basic data part of the entire risk assessment process, covering various types of information related to specific financial entities (such as listed companies). This information comes from multiple sources, including but not limited to financial data platforms such as Caihui, DM, Wind, YY, and public opinion data sources such as news reports and social media. The specific content of subject information includes but is not limited to the basic information of listed companies (such as company name, unified social credit code, industry), financial indicators (such as asset-liability ratio, revenue growth rate, profit margin), business operation data, and public opinion information collected from news reports and social media channels.
[0087] Large Language Model Risk Identification: Large language model risk identification is one of the core analysis modules of the system. This module uses the powerful natural language processing ability of the large language model to deeply analyze the preprocessed public opinion data and financial indicators. Specifically, the large language model will perform preliminary screening, polarity judgment, and classification and rating operations. Through the above operations, the large language model risk identification module can accurately identify potential financial risk signals and provide basic data support for subsequent comprehensive evaluation.
[0088] Collection of Analysis Results and Recommendations: The collection module of analysis results and recommendations is responsible for summarizing various risk rating results output by the large language model risk identification module and organizing them into a structured data format for subsequent comprehensive evaluation and decision support. These analysis results include but are not limited to the specific rating scores of different risk dimensions (such as strategic risk rating, market risk rating), risk type identifiers, and detailed description information related to risks. At the same time, this module will also combine the preset rules and risk assessment models of the system to generate corresponding risk warning suggestions and management strategies based on the analysis results. These suggestions and strategies are designed to help financial institutions or investors timely understand potential risks and take corresponding measures for risk prevention and control.
[0089] Event Association: The main function of the event association module is to identify the relevance between financial events and the potential connections between these events and financial entities by integrating and analyzing information from different data sources. The specific operations are as follows: Data Integration and Analysis: This module will collect and integrate data from multiple sources, including but not limited to news reports, social media, industry reports, and market quotation data. By analyzing and processing these massive amounts of data, the hidden associated information is mined.
[0090] Event Association Identification: Based on the understanding and reasoning ability of the large language model, the integrated data is deeply analyzed to identify the association relationships between different financial events.
[0091] Risk conduction analysis: Further analyze the impact of the correlation between events on the risk status of financial entities, and identify potential risk conduction paths. Through this analysis, the system can more comprehensively evaluate the overall risk environment faced by financial entities, providing more accurate risk warnings and decision-making support for financial institutions and investors.
[0092] The specific architecture diagram of this system is as attached Figure 2 shown as follows: Data integration module 1. Third-party data source docking for Caixin, WIND, DM, and YY: Responsible for docking mainstream financial data platforms such as Caixin, DM, Wind, and YY, collecting multi-dimensional data such as financial indicators of listed companies, industry white papers, and public opinion data, ensuring the wide range and diversity of data sources.
[0093] Data preprocessing: 1. Data preprocessing process: Perform a series of processing operations on the collected data, including data cleaning, deduplication, format conversion, and data association integration, to improve data quality.
[0094] Data processing tasks: 1. Scheduled execution: Through the scheduled execution function, the system automatically starts the data processing process at preset time intervals to ensure the regular update and analysis of data, and guarantee the system's real-time monitoring ability of the financial market dynamics.
[0095] 2. Data synchronization: After docking multiple data sources, the system executes the data synchronization task to ensure the consistency and timeliness of data from different sources, providing a unified data basis for subsequent analysis.
[0096] 3. Synchronization scheduling: Responsible for coordinating the execution order and time of multiple data synchronization tasks, optimizing data processing efficiency, and ensuring the orderly progress of tasks.
[0097] 4. Data cleaning: Remove irrelevant noise data and invalid information, such as identifying and correcting error data, filling in missing values, and removing redundant parts in duplicate records, to improve data quality.
[0098] 5. Index scheduling: Schedule and manage the calculation tasks of various risk indicators to ensure that the indicators can be updated and applied in a timely manner, providing an accurate basis for risk assessment.
[0099] Index development (including public opinion): 1. Public opinion content rating: Rate the content of public opinion information, judge its potential impact on financial entities, and extract risk signals in public opinion.
[0100] 2. Public opinion entity rating: Rate the entities involved in public opinion, evaluate their risk status, and determine the scope and degree of the impact of public opinion on specific financial entities.
[0101] Model Training: 1. Model Training Scheduling: Manage the execution plan and resource allocation of model training tasks to ensure that the model can be trained efficiently and updated in a timely manner to adapt to market changes.
[0102] 2. Training Set Generation: Generate a data set for model training, ensuring the quality and representativeness of the training data to provide strong support for model training.
[0103] 3. Model Training: Train the training set through machine learning algorithms to generate a model that can be used for risk assessment, continuously improving the performance and accuracy of the model.
[0104] Model Training Implementation Method Data Preparation: Obtain data that has been cleaned, de-duplicated, and format-converted from the data preprocessing unit, including multi-dimensional information such as financial indicators and public opinion data of listed companies. This data will serve as the basic material for model training.
[0105] Training Set Generation: Randomly select representative samples from the prepared data according to a predetermined data sampling strategy to construct a training set. Ensure that the training set covers data of financial entities with different industries, scales, and risk characteristics to improve the generalization ability of the model.
[0106] Model Selection and Initialization: According to the characteristics and requirements of financial risk assessment, select a suitable basic model architecture, such as logistic regression, decision tree, random forest, neural network, etc. Perform initialization settings on the selected model, including parameter initialization, definition of loss function and optimizer, etc.
[0107] Model Training Process: Input the training set data into the model, and calculate the output result of the model through forward propagation. Calculate the loss value between the model output result and the actual label, and use the optimizer to update the parameters of the model through backpropagation according to the loss value. Repeat the above steps for multiple training epochs until the performance of the model on the training set reaches the preset convergence condition, such as the loss value is lower than a certain threshold or the performance metrics (such as accuracy, recall, etc.) no longer improve significantly.
[0108] Model Validation and Adjustment: During the model training process, regularly use the validation set to validate the model and evaluate the performance of the model on unseen data. According to the validation results, adjust the hyperparameters of the model, such as learning rate, regularization parameter, model structure parameter, etc., to optimize the performance of the model and prevent overfitting or underfitting.
[0109] Core User Functions: Daytime Process (Credit Risk): The general term for a series of automated analysis and early warning operations performed by the system for credit risk during normal working hours. Its purpose is to monitor and evaluate the credit risk status of financial entities in real time, so as to timely detect potential risks and take corresponding measures.
[0110] Specifically, this process covers multiple links from data collection, processing, to risk identification, analysis, comprehensive evaluation, and then to early warning notification. By docking with multiple financial data sources, multi-dimensional data including financial indicators and public opinion information is obtained, and the large language model is used to deeply analyze this data. Achieve precise identification and quantitative evaluation of credit risk, generate a comprehensive risk score, and provide timely and accurate risk early warning and decision-making support for financial institutions or investors.
[0111] 1. Default Probability Prediction: Predict the likelihood of a financial entity defaulting, providing a key indicator for risk assessment and helping users take preventive measures in advance.
[0112] 2. Risk Identification: Identify potential financial risk factors, mine risk characteristics in the data, and help users identify risk points in advance.
[0113] 3. Comprehensive Evaluation: Comprehensively consider multiple risk factors, conduct a comprehensive risk assessment of the financial entity, and generate a comprehensive risk score.
[0114] 4. TopN Generation: Generate a list of the top N financial entities with higher risk ratings, facilitating users to focus on high-risk entities and optimizing the user's risk management process.
[0115] 5. Early Warning: Trigger the early warning mechanism according to the risk assessment results, timely notify users to take measures, and reduce potential losses.
[0116] 6. Entity Analysis: Conduct in-depth analysis of a specific financial entity, provide a detailed risk assessment report, and help users deeply understand the risk status of a specific entity.
[0117] Large Language Model: 1. Invocation: Invoke the large language model during the process of public opinion content rating and public opinion entity rating, utilize its powerful natural language processing ability for risk analysis, and improve the accuracy and depth of analysis.
[0118] Report: 1. Report Template: Provide a standardized report format, facilitating users to quickly generate risk assessment reports and ensuring the standardization and consistency of reports.
[0119] 2. Chapter Template: Define the chapter structure of the report, ensure the integrity and logic of the report content, and help users clearly present the risk assessment results.
[0120] 3. Material Template: Provide the material formats required for reports, facilitating users to collect and organize data and improving the report generation efficiency.
[0121] 4. Report Generation: Generate the final risk assessment report based on the analysis results and preset templates, providing users with a high-quality way to display risk assessment results.
[0122] Model Iteration: 1. Model Version Management: Manage different versions of the model to ensure the traceability and stability of the model, facilitating users to trace back and compare the performance of different version models.
[0123] 2. Model Training Scheduling: Coordinate the execution of model training tasks, optimize the model training process, and ensure the efficient execution of model training tasks.
[0124] 3. Model Creation: Create new models to adapt to the changing market environment and risk characteristics, maintaining the timeliness and effectiveness of the models.
[0125] 4. Model Backtesting: Conduct historical data testing on the model, evaluate the effectiveness and reliability of the model, and verify the performance of the model in actual applications.
[0126] 5. Model Evaluation: Comprehensively evaluate the performance indicators of the model to ensure that the model can meet the actual application requirements and provide users with a comprehensive model performance evaluation result.
[0127] Model Iteration Implementation Method Model Version Management: Establish a dedicated model version control system, assign a unique version number and identification information to each newly created or updated model. Version-store relevant materials such as the model's code, configuration files, training parameters, and performance indicators, and record information such as the change content, change time, and change reason for each version. Provide functions for querying, comparing, and tracing back model versions, facilitating users to view the historical version evolution of the model at any time and be able to quickly restore to a previous stable version when necessary.
[0128] Model Training Scheduling: Based on task scheduling tools (such as Apache Airflow, Cron), set the timed scheduling plan for model training tasks, and regularly trigger the model training process according to business requirements and data update frequencies. Reasonably allocate computing resources (such as CPU, GPU) for training tasks according to the priority and resource requirements of the model to ensure the efficient and orderly execution of multiple model training tasks. Support the manual triggering function of model training tasks so that in case of special needs or emergencies, the model update operation can be started in a timely manner.
[0129] Model Creation and Update: Based on the changes in the financial market, new business development requirements, and the evaluation results of the performance of existing models, determine whether new models need to be created or existing models need to be updated.
[0130] When creating new models, refer to the design ideas and lessons learned from existing models, and combine new data features and risk assessment indicators to design a more reasonable model architecture and algorithm logic. When updating existing models, based on the records of version management, analyze the improvement points and optimization directions of the models, and adopt appropriate model transfer learning or incremental learning methods to combine the existing model knowledge with new training data to improve the efficiency and effectiveness of model iteration.
[0131] Model Backtesting: Establish a model backtesting mechanism, select historical data as the backtesting dataset, and ensure that the time range, data features, etc. of the backtesting data are consistent with the actual application scenarios of the model.
[0132] Apply the trained model to the backtesting dataset, simulate the risk assessment and early warning performance of the model in the historical market environment, and record the comparison between the prediction results of the model and the actual risk events that occurred.
[0133] Comprehensively evaluate the performance of the model using multiple backtesting metrics (such as accuracy, recall rate, F1 score, area under the ROC curve, etc.), analyze the stability and reliability of the model under different market conditions, and identify the problems and risk points existing in the model.
[0134] Model Evaluation and Optimization: Comprehensively consider the performance of the model in the training set, validation set, and backtesting, as well as practical application factors such as the interpretability, computational cost, and deployment difficulty of the model, and conduct a comprehensive evaluation of the model. According to the evaluation results, determine whether the model meets the requirements for online deployment. For models that do not meet the expected performance, analyze the reasons and take corresponding optimization measures, such as further adjusting model parameters, optimizing model structure, expanding the scale of training data, etc. Regularly re-evaluate and optimize the models that have been deployed online to adapt to the changes in the market environment and the continuous development of the business, and ensure that the models always maintain good performance and risk early warning capabilities in actual applications.
[0135] Explanation of the Operating Principle of the System Architecture: When the system is running, the data integration module collects multi-dimensional data from third-party data sources such as Caihui, WIND, DM, and YY to ensure a wide range of data sources. The collected data enters the data preprocessing module, and the timing execution function of the data processing task module triggers the data synchronization task. The synchronization scheduler coordinates the execution order of the synchronization tasks to ensure efficient data synchronization.
[0136] After synchronization is completed, the system performs data cleaning operations to remove noise and invalid information, improving data quality. The cleaned data enters the indicator development module through the indicator scheduling function. The indicator development module receives the preprocessed data and conducts sentiment content rating and sentiment subject rating. During this process, the large language model is called to deeply analyze the sentiment information using its natural language processing capabilities, extract risk signals in the sentiment, and conduct risk assessment on the subjects involved in the sentiment.
[0137] The sentiment rating results and risk indicator data generated by the indicator development module are transmitted to the model training module. The model training scheduling sub-module manages the execution plan of the model training task, the training set generation sub-module generates high-quality training data, and the model training sub-module uses machine learning algorithms to perform model training and generate a model for risk assessment. The trained model is stored in the model version management sub-module, facilitating model iteration update and performance comparison.
[0138] The trained model is applied to the core user function module to implement functions such as default probability prediction, risk identification, comprehensive evaluation, TopN generation, early warning, and subject analysis. The core user function module provides users with personalized risk assessment services and timely risk warnings based on the analysis results of the model, helping users identify and prevent financial risks in advance.
[0139] At the same time, the system provides a report module to support users in generating standardized and professional risk assessment reports according to preset report templates, chapter templates, and material templates. The report generation sub-module generates a detailed assessment report based on the analysis results and preset templates, facilitating internal reporting and external communication for users.
[0140] In terms of model management, the model iteration module ensures that the model can continuously adapt to the dynamic changes of the market and maintain good performance and reliability through the close cooperation of sub-modules such as model version management, model training scheduling, model creation, model backtesting, and model evaluation. The model backtesting sub-module conducts historical data testing on the model to verify the effectiveness and stability of the model; the model evaluation sub-module comprehensively evaluates various performance indicators of the model and provides users with a more comprehensive model performance evaluation report to help users select the optimal model for risk assessment.
[0141] The design focus of the present invention lies in 1. Efficiently processing massive data and multi-dimensional risk identification The application of large language models in the field of financial risk demonstrates their excellent data processing capabilities, enabling rapid analysis and processing of massive amounts of unstructured data from sources such as news reports, social media content, and policy announcements. This efficient data processing capability allows financial institutions to capture dynamic changes in real-time in a rapidly changing market environment, breaking free from the constraints of the inefficiency and limitations of traditional data processing methods. In this way, financial institutions can not only obtain the latest market information in a timely manner but also quickly respond to market changes and optimize investment strategies.
[0142] 2. Multi-dimensional Risk Identification and In-depth Understanding Different from traditional risk assessment methods that only rely on structured data (such as financial statements), the large language model in this invention can extract risk information from various types of unstructured data, achieving true multi-dimensional risk identification. For example, by analyzing factors such as public opinion, market sentiment, and industry dynamics, the model can identify potential market fluctuations, policy changes, or enterprise-specific risks, providing a comprehensive and detailed risk assessment framework. This all-round risk identification method helps to reveal subtle trends and signals hidden in a large amount of data, providing more accurate risk predictions for investors.
[0143] 3. Real-time and Dynamic Response Mechanism Large language models have the ability to process and analyze continuously updated data in real-time, enabling rapid capture of sudden market changes or fluctuations in public opinion. In the face of a rapidly changing market environment, the model can immediately provide risk warnings to help investors and decision-makers respond quickly. This real-time and dynamic response mechanism ensures that financial institutions can respond to market changes within the shortest possible time, thereby reducing investment risks caused by information lag.
[0144] 4. In-depth Understanding and Context Analysis Compared with traditional models, large language models have stronger context understanding and context analysis capabilities, and can handle complex factors in text, such as context, emotion, and tone. This enables it to more accurately extract potential risk signals. For example, through sentiment analysis of news reports or social media content, the model can identify signals such as negative emotions or market panic, and predict potential crisis events in the future in advance. This in-depth understanding ability greatly improves the accuracy of risk identification.
[0145] 5. Automated and Personalized Risk Control Large language models can automatically generate personalized risk warnings based on real-time data and specific investment needs, which is a highly intelligent risk control mechanism. It not only improves the efficiency of risk management but also provides customized suggestions according to different market conditions and risk preferences. In this way, financial institutions can more flexibly adjust their risk management strategies to adapt to the changing market conditions.
[0146] 6. Capturing Subtle Changes and Proactiveness With large-scale pre-training and continuous learning capabilities, large language models can capture subtle changes and potential risks in the market, which are often difficult to detect by traditional analysis methods. For example, the model can identify potential market sentiment changes or industry crises from minor market fluctuations and issue early warnings. This proactiveness analysis ability is crucial for preventing problems before they occur.
[0147] 7. Scalability and Flexibility Design Large language models have good scalability and flexibility, and can adapt to the changing market environment and new types of risks. Whether integrating new data sources or meeting the needs of different fields, the model can be trained and optimized accordingly to flexibly handle various risk analysis tasks. In addition, the model can integrate data from multiple channels such as social media, news websites, financial reports, and forum comments across platforms for comprehensive analysis, further enhancing the comprehensiveness and accuracy of risk identification and decision support.
[0148] The main applications of this system are as follows: Application of financial risk information extraction: Using large language models to extract key information related to financial risks from multi-source data such as social media and news reports. Application scenarios include market sentiment analysis, policy change monitoring, and industry trend tracking, providing decision support for financial institutions.
[0149] Real-time risk warning system: Real-time processing of public opinion data and rapid identification of potential risk signals, applied to financial market volatility prediction and policy change response. Through historical data analysis and real-time monitoring, timely risk warnings are issued to investors.
[0150] Application of deep semantic understanding and context correlation: Specific applications in financial risk assessment, such as identifying implicit risks in financial statements or analyzing risk factors in complex texts. Emphasize how to improve the accuracy of risk prediction by understanding the deep meaning of texts.
[0151] Application of personalized risk warning services: Risk warning services customized for different investment entities, such as automatically identifying negative public opinions for specific industries or enterprises and generating personalized risk reports.
[0152] Support investors to flexibly adjust strategies according to the latest market dynamics.
[0153] Application of large language models in the field of financial risks: By combining public opinion data with large language models, key risk information in massive financial data can be intelligently sorted out and extracted, key words related to risks can be accurately identified, and targeted risk warnings can be provided.
[0154] Compared with traditional risk analysis methods, this method avoids the problem of difficultly and accurately extracting key risk information from massive data, reducing the complexity and uncertainty of investment decisions. The traditional manual information processing method is inefficient and has limited coverage, which is very likely to lead to the neglect or delayed identification of potential risk information. The financial market is highly volatile and requires high timeliness for risk early warning. The present invention uses large language models to quickly and accurately identify risk information, improving the timeliness and accuracy of risk early warning, helping investors to react faster and reducing potential losses.
[0155] Risk identification and analysis based on large language models: By combining the powerful natural language processing ability of large language models and the unique advantages of the Chain of Thought (COT) method, in-depth understanding and analysis of public opinion data and financial indicators are realized. The system uses large language models for preliminary screening to judge whether public opinion information is related to the designated entity, and further uses large language models for polarity analysis to distinguish positive public opinion from negative public opinion. Separate and detailed ratings are given to the main risk types such as strategic risk, compliance risk, financial risk, market risk, operational risk and legal risk, and the ratings under each risk category generate specific results based on specific Prompts.
[0156] Comprehensive evaluation and generation of professional suggestions: The system uses large language models to conduct a comprehensive rating on the results of financial indicators and classification ratings, summarizes the risk rating results of each dimension and integrates and analyzes them in combination with financial indicators. The large language model generates a comprehensive risk score, considering the direct impact of different risk types and their interactions, and provides accurate risk prediction and early warning. According to the comprehensive rating results and the current position status of the entity, professional risk management suggestions are generated to help decision-makers take timely measures to cope with the challenges brought by uncertainties.
[0157] In summary, by constructing an efficient, intelligent, and real-time financial risk early warning system. The data integration module connects to multiple financial data sources, collects financial indicators of listed companies and multi-dimensional data from authoritative industry white papers, and uses a real-time task scheduling engine to integrate the data, constructing a 7×24-hour intelligent monitoring network to ensure the timeliness and integrity of the data. The data preprocessing unit cleans, de-duplicates, and converts the format of the collected data, extracts effective information closely related to the specific subject company, and provides high-quality data support for subsequent analysis. The large language model analysis module combines the chain of thought method to deeply analyze public opinion data and financial indicators, including steps such as preliminary screening, polarity judgment, and classification rating, and can accurately identify the main risk types of strategic risks, compliance risks, financial risks, market risks, operational risks, and legal risks. The comprehensive evaluation module summarizes the classification rating results, conducts a comprehensive rating in combination with financial indicators, generates a comprehensive risk score, and dynamically adjusts the threshold according to the industry β coefficient and market volatility. At the same time, it considers the liquidity factor to adjust the market risk weight to determine the final risk weight. The early warning module triggers the early warning mechanism according to the comprehensive risk score, sends an alarm notification to relevant decision-makers, and realizes the real-time early warning and response of risks. Through the collaborative work of each module, the entire system gives full play to the natural language processing ability and in-depth understanding advantages of the large language model, and realizes the intelligent early warning of financial risks.
[0158] The above are only the preferred embodiments of the present invention, and do not impose any limitation on the technical scope of the present invention. Therefore, any minor modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. An intelligent financial risk early warning system based on large language models, characterized in that, Including: A data integration module for connecting to multiple financial data sources and collecting financial indicators of listed companies and multi-dimensional data of industry white papers; A real-time task scheduling engine for integrating the collected data, supporting dual-mode collection of scheduled polling and event triggering, and building a 7×24-hour intelligent monitoring network; A data preprocessing unit for cleaning, deduplicating, and format-converting the collected data, and extracting effective information closely related to a specific subject company; A large language model analysis module for analyzing the preprocessed public opinion data and financial indicators, including steps of preliminary screening, polarity judgment, and classification rating; A comprehensive evaluation module for summarizing the classification rating results, combining financial indicators for comprehensive rating, and generating a comprehensive risk score; An early warning module for triggering an early warning mechanism based on the comprehensive risk score and sending an alarm notification to decision-makers; The data collected by the data integration module is integrated by the real-time task scheduling engine and then transmitted to the data preprocessing unit. The data processed by the data preprocessing unit is transmitted to the large language model analysis module. The analysis results of the large language model analysis module are transmitted to the comprehensive evaluation module. The comprehensive risk score of the comprehensive evaluation module is transmitted to the early warning module.
2. The intelligent financial risk early warning system based on the large language model according to claim 1, wherein The data integration module adopts standardized interface technology and supports connection to financial data platforms such as Caihui, DM, Wind, and YY.
3. The intelligent financial risk early warning system based on large language models according to claim 1, wherein The real-time task scheduling engine is developed based on the SpringBoot framework. Task registration is achieved through the @EnableScheduling annotation, and task policies are dynamically loaded using Spring Cloud Config. It adopts a master-slave architecture design, and the health status of nodes is monitored in real time based on the Spring Boot Actuator health check endpoint. The dynamic switching of master and standby roles is achieved through the @ConditionalOnProperty conditional annotation. The task allocation strategy relies on Spring Boot Metrics to collect CPU load and memory usage indicators of each node in real time, and the node weight value is dynamically calculated through the weighted round-robin algorithm.
4. The intelligent financial risk early warning system based on the large language model according to claim 1, characterized in that, The processing steps of the data preprocessing unit include: Data cleaning to remove irrelevant noise data and invalid information; Deduplication processing, using efficient algorithms and technologies to identify and eliminate duplicate records; Format conversion to unify data from different sources into a standard format; data association and integration, using the unified social credit code of enterprises as the core association identifier to achieve entity alignment and information integration across data sources.
5. The intelligent financial risk early warning system based on the large language model according to claim 1, characterized in that, When conducting classification rating, the large language model analysis module includes detailed ratings of strategic risk, compliance risk, financial risk, market risk, operational risk, and legal risk. The ratings under each risk category generate specific results based on specific preset instructions or question templates.
6. The intelligent financial risk early warning system based on the large language model according to claim 1, characterized in that, In the deduplication process of the data preprocessing unit, the SimHash algorithm is used to calculate the text similarity, and the redundancy of the data is evaluated based on the text similarity. When the similarity threshold meets the following formula, it is determined as redundant content and corresponding processing is carried out: Similarity = 1 - Hamming distance / fingerprint length ≥ 95%; where the fingerprint length is the length of the binary fingerprint generated by SimHash.
7. The intelligent financial risk early warning system based on the large language model according to claim 1, characterized in that, During the task assignment process of the real-time task scheduling engine, a specific task assignment strategy is adopted to optimize task assignment. The strategy is expressed as: Node weight = α · CPU load + β · memory usage rate + γ · IO throughput; where α, β, and γ are preset coefficients, and α + β + γ = 1.
8. The intelligent financial risk early warning system based on a large language model according to claim 1, characterized in that, When the comprehensive evaluation module conducts a comprehensive rating, it considers the impact of market fluctuations on risks and adjusts the dynamic threshold based on the industry β coefficient. The adjustment formula is: Comprehensive risk score threshold = benchmark threshold × (1 + industry β coefficient × market volatility); where the industry β coefficient is the covariance / variance of the industry index and the market index in the most recent 12 months.
9. The financial risk intelligent early warning system based on a large language model according to claim 1, characterized in that, When the comprehensive evaluation module conducts a comprehensive rating, it adjusts the market risk weight by combining liquidity factors to determine the final risk weight. The weight adjustment follows the following formula: Final risk weight = market risk weight × [1 + (1 - average trading volume in the past 5 days / benchmark trading volume)]; where the benchmark trading volume is the average daily trading volume in the past year; 1 + (1 - average trading volume in the past 5 days / benchmark trading volume) is the liquidity adjustment factor.
10. The intelligent financial risk early warning system based on the large language model according to claim 1, characterized in that, When the comprehensive evaluation module generates a comprehensive risk score, the following formula is used for calculation: ; where i is the risk dimension serial number, n is the total number of risk dimensions, and the normalized weight is adjusted by the dynamic calibration rule engine.
11. A method applied to the financial risk intelligent early warning system based on the large language model as described in any one of claims 1-10, characterized in that, Including the following steps: S1. Connect to multiple financial data sources through the data integration module, and collect multi-dimensional data such as financial indicators of listed companies and industry white papers. S2. Use the real-time task scheduling engine to integrate the collected data, support dual-mode collection of timed polling and event triggering, and build a 7×24-hour intelligent monitoring network. S3. The data preprocessing unit cleans, de-duplicates, and converts the format of the collected data, and extracts effective information closely related to a specific subject company. S4. Analyze the preprocessed public opinion data and financial indicators through the large language model analysis module, including steps of preliminary screening, polarity judgment, and classification rating. S5. The comprehensive evaluation module summarizes the classification rating results, combines the financial indicators for a comprehensive rating, and generates a comprehensive risk score. S6. If a major risk or abnormal signal is detected, trigger the warning module and send an alarm notification to the decision maker.
Citation Information
Patent Citations
Enterprise negative public opinion intelligent risk identification and index construction method
CN116227909A
Financial calculation system and method based on financing project
CN119130656A
Enterprise financial risk early warning system and method based on big data
CN119151694A
Multi-dimensional financial pressure testing and financial early warning method and system
CN119941410A
Methods and systems for risk mining and for generating entity risk profiles
US20120221485A1
Cited By
Intelligent financial risk early warning method, system and device and storage medium
CN121032676A
National risk intelligent early warning system based on multi-source data fusion
CN121190215A
Credit rating report generation method and device based on multi-modal data and storable medium
CN121257479A
System and method for intelligently acquiring, extracting and analyzing computing power center information based on large model
CN122019656A