Enterprise financial risk early warning method based on large language model and Logistic model

Through a method based on the large language model and Logistic model, the performance briefing text data is analyzed from multiple dimensions and combined with financial data to predict, the problem of difficulty in comprehensively and accurately predicting the financial risks of listed companies in the existing technology is solved, and the prediction capability of the credit risk warning model is improved.

CN119990752APending Publication Date: 2025-05-13YANSHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063935.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

It is difficult for the existing technology to comprehensively and accurately predict the financial risks of listed companies from multiple dimensions, especially in the application of performance briefing text data.

Method used

The enterprise financial risk warning method based on large language model and Logistic model is used to analyze the performance briefing text data from multiple dimensions such as text sentiment value, text readability, text similarity and question-and-answer correlation, and use the Logistic model to predict it in combination with financial data.

Benefits of technology

It has achieved a more comprehensive and accurate prediction of the financial risks of listed companies, improved the predictive capabilities of the credit risk warning model, and helped protect the interests of investors and maintain the stability of the securities market.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990752A_ABST
    Figure CN119990752A_ABST
Patent Text Reader

Abstract

The invention discloses an enterprise financial risk early warning method based on a large language model and a Logistic model, and belongs to the technical field of computer natural language processing, and the method comprises the following steps: S1, respectively collecting text data of an annual performance description of a (t-2) th year and one-quarter, half-year and three-quarter performance description of a (t-1) th year of a listed company, collecting financial data of the (t-1) th year; s2, respectively calculating a text sentiment value, text readability and text similarity for a management layer text on a performance description meeting, and calculating question and answer correlation for a dialogue text between a management layer and an investor; and S3, predicting whether the listed companies are subjected to ST processing or not by using a Logistic model by using company financial data in combination with the text sentiment values, the text readability, the text similarity and the question and answer correlation. According to the method, the financial risk of the listed company can be predicted more comprehensively and accurately from multiple dimensions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer natural language processing, and in particular relates to an enterprise financial risk early warning method based on a large language model and a Logistic model. Background Art

[0002] Selecting reasonable risk assessment indicators and building a reasonable credit risk early warning model are of great significance to help operators timely discover company financial problems and respond and prevent them, as well as to protect the interests of investors and maintain the stability of the securities market. Traditional early warning models are usually based on the financial indicator data of listed companies. In recent years, some studies have also included corporate information disclosure texts in the research objects, such as the "management discussion and analysis (MD&A) text" in the company's annual report. First, the emotional tone of the text is analyzed. By introducing the emotional dictionary, the emotional information of the text sentences is calculated and analyzed. The risk of the company's stock will increase with the increase of negative tone, and the return will increase with the increase of positive tone; secondly, the impact of text similarity on corporate risk is examined. A large amount of similar information will conceal the negative information of the company, and the internal risk of the company will increase; then the impact of text readability on corporate risk is explored. By increasing the complexity of the text, the difficulty of extracting text information is increased, and the information content that the company is unwilling to disclose is hidden.

[0003] As an important part of the company's information disclosure, the performance briefing is an important channel for investors to understand the company's production and operation conditions. It is also an important platform for the company's management to communicate with investors. The performance briefing is usually held after the annual report is disclosed. The company's executives will interpret the company's annual report, explain and answer the industry conditions, development strategies, financial conditions and operating performance, and other issues that investors are concerned about. However, there is a relative lack of research on the text of the performance briefing. The lower the relevance of the management's answers to the questions asked by the questioner in the performance briefing, the more likely it is that there are certain hidden dangers in the company's production and operation. With the development of large language model technology, as a large language model, the recognition of text features such as text sentiment and question-answer relevance (the relevance between the management's answers and the questions asked) has been significantly improved compared with traditional models. The large language model is a deep learning model, especially in the field of natural language processing (NLP). It generally refers to a language model containing hundreds of billions or more parameters. These parameters are trained on a large amount of text data. The purpose of the large language model is to understand and generate natural language, and to predict the next word or generate content related to a given text by learning a large amount of text data. Based on this, the performance briefing text disclosure indicators can be quantified from four perspectives: text sentiment value, text readability, text similarity, and question-answer relevance, and introduced into the enterprise credit risk early warning model. The multi-dimensional examination of performance briefing text disclosure indicators can improve the accuracy of credit risk early warning. A large number of experiments have shown that the performance briefing text characteristics have a significant effect on improving the effect of the credit risk early warning model.

[0004] Therefore, there is a need for an enterprise financial risk warning method based on a large language model and a logistic model that can more comprehensively and accurately predict the financial risks of listed companies from multiple dimensions such as text sentiment value, text readability, text similarity, and question-answer relevance. Summary of the invention

[0005] The purpose of the present invention is to provide an enterprise financial risk early warning method based on a large language model and a logistic model, which can more comprehensively and accurately predict the financial risks of listed companies from multiple dimensions such as text sentiment value, text readability, text similarity and question-answer correlation, and help protect the interests of investors and maintain the stability of the securities market.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is:

[0007] An enterprise financial risk early warning method based on a large language model and a Logistic model comprises the following steps:

[0008] Step S1: Collect the text data of the annual performance briefing of the listed company in the t-2th year and the first quarter, half-year, and third quarter performance briefings in the t-1th year, and collect its financial data in the t-1th year;

[0009] Step S2: For the management text in the performance briefing, the text sentiment value, text readability, and text similarity are calculated respectively; for the dialogue text between the management and investors, the question-answer correlation is calculated;

[0010] Step S3: Using the company's financial data and combining text sentiment value, text readability, text similarity, and question-answer relevance, a Logistic model is used to predict whether a listed company will be subject to ST treatment.

[0011] A further improvement of the technical solution of the present invention is that in step S1, the text data of the performance briefing of the target company is collected. This process can be described as follows:

[0012] document=html.xpath(′ / / div[@id="main_content"] / / text()′)

[0013] Among them, html is the webpage where the target company's performance briefing is held, and main_content is the webpage element id of the dialogue text between the company's management and investors in the webpage.

[0014] A further improvement of the technical solution of the present invention is that in step S2, the sentiment value of the text is calculated, and this process can be described as:

[0015]

[0016] Among them, S is the sentiment value of the text, C i is the sentiment category of the i-th company management statement, with values ​​of -1 negative, 0 neutral, or 1 positive, and N is the total number of statements.

[0017] A further improvement of the technical solution of the present invention is that in step S2, the question-answer correlation is calculated. This process can be described as:

[0018]

[0019] Among them, R represents the question-answer relevance, R i It represents the relevance category of the i-th question-answer pair, with values ​​of 0 (irrelevant), 1 (generally relevant), or 2 (highly relevant), and M is the total number of all question-answer pairs.

[0020] A further improvement of the technical solution of the present invention is that in step S2, the text similarity is calculated, and this process can be described as:

[0021] ω i=(ω i1 ,ω i2 ,...,ω in )

[0022] Among them, ω i Represents a sentence, ω in Indicates the frequency of occurrence of a word in a sentence;

[0023]

[0024] Among them, n means there are n text pairs in the tth year, Sim t Represents the similarity of the text of the company's performance briefing in year t.

[0025] A further improvement of the technical solution of the present invention is that in step S2, the text readability is calculated, and this process can be described as follows:

[0026]

[0027] Among them, Rab is the readability of the text, Cpx is the number of complex words in the sentence, and Level C and Level D words consisting of more than three morphemes are selected as complex words. Pfl is the number of professional words in the sentence, and Long represents the length of the sentence, that is, the total number of words.

[0028] A further improvement of the technical solution of the present invention is that in step S3, the company's financial data is used in combination with text sentiment value, text readability, text similarity, and question-answer correlation, and a Logistic model is used to predict whether a listed company will be subject to ST treatment. This process can be described as follows:

[0029]

[0030] Among them, pre is the probability that the predicted result is processed by ST, x1, x2, ..., x n They are text sentiment value, text readability, text similarity, question-answer relevance, and financial indicators.

[0031] Due to the adoption of the above technical solution, the technical progress achieved by the present invention is:

[0032] The enterprise financial risk early warning method based on the large language model and the Logistic model in the present invention supplements the lack of research on the performance briefing text of listed companies and enterprise financial risk early warning. It can predict the financial risks of listed companies more comprehensively and accurately from multiple dimensions such as text sentiment value, text readability, text similarity and question-answer correlation, which is helpful to protect the interests of investors and maintain the stability of the securities market.

[0033] The enterprise financial risk early warning method based on the large language model and the Logistic model of the present invention adopts text sentiment value, text similarity and text readability which can not only be used for the research of credit risk early warning on MD&A text, but also can be used on the performance briefing text to improve the prediction ability of the credit risk early warning model. The question-answer correlation text features unique to the performance briefing text can significantly improve the prediction ability of the credit risk early warning model.

[0034] The enterprise financial risk early warning method of the present invention is based on a large language model and a Logistic model. The sentiment analysis ability of the large language model adopted by the invention, such as ChatGLM3-6B, after fine-tuning is better than that of the Bert model before and after fine-tuning. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a model flow chart of the present invention. DETAILED DESCRIPTION

[0036] The present invention is further described in detail below in conjunction with embodiments:

[0037] like Figure 1 As shown, the present invention provides an enterprise financial risk early warning method based on a large language model and a Logistic model, comprising the following steps:

[0038] Step S1: Collect the text data of the annual performance briefing of the listed company in the t-2th year and the first quarter, half-year, and third quarter performance briefings in the t-1th year, and collect its financial data in the t-1th year;

[0039] Specifically, the process of collecting the text data of the target company's performance briefing can be described as follows:

[0040] document=html.xpath(′ / / div[@id="main_content"] / / text()′)

[0041] Among them, html is the webpage where the target company's performance briefing is held, and main_content is the webpage element id of the dialogue text between the company's management and investors in the webpage;

[0042] Step S2: For the management text in the performance briefing, the text sentiment value, text readability, and text similarity are calculated respectively. For the dialogue text between management and investors, the question-answer correlation is calculated. Specifically, the text sentiment value is calculated. This process can be described as:

[0043]

[0044] Among them, S is the sentiment value of the text, C iis the sentiment category of the i-th company management statement, taking values ​​of -1 (negative), 0 (neutral), or 1 (positive), and N is the total number of statements;

[0045] Calculate the question-answer relevance. This process can be described as:

[0046]

[0047] Among them, R represents the question-answer relevance, R i represents the relevance category of the i-th question-answer pair, and its value is 0 (irrelevant), 1 (moderately relevant) or 2 (highly relevant), and M is the total number of all question-answer pairs;

[0048] Calculate text similarity. This process can be described as:

[0049] ω i =(ωi i1 ,ω i2 ,...,ω in )

[0050] Among them, ω i Represents a sentence, ω in Indicates the frequency of occurrence of a word in a sentence;

[0051]

[0052] Among them, n means there are n text pairs in the tth year, Sim t represents the similarity of the text of the company's performance briefing in year t;

[0053] Calculate the text readability. This process can be described as:

[0054]

[0055] Among them, Rab is the readability of the text, Cpx is the number of complex words in the sentence, and the "Outline of Chinese Proficiency Vocabulary and Chinese Character Levels" issued by the National Chinese Proficiency Test Committee selects Level C and Level D words consisting of more than three morphemes as complex words, Pfl is the number of professional words in the sentence, and professional words refer to the "Accounting Professional Vocabulary Dictionary", and Long represents the length of the sentence, that is, the total number of words;

[0056] Step S3: Using the company's financial data and combining text sentiment value, text readability, text similarity, and question-answer relevance, the Logistic model is used to predict whether the listed company will be ST-treated. This process can be described as:

[0057]

[0058] Among them, pre is the probability that the predicted result is processed by ST, x1, x2, ..., x n They are text sentiment value, text readability, text similarity, question-answer relevance, and financial indicators.

[0059] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.

Claims

1. A corporate financial risk early warning method based on a large language model and a logistic model, characterized in that The following steps are involved: Step S1: Collect the text data of the annual performance briefing of the listed company in the t-2th year and the first quarter, half-year, and third quarter performance briefings in the t-1th year, and collect its financial data in the t-1th year; Step S2: For the management text in the performance briefing, the text sentiment value, text readability, and text similarity are calculated respectively; for the dialogue text between the management and investors, the question-answer correlation is calculated; Step S3: Using the company's financial data and combining text sentiment value, text readability, text similarity, and question-answer relevance, a Logistic model is used to predict whether a listed company will be subject to ST treatment.

2. The enterprise financial risk early warning method based on a large language model and a logistic model according to claim 1 is characterized by: In step S1, the text data of the target company's performance briefing is collected. This process can be described as follows: document=html.xpath(′ / / div[@id="main_content"] / / text()′) Among them, html is the webpage where the target company's performance briefing is held, and main_content is the webpage element id of the dialogue text between the company's management and investors in the webpage.

3. The enterprise financial risk early warning method based on large language model and logistic model according to claim 1 is characterized by: In step S2, the sentiment value of the text is calculated. This process can be described as: Among them, S is the sentiment value of the text, C i is the sentiment category of the i-th company management statement, with values ​​of -1 negative, 0 neutral, or 1 positive, and N is the total number of statements.

4. The enterprise financial risk early warning method based on large language model and logistic model according to claim 3 is characterized by: In step S2, the question-answer correlation is calculated. This process can be described as: Among them, R represents the question-answer relevance, R i It represents the relevance category of the i-th question-answer pair, with values ​​of 0 (irrelevant), 1 (generally relevant), or 2 (highly relevant), and M is the total number of all question-answer pairs.

5. The enterprise financial risk early warning method based on large language model and logistic model according to claim 4 is characterized by: In step S2, the text similarity is calculated. This process can be described as: oh i =(ω i1 ,oh i2 ,...,oh in ) Among them, ω i Represents a sentence, ω in Indicates the frequency of occurrence of a word in a sentence; Among them, n means there are n text pairs in the tth year, Sim t Represents the similarity of the text of the company's performance briefing in year t.

6. The enterprise financial risk early warning method based on large language model and logistic model according to claim 5 is characterized by: In step S2, the text readability is calculated. This process can be described as: Among them, Rad is the text readability, Cpx is the number of complex words in the sentence, and Level C and Level D words consisting of more than three morphemes are selected as complex words. Pfl is the number of professional words in the sentence, and Long represents the length of the sentence, that is, the total number of words.

7. The enterprise financial risk early warning method based on large language model and logistic model according to claim 1 is characterized by: In step S3, the company's financial data is combined with text sentiment value, text readability, text similarity, and question-answer relevance to use the Logistic model to predict whether the listed company will be subject to ST treatment. This process can be described as follows: Among them, pre is the probability that the predicted result is processed by ST, x1, x2, ..., x n They are text sentiment value, text readability, text similarity, question-answer relevance, and financial indicators.