An optimization and invocation method, apparatus, device, and storage medium for a local large language model suitable for financial statement analysis.

By building an industry indicator library and optimizing the local large language model, the problems of sensitive information leakage and industry differences in financial statement analysis have been solved, achieving highly accurate financial analysis and meeting the actual needs of financial institutions.

CN119294382BActive Publication Date: 2025-10-28GUANGDONG YUECAI FINANCIAL CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411058875.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2025-10-28
Estimated Expiration
2044-08-02

AI Technical Summary

Technical Problem

Existing large language models suffer from risks of leaking sensitive internal information and insufficient accuracy due to industry differences in financial statement analysis, making it difficult to meet the actual application needs of financial institutions.

Method used

By building an industry indicator library, generating industry-specific prompts, using external large language models for analysis, and optimizing the local large language model through manual annotation, we provide industry feature support and generate financial analysis reports.

Benefits of technology

It improves the accuracy and reliability of local large language models in financial statement analysis, avoids the leakage of sensitive information, and meets the actual application needs of financial institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119294382B_ABST
    Figure CN119294382B_ABST
Patent Text Reader

Abstract

This invention provides an optimization and invocation method for a local large language model suitable for financial statement analysis. It constructs an industry indicator library by categorizing financial statement data from the internet by industry and calculating their statistical values, providing data with industry characteristics. Then, a prompt is generated based on this data, and an external large language model is used to output the corresponding financial analysis report. These financial analysis reports are used as training samples to train the local large language model. Thus, the trained local large language model can be invoked to output industry-specific analysis results for the financial statement data to be analyzed within the enterprise. This solves the problems of directly using external large language models for financial statement analysis easily leading to the leakage of sensitive internal information, and the difficulty of external large language models in combining industry characteristics to output different financial statement analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of application deployment technology, specifically relating to an optimization and invocation method, apparatus, device, and storage medium for a local large language model suitable for financial statement analysis. Background Technology

[0002] Financial statements are crucial documents that allow companies to present their financial condition and operating results to the public. Analyzing financial statements is essential for investors, management, creditors, and other stakeholders. On one hand, financial reports provide historical data, which companies can use to develop budgets, make financial forecasts, and plan for the long term. On the other hand, investors use financial reports to assess a company's financial health, profitability, risk level, and future growth potential, thereby making investment decisions to buy, hold, or sell.

[0003] With the rapid development of information technology, information technologies such as the Internet of Things, big data, cloud computing, and mobile internet have emerged one after another. The financial industry is also undergoing in-depth digital transformation, and financial institutions' demand for intelligent financial statement analysis is becoming increasingly prominent. Using large language models for financial statement analysis brings many advantages, such as the ability to quickly process large amounts of data and identify patterns and trends.

[0004] However, despite the progress made by large language models in financial statement analysis, significant constraints and challenges remain. On the one hand, directly using external large language models for financial statement analysis poses a potential risk of leakage of sensitive internal information to financial institutions due to the data transmission and processing involved. This risk could lead to the illegal acquisition or misuse of core business data, threatening trade secrets and customer privacy. On the other hand, while localized large language models offer advantages in data security and privacy protection, their performance often cannot match that of advanced external models. This is because, although financial statement data is abundant online, high-quality data samples suitable for training local large language models are scarce. This imbalance in data resources severely restricts financial institutions' ability to train localized large language models suitable for financial statement analysis, making it difficult for the models to learn effective analytical patterns and rules, thus affecting their accuracy and reliability in financial statement analysis. This makes it difficult for localized large language models to meet the actual application needs of financial institutions, limiting their promotion and application in actual business. In particular, financial statement data from companies in different industries differ significantly in content and characteristics. Therefore, when using a general large model to analyze financial statements of different industries, it is often difficult to achieve the desired results. This industry difference makes it difficult for large language models to be applied across industries, and further optimization is needed for different industry characteristics. Summary of the Invention

[0005] The purpose of this invention is to solve the above-mentioned technical problems and provide an optimization and calling method for a local large language model suitable for financial statement analysis. This method can avoid the leakage of sensitive internal information caused by directly using external large language models to analyze internal financial statements, and can output different financial statement analysis results for different industries.

[0006] To solve the above problems, the present invention is implemented according to the following technical solution:

[0007] In a first aspect, the present invention provides an optimization and invocation method for a local large language model suitable for financial statement analysis, the method comprising:

[0008] Calculate statistical values ​​of financial indicators from multiple financial statements by industry dimension to form an industry indicator library;

[0009] Based on the industry indicator library, obtain the statistical values ​​of the financial indicators corresponding to the industry to which each of the financial statement data belongs;

[0010] Obtain the first raw value for each item in each of the financial statement data, wherein the first raw value is the unadjusted value reported by each item in the financial statement data;

[0011] Based on each of the financial statement data, a first financial indicator value is obtained, wherein the first financial indicator value is obtained by calculating the financial indicator value of each company corresponding to the financial statement data.

[0012] Based on the statistical values ​​of financial indicators, the first original value, and the first financial indicator value corresponding to the industry to which the financial statement data belongs, the first prompt is generated in batches. The first prompt is a prompt or instruction to request an external large language model to output analysis results for the financial statement data.

[0013] Based on the first prompt, a first financial analysis report is obtained, wherein the first financial analysis report is the result of multiple analyses of the financial statement data by multiple external large language models;

[0014] The local large language model is optimized by training samples, wherein the training samples include manually annotated first financial analysis reports;

[0015] A second prompt is generated for the financial statement data to be analyzed, wherein the second prompt is a prompt or instruction to request the local large language model to output the analysis results for the financial statement data to be analyzed;

[0016] Based on the second prompt, a second financial analysis report is obtained, wherein the second financial analysis report is the analysis result of the local large language model on the financial statement data to be analyzed.

[0017] Preferably, the data in the industry indicator library is panel data.

[0018] Preferably, the local large language model is optimized through supervised fine-tuning.

[0019] Preferably, the step of generating a second prompt for the financial statement data to be analyzed includes: obtaining statistical values ​​of financial indicators corresponding to the financial statement data to be analyzed from the industry indicator library; obtaining second raw values ​​for each item in the financial statement data to be analyzed, wherein the second raw values ​​are the unadjusted values ​​reported by each item in the financial statement data to be analyzed; obtaining second financial indicator values, wherein the second statistical values ​​of financial indicators are obtained by calculating the financial indicator values ​​of the company corresponding to the financial statement data to be analyzed; and generating a second prompt based on the statistical values ​​of financial indicators, the second raw values, and the second financial indicator values ​​corresponding to the financial statement data to be analyzed.

[0020] Preferably, the step of obtaining the statistical values ​​of the financial indicators corresponding to the financial statement data to be analyzed from the industry indicator library includes: obtaining the industry to which the financial statement data to be analyzed belongs; obtaining the financial reporting time point of the financial statement data to be analyzed; and obtaining the statistical values ​​of the financial indicators corresponding to the financial statement data to be analyzed from the industry indicator library based on the industry to which the financial statement data to be analyzed belongs and the financial reporting time point.

[0021] Preferably, the step of obtaining the second original value of each item in the financial statement data to be analyzed includes: performing data cleaning on the financial statement data to be analyzed; and obtaining the second original value of each item in the financial statement data to be analyzed based on the data-cleaned financial statement data.

[0022] Preferably, the method for obtaining the second financial indicator value includes the following steps: obtaining the second financial indicator value based on the cleaned financial statement data to be analyzed, wherein the second financial indicator statistical value is obtained by calculating the financial indicator value of the company corresponding to the financial statement data to be analyzed.

[0023] Secondly, the present invention provides an optimization and invocation apparatus for a local large language model suitable for financial statement analysis. The apparatus is configured to execute the optimization and invocation method for the local large language model suitable for financial statement analysis. The apparatus includes:

[0024] An industry indicator library construction module is used to calculate the statistical values ​​of financial indicators from multiple financial statement data according to industry dimensions, thereby forming an industry indicator library.

[0025] The first prompt generation module is used to: obtain statistical values ​​of financial indicators corresponding to the industry to which each set of financial statement data belongs, based on the industry indicator library; obtain the first original value of each item in each set of financial statement data, wherein the first original value is the value reported by each item in the financial statement data and has not been adjusted; obtain a first financial indicator value based on each set of financial statement data, wherein the first financial indicator value is obtained by calculating the financial indicator value of each company corresponding to the financial statement data; and generate a first prompt in batches based on the statistical values ​​of financial indicators corresponding to the industry to which the financial statement data belongs, the first original value, and the first financial indicator value, wherein the first prompt is a prompt or instruction requesting an external large language model to output analysis results for the financial statement data.

[0026] An external large language model calling module is used to obtain a first financial analysis report based on the first prompt, wherein the first financial analysis report is the result of multiple analyses of the financial statement data by multiple external large language models;

[0027] A local large language model optimization module is used to optimize a local large language model through training samples, wherein the training samples include manually annotated first financial analysis reports.

[0028] The local large language model invocation module is used to generate a second prompt for the financial statement data to be analyzed, wherein the second prompt is a prompt or instruction requesting the local large language model to output analysis results for the financial statement data to be analyzed; and according to the second prompt, a second financial analysis report is obtained, wherein the second financial analysis report is the analysis result of the local large language model on the financial statement data to be analyzed.

[0029] Thirdly, the present invention provides an electronic device, the electronic device comprising:

[0030] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform an optimization and invocation method for a local large language model suitable for financial statement analysis, as described in any of the first aspects above.

[0031] Fourthly, the present invention provides a computer-readable storage medium storing a computer program for causing a processor to execute an optimization and invocation method for a local large language model suitable for financial statement analysis, as described in any one of the first aspects above.

[0032] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides an optimization and invocation method for a local large language model suitable for financial statement analysis. By classifying unlabeled financial statement data by industry and calculating statistical values ​​of financial indicators, an industry indicator library is constructed, providing industry indicator characteristics. This enables the local large language model to combine industry characteristics to output financial analysis reports during subsequent training, solving the problem that the existence of industry differences poses a significant challenge to the cross-industry application of large language models. Combining industry indicator characteristics, a prompt is generated for each financial statement data. An external large language model is then used to output financial analysis reports from these prompts. This transforms unlabeled financial statement data from the internet into training samples that can be used to train the local large language model, providing broader and more comprehensive data support for its training. This improves the accuracy and reliability of the data output by the local large language model in financial statement analysis, meeting the practical application needs of various financial institutions. A prompt is generated for the financial statements to be analyzed. This prompt is then input into a pre-trained local large language model to obtain the corresponding financial analysis report. This avoids the leakage of sensitive internal information that can occur when using an external large language model for internal financial statement analysis, thus ensuring data security.

[0033] Therefore, the optimization and invocation method for a local large language model suitable for financial statement analysis provided by this invention can combine with an external large language model to transform unlabeled financial statement data into training samples that can be used to train the local large language model. This allows internal financial statements to be directly analyzed through the local large language model without calling an external large language model. Furthermore, the local large language model can also output financial statement analysis results based on industry characteristics. This solves the problems that directly using an external large language model for financial statement analysis can easily lead to the leakage of sensitive internal information, and that external large language models are difficult to combine with industry characteristics to output different financial statement analysis results.

[0034] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0035] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings, wherein:

[0036] Figure 1 This is a structural diagram of an optimization and invocation method for a local large language model applicable to financial statement analysis, according to an embodiment of the present invention.

[0037] Figure 2 This is a specific implementation example diagram of constructing an industry indicator library in this invention embodiment;

[0038] Figure 3 This is a specific implementation example diagram of batch generating the first prompt in an embodiment of the present invention;

[0039] Figure 4 This is a specific implementation example diagram of obtaining local large language model training samples in an embodiment of the present invention;

[0040] Figure 5 This is a specific implementation example diagram of calling a local large language model for financial statement analysis in an embodiment of the present invention;

[0041] Figure 6 This is a module diagram of an optimization and calling device for a local large language model applicable to financial statement analysis, according to an embodiment of the present invention.

[0042] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0043] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0044] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. Unless otherwise defined, the technical or scientific terms used in this specification should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar words used in this specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different technical features.

[0045] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0046] Figure 1 This is a structural diagram of a method for optimizing and calling a local large language model suitable for financial statement analysis, as described in this invention. This method can be executed by a device for optimizing and calling a local large language model suitable for financial statement analysis. This device can be implemented in hardware and / or software and can be configured in an electronic device. The method includes:

[0047] Step 101: Calculate the statistical values ​​of financial indicators from multiple financial statements according to industry dimensions to form an industry indicator library.

[0048] It is worth noting that many of the financial statement data are financial statement data that can be collected on the Internet, such as financial statement data of listed companies.

[0049] It is worth noting that the data in the industry indicator library is panel data. Panel data, also known as longitudinal or tracking data, is a type of statistical data that records observations of multiple individuals (such as multiple financial statements) at multiple points in time. Panel data is characterized by possessing both time-series (spanning multiple time points) and cross-sectional (multiple individuals) characteristics, making it highly useful in data analysis and statistical modeling. In this invention, using panel data as the statistical data type for the industry indicator library can reduce estimation bias caused by omitted variables by controlling for individual fixed effects. Panel data can provide more information to infer causal relationships between variables; furthermore, compared to pure time-series or cross-sectional data, panel data provides more information, thereby improving estimation efficiency.

[0050] It is worth noting that the industry dimension refers to organizing and interpreting data according to the standards and characteristics of different industries during the analysis and research process. In this invention, calculating the statistical values ​​of financial indicators by industry dimension means classifying financial statement data according to different industries and calculating the statistical values ​​of financial indicators for each industry separately. Financial indicators include debt-to-equity ratio, current ratio, profit margin, return on equity, and quick ratio, etc., and the statistical values ​​refer to the minimum, average, and maximum values, etc.

[0051] It is worth noting that the data time in the industry indicator library of this invention needs to correspond to the data time of the financial statements.

[0052] Step 102: Based on the industry indicator library, obtain the statistical values ​​of the financial indicators corresponding to the industry to which each financial statement data belongs.

[0053] Step 103: Obtain the first raw value of each item in each financial statement data, wherein the first raw value is the unadjusted value reported by each item in the financial statement data.

[0054] Understandably, in financial statements, "accounts" refers to accounting subjects. These are names used in the accounting information system to classify, record, and summarize economic transactions. Accounting subjects are categorized according to their nature and purpose, such as asset accounts, liability accounts, owner's equity accounts, cost accounts, and profit and loss accounts. When preparing financial statements, accountants record each transaction under the corresponding accounting subject based on the actual economic transactions that occur. The summarization and analysis of data from these "accounts" are crucial for understanding a company's financial position, operating results, and cash flow. By analyzing the changes and relationships between different accounts, an assessment of a company's financial health, profitability, liquidity, and risk profile can be made.

[0055] Step 104: Based on the data from each financial statement, obtain the first financial indicator value, where the first financial indicator value is obtained by calculating the financial indicator value of each company corresponding to the financial statement data.

[0056] Understandably, financial metrics are crucial tools for every company, serving as indicators of its financial health, operational efficiency, and profitability. These metrics help investors, management, creditors, and other stakeholders assess a company's health and future potential. Examples of financial metrics include: debt-to-equity ratio, current ratio, quick ratio, equity ratio, profit margin, and return on equity.

[0057] Step 105: Based on the statistical values ​​of financial indicators corresponding to the industry to which the financial statement data belongs, the first original value, and the first financial indicator value, generate the first prompt in batches. The first prompt is a prompt or instruction to request an external large language model to output analysis results for the financial statement data.

[0058] It should be noted that in the use of large language models, "prompt" refers to the text input by the user, used to guide the large language model to make a specific answer or perform a specific task. A prompt can be a question, instruction, topic, or any form of text input used to stimulate the large language model to generate a response or perform an operation. In this invention, the first prompt is a prompt or instruction requesting the external large language model to output analysis results for financial statement data.

[0059] In this embodiment, the first prompt example is: "You are a professional financial analyst. Please conduct a financial statement analysis based on the following data: ① The company belongs to the XX industry, and the industry indicators are as follows: XXXX; ② The company's original financial statements are as follows: XXXX; ③ The company's commonly used financial indicators are as follows: XXXX."

[0060] It's important to note that external large language models refer to large machine learning models developed by other organizations or companies for processing natural language. Using external large language models to analyze a company's internal financial statement data can easily lead to the leakage of sensitive internal information. This could result in the unauthorized acquisition or misuse of an organization's core business data, thereby threatening its trade secrets and customer privacy.

[0061] Step 106: Based on the first prompt, obtain the first financial analysis report, which is the result of multiple analyses of the financial statement data by multiple external large language models.

[0062] Step 107: Optimize the local large language model using training samples, where the training samples include manually annotated first financial analysis reports.

[0063] It should be noted that a local large language model refers to a large language model deployed in the user's local environment (such as a company's internal computer, server, or local network). Because the model runs locally, sensitive data does not need to be uploaded to an external server, enhancing data privacy and security.

[0064] It should be noted that, in this invention, manual annotation refers to financial experts sorting or scoring multiple financial analysis reports of the same financial statement, which allows for fine-tuning of the local large language model based on the manual annotation.

[0065] Specifically, by manually screening and labeling the primary financial analysis reports, higher-quality samples can be identified from these reports and used as training samples for the local large-scale language model. This further improves the accuracy and reliability of the financial statement analysis results output by the localized large-scale model.

[0066] Specifically, supervised fine-tuning optimizes a local large language model. Supervised fine-tuning is a method for optimizing machine learning models. In supervised fine-tuning, a pre-trained machine learning model (called a pre-trained model) is used as the initial state, and then fine-tuned on the training set (training samples) of the target task, allowing the model to better adapt to the target task. In supervised fine-tuning, the model learns from labeled training data to predict or classify new, unseen data. Supervised fine-tuning can employ methods such as P-tuning and LoRA. In this invention, supervised fine-tuning optimizes and improves the overall performance of a pre-trained local large language model. The local large language model optimized through supervised fine-tuning can support data-driven decision-making, helping enterprises and organizations extract valuable insights from historical data.

[0067] Step 108: Generate a second prompt for the financial statement data to be analyzed, wherein the second prompt is a prompt or instruction to request the local large language model to output the analysis results for the financial statement data to be analyzed.

[0068] Specifically, the steps for generating a second prompt for the financial statement data to be analyzed include: (1) obtaining the industry to which the financial statement data to be analyzed belongs and the financial reporting time; based on the industry and the financial reporting time, obtaining the statistical values ​​of the financial indicators corresponding to the industry of the financial statement data to be analyzed from the industry indicator library. (2) performing data cleaning on the financial statement data to be analyzed, and obtaining the second original value of each item in the financial statement data to be analyzed based on the data cleaned financial statement data to be analyzed, wherein the second original value is the value reported by each item in the financial statement data to be analyzed and has not been adjusted. (3) calculating the financial indicator value of the company corresponding to the financial statement data to be analyzed based on the data cleaned financial statement data to be analyzed. (4) generating a second prompt based on the statistical value of the financial indicator, the second original value and the second financial indicator value corresponding to the financial statement data to be analyzed.

[0069] It should be noted that the format of the second prompt is the same as that of the first prompt. An example of the second prompt is: "You are a professional financial analyst. Please conduct a financial statement analysis based on the following data: ① The company belongs to the XX industry, and the industry indicators are as follows: XXXX; ② The company's original financial statements are as follows: XXXX; ③ The company's commonly used financial indicators are as follows: XXXX."

[0070] It should be noted that the financial reporting point in time refers to a specific date or time within the financial reporting cycle, and these points in time are crucial for determining the accuracy and completeness of financial data.

[0071] It's important to note that data cleaning refers to operations such as removing duplicate data, filling in missing values, handling outliers, and standardizing data formats from collected data. This reduces errors and biases in data analysis, modeling, and model training, improving data quality and reliability. During data mining collection, massive amounts of raw data contain numerous incomplete (missing values), inconsistent, and anomaly-laden data points, severely impacting the efficiency of data mining modeling and potentially leading to biased results. Therefore, data cleaning is crucial. Data cleaning aims to improve data quality and better adapt the data to specific mining techniques or tools.

[0072] Step 109: Based on the second prompt, obtain the second financial analysis report, which is the analysis result output by the local large language model of the financial statement data to be analyzed.

[0073] In summary, this invention provides an optimization and invocation method for a local large language model suitable for financial statement analysis. By classifying unlabeled financial statement data by industry and calculating statistical values ​​of financial indicators, an industry indicator library is constructed, providing industry indicator features. This enables the local large language model to combine industry features to output financial analysis reports during subsequent training, solving the problem that the existence of industry differences poses a huge challenge to the cross-industry application of large language models.

[0074] Secondly, by combining industry indicator characteristics, a prompt is generated for each financial statement data. An external large language model is then used to output financial analysis reports from these prompts. This transforms the original financial statement data from the Internet into training samples that can be used to train the local large language model. This provides broader and more comprehensive data support for the training of the local large language model, improves the accuracy and reliability of the data output by the local large language model in financial statement analysis, and meets the actual application needs of various financial institutions.

[0075] Finally, a prompt is generated for the financial statements to be analyzed. This prompt is then input into the trained local large language model to obtain the corresponding financial analysis report. This avoids the leakage of sensitive internal information that would occur when using an external large language model for internal financial statement analysis, thus ensuring data security.

[0076] Therefore, the optimization and invocation method for a local large language model suitable for financial statement analysis provided by this invention can combine with an external large language model to transform unlabeled financial statement data into training samples that can be used to train the local large language model. This allows internal financial statements to be directly analyzed through the local large language model without calling an external large language model. Furthermore, the local large language model can also output financial statement analysis results based on industry characteristics. This solves the problems that directly using an external large language model for financial statement analysis can easily lead to the leakage of sensitive internal information, and that external large language models are difficult to combine with industry characteristics to output different financial statement analysis results.

[0077] Figure 6 This is a module diagram of an optimization and calling device for a local large language model suitable for financial statement analysis, as described in this invention. Figure 6 As shown, this module includes:

[0078] 301. Industry Indicator Library Construction Module: This module is used to calculate the statistical values ​​of financial indicators from multiple financial statements according to industry dimensions, forming an industry indicator library.

[0079] 302. A first prompt generation module, which is used to obtain the statistical values ​​of financial indicators corresponding to the industry to which each financial statement data belongs, based on an industry indicator library; obtain the first original value of each item in each financial statement data, wherein the first original value is the value reported by each item in the financial statement data and has not been adjusted; obtain the first financial indicator value based on each financial statement data, wherein the first financial indicator value is obtained by calculating the financial indicator value of each company corresponding to the financial statement data; and generate a first prompt in batches based on the statistical values ​​of financial indicators corresponding to the industry to which the financial statement data belongs, the first original value, and the first financial indicator value, wherein the first prompt is a prompt or instruction requesting an external large language model to output analysis results for the financial statement data.

[0080] 303. External large language model calling module, the external large language model calling module is used to obtain the first financial analysis report according to the first prompt, wherein the first financial analysis report is the result of multiple analysis of the financial statement data by multiple external large language models respectively.

[0081] 304. Local Large Language Model Optimization Module: This module optimizes the local large language model using training samples, which include manually labeled first financial analysis reports.

[0082] 305. Local Large Language Model Calling Module: This module generates a second prompt for the financial statement data to be analyzed. The second prompt is a prompt or instruction requesting the local large language model to output analysis results for the financial statement data to be analyzed. Based on the second prompt, a second financial analysis report is obtained, which is the analysis result of the local large language model for the financial statement data to be analyzed.

[0083] Figure 7 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0084] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0085] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0086] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as solving an optimization and invocation method for a local large language model suitable for financial statement analysis.

[0087] In some embodiments, a method for optimizing and invoking a local large language model suitable for financial statement analysis can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for optimizing and invoking a local large language model suitable for financial statement analysis described above can be performed. Alternatively, in other embodiments, processor 11 can be configured by any other suitable means (e.g., by means of firmware) to execute a method for optimizing and invoking a local large language model suitable for financial statement analysis.

[0088] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0089] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0090] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0091] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0092] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0093] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0094] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements an optimization and invocation method for a local large language model suitable for financial statement analysis, as provided in this invention.

[0095] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0096] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0097] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An optimization and invocation method for a local large language model suitable for financial statement analysis, characterized in that, include: Calculate statistical values ​​of financial indicators from multiple financial statements by industry dimension to form an industry indicator library; Based on the industry indicator library, obtain the statistical values ​​of the financial indicators corresponding to the industry to which each of the financial statement data belongs; Obtain the first raw value for each item in each of the financial statement data, wherein the first raw value is the unadjusted value reported by each item in the financial statement data; Based on each of the financial statement data, a first financial indicator value is obtained, wherein the first financial indicator value is obtained by calculating the financial indicator value of each company corresponding to the financial statement data. Based on the statistical values ​​of financial indicators, the first original value, and the first financial indicator value corresponding to the industry to which the financial statement data belongs, the first prompt is generated in batches. The first prompt is a prompt or instruction to request an external large language model to output analysis results for the financial statement data. Based on the first prompt, a first financial analysis report is obtained, wherein the first financial analysis report is the result of multiple analyses of the financial statement data by multiple external large language models; The local large language model is optimized by training samples, wherein the training samples include manually annotated first financial analysis reports; A second prompt is generated for the financial statement data to be analyzed, wherein the second prompt is a prompt or instruction to request the local large language model to output the analysis results for the financial statement data to be analyzed; Based on the second prompt, a second financial analysis report is obtained, wherein the second financial analysis report is the analysis result of the local large language model on the financial statement data to be analyzed; The step of generating a second prompt from the financial statement data to be analyzed includes: Obtain the statistical values ​​of the financial indicators corresponding to the financial statement data to be analyzed from the industry indicator library; Obtain the second raw value of each item in the financial statement data to be analyzed, wherein the second raw value is the unadjusted value reported by each item in the financial statement data to be analyzed; Obtain a second financial indicator value, wherein the second financial indicator value is obtained by calculating the financial indicator value of the company corresponding to the financial statement data to be analyzed; Based on the statistical values ​​of financial indicators, the second original values, and the second financial indicator values ​​corresponding to the financial statement data to be analyzed, a second prompt is generated.

2. The optimization and invocation method for a local large language model suitable for financial statement analysis according to claim 1, characterized in that: The data in the industry indicator library is panel data.

3. The optimization and invocation method for a local large language model suitable for financial statement analysis according to claim 1, characterized in that: Optimize the local large language model through supervised fine-tuning.

4. The optimization and invocation method for a local large language model suitable for financial statement analysis according to claim 1, characterized in that, The step of obtaining the statistical values ​​of financial indicators corresponding to the financial statement data to be analyzed from the industry indicator library includes: Obtain the industry to which the financial statement data to be analyzed belongs; The financial reporting point in time at which the financial statement data to be analyzed is obtained; Based on the industry and the financial reporting time, the statistical values ​​of the financial indicators corresponding to the financial statement data to be analyzed are obtained from the industry indicator library.

5. The optimization and invocation method for a local large language model suitable for financial statement analysis according to claim 1, characterized in that, The step of obtaining the second original value of each item in the financial statement data to be analyzed includes: Data cleaning is performed on the financial statement data to be analyzed; Based on the cleaned financial statement data to be analyzed, the second original value of each item in the financial statement data to be analyzed is obtained.

6. The optimization and invocation method for a local large language model suitable for financial statement analysis according to claim 5, characterized in that, The steps to obtain the second financial indicator value include: Based on the cleaned financial statement data to be analyzed, the second financial indicator value is obtained.

7. An optimization and invocation device for a local large language model suitable for financial statement analysis, characterized in that, include: The industry indicator library construction module is used to calculate the statistical values ​​of financial indicators from multiple financial statement data according to industry dimensions, and form an industry indicator library. The first prompt generation module is used to: obtain statistical values ​​of financial indicators corresponding to the industry to which each set of financial statement data belongs, based on the industry indicator library; obtain the first original value of each item in each set of financial statement data, wherein the first original value is the value reported by each item in the financial statement data and is not adjusted; obtain a first financial indicator value based on each set of financial statement data, wherein the first financial indicator value is obtained by calculating the financial indicator value of each company corresponding to the financial statement data; and generate a first prompt in batches based on the statistical values ​​of financial indicators corresponding to the industry to which the financial statement data belongs, the first original value, and the first financial indicator value, wherein the first prompt is a prompt or instruction to request an external large language model to output analysis results for the financial statement data. The external large language model calling module is used to obtain a first financial analysis report based on the first prompt, wherein the first financial analysis report is the result of multiple analyses of the financial statement data by multiple external large language models respectively; A local large language model optimization module is used to optimize the local large language model through training samples, wherein the training samples include manually annotated first financial analysis reports; The local large language model calling module is used to generate a second prompt for the financial statement data to be analyzed, wherein the second prompt is a prompt or instruction to request the local large language model to output analysis results for the financial statement data to be analyzed; and according to the second prompt, a second financial analysis report is obtained, wherein the second financial analysis report is the analysis result of the local large language model on the financial statement data to be analyzed; The generation of the second prompt from the financial statement data to be analyzed includes: Obtain the statistical values ​​of the financial indicators corresponding to the financial statement data to be analyzed from the industry indicator library; Obtain the second raw value of each item in the financial statement data to be analyzed, wherein the second raw value is the unadjusted value reported by each item in the financial statement data to be analyzed; Obtain a second financial indicator value, wherein the second financial indicator value is obtained by calculating the financial indicator value of the company corresponding to the financial statement data to be analyzed; Based on the statistical values ​​of financial indicators, the second original values, and the second financial indicator values ​​corresponding to the financial statement data to be analyzed, a second prompt is generated.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is... The at least one processor executes such that it is capable of performing an optimization and invocation method for a local large language model suitable for financial statement analysis, as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program. The computer program is used to enable the processor to implement, when executed, an optimization and invocation method for a local large language model suitable for financial statement analysis as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Enterprise financial fraud identification method, device, equipment, medium and program product

    CN117994055A

  • Intelligent session method and system suitable for vertical field based on large model

    CN118035408A