Financial research report question answering and analysis method and device based on knowledge graph and large model

Through the financial research report Q&A method based on knowledge graphs and big models, the shortcomings of the existing technology in understanding user semantics and data dependence are solved, and efficient and accurate personalized Q&A is achieved to adapt to the rapidly changing financial market.

CN120492568AInactive Publication Date: 2025-08-15SHANGHAI JIAJIAN SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510525382.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing financial research report Q&A technology has shortcomings in understanding user semantics, responding to complex problems, data dependence, computing costs and generalization capabilities, and it is difficult to meet the rapidly changing and diversified financial market needs.

Method used

Using a question-and-answer method based on knowledge graphs and big models, we identify research report content through OCR, use big models to generate rewrite questions and combine knowledge graph query to obtain key information on research reports, and realize personalized question-and-answer.

Benefits of technology

It improves the efficiency and accuracy of Q&A, reduces data demand and computing costs, enhances the flexibility and adaptability of the system, and meets the rapidly changing financial market needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492568A_ABST
    Figure CN120492568A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing, in particular to a financial research report question answering and analysis method and device based on a knowledge graph and a large model. The question answering and analysis method comprises the following steps: S1, when a user uploads a financial research report, reading the research report, and returning the document content and the document format of the research report; s2, putting the document content and the document format obtained by the financial research report reading module into a large model for processing; s3, after a user inputs a question, generating a similar question rewritten by the question, a question keyword and a knowledge graph query statement corresponding to the question according to the question by utilizing the large model, querying graph knowledge and research and report keywords, an overall abstract, a segmented abstract and document contents related to research and report from the graph by utilizing the knowledge graph query statement, and after packaging, storing the graph knowledge and the research and report keywords, the overall abstract, the segmented abstract and the document contents in a database; and inputting the result into the large model, and outputting a final answer after the result is processed by the large model. The question and answer efficiency is improved, the question and answer accuracy is enhanced, personalized customization can be achieved, the data requirement is reduced, and the system flexibility is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a method and device for question-answering and analyzing financial research reports based on knowledge graphs and large models. Background Art

[0002] In the financial field, financial research reports are an important reference material for investors. Financial research reports are usually rich in content and information-intensive. For ordinary investors, understanding and extracting key information is a challenging task. Financial research report question-answering technology aims to help investors or analysts automatically generate answers that meet their needs, thereby improving the understanding and utilization efficiency of financial research reports. Current technical methods are mainly divided into the following four categories: full-text search and keyword matching methods, rule-based question-answering systems, traditional machine learning-based question-answering systems, and pre-trained language model-based question-answering systems. The shortcomings of these four methods are:

[0003] 1. Full-text search and keyword matching: While simple to implement, this method relies solely on keyword matching and fails to understand the semantics of user questions. Consequently, search results contain a large amount of irrelevant content, making it difficult to provide accurate answers. Furthermore, this method has poor handling of synonyms and semantic variations, making it incapable of addressing complex and diverse user questions.

[0004] 2. Rule-based question-answering systems: Rules and templates require manual definition and maintenance, which is time-consuming and labor-intensive. Furthermore, the rules have limited coverage and are unable to cope with the volatile financial markets and the ever-changing content of financial research reports. When the financial market environment changes or new research report formats emerge, the rules need to be redefined and adjusted, increasing system maintenance costs. Furthermore, the answering capabilities are limited by predefined rules, making it difficult to handle questions outside the scope of the rules.

[0005] 3. Question-answering systems based on traditional machine learning: They are highly data-dependent. Model training requires a large amount of high-quality labeled data, which is time-consuming, labor-intensive, and costly to obtain and label. Feature engineering is complex, requiring manual design and extraction of features, extensive domain knowledge, extensive experimentation, and debugging, which increases the difficulty and cost of model development. Generalization capabilities are limited, and models are prone to overfitting. They perform well on training data but are ineffective in real-world applications.

[0006] 4. Question-answering system based on pre-trained language models: Training and inference of pre-trained language models require a lot of computing resources and are computationally expensive. Although pre-trained models can be initially trained on general data, they still require a large amount of labeled data for fine-tuning in specific fields (such as finance) to ensure that the model has sufficient domain knowledge. Pre-trained models rely on semantic understanding of training data, and their performance depends on the coverage and quality of their training data. When encountering contexts or fields not covered by the training data (such as the financial field), the model performs poorly.

[0007] In summary, the shortcomings of the above-mentioned existing technologies are: strong data dependence, complex feature engineering, high computational cost and limited generalization ability. Summary of the Invention

[0008] The purpose of the present invention is to address the problems existing in the background technology and propose a financial research report question-answering and analysis method and device based on knowledge graphs and large models, which improves the efficiency of question-answering, enhances the accuracy of question-answering, can be customized, reduces data requirements and computing costs, and improves system flexibility.

[0009] On the one hand, the present invention proposes a financial research report question-answering and analysis method based on a knowledge graph and a large model, comprising the following steps:

[0010] S1. When a user uploads a financial research report, read the report and return the document content and format of the report;

[0011] S2. The document content and format obtained by the financial research report reading module are processed in the large model to obtain possible entities, relationships, and attributes, and then put into the knowledge graph;

[0012] S3. After the user inputs a question, the big model is used to generate similar questions that are rewritten based on the question, question keywords, and knowledge graph query statements corresponding to the question. The knowledge graph query statements are used to query the graph to obtain graph knowledge and research report keywords, overall summaries, segmented summaries, and document content related to the research report. After being encapsulated through the corresponding prompt, they are input into the big model together, and the final answer is output after being processed by the big model.

[0013] Preferably, in step S1, after the user uploads the financial research report, the research report is uniformly converted into an image format, and the OCR model (Optical Character Recognition) is used to identify the text content of the research report, as well as the format-related information of the font, size, table, and text block coordinates. The obtained research report document content and document format information are encapsulated in a fixed prompt and placed in the large model to obtain information related to the research report, including research report keywords, overall summary, segmented summary, document content and document format.

[0014] Preferably, in step S2, the obtained entities, relationships and attributes are put into the knowledge graph, and the knowledge graph is added, deleted and modified.

[0015] Preferably, in step S3, the large model converts the user question into a graph query statement for searching related questions in the knowledge graph.

[0016] On the other hand, the present invention proposes a financial research report question-answering and analysis device based on a knowledge graph and a big model, comprising a financial research report reading module, a knowledge graph updating module and a user question-answering module; the financial research report reading module is used to read the financial research report when the user uploads it, and return the document content and document format of the research report; the knowledge graph updating module is used to put the document content and document format obtained by the financial research report reading module into the big model for processing, obtain possible entities, relationships and attributes, and then put them into the knowledge graph; the user question-answering module is used to use the big model to generate similar questions rewritten from the question, question keywords, and knowledge graph query statements corresponding to the question after the user inputs a question, use the knowledge graph query statement to query the graph to obtain graph knowledge and research report keywords, overall summary, segmented summary and document content related to the research report, and after being encapsulated by the corresponding prompt, input them into the big model together, and output the final answer after being processed by the big model.

[0017] Preferably, after the user uploads the financial research report, the financial research report reading module uniformly converts the research report into an image format, uses the OCR model to identify the text content of the research report, as well as the format-related information of the font, size, table, and text block coordinates, and encapsulates the obtained research report document content and document format information in a fixed prompt and places it into the large model to obtain information related to the research report, including research report keywords, overall summary, segmented summary, document content, and document format.

[0018] Preferably, the knowledge graph updating module places the entities, relationships and attributes obtained by the financial research report reading module into the knowledge graph, and performs addition, deletion and modification of the knowledge graph.

[0019] Preferably, in the user question-answering module, the large model converts user questions into graph query statements, which are used to query related questions in the knowledge graph.

[0020] Compared with the prior art, the present invention has the following beneficial technical effects:

[0021] This invention can effectively improve the efficiency of interpreting financial research reports, enhance the efficiency and accuracy of question-answering, improve user experience, reduce data requirements and computing costs, improve system flexibility and generalization capabilities, and meet the rapidly changing and diverse needs of the financial market. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1This is an overall flow chart of the financial research report question-answering and analysis method based on the knowledge graph and large model according to an embodiment of the present invention;

[0023] Figure 2 Flowchart for reading financial research reports;

[0024] Figure 3 Flowchart for knowledge graph updates;

[0025] Figure 4 Provide a user question and answer flow chart. DETAILED DESCRIPTION

[0026] Example 1

[0027] like Figure 1-Figure 4 As shown, this embodiment proposes a financial research report question-answering and analysis method based on a knowledge graph and a large model, including the following steps:

[0028] S1. When a user uploads a financial research report, the report is read and the document content and format of the report are returned. Specifically, after the user uploads the financial research report, the report is uniformly converted into an image format, and the OCR model is used to recognize the text content of the report, as well as format-related information such as font, size, table, and text block coordinates (OCR recognition is a technology that converts text in an image into editable and searchable text). The obtained research report document content and document format information are encapsulated in a fixed prompt and placed in the large model to obtain information related to the research report, including research report keywords, overall summary, segmented summary, document content, and document format; the prompt of the large model refers to the initial text or question input into the model when using a large machine learning model (such as GPT-3, etc.);

[0029] S2. The document content and format obtained by the financial research report reading module are processed in the large model to obtain possible entities, relationships, and attributes, and then put into the knowledge graph to add, delete, and modify the knowledge graph;

[0030] S3. After the user inputs a question, the big model is used to generate similar questions that are rewritten based on the question, question keywords, and knowledge graph query statements corresponding to the question. That is, after the user inputs the question, the big model processes it and rewrites it into several similar questions. At the same time, the keywords of the question are extracted to allow the big model to better answer it. Through this process, the question itself, similar questions after the question is rewritten, and question keywords are obtained; the big model converts the user question into a graph query statement, and uses the knowledge graph query statement to query the graph to obtain graph knowledge and research report keywords, overall summary, segment summary and document content related to the research report. After being encapsulated through the corresponding prompt, they are input into the big model together, and the final answer is output after being processed by the big model; the big model converts the user question into a graph query statement, which is used to query related questions in the knowledge graph. Since the knowledge contained in the knowledge graph is a priori, this process can well help the big model alleviate the big model illusion problem when answering questions.

[0031] Example 2

[0032] This embodiment proposes a financial research report question-answering and analysis device based on a knowledge graph and a large model, including a financial research report reading module, a knowledge graph updating module and a user question-answering module.

[0033] The financial research report reading module is used to read financial research reports when users upload them, and return the report's document content and document format. Specifically, after users upload financial research reports, the module converts the reports into image formats and uses the OCR model to identify the report's text content, as well as information related to the font, size, table, and text block coordinate format. This module can more quickly obtain key information from the research report. The obtained report document content and document format information are encapsulated in a fixed prompt and placed in the large model to obtain information related to the research report, including research report keywords, overall summary, segmented summary, document content, and document format. Through the processing of OCR technology and the large model, key information of the research report can be quickly obtained, greatly improving the utilization efficiency of financial research reports.

[0034] The knowledge graph update module is used to process the document content and document format obtained by the financial research report reading module into a large model, obtain possible entities, relationships and attributes, and then put them into the knowledge graph to add, delete and modify the knowledge graph, so that the knowledge graph can be updated in real time, so that the information in the knowledge graph is always kept up to date, thereby improving the accuracy of questions and answers.

[0035] The user question-answering module is used to use the big model to generate similar questions of the rewritten questions, question keywords, and knowledge graph query statements corresponding to the questions after the user inputs the question. That is, after the user inputs the question, the big model processes it and rewrites it into several similar questions. At the same time, the keywords of the questions are extracted to better allow the big model to answer. Through this process, the question itself, similar questions after the question is rewritten, and question keywords are obtained; the big model converts the user question into a graph query statement, and uses the knowledge graph query statement to query the graph from the graph to obtain graph knowledge and research report keywords, overall summary, segment summary and document content related to the research report. After being encapsulated by the corresponding prompt, they are input into the big model together and the final answer is output after being processed by the big model; the big model converts the user question into a graph query statement, which is used to query related questions in the knowledge graph. Since the knowledge contained in the knowledge graph is a priori, this process can well help the big model alleviate the big model illusion problem when answering questions. Users can directly input questions, and the large model will automatically generate similar questions that are rewritten, question keywords, and knowledge graph query statements corresponding to the questions, allowing users to obtain the required accurate information more conveniently and quickly, greatly improving the user experience, generating personalized answers, meeting the specific needs of different users, and enhancing the adaptability of the system.

[0036] The present invention can effectively improve the efficiency of interpreting financial research reports, enhance the efficiency and accuracy of question-answering, improve user experience, reduce data requirements and computing costs, and improve system flexibility and generalization capabilities. Through automated reading of research report content and updating of knowledge graphs, the system can quickly respond to user question-answering requests, reduce manual participation, and improve question-answering efficiency. By utilizing the powerful semantic understanding capabilities of large models and the structured information of knowledge graphs, more accurate and relevant answers are provided to enhance user experience. Based on the questions input by users, personalized answers are generated to meet the specific needs of different users and enhance the adaptability of the system. By combining knowledge graphs and large models, dependence on large amounts of high-quality training data is reduced, and the development and maintenance costs of the system are reduced. The knowledge graph can be dynamically updated and expanded, and can adapt to the ever-changing financial market and research report content, thereby improving the flexibility and scalability of the system.

[0037] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.

Claims

1. A financial research report question answering and analysis method based on knowledge graph and large model, characterized by: The following steps are involved: S1. When a user uploads a financial research report, read the report and return the document content and format of the report; S2. The document content and format obtained by the financial research report reading module are processed in the large model to obtain possible entities, relationships, and attributes, and then put into the knowledge graph; S3. After the user inputs a question, the big model is used to generate similar questions that are rewritten based on the question, question keywords, and knowledge graph query statements corresponding to the question. The knowledge graph query statements are used to query the graph to obtain graph knowledge and research report keywords, overall summaries, segmented summaries, and document content related to the research report. After being encapsulated through the corresponding prompt, they are input into the big model together, and the final answer is output after being processed by the big model.

2. The financial research report question-answering and analysis method based on knowledge graph and large model according to claim 1 is characterized in that: In step S1, after the user uploads the financial research report, the research report is uniformly converted into an image format, and the OCR model is used to identify the text content of the research report, as well as the format-related information of the font, size, table, and text block coordinates. The obtained research report document content and document format information are encapsulated in a fixed prompt and placed in the large model to obtain information related to the research report, including research report keywords, overall summary, segmented summary, document content and document format.

3. The financial research report question-answering and analysis method based on knowledge graph and large model according to claim 2 is characterized in that: In step S2, the obtained entities, relationships, and attributes are put into the knowledge graph to add, delete, and modify the knowledge graph.

4. The financial research report question-answering and analysis method based on knowledge graph and large model according to claim 3 is characterized in that: In step S3, the large model converts the user question into a graph query statement, which is used to query related questions in the knowledge graph.

5. A financial research report question-answering and analysis device based on knowledge graph and large model, characterized by: include: The financial research report reading module is used to read the financial research report uploaded by the user and return the document content and document format of the research report; The knowledge graph update module is used to process the document content and format obtained by the financial research report reading module into the large model, obtain possible entities, relationships and attributes, and then put them into the knowledge graph; The user question-and-answer module is used to generate similar questions that are rewritten based on the question, question keywords, and knowledge graph query statements corresponding to the question after the user inputs a question. The knowledge graph query statement is used to query the graph to obtain graph knowledge and research report keywords, overall summaries, segmented summaries, and document content related to the research report. After being encapsulated through the corresponding prompt, they are input into the big model together, and the final answer is output after being processed by the big model.

6. The financial research report question-answering and analysis device based on knowledge graph and large model according to claim 5 is characterized in that: After the user uploads the financial research report, the financial research report reading module converts the report into an image format and uses the OCR model to identify the text content of the research report, as well as format-related information such as font, size, table, and text block coordinates. The obtained research report document content and document format information are encapsulated in a fixed prompt and placed in the large model to obtain information related to the research report, including research report keywords, overall summary, segmented summary, document content, and document format.

7. The financial research report question-answering and analysis device based on knowledge graph and large model according to claim 6 is characterized in that: The knowledge graph update module puts the entities, relationships and attributes obtained by the financial research report reading module into the knowledge graph, and adds, deletes and modifies the knowledge graph.

8. The financial research report question-answering and analysis device based on knowledge graph and large model according to claim 7 is characterized in that: In the user question-answering module, the large model converts user questions into graph query statements, which are used to query related questions in the knowledge graph.