A method for automatically generating research reports based on knowledge graphs and large models
By constructing enterprise knowledge graphs and large model technology, industry research report templates are generated and the text is polished, solving the problems of limited information processing capabilities and low content credibility in existing research report generation methods, and realizing efficient and reliable research report generation.
Patent Information
- Application Number
- CN202310811921.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-04
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-07-04
AI Technical Summary
Existing research report generation methods rely on human labor and web scraping technology, which have problems such as limited information processing capabilities, insufficient language fluency, poor data accuracy, and legal risks.
By constructing a knowledge graph of enterprise multidimensional data and combining it with big data modeling technology, industry research report templates are generated. The big data model is then used to refine the basic text, generate semantically coherent research report content, and supplement the chart generation program.
It enables the generation of efficient and reliable research reports, reduces labor costs, improves information processing capabilities and content credibility, and avoids problems such as rigidity and data fabrication.
Smart Images

Figure CN116842195B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge graphs, and more particularly to a method for automatically generating research reports based on knowledge graphs and large models. Background Technology
[0002] Knowledge graphs are structured semantic knowledge bases. Unlike traditional relational databases, they can represent entity concepts and relationships in the physical world in symbolic form. The basic building block of a knowledge graph is the "entity-relation-entity" triple, with entities interconnected using relations to form a network-like knowledge structure. The main difference between knowledge graphs and traditional relational databases lies in their ability to store knowledge in a graph format, facilitating knowledge representation, retrieval, and computation.
[0003] Large language models (LLMs) refer to language models containing hundreds of billions (or more) parameters trained on massive amounts of text data, such as models GPT-3, PalM, Galactica, and LLaMA. Specifically, LLMs are built on the Transformer architecture, where multi-head attention layers are stacked within a very deep neural network. Existing LLMs primarily employ a similar model architecture (i.e., Transformer) and pre-training objective (i.e., language modeling) to smaller language models. The difference lies in the fact that LLMs significantly expand the model size, pre-training data, and total computational cost.
[0004] Research reports generally refer to securities firm research reports, often called "sell-side research reports." These reports are primarily prepared by securities firm researchers who use securities and related data, research information, facts, and industry development trends as a basis to conduct a comprehensive analysis of the value of specific securities and summarize their views on listed companies, industries, and macroeconomic policies.
[0005] Current research reports are mainly divided into the following types:
[0006] 2.1 Manual writing: Research reports are written by securities analysts based on their comprehensive judgment of data analysis, indicator calculation, and industry experience across various dimensions.
[0007] 2.2 Research Report Crawler: Relies on crawling research reports written and published by analysts on various public websites to obtain data;
[0008] 2.3 Machine writing:
[0009] 2.3.1 Based on preset questions and a preset knowledge base, the matching degree between the questions and each paragraph is determined, and the most matching answers from each paragraph are selected from the knowledge base and spliced together.
[0010] 2.3.2 Automatic generation based on language models. For example, the traditional sequence generation model Seq2Seq and the large-scale language model LLaMa.
[0011] Existing technologies rely heavily on human resources, placing high demands on securities analysts' information gathering abilities and industry judgment experience, and limiting their ability to process information across the entire industry and various dimensions.
[0012] Web scraping is merely a means of knowledge transfer and cannot truly assess the value of analytical reports. For some paid research reports, blindly scraping them may carry legal risks.
[0013] For the first type: Since the focus of knowledge base and question-answer pairs varies across industries, significant resources are often required to build these systems. Furthermore, the accuracy of matching the question-answer pairs with the knowledge base needs to be considered, and the generated language may lack naturalness, leaning towards procedural language. For the second type: Larger models offer significantly better fluency and self-consistency compared to traditional generative sequence models. However, their biggest drawback is the inability to provide accurate data, easily giving the impression of seemingly nonsensical statements, thus compromising trust in real-world production applications. Summary of the Invention
[0014] In view of the above problems, the present invention is proposed to provide a method for automatically generating research reports based on knowledge graphs and large models to overcome or at least partially solve the above problems.
[0015] According to one aspect of the present invention, a method for automatically generating research reports based on knowledge graphs and large models is provided, the method comprising:
[0016] Step S1: Construct a company intelligence knowledge graph based on the company's multidimensional data;
[0017] Step S2: Develop industry research report templates based on business experts;
[0018] Step S3: Obtain the corresponding industry and determine the corresponding industry research report template based on the specified company;
[0019] Step S4: Perform corresponding data queries, calculations, and basic text generation for each paragraph.
[0020] Step S5: Generate content based on the basic text of each paragraph, and refine and adapt it using the hinting engineering of the large model;
[0021] Step S6: Supplement relevant charts as needed using a chart generation program;
[0022] Step S7: Unify and polish the generated content of each paragraph;
[0023] Step S8: Output the final result as the result of the research report generation program.
[0024] Optionally, the enterprise's multidimensional data may specifically include: financial data, subsidiary data, director and senior executive data, industry development data, and relevant policy data.
[0025] Optionally, the charts may specifically include: bar charts, line charts, and tables.
[0026] Optionally, for a specific company, forming basic content specifically includes: firstly, querying the industry in which the company is located, and then performing relevant knowledge queries and calculations according to each paragraph to form basic content.
[0027] Optionally, after refining the text using large model technology based on the basic content and prompt words, the process further includes: if a chart is needed, generating the relevant chart through a chart generation program.
[0028] This invention provides a method for automatically generating research reports based on knowledge graphs and large-scale models. The method includes: constructing a company intelligence knowledge graph based on multi-dimensional enterprise data; building industry research report templates for different industries; forming basic content for a specific company; refining the basic content using large-scale model technology in conjunction with prompt words to obtain the refined text; assembling and integrating the content generated from each paragraph; and finally outputting a complete research report. This method combines the advantages of knowledge graphs (excelling in precise querying and calculation) and large-scale models (generating semantically coherent text) with the disadvantages of knowledge graph templates (being overly mechanical and rigid) and large-scale models (fabricating data), resulting in credible and usable content, reducing labor costs, and providing decision support.
[0029] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0030] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 A flowchart of a method for automatically generating research reports based on knowledge graphs and large models, provided in an embodiment of the present invention;
[0032] Figure 2 An example of a company intelligence knowledge graph provided for an embodiment of the present invention. Detailed Implementation
[0033] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0034] The terms "comprising" and "having," and any variations thereof, in the specification, embodiments, claims, and drawings of this invention are intended to cover non-exclusive inclusion, such as including a series of steps or units.
[0035] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0036] This invention discloses a method for generating research reports using knowledge graph technology and large-scale modeling technology. First, a company intelligence knowledge graph is constructed based on multi-dimensional corporate data (such as financial data, subsidiary data, board and executive information, industry development, and relevant policies). Then, industry research report templates are created for different industries. For a specific company, the industry is first queried, and relevant knowledge queries and calculations are performed sequentially according to each paragraph to form basic content. Then, based on the basic content and prompts, large-scale modeling technology is used to refine the text, resulting in the refined text. If charts are needed, relevant charts (bar charts, line charts, tables, etc.) are generated using a chart generation program. Finally, the content generated from each paragraph is assembled and integrated, and the complete research report result is output.
[0037] like Figure 1 As shown, a method for automatically generating research reports based on knowledge graphs and large models is proposed:
[0038] Step S1: Construct a company intelligence knowledge graph based on multi-dimensional enterprise data (financial data, executive data, industry data, policy data, etc.);
[0039] Step S2: Develop industry research report templates based on business experts;
[0040] Step S3: Obtain the corresponding industry and determine the corresponding industry research report template based on the specified company;
[0041] Step S4: Perform corresponding data queries, calculations, and basic text generation for each paragraph.
[0042] Step S5: Generate content based on the basic text of each paragraph, and refine and adapt it using the hinting engineering of the large model;
[0043] Step S6: Supplement relevant charts as needed using a chart generation program;
[0044] Step S7: Unify and polish the generated content of each paragraph;
[0045] Step S8: Output the final result as the result of the research report generation program.
[0046] A method for automatically generating research reports based on knowledge graphs and large models requires, in step S1, multi-dimensional enterprise data (financial data, executive data, industry data, policy data, etc.) to construct a company intelligence knowledge graph, such as... Figure 2 As shown.
[0047] A method for automatically generating research reports based on knowledge graphs and large models is proposed. In step S3, based on the company intelligence knowledge graph ER diagram, a specified company's industry is queried, and the corresponding industry research report template is determined. Taking the 2022 research report generation task of Bank of Beijing as an example, the industry to which "Bank of Beijing" belongs is "Banking". The research report template for the banking industry is as follows:
[0048]
[0049]
[0050] A method for automatically generating research reports based on knowledge graphs and large models, step S4 being knowledge graph calculation. The knowledge calculation is described in the first paragraph of the template; similar calculations will not be elaborated further. The first paragraph's theme is industry analysis, requiring a query of industry asset size. By calculating the asset size and total asset growth rate of all commercial banks, the following table can be obtained:
[0051]
[0052] Based on the above table, the programmed text is as follows: "In 2014, the total assets of the commercial banking industry were 135.24 trillion yuan, with a year-on-year growth rate of 13.5%; in 2015, the total assets of the commercial banking industry were 151.21 trillion yuan, with a year-on-year growth rate of 11.8%; in 2016, the total assets of the commercial banking industry were 172.34 trillion yuan, with a year-on-year growth rate of 13.9%;..."
[0053] A method for automatically generating research reports based on knowledge graphs and large models is proposed. In step S5, the obtained basic text is combined with the large model for prompting engineering polishing. The prompt is as follows: "Please perform reasoning analysis and polishing on the following text, while keeping the data content unchanged. {In 2014, the total assets of the commercial banking industry were 135.24 trillion yuan, with a year-on-year growth rate of 13.5%; in 2015, the total assets of the commercial banking industry were 151.21 trillion yuan, with a year-on-year growth rate of 11.8%; in 2016, the total assets of the commercial banking industry were 172.34 trillion yuan, with a year-on-year growth rate of 13.9%;...}". The final, polished text for this paragraph can be: "In recent years, commercial banks have actively embraced financial technology and promoted digital transformation, resulting in a continuous expansion of the overall industry's asset scale. From 2014 to 2021, the total assets of my country's commercial banks increased from RMB 135.4 trillion to RMB 281.77 trillion, maintaining steady growth. As of the end of 2022, the assets of China's commercial banks had grown to RMB 312.75 trillion, a year-on-year increase of 10.99%, showing a generally positive development trend."
[0054] A method for automatically generating research reports based on knowledge graphs and large models. In step S6, chart generation is performed. If a chart is required, the data from step S4 is automatically converted into the corresponding chart by the program according to the paragraph requirements in the template.
[0055] A method for automatically generating research reports based on knowledge graphs and large models is presented. Step S7 generates the research report result. Steps S4 to S6 are repeated to generate all paragraph content. The generated content is then spliced together, and manual review and polishing are performed to obtain the final research report result.
[0056] Beneficial effects: The model of generating research reports by combining knowledge graphs and large-scale models. This model can combine the advantages of knowledge graphs in terms of accurate querying and calculation, and large-scale models in terms of generating semantically coherent text, while avoiding the shortcomings of knowledge graph templates in generating text that is too mechanical and rigid, and large-scale models in terms of fabricated data. As a result, the generated content is credible and usable, reduces labor costs, and assists in providing decision support.
[0057] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for automatically generating research reports based on knowledge graphs and large models, characterized in that, The automated research report generation method includes: Step S1: Construct a company intelligence knowledge graph based on the company's multidimensional data; Step S2: Develop industry research report templates based on business experts; Step S3: Obtain the corresponding industry and determine the corresponding industry research report template based on the specified company; Step S4: Perform corresponding data queries, calculations, and basic text generation for each paragraph. Step S5: Generate content based on the basic text of each paragraph, and refine and adapt it using the hinting engineering of the large model; Step S6: Supplement relevant charts as needed using a chart generation program; Step S7: Unify and polish the generated content of each paragraph; Step S8: Output the final result as the result of the research report generation program.
2. The method for automatically generating research reports based on knowledge graphs and large models according to claim 1, characterized in that, The enterprise's multidimensional data specifically includes: financial data, subsidiary data, board members and senior executives, industry development, and relevant policy data.
3. The method for automatically generating research reports based on knowledge graphs and large models according to claim 1, characterized in that, The charts specifically include: bar charts, line charts, and tables.
4. The method for automatically generating research reports based on knowledge graphs and large models according to claim 1, characterized in that, Step S4: Based on each paragraph, perform the corresponding data query, calculation, and basic text generation content. Specifically, this includes: firstly, querying the industry, and then performing relevant knowledge queries and calculations according to each paragraph to generate basic text content.
5. The method for automatically generating research reports based on knowledge graphs and large models according to claim 1, characterized in that, The process of generating content based on the basic text of each paragraph, and then refining and adapting it using the prompting engineering of the large model, also includes: if charts are needed, generating relevant charts through a chart generation program.
Citation Information
Patent Citations
Financial document intelligent verification method, device and storage medium
CN109117479A
Work ticket process management system
CN113393084A