Intelligent data report generation system and method based on large language model

By introducing an intelligent data report generation system based on large language models in the intelligent assistant, combined with Text_to_sql and Retrieval-Augmented Generation technology, the shortcomings of traditional intelligent assistants in complex user queries and data retrieval are solved, and efficient and accurate data report generation and query are achieved.

CN120146016APending Publication Date: 2025-06-13ZHEJIANG GONGSHANG UNIVERSITY

Patent Information

Application Number
CN202510629180.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional conversational intelligent assistants have shortcomings in accurately understanding complex user queries and efficient data retrieval, especially in application scenarios involving a large number of knowledge bases or database backgrounds, and it is difficult to support free question-and-answer and precise extraction of information from knowledge documents and databases based on the context of user queries.

Method used

An intelligent data report generation system based on a large language model is adopted, which includes a natural language processing module, a database query module, a knowledge document search module and a report generation module. Database query is carried out through Text_to_sql technology, knowledge base search is used to search through Retrieval-Augmented Generation technology, and interactive data reports are generated in combination with report generation technology.

Benefits of technology

It significantly improves the efficiency and accuracy of data report production and query, allows users to query and report generation through natural language, and provides intuitive, simple and accurate intelligent data services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146016A_ABST
    Figure CN120146016A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent data report generation system and method based on a large language model, and relates to the field of computer software and artificial intelligence. A natural language processing module in the system is used for analyzing user requirements and judging whether a data source of the user requirements is a data storage module or an enterprise knowledge base module; the database query module is used for generating structured data of the SQL statement query data storage module; the database query module queries a database; the knowledge document retrieval module retrieves unstructured data of the enterprise knowledge base module; the knowledge document retrieval module retrieves a knowledge base by using an RAG technology; the report generation module is used for generating a visual chart according to a user demand and a query result; the data storage module is used for storing the structured data and the generated visual chart; and the enterprise knowledge base module stores the vectorized unstructured knowledge document. According to the method, the data report making and querying efficiency and accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computer software and artificial intelligence, and in particular to an intelligent data report generation system and generation method based on a large language model. Background Art

[0002] With the rapid development of artificial intelligence technology, especially the breakthrough of natural language processing (NLP) technology, the application of intelligent assistants in various industries has gradually deepened. Traditional data analysis and report generation methods often rely on manual operations, which are inefficient and prone to human errors. In order to improve the data decision-making efficiency of enterprises and individuals, more and more intelligent assistants are beginning to be used in scenarios such as automatic generation of data reports, data analysis, and processing of business queries. Conversational data reporting intelligent assistants, especially those that can interact with users in natural language, have become an important tool in the field of data analysis. It can understand the natural language queries entered by users, quickly extract relevant information from large amounts of data, and present it to users in an easy-to-understand form.

[0003] However, in traditional conversational intelligent assistants, how to accurately understand complex user queries and perform efficient data retrieval is still a difficult problem that needs to be solved. For application scenarios with a large knowledge base or database background, how to make conversational intelligent assistants not only support free question and answer, but also accurately extract information from knowledge documents and databases according to the context of user queries and generate accurate reports has always been an important direction in technology research and development. Summary of the invention

[0004] The purpose of this application is to provide an intelligent data report generation system and generation method based on a large language model, which can improve the efficiency and accuracy of data report production and query.

[0005] To achieve the above objectives, this application provides the following solutions: In a first aspect, the present application provides an intelligent data report generation system based on a large language model, comprising: Enterprise knowledge base, data storage module, natural language processing module, knowledge document retrieval module, database query module and report generation module.

[0006] The natural language processing module is used to analyze user needs and determine whether the data source of the user needs is the data storage module or the enterprise knowledge base module; when the data source is the data storage module, a database query request is issued to the database query module; when the data source is the enterprise knowledge base module, a knowledge base retrieval request is issued to the knowledge document retrieval module.

[0007] The database query module generates an SQL statement to query the structured data in the data storage module in response to the database query request of the natural language processing module, and obtains the query result. The database query module uses the Text_to_sql technology to query the database.

[0008] The knowledge document retrieval module retrieves the unstructured data in the enterprise knowledge base module in response to the knowledge base retrieval request of the natural language processing module. The knowledge document retrieval module uses the Retrieval-Augmented Generation technology to retrieve the knowledge base.

[0009] The report generation module is used to generate a visual chart according to the user's needs and the query result.

[0010] The data storage module is used to store the structured data and the generated visual chart.

[0011] The enterprise knowledge base module is used to store the vectorized unstructured knowledge documents.

[0012] Optionally, the natural language processing module determines the type of user needs through a large language model: if the type of user needs contains a report generation instruction, the query result is input into the report generation module; otherwise, the data is directly returned.

[0013] Optionally, the way for the enterprise knowledge base module to store unstructured knowledge documents is as follows: Use the BERT model to vectorize the knowledge documents.

[0014] Convert the preprocessed text content into high-dimensional vectors and store them in the database table of the enterprise knowledge base module.

[0015] In a second aspect, the present application provides an intelligent data report generation method based on an intelligent data report generation system based on a large language model, including: Obtain the user's needs.

[0016] Based on the user's needs processed by the natural language processing module, determine the data source queried in the user's needs. The data source is the data storage module or the enterprise knowledge base.

[0017] If the data queried in the user's needs is in the data storage module, enter the database query module to query and obtain the query result.

[0018] If the queried data is in the enterprise knowledge base, enter the knowledge document retrieval module to query and obtain the query result.

[0019] Based on the query result and the user's needs, generate a visual report or directly return the data based on the report generation module.

[0020] Optionally, before obtaining user requirements, it further includes: Configuring the database through a large language model, and inputting the database structure information text as a prompt into the large language model.

[0021] Optionally, configuring the database through a large language model, and inputting the database structure information text as a prompt into the model, specifically including: Writing the database structure information text corresponding to the enterprise database.

[0022] Uploading the database structure information text as a prompt to the large language model; the large language model queries the database tables based on the prompt.

[0023] Optionally, if the data queried in the user requirements is in the data storage module, it enters the database query module for querying to obtain the query result, specifically including: If the data queried in the user requirements is in the data storage module, use the Text_to_sql technology to enter the database query module for querying.

[0024] Optionally, if the queried data is in the enterprise knowledge base, it enters the knowledge document retrieval module for querying to obtain the query result, specifically including: Vectorize the user requirements using an embedding model to obtain a vectorized representation of the user requirements.

[0025] Calculate the cosine similarity between the vectorized representation of the user requirements and each knowledge vector in the enterprise knowledge base to obtain several vector similarity values.

[0026] Use the knowledge vector corresponding to the maximum value among the vector similarity values as a prompt to input into the large language model.

[0027] Based on the user requirements and the prompt, obtain the query result based on the large language model.

[0028] Optionally, the calculation formula for cosine similarity is: .

[0029] Where a represents the user's question vector, b represents the knowledge vector of the enterprise knowledge base, and s represents the vector similarity value.

[0030] Optionally, generating a visual report based on the report generation module, specifically including: Combining the user requirements and the query data as a prompt to input into the large language model.

[0031] The model generates an ECharts script and renders it as a chart.

[0032] According to the specific embodiments provided in this application, the following technical effects are disclosed in this application: This application provides an intelligent data report generation system and a generation method based on a large language model. In this system, first, a knowledge document retrieval module based on the RAG (Retrieval-Augmented Generation) technology is used to convert the uploaded enterprise knowledge documents into vectors for storage, and calculate the document similarity in a vectorized manner, and extract relevant documents therefrom to provide support for generating accurate answers; secondly, based on the Text_to_SQL technology, the system can automatically generate SQL query statements according to the natural language questions input by the user, and efficiently extract relevant data from the database; finally, through the integrated report generation technology, the system generates an interactive data report according to the user requirements and the SQL query results, realizing automated and intelligent report production. This system allows users to query and generate reports through natural language, providing intuitive, simple, and accurate intelligent data services, greatly improving the efficiency and accuracy of data report production and query. Brief Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 It is a schematic diagram of the functional modules of an intelligent data report generation system based on a large language model provided in an embodiment of this application.

[0035] Figure 2 It is a schematic diagram of the process of an intelligent data report generation method based on a large language model provided in an embodiment of this application. Detailed Embodiments

[0036] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0037] Currently, many existing conversational intelligent assistants mainly rely on preset rules, templates, or simple retrieval algorithms to answer user queries. However, these technologies have certain limitations.

[0038] Insufficient knowledge retrieval: For queries involving a large amount of domain-specific knowledge or external document data, existing intelligent assistants often rely on simple keyword matching and lack flexible context understanding and in-depth reasoning.

[0039] Complexity of database queries: Traditional database queries usually require users to have a certain level of SQL query ability or rely on developers to perform complex query customization. Most existing intelligent assistants only support basic retrieval functions in automated database queries and lack flexible natural language to SQL capabilities.

[0040] Challenges in natural language understanding: The query habits, query contexts, and semantic complexities of different users can vary greatly. Existing technologies have not fully addressed the challenge of accurately mapping natural language to database queries or complex document retrievals.

[0041] Therefore, how to combine the latest natural language processing technologies, especially by introducing Retrieval-Augmented Generation (RAG) and Text-to-SQL technologies, to enhance the capabilities of conversational intelligent assistants in knowledge document retrieval and database queries has become a hot topic in the current innovation of intelligent assistant technologies.

[0042] The purpose of this application is to provide an intelligent data report generation system and method based on a large language model, which can improve the efficiency and accuracy of data report production and query.

[0043] To make the above objects, features, and advantages of this application more obvious and understandable, the following further details this application in conjunction with the accompanying drawings and specific embodiments.

[0044] Embodiment 1 As Figure 1 shown, this embodiment provides an intelligent data report generation system based on a large language model, including: An enterprise knowledge base, a data storage module, a natural language processing module, a knowledge document retrieval module, a database query module, and a report generation module.

[0045] The natural language processing module is used to parse the user's requirements and determine whether the data source of the user's requirements is the data storage module or the enterprise knowledge base module; when the data source is the data storage module, it sends a database query request to the database query module; when the data source is the enterprise knowledge base module, it sends a knowledge base retrieval request to the knowledge document retrieval module.

[0046] The database query module generates an SQL statement to query the structured data in the data storage module in response to the database query request from the natural language processing module; the database query module uses the Text_to_sql technology to query the database.

[0047] The knowledge document retrieval module retrieves the unstructured data in the enterprise knowledge base module in response to the knowledge base retrieval request from the natural language processing module; the knowledge document retrieval module uses the Retrieval-Augmented Generation technology to retrieve the knowledge base.

[0048] The report generation module is used to generate visual charts according to user requirements and query results.

[0049] The data storage module is used to store structured data and the generated visual charts.

[0050] The enterprise knowledge base module stores the vectorized unstructured knowledge documents.

[0051] Specifically, the report generation module is used to generate various reports according to user requirements, such as line charts, bar charts, pie charts, etc. For example: after the user inputs the data report requirement for generating a line chart, the report generation module can generate the corresponding line chart according to the data generated by the database query module and the user's requirements.

[0052] Specifically, the natural language processing module is used to process the user's requirements, analyze whether the data queried by the user is in the data storage module or the enterprise knowledge base. If it is in the data storage module, it enters the database query module to query the enterprise database. If the user also has a report requirement, after querying the data, it enters the report generation module to draw the report. If it is the enterprise knowledge base, it enters the knowledge document retrieval module to query the enterprise knowledge base information.

[0053] Specifically, the database query module is used to query the data in the enterprise database. After receiving the database query requirement from the natural language processing module, this module will generate the corresponding query statement to query the information in the data storage module.

[0054] Specifically, the data storage module is used to store the data in the database tables in the enterprise business system. This module provides structured data for the database query module.

[0055] Specifically, the knowledge document retrieval module is used to retrieve the data in the enterprise knowledge base. After receiving the enterprise knowledge base query requirement from the natural language processing module, this module will retrieve the enterprise knowledge base information.

[0056] Specifically, the enterprise knowledge base module is used to store data of enterprise knowledge documents. This module provides unstructured data for the knowledge document retrieval module. Among them, the BERT model is used to vectorize the knowledge documents; the content after text preprocessing is converted into high-dimensional vectors and stored in the database table of the enterprise knowledge base module.

[0057] Embodiment 2 As Figure 2 shown, the present application provides an intelligent data report generation method based on a large language model-based intelligent data report generation system, including: Step 201: Obtain user requirements.

[0058] Step 202: Based on the user requirements processed by the natural language processing module, judge the data source queried in the user requirements; the data source is the data storage module or the enterprise knowledge base.

[0059] Step 203: If the data queried in the user requirements is in the data storage module, enter the database query module to query and obtain the query result.

[0060] Step 204: If the queried data is in the enterprise knowledge base, enter the knowledge document retrieval module to query and obtain the query result.

[0061] Step 205: Based on the query result and user requirements, generate a visual report or directly return data based on the report generation module.

[0062] Among them, in some embodiments, before obtaining user requirements, it further includes: Configuring the database through the large language model, and taking the database structure information text as the prompt and inputting it into the large language model, specifically including: Step 1: Compile the database structure information text corresponding to the enterprise database; that is, prepare data in the enterprise database, which is presented in the form of a structured data table and can be queried by sql statements.

[0063] Step 2: Upload the database structure information text as the prompt to the large language model; the large language model queries the database table based on the prompt.

[0064] In some embodiments, when executing step 202, it can be specifically as follows: Analyze whether the data queried by the user is in the data storage module or the enterprise knowledge base. If it is in the data storage module, enter the database query module to query the enterprise database. If the user also has a report requirement, enter the report generation module to draw the report after querying the data. If it is the enterprise knowledge base, enter the knowledge document retrieval module to query the enterprise knowledge base information.

[0065] In some embodiments, when performing step 203, it can be specifically as follows: Step 1: The user puts forward a requirement for database query and uploads a large language model.

[0066] Step 2: The large language model generates corresponding SQL statements according to the database structure information text provided by the data storage module and the user's requirements.

[0067] Step 3: The system queries corresponding data from the database according to the SQL statements generated by the large language model.

[0068] Step 4: The large language model makes different responses according to the user's requirements. If there is no report requirement, the front end presents it in the form of a table. If there is a report requirement, it enters the report generation module to generate a report.

[0069] In some embodiments, when performing step 204, it can be specifically as follows: The user requirements are vectorized using an embedding model to obtain a vectorized representation of the user requirements.

[0070] The cosine similarity is calculated between the vectorized representation of the user requirements and each knowledge vector in the enterprise knowledge base to obtain several vector similarity values.

[0071] The knowledge vector corresponding to the maximum value among the vector similarity values is used as a prompt and input into the large language model.

[0072] Based on the user requirements and the prompt, a query result is obtained based on the large language model.

[0073] Among them, the calculation formula for cosine similarity is: .

[0074] Among them, a represents the user's question vector, b represents the knowledge vector in the enterprise knowledge base, and s represents the vector similarity value.

[0075] Among them, the user requirements are vectorized using an embedding model to obtain a vectorized representation of the user requirements, which can be specifically as follows: The knowledge documents are vectorized using an embedding model and stored in a database table.

[0076] The detailed process of vectorization by the embedding model is: 1. Text preprocessing: Before vectorizing the user question, the input text needs to be preprocessed to ensure that the text format is suitable for input into the embedding model. Common preprocessing steps include: Word segmentation: Word segmentation is necessary for text, especially for languages ​​without space separators such as Chinese.

[0077] Remove stop words: Remove some common words without actual semantics, such as "的", "是", etc.

[0078] Lowercase: For languages ​​such as English, all letters are converted to lowercase for uniform processing.

[0079] Remove special symbols and punctuation: Clean the text of special symbols and punctuation marks unless they have a specific role to play in the task.

[0080] These steps can help the embedding model effectively extract important semantic information from text.

[0081] 2. Select the embedding model: Choosing a suitable embedding model is the key to vectorization. We use the BERT model, which can provide richer and more accurate word vector representation in semantics.

[0082] 3. Convert text into vector: Each piece of knowledge in the enterprise knowledge base is input into the selected embedding model, and the model converts the text into a high-dimensional vector based on the trained weights. Specifically: the text is passed through the input layer of the model, and the context is encoded through multiple layers of Transformers, and finally a fixed-dimensional vector representation is generated.

[0083] In some embodiments, when executing step 204, the specific steps may be as follows: Combine user needs and query data into prompts and input them into a large language model.

[0084] The model generates ECharts scripts and renders them as charts.

[0085] Specifically, the user inputs the report requirements and obtains the corresponding data in the database query module. The user requirements and data are uploaded to the big language model as prompts. The big language model generates the corresponding echarts script based on the prompt, and the script can be used to generate the corresponding data report.

[0086] This application also provides usage scenarios of the intelligent data report generation system based on the large language model: For example, a large enterprise has a large amount of business data and internal knowledge documents. To better perform data analysis and decision support, the enterprise decides to introduce an intelligent data report generation system based on large language models.

[0087] I. Daily report generation.

[0088] Salespersons need to generate sales reports regularly to monitor sales progress and performance. They input requirements in natural language, such as "total sales and monthly distribution in the past three months". After the system analyzes the requirements, it automatically determines that the data source is the data storage module and sends a query request to the database query module. The database query module uses Text_to_sql technology to generate SQL statements, query structured data, and pass the results to the report generation module. The report generation module generates bar charts or line charts based on the query results to clearly display the sales data.

[0089] II. Complex queries and knowledge integration.

[0090] The marketing department hopes to analyze the market dynamics of competitors and the market share of the enterprise. They input the requirement: "changes in the market share of competitor A in the past year and comparison with our market share". After the system analyzes the requirements, it determines that the data sources include both the data storage module (market share data) and the enterprise knowledge base module (market dynamics of competitors). The system sends requests to the database query module and the knowledge document retrieval module respectively. The database query module queries structured data, and the knowledge document retrieval module uses Retrieval-Augmented Generation technology to retrieve unstructured knowledge documents. The report generation module integrates the query results and generates comparison charts to help the marketing department quickly understand the market dynamics and competitive situation.

[0091] III. Dynamic reports and decision support.

[0092] Senior management needs to understand the enterprise's operating conditions in real time to make quick decisions. They input the requirement: "current financial status of the enterprise, including revenue, cost, and profit". The system analyzes the requirements in real time, queries the latest data, and generates a dynamic report. The report includes visual charts and data tables, intuitively showing the financial status. Managers can make decisions to adjust strategies or optimize operations quickly based on the report data.

[0093] By introducing the intelligent data report generation system, the enterprise has achieved efficient utilization of data and rapid integration of knowledge, providing strong support for business decisions.

[0094] In summary, this application has the following technical effects: 1) Improved query efficiency.

[0095] This application generates SQL query statements through a large language model. Combining with the database structure information text, the system can quickly generate accurate SQL queries without the need for manual writing of complex query logic. In addition, the combination of the database query module and natural language processing enables users to perform queries through simple natural language requirements, greatly reducing the query time and operation complexity.

[0096] Specifically, the Text-to-SQL technology of the database query module can automatically convert natural language requirements into SQL query statements, avoiding manual intervention and saving query time.

[0097] By uploading the structure information of the database in advance as a prompt to the large language model, the system can directly generate targeted SQL queries based on this structure information, improving the accuracy and efficiency of the queries.

[0098] 2) Improved the accuracy and quality of data analysis.

[0099] In terms of data analysis and report generation, this application uses a large language model to generate corresponding echarts scripts according to user requirements, thus ensuring that the presentation of reports and the interpretation of data meet the expectations and requirements of users. The natural language understanding ability of the large language model can effectively identify the report requirements of users, making report generation more accurate.

[0100] Specifically, the generation of echarts scripts in the report generation module: The large language model generates dynamic chart scripts according to user requirements, which can intelligently identify user report requirements, avoiding the inefficiency and errors of manually generating scripts.

[0101] The combination of natural language processing and data: The large language model understands user requirements and combines with the accurate data provided by the database query module to generate report results that meet user expectations.

[0102] 3) Supports efficient retrieval of enterprise knowledge bases.

[0103] Through the RAG (Retrieval-Augmented Generation) technology, this application enables the system to efficiently retrieve relevant information from the enterprise knowledge base. Using the embedding model to vectorize knowledge documents makes the retrieval process no longer rely on traditional keyword matching, but perform efficient matching based on semantic similarity, thereby improving the accuracy and relevance of knowledge retrieval.

[0104] Specifically, the Embedding model and RAG technology: By vectorizing enterprise documents, the system can quickly find the knowledge most relevant to user requirements from a large number of documents by calculating cosine similarity, thus providing accurate answers.

[0105] Preprocessing and Embedding of Vectorization Technology: Through text preprocessing and the adoption of efficient embedding models (such as BERT), the system can accurately capture the semantic information of the text, improving the accuracy of document retrieval.

[0106] 4) Flexible report display methods.

[0107] This application can generate reports in various display forms according to different needs, including tables and charts, meeting the needs of different users. For data that requires visual presentation, the system creates dynamic charts by generating chart scripts, while for simple data requirements, it can be quickly presented in tabular form.

[0108] Specifically, the generation of echarts scripts: Through the echarts scripts generated by the large language model, the system can automatically generate charts, thus providing a more rich and intuitive data display.

[0109] 5) Improved the intelligence and simplicity of user interaction.

[0110] Compared with the traditional manual query and report-making methods, in this application, users only need to put forward their needs in natural language, and the system can automatically generate corresponding query results or reports, greatly reducing the operation difficulty of users and improving the interaction efficiency at the same time. Through the natural language processing module in the system design, the interaction between users and the database or knowledge base becomes more intelligent.

[0111] Specifically, the natural language processing module: Through natural language processing technology, the system can intelligently parse the needs input by users, automatically decide whether to query the database or the knowledge base, and then generate corresponding queries or reports, reducing manual operations.

[0112] Combination of Text-to-SQL Technology and Natural Language Query: The natural language needs of users can be automatically converted into SQL query statements, simplifying the user operation process.

[0113] 6) Enhanced the intelligence and automatic update ability of the enterprise knowledge base.

[0114] Due to the adoption of the embedding model to vectorize and store knowledge documents in this application, the enterprise knowledge base can be dynamically updated according to the growing document data without manual intervention. As the enterprise knowledge volume increases, the system can continue to provide accurate retrieval services.

[0115] Specifically, vectorization and embedding technology: This technology can not only improve the retrieval accuracy, but also support the automatic update of the vector representation of documents as the enterprise knowledge base expands continuously, ensuring that the knowledge base can continuously provide accurate answers in a changing environment.

[0116] RAG Technology: By combining vectorized documents and retrieval techniques, RAG technology enables efficient retrieval of knowledge bases, automatically obtaining the most relevant knowledge from large-scale documents to ensure the timeliness and accuracy of information.

[0117] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0118] Specific examples are used in this article to elaborate on the principles and implementation methods of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, based on the idea of this application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. An intelligent data report generation system based on a large language model, characterized in that: The intelligent data report generation system includes: an enterprise knowledge base, a data storage module, a natural language processing module, a knowledge document retrieval module, a database query module and a report generation module; A natural language processing module is used to analyze user needs and determine whether the data source of the user needs is a data storage module or an enterprise knowledge base module; when the data source is the data storage module, a database query request is sent to the database query module; when the data source is the enterprise knowledge base module, a knowledge base search request is sent to the knowledge document search module; A database query module, in response to a database query request from a natural language processing module, generates an SQL statement to query the structured data of the data storage module and obtains a query result; the database query module uses a text_to_sql technology to query the database; A knowledge document retrieval module, in response to a knowledge base retrieval request from a natural language processing module, retrieves unstructured data from an enterprise knowledge base module; the knowledge document retrieval module uses a Retrieval-Augmented Generation technology to retrieve the knowledge base; Report generation module, used to generate visual charts based on user needs and query results; Data storage module, used to store structured data and generated visualization charts; The enterprise knowledge base module is used to store vectorized unstructured knowledge documents.

2. According to claim 1, the intelligent data report generation system based on the large language model is characterized in that: The natural language processing module determines the user demand type through a large language model: if the user demand type includes a report generation instruction, the query result is input into the report generation module; otherwise, the data is directly returned.

3. The intelligent data report generation system based on a large language model according to claim 1, characterized in that: The enterprise knowledge base module stores unstructured knowledge documents in the following manner: Use the BERT model to vectorize knowledge documents; The preprocessed text content is converted into high-dimensional vectors and stored in the database table of the enterprise knowledge base module.

4. An intelligent data report generation method based on an intelligent data report generation system based on a large language model according to any one of claims 1 to 3, characterized in that: The intelligent data report generation method comprises: Obtain user needs; Based on the user needs processed by the natural language processing module, determine the data source queried in the user needs; the data source is a data storage module or an enterprise knowledge base; If the data queried in the user's needs is in the data storage module, the query is performed in the database query module to obtain the query result; If the queried data is in the enterprise knowledge base, the knowledge document retrieval module is entered to perform the query and obtain the query result; Based on the query results and user needs, the report generation module generates visual reports or directly returns data.

5. The intelligent data report generation system based on a large language model according to claim 4 is characterized in that: Before obtaining user needs, it also includes: The database is configured through a large language model, and the database structure information text is input into the large language model as a prompt.

6. The intelligent data report generation system based on a large language model according to claim 5, characterized in that: Configure the database through a large language model and input the database structure information text as a prompt into the model, including: Write database structure information text corresponding to the enterprise database; The database structure information text is uploaded to the large language model as a prompt; and the large language model performs a database table query based on the prompt.

7. The intelligent data report generation system based on a large language model according to claim 4, characterized in that: If the data queried in the user's needs is in the data storage module, the query module is entered to perform the query to obtain the query result, which specifically includes: If the data queried in the user's needs is in the data storage module, the Text_to_sql technology is used to enter the database query module for query.

8. The intelligent data report generation system based on a large language model according to claim 4, characterized in that: If the queried data is in the enterprise knowledge base, the knowledge document retrieval module is entered for query to obtain the query results, including: Vectorize user needs using the embedding model to obtain a vectorized representation of user needs; Calculate the cosine similarity between the vectorized representation of user needs and each knowledge vector in the enterprise knowledge base to obtain several vector similarity values; The knowledge vector corresponding to the maximum value of the similarity of each vector is input into the large language model as a prompt; According to user needs and prompts, query results are obtained based on the large language model.

9. The intelligent data report generation system based on a large language model according to claim 8, characterized in that: The calculation formula for cosine similarity is: ; Among them, a represents the user's question vector, b represents the knowledge vector of the enterprise knowledge base, and s represents the value of vector similarity.

10. The intelligent data report generation system based on a large language model according to claim 4, characterized in that: Generate visual reports based on the report generation module, including: Combine user needs and query data into prompts and input them into a large language model; The model generates ECharts scripts and renders them as charts.

Citation Information

Patent Citations

  • Intelligent question answering method, system and equipment based on large language model and database

    CN118606348A

  • Enterprise analysis report generation method and system based on large language model

    CN118607480A

Cited By

  • Intelligent report generation method and system

    CN120893414A

  • Intelligent gas customer service response method

    CN121092754A