Report generation method, agent module, report generation system, program product and storage medium

By configuring a specific domain knowledge base and rule base, combining a large language model and an agent module, users can generate adaptive and accurate reports without operations, solving the complex operation of existing report generation solutions and improving user experience.

CN119962498APending Publication Date: 2025-05-09CHINA UNIONPAY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411474269.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-21
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing report generation solution has high requirements for user operations, making it difficult to achieve simple and accurate self-service report generation.

Method used

By configuring a domain-specific knowledge base and rule base, combining large language models and agent modules, reports are generated automatically, including information fusion, plug-in and unplugging of visual components and adaptive adjustment.

Benefits of technology

Users can generate accurate adaptive reports without complex operations, which improves the convenience and accuracy of report generation and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962498A_ABST
    Figure CN119962498A_ABST
Patent Text Reader

Abstract

The invention provides a report generation method which comprises the steps that in response to a user request, information associated with the user request is retrieved from a database, and the database is pre-configured to comprise a specific domain knowledge base and a rule base; the user request and the retrieved information are fused, first intermediate information is generated, and compared with the user request, the first intermediate information can express the user request more accurately; transmitting the first intermediate information to a large language model; processing the first intermediate information by the large language model to generate an executable statement; generating a new prompt word based on the executable statement; the big data model generates summary and report style recommendation according to the new cue word; and generating a report matched with the user request according to the user request, the summary and the report style recommendation, wherein the generated report is a report based on the information of the specific field. The invention further provides a system, a storage medium, a program product and the like corresponding to the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to data processing technology, and more specifically, to intelligent report generation technology. Background Art

[0002] Data reports are often used in data analysis. In order to generate appropriate and targeted data reports, the industry provides a variety of report generation methods.

[0003] The Chinese patent application with publication number CN103218448A provides a self-service report generation method. According to the patent application, after receiving the report attribute parameters, report style maintenance parameters, and report area maintenance parameters input by the user, a data table will be automatically created, and the result data will be obtained according to the report formula input by the user for the data cell in the data area, and finally the result data will be displayed according to the report style.

[0004] The Chinese patent application with publication number CN118052214A provides a method for automatically generating reports. In the solution of the patent application, initial data is obtained from a data source table, the initial data is calculated and processed to obtain a target data set corresponding to a data indicator; a report template is configured and report subjects are set; an association relationship between report subjects and data indicators is established; according to the association relationship, target data is selected from the target data set corresponding to the data indicator and imported into the report subject to obtain a report of the target data.

[0005] Although there are different report generation solutions, each solution has certain operation requirements for users, and it is necessary to provide an improved report generation solution. Summary of the invention

[0006] According to one aspect of the present application, a report generation method is provided to provide improvement at least in terms of ease of use.

[0007] According to one aspect of the present application, a report generation method is provided, which includes: in response to a user request, retrieving information associated with the user request from a database,

[0008] The database is pre-configured to include a domain-specific knowledge base and a rule base; the user request and the retrieved information are integrated to generate first intermediate information, which can more accurately express the user request than the user request; the first intermediate information is transmitted to a large language model; the large language model processes the first intermediate information to generate an executable statement; based on the executable statement, a new prompt word is generated; the big data model generates a summary and a report style recommendation based on the new prompt word; and a report matching the user request is generated based on the user request, the summary and the report style recommendation; wherein the generated report is a report based on the information in the specific domain.

[0009] According to another aspect of the present application, a report generation method is provided, which includes: receiving a user request; retrieving information associated with the user request from a database, wherein the database is pre-configured to include a specific domain knowledge base and a rule base; fusing the user request and the retrieved information to generate first intermediate information, and compared with the user request, the first intermediate information can more accurately express the user request; transmitting the first intermediate information to the large language model; based on the executable statement sent by the large language model, generating a new prompt word and sending it to the large language model via the new prompt word; and generating a report matching the user request according to the user request, summary and report style recommendation, wherein the summary and report style recommendation are generated by the large data model according to the new prompt word; wherein the generated report is a report based on the information of the specific domain.

[0010] In the provided report generation method, as an example or supplement, in response to a user request, information associated with the user request is retrieved from a database, including: forming the user request into multiple sub-questions, the sub-questions including the original question in the user request or the decomposed question of the original question, and the sub-questions also including new questions related to the user request; using multiple retrieval techniques to search in the specific field knowledge base to obtain knowledge for the multiple sub-questions from the knowledge base.

[0011] In the report generation method provided, as an example or supplement, the user request and the retrieved information are fused to generate first intermediate information, including: according to the user intention expressed in the user request and the semantic overlap in the knowledge, the knowledge used for the multiple sub-problems is processed to generate the first intermediate information which is more accurate than the user request and includes more knowledge, as a prompt word to be input into the large language model.

[0012] In the report generation method provided, as an example or supplement, new prompt words are generated based on the executable statement, including: sending the executable statement to the specific domain knowledge base to obtain execution results for subsequent data analysis; combining the execution results with the report rules in the specific domain knowledge base and the rule base to generate new prompt words.

[0013] In the report generation method provided, as an example or supplement, the specific domain knowledge base includes metadata, data lineage and user preferences. The method further includes when configuring the specific domain knowledge database: analyzing the data in the specific domain knowledge base through one or more methods of word frequency statistics, keyword extraction, and topic clustering; and constructing the analyzed data into structured data.

[0014] In the provided report generation method, as an example or supplement, generating a report matching the user request includes: using a visualization component required for generating the report from a component library in a pluggable manner.

[0015] In the report generation method provided, as an example or supplement, the method also includes: continuously making adaptive adjustments by collecting user feedback and / or capturing user intentions; and continuously acquiring new knowledge to improve the knowledge base.

[0016] According to another aspect of the present application, an intelligent agent module is also provided. The intelligent agent module is configured to: receive a user request; retrieve information associated with the user request from a database, wherein the database is pre-configured to include a specific domain knowledge base and a rule base; fuse the user request and the retrieved information to generate first intermediate information, which can more accurately express the user request than the user request; transmit the first intermediate information to the large language model; generate new prompt words based on the executable statements sent by the large language model and send them to the large language model via the new prompt words; and generate a report matching the user request according to the user request, summary and report style recommendation, wherein the summary and report style recommendation are generated by the large data model according to the new prompt words; wherein the generated report is a report based on the information of the specific domain.

[0017] The provided intelligent agent module may, optionally or additionally, be further configured to retrieve information associated with the user request from a database through the following process: forming the user request into a plurality of sub-problems, wherein the sub-problems include the original problem in the user request or the decomposed problem of the original problem, and the sub-problems also include new problems related to the user request; and using a variety of retrieval techniques to search in the domain-specific knowledge base to obtain knowledge for the plurality of sub-problems from the knowledge base.

[0018] The provided intelligent agent module may, optionally or additionally, be further configured to fuse the user request and the retrieved information to generate first intermediate information through the following process: based on the user intention expressed in the user request and the semantic overlap in the knowledge, the knowledge used for the multiple sub-problems is processed to deduplicate and generate the first intermediate information which is more accurate than the user request and includes more knowledge, as a prompt word to be input into the large language model.

[0019] The provided intelligent agent module, optionally or supplementally, the specific domain knowledge base includes metadata, data representing data lineage and data representing user preferences, and the intelligent agent module is further configured when configuring the specific domain knowledge database: analyzing the data in the specific domain knowledge base through one or more methods of word frequency statistics, keyword extraction, and topic clustering; and constructing the analyzed data into structured data.

[0020] The provided intelligent agent module can be optionally or supplementally configured to use the visualization components required for generating the report from a component library in a pluggable manner.

[0021] The provided intelligent agent module can be optionally or supplementally further configured to: continuously perform adaptive adjustments by collecting user feedback and / or capturing user intentions; and continuously acquire new knowledge to improve the knowledge base.

[0022] According to another aspect of the present application, a report generation system is also provided. The system includes: a database, including a domain-specific knowledge base and a rule base; a component library, including multiple visual components; an intelligent agent module, the intelligent agent module is connected to the database, the component library, and the large language data model; wherein the intelligent agent module adopts any one of the above-mentioned intelligent agent modules, or the intelligent agent module is configured to execute any one of the methods described herein; wherein the intelligent agent module obtains the required visual components from the component library when generating a report matching the user request according to the user request, summary, and report style recommendation.

[0023] A program product is also provided, which includes program instructions, and when the program instructions are executed, any one of the methods described above is implemented.

[0024] A non-transitory storage medium is also provided, which stores program instructions, and when the program instructions are executed, any one of the methods described above is implemented.

[0025] By executing the method according to the example of this application or adopting the system or intelligent agent module according to the example of this application, reports can be "adaptively" generated for non-technical users according to their needs. Users only need to input the needs of generating reports in a convenient way, such as voice, text or video. According to the scheme of this application, the needs will be analyzed and more accurate and comprehensive prompt words will be generated for processing by the large language model. Based on the feedback of the large language model, the intelligent agent module in the example of this application will further combine user needs, domain-specific knowledge bases and rule bases, and obtain the required visual components from the component library in a drag-and-drop manner, and finally run the report. Throughout the process, users can obtain adaptively generated visual reports without having to design report generation, which greatly improves the convenience of report generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The following will describe the embodiments of the present application in detail with reference to the accompanying drawings so that the present application can be more fully understood, wherein:

[0027] Figure 1 is a flow chart of a report generation method according to some embodiments of the present application;

[0028] Figure 2 is an executable application according to Figure 1 The schematic diagram of the method system structure shown;

[0029] Figure 3 A flow chart of the report generation method according to this embodiment is illustrated;

[0030] Figure 4 is a schematic diagram of the structure of an intelligent agent module according to an example of the present application;

[0031] Figure 5 It is a structural diagram of a report making system according to an example of this application;

[0032] Figure 6 It is a schematic diagram of the architecture of a report making system according to some examples of the present application;

[0033] Figure 7 It is a schematic diagram of the architecture of data processing implementation according to some examples of the present application. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the accompanying drawings to clearly and completely describe the implementation methods of the present application. It should be noted that the described implementation methods are only part of the implementation methods of the technical solution of the present application, not all. All other implementation methods obtained by ordinary technicians in this field based on the implementation methods recorded in the application documents without paying creative work are covered by the protection scope of the present application.

[0035] Figure 1 Flowchart of a report generation method according to some embodiments of the present application. In some examples, the method can be executed by an electronic device such as a computer, a tablet, etc. In other examples, the method is executed by an electronic device and another electronic device or a cloud device. The hardware environment for executing the present invention is diverse and is not limited to one or two hardware devices, which will be explained in conjunction with examples below.

[0036] In step S100, a user request is received. The user request may be input by the user through the interactive interface of the electronic device, and the input may be one or more of text input, voice input, and video input. To generate a data report, the user only needs to input the requirements in a simple and direct manner without performing more complicated operations.

[0037] In step S102, in response to the user request, information associated with the user request is retrieved from the database, wherein the database is pre-configured to include a domain-specific knowledge base and a rule base. Here, the database includes general knowledge, and is further configured to include a domain-specific knowledge base. The database also includes a rule base, which is also a rule base for a specific domain. The specific domain refers to the domain corresponding to the data required to generate a report. For example, the data for generating a report should be financial data, and the specific domain becomes the financial domain. The database is pre-configured with a financial domain knowledge base and a rule base for financial domain reports. Accordingly, in order to meet the needs of different reports, the pre-configured rule base includes rule models for multiple rules such as data filtering, aggregation, and sorting, and each model has corresponding attribute constraints, thereby constituting a rule base.

[0038] By way of example and not limitation, the domain-specific knowledge base includes metadata, data lineage, and user preferences. Metadata describes various attributes of data, such as data structure, data type, etc. Data lineage records information such as the source of data, conversion process, etc. User preferences record the user's personal settings and preferences, such as settings and preferences when using the method according to the example of this application or the system and agent model to be described below, such as report theme color, font size, layout style, data analysis method, etc.

[0039] According to some examples of the present application, when pre-configuring a specific domain knowledge base, the data in the specific domain database can be analyzed by one or more methods such as word frequency statistics, keyword extraction, and topic clustering, and the analyzed data can be constructed as structured data. For example, the intelligent agent performs text mining and conducts in-depth exploration of a large amount of unstructured text in various data sources such as documents and reports. Through word frequency statistics, keyword extraction, and topic clustering, deep information with use and reference value in the text is mined. Through operations such as named entity recognition, relationship extraction, and event extraction, the key knowledge points in this information (including deep information) are converted into structured data. Based on ontology and knowledge graph technology, a structured knowledge graph containing entities, relationships, and events is constructed. In addition, the embedding technology is used to convert unstructured text data into vector form and construct a vector library, which is convenient for subsequent semantic search and knowledge representation.

[0040] In practical applications, user expressions have personalized characteristics, which leads to the diversity and ambiguity of user input requests, including questions, etc. Therefore, it is necessary to optimize user requests, on the one hand to standardize them, and on the other hand to enrich their content, so that the intelligent agent can more quickly and accurately retrieve relevant information in the knowledge base and rule base. According to this embodiment of the present application, user requests will be optimized. For example, the user request is formed into multiple sub-problems, which may include the original questions in the user request, or these sub-problems include decomposed questions after decomposing the original questions; as an example, these sub-problems also include new questions related to the user request. A variety of retrieval techniques are used to search in the specific field knowledge base to obtain knowledge for multiple sub-problems from the knowledge base. For example, the optimization and enrichment process can be reflected in rewriting the user's question to generate multiple new questions or sub-questions, and at the same time, using multiple retrieval technologies including vector retrieval (capturing semantic similarity), full-text retrieval (based on keyword matching) and graph retrieval (using entities and relationships in the knowledge graph) to construct a multi-dimensional and comprehensive retrieval strategy to search in the knowledge base and / or rule base to improve recall and accuracy and cover the knowledge related to the original question as comprehensively as possible. Recall refers to the situation in which knowledge related to the user's request is retrieved from the database, and the retrieved knowledge includes data structure, samples, user preferences, etc.

[0041] As an example, in the method according to this example, when retrieving information associated with a user request from a database, a Retrieved-augmented Generation (RAG) model can be used. Simply put, the RAG model combines language models and information retrieval technology. When the model needs to generate text or answer questions, it can retrieve relevant information from a large document collection, and then use the retrieved information to guide text generation.

[0042] In step S104, the user request and the detected information are fused to generate first intermediate information. Compared with the user request, the first intermediate information can more accurately express the user request and the content it includes is richer and more comprehensive. Specifically, according to the user intention expressed in the user request and the semantic overlap in the retrieved knowledge, the knowledge used for multiple sub-problems is processed, such as removing overlaps and performing fusion processing such as sorting. Thus, the first intermediate information including more knowledge is formed, which is used as the prompt word to be input to the large language model. The first intermediate information generated here removes overlapping semantics and includes a more comprehensive knowledge representation, and the prompt words generated accordingly are of higher quality.

[0043] In one example of this application, in the knowledge fusion stage, technical means such as data alignment and conflict resolution can be used to compare and fuse the newly extracted knowledge with the existing knowledge to eliminate ambiguity and redundant information. To ensure the accuracy and reliability of knowledge, multiple methods such as expert systems, rule engines, manual review and user feedback can be used for verification.

[0044] In step S106, a large language model (LLM) processes the first intermediate information to generate executable statements. Specifically, the large language model generates a reasonable SQL (Structured Query Language) statement according to the prompt word.

[0045] In step S108, a new prompt word is generated based on the executable statement. Specifically, the SQL statement generated in step S106 is sent to the database for execution, thereby generating an execution result for subsequent data analysis and summary by the user. Further, a new prompt word is generated in combination with the execution result, the relevant knowledge searched in the specific knowledge base, and the relevant rules for generating the report searched.

[0046] According to the example of this application, the rule base in the database is intended to manage the rules related to report production, including report rules, API rules, etc., covering many forms of report production, such as table reports (complex Chinese-style reports, pivot table functions, etc.), graphic reports (line charts, bar charts and other statistical charts), U chat reports (illustrated briefing reports). The initial value of the rule is generally set manually based on business experience, and the content of the rule is maintained according to the actual effect during use through a combination of autonomous learning and manual confirmation and correction.

[0047] In step S110, the big data model generates a summary and report style recommendation based on the new prompt word.

[0048] In step S112, a report matching the user request is generated based on the user request, the summary generated in step S110, and the style recommendation. Specifically, in the process of generating the report, the data features, user intent, user preferences, and other information are cleaned, formatted, standardized, and other processing operations are performed to ensure the accuracy and consistency of the information, and a logical framework between the information is constructed for targeted fusion. For example, based on the relevance and importance of the information, weighted average, Bayesian inference, and other methods are used for fusion.

[0049] In addition, in the process of generating a report, the required visualization components can be selected from a component library including multiple visualization components. The visualization components in the component library are configured to be used in a pluggable manner. In some examples, the component library encapsulates open source components such as ECharts, Bootstrap, and D3 (Data-Driven Documents), and constructs visualization components for report presentation, such as ECharts chart components, Bootstrap layout style components, and D3 complex data visualization components. In the process of establishing a report, according to the report type, report style, and result data set in the rule library, combined with actual business needs, different visualization components are plugged in to make a report. For example, when a user needs a visualization component, the visualization component can be directly used by dragging and dropping. If the component is not needed during the report making process, it can be removed by dragging and dropping. In this way, the report making process does not require cumbersome custom development, and the intuitive display of data can be achieved. The pluggable components, or the implementation of the component library, are based on the "pluggable use" adaptation of the visualization chart, which improves the flexibility of data visualization, while also ensuring that the chart component can adapt to different data sources and display environments. In each example of the present application, the visualization components of the component library can be developed in advance. These visualization components are classified and organized to form a visualization report component library.

[0050] In summary, the most appropriate presentation method is intelligently selected according to the needs of the user, that is, the most appropriate visual components (such as tables, charts, texts, etc.) are assembled to achieve "adaptive" report generation for non-technical personnel, that is, non-technical users do not need to operate the report design software to generate visual reports themselves, but after the user inputs the user request, the electronic device executing the method performs a series of operations by itself to adaptively generate visual reports. Finally, the generated results can be presented to the user so that the user can quickly understand and use the information.

[0051] According to the example of the present application, the method also includes continuously collecting user feedback and / or capturing user intent for adaptive adjustment. The method may also include continuously acquiring new knowledge to improve the knowledge base. For example, regularly updating and optimizing the content and structure of the knowledge base, continuously collecting information such as user feedback and usage data, capturing user intent and performing adaptive adjustment and optimization, continuously integrating new knowledge, and continuously improving the knowledge base.

[0052] Figure 2 is an executable application according to Figure 1 The schematic diagram of the method system structure is shown in FIG. Figure 2 As shown, the system includes an electronic device 1 and an electronic device 2, and the electronic device 1 is communicatively connected with the electronic device 2. The electronic device 1 is, for example, any one of electronic devices such as a smart phone, a tablet, a wearable electronic device, a laptop computer, a desktop computer, etc. Similar to the electronic device 1, the electronic device 2 can be one of a laptop computer, a desktop computer, a tablet, a smart phone, and a wearable electronic device.

[0053] The user inputs a user request through the electronic device 1, for example, through text input, voice input or video input through the interactive interface of the electronic device 1. The user request is transmitted to the electronic device 2. The electronic device 2 executes steps S102 to S112 to generate a report. Further, the electronic device 2 can display the report or transmit the generated report to the electronic device 1 for display. The database 3 is communicatively connected with the electronic device 2 so that the electronic device 2 can configure, access, store and other operations on the data in the database during the execution of the method according to the example of the present application. In some cases, the system may also include a cloud. For example, when the large language model is not set in the electronic device 2, the electronic device 2 accesses the cloud during the execution of the method according to the example of the present application to interact with the large language model.

[0054] In other cases, executing Figure 1 The system of the method shown may include only the electronic device 2 but not the electronic device 1. In this example, the user inputs the user request through the interactive interface of the electronic device 2.

[0055] According to the present application, another report generation method is also provided, which can be, for example, Figure 2 Executed by electronic device 2 in. Figure 3 FIG. 4 is a flowchart of the report generation method according to this embodiment. Figure 3 As shown, in step S300, a user request is received. The user request is transmitted through the electronic device 2 ( Figure 2 ) is received by an interactive interface, which may be a text interactive interface, an audio input port, or a video capture component. The user request includes the user's intention, that is, the information that the user wants to generate a report. In step S302, information associated with the user request is retrieved from a database, wherein a specific domain knowledge base including specific domain knowledge and a specific domain rule base are pre-configured in the database. The rule base is a rule base for rules for generating reports. Information associated with the user request is retrieved from the database, and combined with the above Figure 1 The described process is similar and will not be described in detail. In step S304, the user request and the retrieved information are merged to generate the first intermediate information. Compared with the user request, the first intermediate information can express the user request more accurately and contains richer and more complete content. In step S306, the first intermediate information is transmitted to the large language model. The large prediction model can be set on the electronic device 2, or it can be set on an electronic device other than the electronic device 2 or set in the cloud. The large language model will generate an executable statement based on the first intermediate information, or the prompt word constructed by the first intermediate information, and return the generated executable statement to the electronic device 2. In step S308, based on the executable statement, a new prompt word is generated and the new prompt word is sent to the large language model. Specifically, the executable statement generated in step S306 is sent to the database for execution, thereby generating the execution result of the user's subsequent data analysis and summary. Further, in combination with the execution result, the relevant knowledge searched in the specific knowledge base, and the relevant rules for generating the report searched, a new prompt word is generated. After receiving the new prompt word, the large language model generates a summary and report style recommendation. The summary and report style recommendation are sent to the electronic device 2, which generates a report matching the user request according to the user request, summary and report style recommendation, as shown in step S310. In the process of generating the report, the required visualization components can be obtained from the preset component library.

[0056] In this example, step S302, step S304, step S308 and step S310 are respectively combined with Figure 1 Step S102, step S104, step S108, and step S112 in the method described are the same or similar, therefore, in combination with Figure 1 When describing the embodiments of the present application, examples related to these steps are applicable to Figure 3 For the sake of brevity, the examples are not repeated here.

[0057] Figure 1 Shown and Figure 3 The report generation method shown can be executed by, for example, electronic device 1 and / or electronic device 2. For example, any example of the report generation method described herein is implemented as a software program through a programming language so as to be executed by an electronic device.

[0058] According to the report generation method of the example of the present application, a program language can be implemented as an intelligent agent model for generating reports, thereby forming an intelligent agent for generating reports. For the intelligent agent model, the intelligent agent model can be continuously optimized by continuously inputting data related to reports in specific fields and data such as generation requirements. In addition, during the use of the intelligent agent model, the model can also continuously collect user feedback for continuous optimization.

[0059] According to an example of the present application, an intelligent agent module is also provided. The intelligent agent module is configured to be able to execute according to Figure 3 In another example, the agent module is configured to perform Figure 1 The agent module can be based on a programming language Figure 3 The software module implemented by the described method flow, or based on Figure 1 The agent module can also be a module implemented by combining software and hardware, for example, a module implemented by a programming language. Figure 1 The illustrated method is incorporated into a component such as a processor to form an agent module.

[0060] Figure 4 4 is a schematic diagram of the structure of an intelligent agent module according to some examples of the present application. The intelligent agent module includes program instructions 40, which are programmed to achieve when executed Figure 3 Each example of the method described above can be implemented Figure 1 The program instructions may be stored in the memory 401 so as to be executed by the processor 402. As an example, the agent module may be a software module constructed by the program instructions 40, which is portable and loadable to a hardware device, such as Figure 2 More specifically, according to the example of the present application, the software module implemented by the program instruction 40 is an intelligent agent model for implementing the method according to the example of the present application.

[0061] In another example, the agent module includes program instructions 40 and a memory 401 storing the program instructions 40. Alternatively, the agent module includes the memory 401, a processor 402, and the program instructions 40 stored in the memory 401.

[0062] Compared with the traditional way of making reports, by executing the report generation method according to the example of this application, or by using the intelligent agent module according to the example of this application, the user who makes the report does not need to have the complex operation skills of database query and report design software, so it is more convenient. For example, in the report making scheme of the example of this application, the knowledge and rule base of the corresponding field of the report are provided, and the big data model is introduced. In this way, the user does not need to search for data by himself, but only needs to briefly describe his needs. After receiving the user's needs, according to the scheme of this application, for example, the intelligent body will automatically search for knowledge related to the user's needs from a specific database and rule base, and the search process can be based on the RAG model. The search process does not require user participation, but is carried out automatically. After the search is completed, according to the scheme of the example of this application, the user does not need to operate the report design software by himself, but based on the searched knowledge, the user request is converted into a prompt word including richer and more complete information, and the prompt word is input into the large language model, and the executable result is generated by the large language model. Because the knowledge of the large language model in the general field has certain limitations, in the example of this application, the knowledge base and rule base of the specific field help the large language model to more accurately generate results related to the generation of reports when used. In addition, the use of RAG can also effectively reduce the hallucination phenomenon of large language models. In addition, when generating a data report for display, a required component can be selected from the visualization components to form a visualization report.

[0063] Figure 5 Schematic diagram of the structure of the report making system according to the example of this application. Figure 5 As shown, the system includes a database 50 , a component library 54 , an intelligent agent module 52 , and an interface 57 .

[0064] In the example of the present application, the database 50 is configured to include a domain-specific knowledge base and a rule base. The domain-specific knowledge base may include metadata, data lineage, and user preferences. Metadata describes various attributes of the data, such as data structure, data type, etc. Data lineage records information such as the source of the data and the conversion process. User preferences record the user's personal settings and preferences, such as the settings and preferences when using the system and intelligent agents, such as report theme color, font size, layout style, data analysis methods, etc. As an example, when pre-configuring the domain-specific knowledge base in the database 50, the knowledge of the specific domain is analyzed by one or more methods such as word frequency statistics, keyword extraction, and topic clustering, and the analyzed data is constructed as structured data for database storage and search. The above is combined with Figure 1 When introducing the method embodiment according to the present application, the description of the database is applicable to this example and will not be described in detail.

[0065] In this example, the agent module 52 is configured to be connected to the large language model, the database, and the component library. The communication connection can be a line connection or a wireless communication connection, as long as the two connected parties can transmit signals. Any of the agent modules described above can be used. Or the agent module 52 is configured to execute according to Figure 3 Alternatively, the agent module 52 is configured to execute Figure 1 As an example, the intelligent agent module 52 in this example can call the required functions, applications, etc. through the interface 57 during the report generation process. Of course, the intelligent agent module 52 can also call other functions, applications, etc. through the interface 57, and this application does not limit this.

[0066] The component library 44 includes a variety of visualization components. Specifically, the component library 54 is configured to encapsulate open source components such as ECharts, Bootstrap, D3, etc., and construct visualization components required for making reports, such as ECharts chart components, Bootstrap layout style components, D3 complex data visualization components, etc. The component library is constructed as the components it includes, which can be used in a plug-in manner. For example, when generating a report, the intelligent body module 52 can obtain the required visualization components from the component library 44 by dragging and dropping, and if the visualization component is not needed, it can also be removed by dragging and dropping.

[0067] Figure 6 Schematic diagram of the architecture of the report making system according to some examples of this application. Figure 6 As shown, the architecture of the report making system mainly includes four parts, namely, a knowledge base and rule base 60, a pluggable component layer 62, an adaptive processing layer 64 and an interface layer 66.

[0068] The knowledge base in the knowledge base and rule base 60 can be the specific domain knowledge base described above, which is mainly used to store and manage metadata, data lineage and user preferences. More specifically, metadata describes various attributes of data, such as data structure, data type, etc.; data lineage records information such as the source and conversion process of data; user preferences record personal settings and preferences of users during the use of the system, such as report theme color, font size, layout style, data analysis method, etc. In the knowledge base configuration stage, including in the subsequent maintenance, update and optimization stages, intelligent agents can be used to perform text mining to conduct in-depth exploration of a large amount of unstructured text in various data sources such as documents and reports. Through means such as word frequency statistics, keyword extraction and topic clustering, the valuable information in the text is revealed. Through operations such as named entity recognition, relationship extraction and event extraction, the key knowledge points in the text are converted into structured data.

[0069] When it comes to knowledge fusion, the intelligent agent can use technical means such as data alignment and conflict resolution to compare and fuse the newly extracted knowledge with the existing knowledge to eliminate ambiguity and redundant information. In order to ensure the accuracy and reliability of knowledge, multiple methods such as expert systems, rule engines, manual review and user feedback can be used for verification. Based on ontology and knowledge graph technology, a structured knowledge graph containing entities, relationships and events is constructed.

[0070] In addition, embedding technology is used to convert unstructured text data into vector form and construct a vector library to facilitate subsequent semantic search and knowledge representation. In order to adapt to business development and changes in demand, the knowledge base can be regularly updated and optimized in terms of content and structure. At the same time, the intelligent agent continuously collects user feedback and usage data, captures user intentions, and makes adaptive adjustments and optimizations, continuously integrates new knowledge, and continuously improves the knowledge base.

[0071] The rule base in the knowledge base and rule base 60 is also configured for a specific field and includes rules for generating reports. Specifically, the rule base manages rules related to report production, including report rules, API rules, etc., covering table reports (complex Chinese-style reports, pivot table functions, etc.), graphic reports (statistical charts such as line charts and bar charts), U chat reports (briefing-style reports with pictures and texts), etc. The initial value of the rule is generally set manually based on business experience. During use, the content of the rule is maintained in a combination of autonomous learning and manual confirmation and correction based on the actual effect.

[0072] By using the knowledge base and rule base, the intelligent agent can better understand and analyze data and make decisions that better match user needs at each stage. For example, in the stage of generating prompt words, more accurate and comprehensive prompt words can be generated because more comprehensive information is obtained from the knowledge base, which helps the large language model to make better processing based on the prompt words. For example, in the stage of generating reports, the rules for report production can be obtained from the rule base. By combining the information in the knowledge base and user needs, better decisions can be made on what kind of visual report to produce that better meets user needs.

[0073] The pluggable component layer 62 encapsulates open source components such as ECharts, Bootstrap, and D3, and constructs data visualization components required for report production, such as ECharts chart components, Bootstrap layout style components, and D3 complex data visualization components. In the process of building reports, users can complete report production by plugging in different chart components and report components in the Report Plugin layer according to the report type, report style, and result data set in the rule library, combined with actual business needs, without tedious custom development, to achieve intuitive display of data. The pluggable component layer can achieve "pluggable use" adaptation based on visual charts, improve the flexibility of data visualization, and ensure that chart components can adapt to different data sources and display environments.

[0074] In this example, the adaptive processing layer 64 includes a large language model LLM and an intelligent agent module according to the example of this application. In short, the intelligent agent module may include four components: storage, planning, tool calling and execution.

[0075] The intelligent agent module stores instant information related to the current task. This information is usually short-lived and changeable, but closely related to the current task or activity. By setting the length of the context window, the instant information related to the current task is retained, which is very critical for processing real-time tasks, making instant reasoning and decisions. The intelligent agent module also stores long-term information such as RAG mechanisms, stores factual knowledge (such as historical information, formula principles, etc.), and stores procedural knowledge (such as skills, strategies, etc.), thereby supporting more complex cognitive tasks. The intelligent agent module here is a module that includes software and hardware, and its memory stores the information mentioned above. However, in actual application, when the adaptive processing layer 64 is configured to an electronic device, the storage component of the electronic device can be used for storage. In addition, some data and information that need to be stored can also be stored in the cloud or other devices, and the intelligent agent module communicates to obtain it.

[0076] The planning module is responsible for both pre-planning and post-improvement. Pre-planning may include breaking down large tasks into smaller, more manageable subtasks. Based on the nature and priority of the subtasks, the planning module allocates resources appropriately, determines the path from the current state to the target state, considers possible obstacles and risks, selects the optimal action plan, and predicts the expected results during task execution so that necessary adjustments can be made during execution.

[0077] Pre-planning not only reduces the complexity of the task, but also improves the efficiency of task processing. Post-improvement refers to self-criticism and reflection on past actions. By analyzing errors and deficiencies, the agent self-learns and modifies, generates new knowledge and experience, and enriches the report generation rule base, from which it learns and improves future action strategies while continuously improving its decision-making ability and task completion quality.

[0078] The tool call part includes the use of external tools and resources in the process of completing tasks, such as database query, file operation, and interface call. This part also includes the use of visual components in the component library. In this way, the plug-in design of the component library realizes the high flexibility and scalability of the use of components in the report generation process, and it is also easy to quickly access different external functions according to business needs to meet different business needs.

[0079] The action part refers to the execution process, such as executing specific actions according to the strategies and plans planned in the early stage, querying the stored information to support the decision-making and execution process of the intelligent agent, and calling external tools and resources through the tool usage module to complete the task.

[0080] As an example, the interface layer 66 can provide five toolkits: agent (Report Agent) development tool (SDK), Uchat development tool (Uchat SDK), file development tool (File SDK), mail development tool (Mail SDK), and browser development tool (Browser SDK). Agent SDK supports the input of user demand description to achieve effective communication with agents; Uchat SDK supports mobile Uchat push and viewing. It uses rich text format dynamic splitting technology to achieve dynamic splitting of Uchat report data, disassembles the content according to the style defined in the rule library, attaches the text to the style, starts from the text, calculates the length of the text and the overall style, and reasonably allocates according to the Uchat length limit rules (from the rule library), reasonably plans the start and end positions of characters, and continuously loops to judge so that the final encapsulation remains segmented and sentence-by-sentence, ensuring that the generated Uchat report conforms to the user's reading habits. File SDK and mail SDK provide corresponding interfaces to generate corresponding reports, supporting multiple file formats such as csv, txt, pdf, png, etc. Users can download or push through emails; Browser SDK supports web reports, allowing users to log in to the front-end page directly through the browser to view the report.

[0081] Figure 7 This is a schematic diagram of the architecture of data processing implementation according to some examples of this application. The architecture mainly includes a rule base, a data synchronization layer, a data adaptation layer and a resource management layer.

[0082] The rule base manages and generates rules for data sets, including visualization (View) rules, resource (Source) rules, data (Data) rules, etc. The initial value of the rule is generally set manually based on business experience. During use, the content of the rule is maintained through a combination of self-learning and manual confirmation and correction based on actual results.

[0083] For the data synchronization layer, according to business rules, synchronization, extraction and conversion of massive data are realized, and finally the data is loaded into TiDB to realize efficient retrieval and calculation of large-scale data. In the process of data synchronization, different synchronization strategies can be adopted for real-time transaction flow data and offline historical data. Offline data implements specific data synchronization and processing tasks by encapsulating datax, and manages a large number of configurable parameters by encapsulating datax-web, such as routing strategy, blocking processing strategy, task failure retry time and number, etc., to realize task scheduling, log viewing and other functions. Scheduled data synchronization tasks (such as scheduled reminders and scheduled file push) are implemented using Quartz; real-time data uses Flink's time rolling window computing capability to receive data pushed by the Kafka interface to realize real-time data calculation.

[0084] For the data source adaptation layer, you can use encapsulated JDBC, File, Https and other docking methods to achieve docking with relevant data sources, such as TiDB, second-generation files, big data platforms, etc., and establish a corresponding data pool.

[0085] The resource management layer manages data resources and computing resources, and creates "data query views" according to the rules in the rule base. Data query views are converted into corresponding SQL query scripts based on the design patterns of different libraries, extract data from the data pool of the interface layer, and process and convert them into the result data required by the report production layer. It mainly consists of five components: Concurrency Control and Server Aggregation dynamically sense the load of data processing and automatically adjust computing resources and algorithm strategies to ensure the efficiency and real-time performance of data processing; Cache and Query are responsible for extracting processing rules from the rule base, combining the characteristics of different data sources, intelligently adjusting processing strategies, realizing effective integration and efficient utilization of heterogeneous source data, and providing a unified data format for report makers; "Data Dependency and Association Management" is responsible for enhancing the system's adaptability to the business and data ends, especially to the complex time series and dependencies in report data. The autocorrelation function (AF) is used to quantify the time series dependencies of time series data. The long short-term memory network (LSTM) can effectively identify and process these complex time series and dependencies, capture the long-term dependencies between data, and optimize the association rules between data. On this basis, data sorting, aggregation, trend analysis and other processing can be further implemented to support the report making module to more intuitively present the time series changes of data.

[0086] The present application also provides a program product, which includes program instructions, and when the program instructions are executed, any one of the methods described above is implemented.

[0087] The present application also provides a non-transitory storage medium, which stores program instructions, and when the program instructions are executed, any one of the methods described above is implemented.

[0088] According to the present application, during the data processing, data can adapt to changes in data source data, obtain data from multiple heterogeneous data sources, and structure it after analysis and store it as a data set for analysis, thereby achieving data source adaptation and adaptability to different database interfaces. For example, the database may include various relational databases (such as UPSQL), big data platform components (such as HIVE), interfaces may include interfaces for big data queries, etc., and customized development is carried out to efficiently generate the data sets required for report production.

[0089] The following are some abbreviations and terminology used herein (including the accompanying drawings):

[0090] ECharts is an open source JavaScript-based data visualization chart library; D3 is a JavaScript graphics library based on Canvas, Svg, and HTML; BootStrap is a powerful, scalable, and feature-rich front-end development toolkit; datax is the open source version of Alibaba Cloud DataWorks data integration and a widely used offline data synchronization tool; datax-web provides a simple and easy-to-use operation interface to reduce the user's learning cost for using datax and shorten the task configuration time; TiDB is a fusion distributed database product that supports both online transaction processing and online analytical processing.

[0091] In the absence of contradiction or conflict, the technical features in the examples described herein may be combined with each other to form implementations not described herein, which should also be covered by the scope of this application.

[0092] Although specific embodiments of the present application have been shown and described in detail to illustrate the principles of the present application, it should be understood that the present application may be implemented in other ways without departing from such principles, for example, the technical features of the various embodiments / examples / examples of the present application may be combined with each other to form new implementation methods.

Claims

1. A report generation method, characterized in that: The method comprises: In response to a user request, retrieving information associated with the user request from a database, wherein the database is pre-configured to include a domain-specific knowledge base and a rule base; fusing the user request and the retrieved information to generate first intermediate information, wherein the first intermediate information can more accurately express the user request than the user request; transmitting the first intermediate information to a large language model; Processing the first intermediate information by the large language model to generate executable statements; Based on the executable statement, generate a new prompt word; The big data model generates a summary and report style recommendation based on the new prompt word; and generating a report matching the user request according to the user request, the summary and the report style recommendation; The generated report is a report based on information in the specific field.

2. The method according to claim 1, characterized in that In response to a user request, retrieving information associated with the user request from a database, including: Forming the user request into a plurality of sub-questions, wherein the sub-questions include the original question in the user request or the decomposed question of the original question, and the sub-questions also include new questions related to the user request; A plurality of retrieval techniques are used to search the domain-specific knowledge base to obtain knowledge for the plurality of sub-problems from the knowledge base.

3. The method according to claim 2, characterized in that The user request and the retrieved information are integrated to generate first intermediate information, including: Based on the user intention expressed in the user request and the semantic overlap in the knowledge, the knowledge used for the multiple sub-problems is processed to generate first intermediate information that is more accurate than the user request and includes more knowledge, as a prompt word to be input into the large language model.

4. The method according to claim 1, characterized in that: Based on the executable statement, a new prompt word is generated, including: Sending the executable statement to the specific domain knowledge base to obtain execution results for subsequent data analysis; The execution result is combined with the report rules in the specific domain knowledge base and the rule base to generate new prompt words.

5. The method according to claim 1, characterized in that The domain-specific knowledge base includes metadata, data lineage, and user preferences, and the method further includes, when configuring the domain-specific knowledge database: Analyze the data in the specific domain knowledge base by one or more methods including word frequency statistics, keyword extraction, and topic clustering; The analyzed data is structured.

6. The method according to claim 1, characterized in that Generating a report matching the user request includes using a visualization component required for generating the report from a component library in a pluggable manner.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Continuously make adaptive adjustments by collecting user feedback and / or capturing user intent; New knowledge is continuously acquired to improve the knowledge base.

8. An intelligent agent module, characterized in that: The intelligent agent module is configured as follows: Receive user requests; Retrieving information associated with the user request from a database, wherein the database is preconfigured to include a domain-specific knowledge base and a rule base; fusing the user request and the retrieved information to generate first intermediate information, wherein the first intermediate information can more accurately express the user request than the user request; transmitting the first intermediate information to the large language model; Based on the executable sentence sent by the large language model, generate a new prompt word and send the new prompt word to the large language model; and Generating a report matching the user request according to the user request, summary and report style recommendation, wherein the summary and report style recommendation are generated by the big data model according to the new prompt word; The generated report is a report based on information in the specific field.

9. The intelligent agent module according to claim 8, characterized in that: The agent module is also configured to retrieve information associated with the user request from a database by: Forming the user request into a plurality of sub-questions, wherein the sub-questions include the original question in the user request or the decomposed question of the original question, and the sub-questions also include new questions related to the user request; A plurality of retrieval techniques are used to search the domain-specific knowledge base to obtain knowledge for the plurality of sub-problems from the knowledge base.

10. The intelligent agent module according to claim 9, characterized in that: The agent module is further configured to fuse the user request and the retrieved information through the following process to generate first intermediate information: According to the user intention expressed in the user request and the semantic overlap in the knowledge, the knowledge used for the multiple sub-problems is processed to regenerate first intermediate information that is more accurate than the user request and includes more knowledge as a prompt word to be input into the large language model.

11. The intelligent agent module according to claim 8, characterized in that: The domain-specific knowledge base includes metadata, data indicating data lineage, and data indicating user preferences. The agent module is further configured to: Analyze the data in the specific domain knowledge base by one or more methods including word frequency statistics, keyword extraction, and topic clustering; The analyzed data is structured.

12. The intelligent agent module according to claim 8, characterized in that: The intelligent agent module is also configured to use the visualization components required for generating the report from a component library in a pluggable manner.

13. The intelligent agent module according to any one of claims 8 to 12, characterized in that: The agent module is further configured as follows: Continuously make adaptive adjustments by collecting user feedback and / or capturing user intent; New knowledge is continuously acquired to improve the knowledge base.

14. A report generation method, characterized in that: The method comprises: Receive user requests; Retrieving information associated with the user request from a database, wherein the database is preconfigured to include a domain-specific knowledge base and a rule base; fusing the user request and the retrieved information to generate first intermediate information, wherein the first intermediate information can more accurately express the user request than the user request; transmitting the first intermediate information to the large language model; Based on the executable sentence sent by the large language model, generate a new prompt word and send the new prompt word to the large language model; and Generating a report matching the user request according to the user request, summary and report style recommendation, wherein the summary and report style recommendation are generated by the big data model according to the new prompt word; The generated report is a report based on information in the specific field.

15. The method according to claim 14, characterized in that Retrieving information associated with the user request from a database, including: Forming the user request into a plurality of sub-questions, wherein the sub-questions include the original question in the user request or the decomposed question of the original question, and the sub-questions also include new questions related to the user request; A plurality of retrieval techniques are used to search the domain-specific knowledge base to obtain knowledge for the plurality of sub-problems from the knowledge base.

16. The method according to claim 15, characterized in that The user request and the retrieved information are integrated to generate first intermediate information, including: According to the user intention expressed in the user request and the semantic overlap in the knowledge, the knowledge used for the multiple sub-problems is processed to regenerate first intermediate information that is more accurate than the user request and includes more knowledge as a prompt word to be input into the large language model.

17. The method according to claim 14, characterized in that The executable statement generates a new prompt word, including: Sending the executable statement to the specific domain knowledge base to obtain execution results for subsequent data analysis; The execution result is combined with the report rules in the specific domain knowledge base and the rule base to generate new prompt words.

18. The method according to claim 14, characterized in that: The domain-specific knowledge base includes metadata, data representing data lineage, and data representing user preferences. The method further includes, when configuring the domain-specific knowledge database: Analyze the data in the specific domain knowledge base by one or more methods including word frequency statistics, keyword extraction, and topic clustering; The analyzed data is structured.

19. The method according to claim 14, characterized in that Generating a report matching the user request includes: using a visualization component required for generating the report from a component library in a pluggable manner.

20. The method according to any one of claims 14 to 19, characterized in that The method further comprises: Continuously make adaptive adjustments by collecting user feedback and / or capturing user intent; New knowledge is continuously acquired to improve the knowledge base.

21. A report generation system, characterized in that: The system comprises: Databases, including domain-specific knowledge bases and rule bases; Component library, including various visualization components; An intelligent agent module, the intelligent agent module is communicatively connected with the database, the component library, and the large language data model; The intelligent agent module adopts the intelligent agent module according to any one of claims 8 to 13, or the intelligent agent module is configured to execute the method according to any one of claims 14 to 20; Wherein, when the intelligent agent module generates a report matching the user request according to the user request, summary and report style recommendation, it obtains the required visualization components from the component library.

22. The system according to claim 21, characterized in that The system further comprises: An interface is used for the intelligent agent module to call at least a required application through the interface during the process of generating the report.

23. The system according to claim 21, characterized in that The component library is configured to encapsulate components to construct the visualization components required to form the report.

24. The system according to claim 21, characterized in that The component library is configured so that the visualization components included therein can be used by the agent in a pluggable manner.

25. A program product, characterized in that The program product includes program instructions, which can implement the method according to any one of claims 14 to 0 when executed; or can implement the method according to any one of claims 1 to 7 when executed.

26. A non-transitory storage medium, characterized in that: The medium stores program instructions, which can implement the method according to any one of claims 14 to 20 when executed; or can implement the method according to any one of claims 1 to 7 when executed.

Citation Information

Patent Citations

  • Self-service report generating method, device and system

    CN103218448A

  • Automatic report generation method and device, computer equipment and storage medium

    CN118052214A