Intelligent data analysis report generation method and device based on intelligent analysis intelligent agent

CN122817291APending Publication Date: 2026-09-25BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611123250.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]相关技术中,针对多源异构数据的分析报告生成,业务人员需要分别查询不同数据源中的异构数据,然后将多源异构数据进行手动整合,形成最终的分析报告,操作复杂,报告生成效率低

Benefits of technology

[0009]通过上述方法,解析对智能分析智能体的报告生成请求,获得结构化的第一任务表示,基于第一任务表示,分别从存储结构化数据的第一数据源获得第一数据以及从存储非结构化数据的第二数据源获得第二数据,进而在第一数据和第二数据中,将对应同一特征对象的数据进行聚合处理得到第一数据集合,基于第一数据集合生成包括至少一个报告片段的第一智能数据分析报告,其中,第一报告片段包括第三数据以及基于第三数据生成的第一内容。采用该方法,通过智能分析智能体将报告生成请求解析为结构化的任务表示,并依据该表示从多源异构数据中召回相关数据,再将召回的数据按照特征对象进行聚合,并生成相应的智能数据分析报告,一方面,结构化的任务表示能够提升数据召回的精确度和效率,为后续生成报告奠定数据基础,另一方面,通过特征对象聚合与内容耦合,避免了手动整合的数据偏差,保证了报告内容的一致性、可靠性和逻辑连贯性,最终有效提升了报告的生成效率以及报告内容的准确性和可读性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817291A_ABST
    Figure CN122817291A_ABST
Patent Text Reader

Abstract

The application discloses an intelligent data analysis report generation method and device based on an intelligent analysis agent, and relates to the technical field of agents.The method comprises the following steps: in response to a report generation request of an intelligent analysis agent, a first task representation is obtained by analyzing the report generation request; first data is obtained from a first data source and second data is obtained from a second data source based on the first task representation, the first data source is used for storing structured data, and the second data source is used for storing unstructured data; in the first data and the second data, data corresponding to the same feature object is aggregated to obtain a first data set; and a first intelligent data analysis report is generated based on the first data set, wherein the first intelligent data analysis report comprises at least one report segment, the at least one report segment comprises a first report segment, and the first report segment comprises third data and first content. Therefore, the report generation efficiency, the accuracy and the readability of the report content are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This article relates to the field of intelligent agent technology, specifically to a method and apparatus for generating intelligent data analysis reports based on intelligent analytical agents. Background Technology

[0002] As data analysis scenarios become increasingly diverse and report structures become more complex, the data supporting these reports exhibits multi-source and heterogeneous characteristics, including structured and unstructured data from different data sources.

[0003] In related technologies, for generating analysis reports on multi-source heterogeneous data, business personnel need to query heterogeneous data from different data sources separately, and then manually integrate the multi-source heterogeneous data to form the final analysis report. This process is complex and the report generation efficiency is low. Summary of the Invention

[0004] This content section is provided to briefly introduce the ideas, which will be described in detail in the examples section later. This content section is not intended to identify key or essential features of the claimed content, nor is it intended to limit the scope of the claimed content.

[0005] Firstly, a method for generating intelligent data analysis reports based on intelligent analytical agents is provided, including: In response to a report generation request from an intelligent analytical agent, the report generation request is parsed to obtain a first task representation, wherein the first task representation is in a structured form; Based on the first task representation, first data is obtained from a first data source and second data is obtained from a second data source, respectively. The first data source is used to store structured data and the second data source is used to store unstructured data. In the first data and the second data, data corresponding to the same feature object are aggregated to obtain a first data set; Based on the first data set, a first intelligent data analysis report is generated. The first intelligent data analysis report includes at least one report segment, the at least one report segment includes a first report segment, the first report segment includes third data and first content generated based on the third data, and the first data set includes the third data.

[0006] Secondly, a device for generating intelligent data analysis reports based on intelligent analytical agents is provided, comprising: The parsing module is used to respond to a report generation request from an intelligent analysis agent, parse the report generation request, and obtain a first task representation, wherein the first task representation is in a structured form; The acquisition module is used to obtain first data from a first data source and second data from a second data source based on the first task representation. The first data source is used to store structured data, and the second data source is used to store unstructured data. The processing module is used to aggregate data corresponding to the same feature object in the first data and the second data to obtain a first data set; A generation module is configured to generate a first intelligent data analysis report based on the first data set. The first intelligent data analysis report includes at least one report segment, the at least one report segment includes a first report segment, the first report segment includes third data and first content generated based on the third data, and the first data set includes the third data.

[0007] Thirdly, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processing device, implements the method described in the first aspect.

[0008] Fourthly, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the method described in the first aspect.

[0009] The above method parses the report generation request to the intelligent analysis agent to obtain a structured first task representation. Based on the first task representation, first data is obtained from a first data source storing structured data, and second data is obtained from a second data source storing unstructured data. Then, data corresponding to the same feature object are aggregated from the first and second data to obtain a first data set. Based on the first data set, a first intelligent data analysis report including at least one report fragment is generated, whereby the first report fragment includes third data and first content generated based on the third data. This method parses the report generation request into a structured task representation by the intelligent analysis agent, retrieves relevant data from multi-source heterogeneous data based on this representation, aggregates the retrieved data according to feature objects, and generates a corresponding intelligent data analysis report. On the one hand, the structured task representation improves the accuracy and efficiency of data retrieval, laying a data foundation for subsequent report generation. On the other hand, by coupling feature object aggregation with content, data biases from manual integration are avoided, ensuring the consistency, reliability, and logical coherence of the report content. Ultimately, this effectively improves the report generation efficiency, accuracy, and readability of the report content.

[0010] Other features and advantages will be described in detail in the following examples section. Attached Figure Description

[0011] The above and other features, advantages, and aspects of this document will become more apparent when viewed in conjunction with the accompanying drawings and the following examples. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a schematic diagram illustrating an implementation environment according to an example.

[0012] Figure 2 This is a schematic diagram of an intelligent data analysis report generation system, as illustrated by an example.

[0013] Figure 3 This is a flowchart illustrating an intelligent data analysis report generation method based on an intelligent analytical agent, according to an example.

[0014] Figure 4 This is a schematic diagram of the structure of an intelligent data analysis report generation device based on an intelligent analysis agent, as illustrated by an example.

[0015] Figure 5 This is a schematic diagram of the structure of an electronic device according to an example. Detailed Implementation

[0016] The following description will be given in more detail with reference to the accompanying drawings. While certain scenarios are shown in the drawings, it should be understood that this document can be implemented in various forms and should not be construed as limited to the scenarios described herein. Rather, these scenarios are provided to provide a more thorough and complete understanding of this document. It should be understood that the accompanying drawings and the scenarios depicted are for illustrative purposes only and are not intended to limit the scope of this document.

[0017] It should be understood that the steps described in the method may be performed in different orders and / or in parallel. Furthermore, the method may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.

[0018] The term "comprising" and its variations can be open-ended, meaning "including but not limited to". The term "based on" can mean "at least partially based on". The term "one case" means "at least one case"; the term "another case" means "at least one additional case"; the term "some cases" means "at least some cases". Definitions of other terms will be given in the following description.

[0019] It should be noted that the concepts of "first" and "second" are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.

[0020] It should be noted that the modifiers “one” and “multiple” can be illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as “one or more”.

[0021] The names of messages or information exchanged between multiple devices are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0022] It is understandable that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant regulations.

[0023] Structured data typically exists in datasets, indicator models, subject tables, or reporting systems, emphasizing numerical calculations, filtering and aggregation, and multi-table joins; unstructured data typically exists in documents, knowledge bases, meeting minutes, program materials, and other carriers, emphasizing semantic retrieval, contextual understanding, and experience reuse.

[0024] In related technologies, the main processing flow of knowledge question-answering agents or knowledge base retrieval systems is as follows: documents are pre-parsed, sliced, and indexed; natural language questions are received to recall document fragments; and then text answers are generated based on the recalled fragments. While capable of answering questions about rules, processes, and business descriptions, they lack the direct computational capability for structured indicators, time-series data, and multidimensional analysis results. The answers remain at the level of textual summaries and cannot provide data analysis results.

[0025] The main processing flow of intelligent query agents or natural language data query systems is as follows: parsing user questions, identifying data indicators and data dimensions, generating query statements for structured data, and generating brief explanations based on the returned query results. While capable of handling indicator queries, trend analysis, and multi-table joins, they rely on pre-built data semantics and field definitions, and have insufficient utilization of unstructured information such as document knowledge and business data.

[0026] Furthermore, solutions based on report generation or analysis tools require users to manually organize document data, select data sources, supplement the analysis framework, and then perform retrieval, data extraction, analysis, charting, and writing processes. This necessitates frequent switching between different data sources and report editing tools, resulting in long operational chains, information fragmentation, poor timeliness, and insufficient quality stability. Moreover, the lack of a unified binding relationship between data definitions, knowledge fragments, and analytical conclusions in the reports makes them difficult to verify or reuse. Additionally, the high coupling between different data sources and output carriers leads to difficulties in system scalability and adaptation to permissions and scenario-based configurations.

[0027] In view of this, this paper provides a method and apparatus for generating intelligent data analysis reports based on intelligent analytical agents to solve the above-mentioned technical problems.

[0028] The intelligent data analysis report generation method based on intelligent analytical agents provided in this paper can be executed by an electronic device, which can be provided as at least one of a terminal and a server. Figure 1 This is an exemplary schematic diagram illustrating an implementation environment; see [link / reference]. Figure 1 The implementation environment includes: intelligent analysis agent 101, first data source 102 and second data source 103.

[0029] For example, the intelligent analysis agent 101 receives and parses the report generation request to obtain a structured task representation. Based on the task representation, it recalls structured data from the first data source 102 and unstructured data from the second data source 103, respectively. It then aggregates the recalled data according to the feature objects to obtain a data set, and finally generates an intelligent data analysis report based on the data set.

[0030] The front-end interface of the intelligent analysis agent 101 can be implemented by a terminal, the back-end processing logic can be implemented by a server, and the first data source 102 and the second data source 103 can be implemented by electronic devices with storage functions, such as servers.

[0031] For example, a terminal can be at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, virtual reality terminal, augmented reality terminal, wireless terminal, and laptop computer. The terminal has communication capabilities and can access wired or wireless networks. "Terminal" can refer to one of multiple terminals; those skilled in the art will understand that the number of terminals can be more or less. A server can be a single physical server, a server cluster consisting of multiple physical servers, or a distributed file system. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and artificial intelligence platforms.

[0032] This article also provides an intelligent data analysis report generation system, such as Figure 2 As shown, the system includes a receiving and parsing unit, a data source access unit, an unstructured data retrieval unit, a structured data analysis unit, a data fusion unit, a source tracing output unit, and a report generation unit. These units can collaborate within a unified task context. The following section provides an example of the intelligent data analysis report generation method based on intelligent analytical agents, as illustrated in this paper, using an intelligent data analysis report generation system as an example.

[0033] The data source access unit is used to access both structured and unstructured data sources. Structured data sources can include datasets, data tables, indicator models, topic tables, preprocessed tabular files, or other computable data. Unstructured data sources can include knowledge base documents from the knowledge engine, uploaded documents, solution materials, meeting minutes, or other text materials. For unstructured data sources, pre-processing can include parsing, slicing, summarizing, extracting keywords, extracting metadata, building indexes, and maintaining access permissions. For structured data sources, information such as field descriptions, indicator definitions, inter-table relationships, access permissions, and descriptions of available analytical capabilities can be maintained. Specific processing can be tailored to requirements and is not limited in this regard.

[0034] Figure 3 This is a flowchart illustrating an exemplary method for generating intelligent data analysis reports based on an intelligent analytical agent. For example... Figure 3 As shown, the method may include the following steps: In step 301, in response to a report generation request from the intelligent analysis agent, the report generation request is parsed to obtain a first task representation, which is in a structured form.

[0035] Among them, the intelligent analysis agent is an automated intelligent execution unit with a large language model (LLM) as its core brain. It has the ability to autonomously schedule data, fuse multi-source data, perform logical reasoning, and generate report content. It can connect to various structured and unstructured heterogeneous data sources and connect to different output carriers to output the generated reports.

[0036] For example, such as Figure 2 As shown, the receiving and parsing unit is used to receive and parse the report generation request to obtain a structured task representation. The report generation request may include at least one of the following: a natural language question, supplementary instructions, historical dialogue information, and optional local files or reference data. Supplementary instructions may be information such as data access addresses that can obtain other data to supplement the report generation requirements. The specific details can be determined according to the needs and are not limited thereto.

[0037] It should be understood that the intelligent analysis agent provided in this paper can generate corresponding intelligent data analysis reports in response to report generation requests, perform data queries or data analysis in response to data processing requests, and answer knowledge questions in response to natural language questions, etc. The specific settings can be configured according to needs, and there are no restrictions on this.

[0038] In one scenario, the first task representation includes at least one of the following: task intent, task object, task theme, time frame, indicator dimension, output format, and constraints.

[0039] For example, the task intent describes the core objectives and business goals to be achieved in this task; the task object clarifies the specific entity objects and / or data objects to be processed in this task; the task theme defines the theme type corresponding to this task, such as a specific industry; the time range constrains the start and end time intervals for data extraction, statistics, and activation in this task; the indicator dimension defines the quantitative evaluation standards for data analysis; and the output format constrains the carrier, format, and display structure of the final output content of this task. Constraints describe the rules that must be followed when executing this task. These can be set according to requirements and are not restricted.

[0040] In addition, when the report generation request is unclear, such as when there is ambiguity in natural language questions, question rewriting, slot completion, or clarification interaction can be performed first to form a structured task representation, so as to ensure the accuracy of the task representation and thus the accuracy of the subsequent report generation.

[0041] In step 302, based on the first task representation, first data is obtained from a first data source and second data is obtained from a second data source. The first data source is used to store structured data, and the second data source is used to store unstructured data.

[0042] In step 303, the data corresponding to the same feature object in the first data and the second data are aggregated to obtain the first data set.

[0043] For example, such as Figure 2 As shown, the data fusion unit maps the retrieved multi-source heterogeneous data into a unified evidence set, forming an evidence chain that supports the final report conclusion and ensures the consistency of the supporting data. The feature objects can be entities, time, indicators, topic tags, etc., and can be set according to requirements without limitation.

[0044] In step 304, a first intelligent data analysis report is generated based on the first data set. The first intelligent data analysis report includes at least one report segment, the at least one report segment includes a first report segment, the first report segment includes third data and first content generated based on the third data, and the first data set includes the third data.

[0045] For example, the output medium for intelligent data analysis reports can be online documents, Word, PDF, or other readable report formats, which can be set according to needs and are not limited thereto. Furthermore, the report content is not simply a patchwork of recalled data, but rather a coupling of the recalled data and the corresponding generated content, ensuring logical coherence and clear hierarchy, effectively improving the readability of the report.

[0046] Thus, knowledge retrieval, data querying, and report generation can be completed through a unified interactive portal, eliminating the need for switching between systems or manually integrating information. Simultaneously, it integrates the dual capabilities of precise calculation of structured data and semantic supplementation of unstructured knowledge, ensuring that the output reports contain both quantitative numerical conclusions and business interpretations, improving the accuracy and completeness of intelligent data analysis reports. The various disparate steps in generating intelligent data analysis reports are linked together into an automated processing flow of the intelligent analysis agent, reducing the overall report production cycle.

[0047] In another scenario, the task execution plan may also include a report generation subtask, which executes the report generation subtask to generate a first intelligent data analysis report based on a first data set.

[0048] For example, the task execution plan may also include a report generation subtask to standardize the report generation process of the intelligent analysis agent. By breaking down the complex report generation process into multiple subtasks and orchestrating them uniformly, these subtasks are executed collaboratively within the same task context. This achieves decoupling of business processes and flexible controllability of execution logic, effectively improving the utilization rate of computing resources.

[0049] It should be understood that the intelligent analysis agent can generate a complete task execution plan at once, or it can first complete the problem rewriting and task parsing, and then trigger the unstructured data retrieval, structured data analysis and report generation processes respectively. The specific settings can be configured according to the requirements, and there are no restrictions on this.

[0050] By employing the above method, the intelligent analysis agent parses the report generation request into a structured task representation, and retrieves relevant data from multi-source heterogeneous data based on this representation. The retrieved data is then aggregated according to feature objects to generate a corresponding intelligent data analysis report. On the one hand, the structured task representation can improve the accuracy and efficiency of data retrieval, laying a data foundation for subsequent report generation. On the other hand, by coupling feature object aggregation with content, the data bias of manual integration is avoided, ensuring the consistency, reliability, and logical coherence of the report content. Ultimately, this effectively improves the report generation efficiency as well as the accuracy and readability of the report content.

[0051] In one scenario, based on a first task representation, obtaining first data from a first data source and second data from a second data source includes: generating a task execution plan based on the first task representation, the task execution plan including a first subtask and a second subtask; generating a query statement and / or analysis strategy based on the data objects corresponding to the first subtask and the relationships between the data objects, executing the query statement and / or analysis strategy to obtain the first data from the first data source; determining at least one search condition based on the second subtask, obtaining data matching the at least one search condition from the second data source to obtain the second data.

[0052] For example, such as Figure 2 As shown, the receiving and parsing unit can also be used to orchestrate a retrieval plan based on the structured task representation. For example, it can determine whether the current request belongs to a knowledge question answering, data query, or report generation task based on the structured task representation. For a report generation task, it can be broken down into several sub-tasks, such as a first sub-task for structured data sources and a second sub-task for unstructured data sources, and assign knowledge retrieval paths, structured analysis paths, or hybrid execution paths to each sub-task. At the same time, a retrieval plan, i.e., a task execution plan, is generated based on the task representation.

[0053] It should be noted that each subtask in the task execution plan can be considered as a step to be executed. For example, subtask 1 is used to query data in field B of data table A, subtask 2 is used to query the related data between data table A and data table C, subtask 3 is used to query knowledge fragments related to object D, subtask 4 is used to query knowledge fragments related to object F, and so on. Furthermore, multiple retrieval subtasks can be executed in parallel to improve retrieval efficiency; the specific settings can be configured according to requirements, and there are no restrictions on this.

[0054] For example, search criteria may include those based on keyword matching, vector matching, summary matching, etc. Figure 2 As shown, the unstructured data retrieval unit can perform keyword retrieval based on keywords in subtasks, vector retrieval based on vectors corresponding to subtasks, and summary retrieval based on task summaries corresponding to subtasks. It can also perform retrieval based on one or more of the above retrieval conditions in unstructured data sources such as knowledge engines or document indexes. Furthermore, it can perform metadata filtering and reordering on the retrieved data that meets the retrieval conditions to obtain the final retrieved unstructured data, such as knowledge fragments that are highly relevant to the task.

[0055] Among them, the data that matches at least one search condition can be data that meets the preset matching conditions. For example, the matching condition for keyword matching can be hitting the preset keyword, the matching condition for vector matching can be that the vector similarity is greater than the first similarity and the first preset number of data with the highest vector similarity can be selected, the matching condition for summary matching can be that the semantic similarity is greater than the second similarity and the first preset number of data with the highest semantic similarity can be selected, etc. The specific settings can be set according to the needs and there are no restrictions on them.

[0056] For example, the structured data analysis unit can identify data objects such as data tables, data fields, and data metrics related to subtasks, as well as the relationships between these data objects. It generates corresponding query statements and / or analysis strategies. For query statements, it can invoke query execution tools; for analysis strategies, it can invoke analysis calculation tools. Alternatively, it can invoke multiple executors, such as query execution tools, analysis calculation tools, or rule calculation engines, to collaboratively process data. This enables operations such as data aggregation, filtering, comparison, trend analysis, multi-table joins, or attribution analysis, yielding structured data such as tabular results, statistical results, and chart results. Furthermore, the retrieval of structured data can be achieved using different methods, including semantic model mapping, metric retrieval, and query plan generation.

[0057] This will standardize the retrieval process of intelligent analytical agents, improve data retrieval efficiency and accuracy, and thus enhance report generation efficiency and content accuracy.

[0058] In one scenario, the method further includes: obtaining a first permission scope, the first permission scope being determined by the requester of the report generation request; generating a task execution plan based on a first task representation, including: generating a task execution plan based on the first task representation and the first permission scope, the first permission scope being used to constrain the retrieval scope of the first subtask and / or the second subtask.

[0059] For example, data access permissions corresponding to the requester of the report generation request can be obtained. Then, a retrieval plan can be generated by combining the topic tags, business objects, time windows, and other content in the structured task representation with the obtained permission scope. For instance, if the requester is allowed to access data in field B of table A, the aforementioned subtask 1 can be generated. Alternatively, the permission scope and subtask 1 can be issued to the corresponding data source. When executing a data query, the permission scope is used to determine whether to execute the query subtask or whether the corresponding data can be returned, thereby ensuring secure data access during the report generation process.

[0060] In one scenario, data corresponding to the same feature object in the first and second data are aggregated to obtain a first data set. This includes: classifying the first and second data according to different feature objects to obtain multiple data subsets, where the data in each subset corresponds to the same feature object and each data has a corresponding confidence level; performing filtering and / or fusion processing on each data subset; wherein the filtering process is used to filter data in the data subset that does not meet a first preset condition, and the fusion process is used to merge data in the data subset whose similarity is greater than a preset threshold; and sorting the data within each data subset according to the confidence level to obtain the first data set.

[0061] For example, such as Figure 2As shown, the data fusion unit aligns the recalled structured and unstructured data corresponding to the same feature object, thereby dividing them into data subsets corresponding to multiple feature objects. Among them, the data within the same data subset is processed according to the task objective, evidence integrity and consistency rules. For example, data with overlapping information or different expressions can be fused, such as fusion of two data with similarity greater than a threshold. Data with content conflicts are filtered out, that is, the report generation does not depend on conflicting content.

[0062] For example, data within the same subset can be sorted according to the confidence level of each data point. This allows subsequent report generation to determine whether to adopt a data point or to assign it a sampling weight based on its confidence level. The specific settings can be customized to meet specific needs and are not restricted. This enables the unified organization of recalled multi-source heterogeneous data into an evidence set, and aligns and fuses it based on multi-dimensional features including entities, time, indicators, and semantic relationships.

[0063] This enables the elimination of differences in data formats and standards from different data sources, efficiently completes unified processing of cross-source evidence, improves data integration efficiency and data consistency, thereby forming an evidence chain that can support the report's conclusions and improves the quality of the report content.

[0064] In one scenario, generating a first intelligent data analysis report based on a first data set includes: determining first structural information based on the first data set, wherein the first structural information is a preset report template or a dynamically generated report outline, and the first structural information includes at least one segment identifier; generating content corresponding to each segment identifier based on the first data set; generating at least one report segment based on at least one segment identifier and the content corresponding to each segment identifier; and obtaining the first intelligent data analysis report, wherein at least one report segment corresponds one-to-one with at least one segment identifier.

[0065] For example, such as Figure 2 As shown, the report generation unit can select a suitable report template based on the chain of evidence or dynamically generate a report outline to generate report structure information. It can also determine the output format based on the request, without restriction. The structure information includes at least one segment identifier, such as sections like summary, background, key findings, analysis process, conclusions and recommendations, and additional explanations. Then, based on the chain of evidence and according to the requirements of each segment, the corresponding report content is generated and populated into the corresponding segments. During the generation process, the recalled data used to generate the content can be embedded into the corresponding segments along with the generated content, so that the final intelligent data analysis report includes both numerical conclusions and textual explanations, improving the readability of the report content.

[0066] In other words, the report generation unit does not simply piece together the recalled content. Instead, it first forms a report based on the evidence set to close the case, then generates the main text by chapter, and backfills the corresponding analysis results and recalled data at the chapter level, thereby improving the stability of the report structure and the interpretability of the results.

[0067] Alternatively, reports can be generated iteratively by segment or chapter, by filling in templates, or by compiling Q&A results. The specific method can be chosen according to the needs, and there are no restrictions on this.

[0068] In one scenario, the first intelligent data analysis report further includes second content, and the method further includes: acquiring fourth data, the fourth data including at least one of fifth data for generating the second content, recall conditions for the fifth data, and generation logic for the second content, and the first data set including the fifth data; wherein the fourth data is displayed in association with the second content in the first intelligent data analysis report.

[0069] For example, such as Figure 2 As shown, the source tracing output unit binds corresponding related data to the content generated in the report (such as key conclusions). This can include the recall data used to generate the content, the recall conditions for the data, the filtering conditions, and the processing procedures for the data (such as execution steps and thought processes). Specific details can be set according to needs and are not limited. The generated content and related content can be displayed in conjunction, for example, as embedded annotations, appendix indexes, metadata panels, or separate evidence lists; there are no restrictions on this as well.

[0070] This allows the generated report content to be linked to the corresponding data results, knowledge fragments, filtering conditions, and execution steps, and displayed in a related manner, thereby improving the report's interpretability and searchability.

[0071] In one scenario, the method further includes: responding to an update request for a first intelligent data analysis report, updating the first intelligent data analysis report based on a first data set and a first context to obtain a second intelligent data analysis report; wherein the first context is intermediate data generated by the intelligent analysis agent during the generation process of the first intelligent data analysis report, and the update process includes modifying and / or adding content to the first intelligent data analysis report.

[0072] For example, the system supports updating report content. For instance, in response to subsequent natural language questions, it can reuse existing task context and evidence sets to perform partial recalculation, content rewriting, or supplementary explanations, thus supporting multi-round report iterations. The context can include the intelligent analysis agent's reasoning process, historical input questions, recommendation questions, template constraints, and permission information, etc., and can be set according to requirements without limitation. This allows the intelligent analysis agent to maintain topic continuity in multi-round dialogues and continuously access multi-source heterogeneous data within the same permission boundaries.

[0073] Using the above method, problem understanding, hybrid retrieval, analytical reasoning, result fusion, and report output can be completed within a unified task chain. It also enables unified arrangement and collaborative processing of structured and unstructured data, generating comprehensive reports that combine data conclusions, knowledge support, and verifiable sources. Furthermore, based on a pluggable data source access and output rendering mechanism, it can flexibly adapt to different data sources, data types, and report output carriers, and can support more data analysis scenarios by incorporating access control, tag filtering, and template constraints.

[0074] This paper also provides a method for generating intelligent data analysis reports based on intelligent analytical agents, including: responding to a report generation request from an intelligent analytical agent, parsing the report generation request to obtain a first task representation, wherein the first task representation is in a structured form; generating a task execution plan based on the first task representation, wherein the task execution plan includes a first subtask and a second subtask; generating a query statement and / or analysis strategy according to the data objects corresponding to the first subtask and the relationships between the data objects, executing the query statement and / or analysis strategy to obtain first data from a first data source; determining at least one retrieval condition according to the second subtask, obtaining data matching the at least one retrieval condition from the second data source to obtain second data, wherein the first data source is used to store structured data and the second data source is used to store unstructured data; classifying the first data and the second data according to different feature objects to obtain multiple data subsets, wherein the data in each data subset corresponds to the same feature object and each data has a corresponding confidence level; and performing filtering processing on each data subset. / or fusion processing; wherein, filtering processing is used to filter data in the data subset that does not meet the first preset condition, and fusion processing is used to merge data in the data subset with a similarity greater than a preset threshold; the data in each data subset are sorted according to confidence level to obtain a first data set; a first structural information is determined based on the first data set, the first structural information being a preset report template or a dynamically generated report outline, the first structural information including at least one segment identifier; content corresponding to each segment identifier is generated based on the first data set, and at least one report segment is generated based on at least one segment identifier and the content corresponding to each segment identifier to obtain a first intelligent data analysis report, at least one report segment corresponding one-to-one with at least one segment identifier, at least one report segment including a first report segment, the first report segment including third data and first content generated based on the third data, and the first intelligent data analysis report displaying associated data of the first content, the first data set including third data, and the associated data including the recall conditions and / or processing logic of the intelligent analysis agent for the third data.

[0075] For example, the processing logic may include the analysis logic of the intelligent analysis agent on the third data, the reasoning process of the intelligent analysis agent generating the first content based on the third data, and so on.

[0076] This method uses an intelligent analysis agent to parse report generation requests into a structured task representation. Based on this representation, relevant data is retrieved from multi-source heterogeneous data. The retrieved data is then aggregated according to feature objects to generate a corresponding intelligent data analysis report, which is then displayed in conjunction with related content. On the one hand, the structured task representation improves the accuracy and efficiency of data retrieval, laying a data foundation for subsequent report generation. On the other hand, through feature object aggregation, content coupling, and content association, the method avoids data biases inherent in manual integration, ensuring the consistency, reliability, logical coherence, and queryability of the report content. Ultimately, this effectively improves the efficiency of report generation as well as the accuracy and readability of the report content.

[0077] Figure 4 This is a schematic diagram of the structure of an intelligent data analysis report generation device based on an intelligent analytical agent, as illustrated by an example. Figure 4 As shown, the intelligent data analysis report generation device 400 based on intelligent analytical agents includes: Parsing module 401 is used to parse the report generation request in response to the intelligent analysis agent to obtain a first task representation, wherein the first task representation is in a structured form; The acquisition module 402 is used to obtain first data from a first data source and second data from a second data source based on the first task representation. The first data source is used to store structured data, and the second data source is used to store unstructured data. The processing module 403 is used to aggregate data corresponding to the same feature object in the first data and the second data to obtain a first data set; The generation module 404 is configured to generate a first intelligent data analysis report based on the first data set. The first intelligent data analysis report includes at least one report segment, the at least one report segment includes a first report segment, the first report segment includes third data and first content generated based on the third data, and the first data set includes the third data.

[0078] Using the aforementioned device, an intelligent analysis agent parses the report generation request into a structured task representation, retrieves relevant data from multi-source heterogeneous data based on this representation, aggregates the retrieved data according to feature objects, and generates a corresponding intelligent data analysis report. On the one hand, the structured task representation can improve the accuracy and efficiency of data retrieval, laying a data foundation for subsequent report generation. On the other hand, by coupling feature object aggregation with content, the data deviation of manual integration is avoided, ensuring the consistency, reliability, and logical coherence of the report content. Ultimately, this effectively improves the report generation efficiency as well as the accuracy and readability of the report content.

[0079] Optionally, the acquisition module 402 is used for: A task execution plan is generated based on the first task representation, and the task execution plan includes a first subtask and a second subtask; Based on the data object corresponding to the first subtask and the relationship between the data objects, generate a query statement and / or an analysis strategy, execute the query statement and / or the analysis strategy, and obtain the first data from the first data source; Based on the second subtask, at least one search condition is determined, and data matching the at least one search condition is obtained from the second data source to obtain the second data.

[0080] Optionally, the intelligent data analysis report generation device 400 based on intelligent analytical agents further includes a permission module, used for: Obtain a first scope of permissions, which is determined by the requester that generated the report; The step of generating a task execution plan based on the first task representation includes: Based on the first task representation and the first permission scope, the task execution plan is generated, wherein the first permission scope is used to constrain the retrieval scope of the first subtask and / or the second subtask.

[0081] Optionally, the generation module 404 is used for: Based on the first data set, first structural information is determined. The first structural information is a preset report template or a dynamically generated report outline. The first structural information includes at least one segment identifier. Based on the first data set, generate content corresponding to each of the fragment identifiers, generate the at least one report fragment based on the at least one fragment identifier and the content corresponding to each of the fragment identifiers, and obtain the first intelligent data analysis report, wherein the at least one report fragment corresponds one-to-one with the at least one fragment identifier.

[0082] Optionally, the processing module 403 is used for: The first data and the second data are classified according to different feature objects to obtain multiple data subsets. The data in each data subset corresponds to the same feature object, and each data has a corresponding confidence level. For each of the data subsets, filtering and / or fusion processing are performed on the data subsets; wherein, the filtering processing is used to filter data in the data subsets that do not meet a first preset condition, and the fusion processing is used to merge data in the data subsets whose similarity is greater than a preset threshold; The data within each subset is sorted according to the confidence level to obtain the first data set.

[0083] Optionally, the first intelligent data analysis report further includes second content, and the intelligent data analysis report generation device 400 based on the intelligent analysis agent further includes an acquisition submodule, used for: Obtain fourth data, the fourth data including at least one of fifth data for generating the second content, recall conditions for the fifth data, and generation logic for the second content, the first data set including the fifth data; The fourth data is displayed in conjunction with the second content in the first intelligent data analysis report.

[0084] Optionally, the intelligent data analysis report generation device 400 based on intelligent analytical agents further includes an update module, used for: In response to the update request for the first intelligent data analysis report, based on the first data set and the first context, the first intelligent data analysis report is updated to obtain the second intelligent data analysis report; Wherein, the first context is the intermediate data of the intelligent analysis agent during the generation process of the first intelligent data analysis report, and the update process includes modifying the content of the first intelligent data analysis report and / or adding content.

[0085] Regarding the intelligent data analysis report generation device 400 based on intelligent analysis agents mentioned above, the method logic executed by each functional module has been explained in detail in the section on methods, and will not be repeated here.

[0086] Based on the same concept, a computer-readable medium is also provided, on which a computer program is stored, which, when executed by a processing device, implements the steps of any of the above-described intelligent data analysis report generation methods based on intelligent analytical agents.

[0087] Specifically, the intelligent analysis agent parses the report generation request into a structured task representation, and retrieves relevant data from multi-source heterogeneous data based on this representation. The retrieved data is then aggregated according to feature objects to generate a corresponding intelligent data analysis report. On the one hand, the structured task representation can improve the accuracy and efficiency of data retrieval, laying a data foundation for subsequent report generation. On the other hand, by coupling feature object aggregation with content, the data deviation of manual integration is avoided, ensuring the consistency, reliability, and logical coherence of the report content. Ultimately, this effectively improves the report generation efficiency as well as the accuracy and readability of the report content.

[0088] Based on the same concept, an electronic device is also provided, which may include: A storage device on which computer programs are stored; A processing device is used to execute a computer program stored in a storage device to implement any of the above-described intelligent data analysis report generation methods based on intelligent analytical agents.

[0089] Specifically, the intelligent analysis agent parses the report generation request into a structured task representation, and retrieves relevant data from multi-source heterogeneous data based on this representation. The retrieved data is then aggregated according to feature objects to generate a corresponding intelligent data analysis report. On the one hand, the structured task representation can improve the accuracy and efficiency of data retrieval, laying a data foundation for subsequent report generation. On the other hand, by coupling feature object aggregation with content, the data deviation of manual integration is avoided, ensuring the consistency, reliability, and logical coherence of the report content. Ultimately, this effectively improves the report generation efficiency as well as the accuracy and readability of the report content.

[0090] Based on the same concept, a computer program product is also provided, including a computer program that, when executed by a processor, implements the above-described intelligent data analysis report generation method based on an intelligent analytical agent.

[0091] Specifically, the intelligent analysis agent parses the report generation request into a structured task representation, and retrieves relevant data from multi-source heterogeneous data based on this representation. The retrieved data is then aggregated according to feature objects to generate a corresponding intelligent data analysis report. On the one hand, the structured task representation can improve the accuracy and efficiency of data retrieval, laying a data foundation for subsequent report generation. On the other hand, by coupling feature object aggregation with content, the data deviation of manual integration is avoided, ensuring the consistency, reliability, and logical coherence of the report content. Ultimately, this effectively improves the report generation efficiency as well as the accuracy and readability of the report content.

[0092] The following is for reference. Figure 5 The diagram illustrates the structure of an electronic device 500 suitable for implementing the above method. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs (Televisions), desktop computers, etc. Figure 5 The electronic device shown is merely an example and should not be construed as limiting its functionality or scope of use.

[0093] like Figure 5As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, the ROM 502, and the RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0094] Typically, the following devices can be connected to the input / output interface 505: input devices 506 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 508 including, for example, magnetic tape, hard disk, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0095] In particular, depending on certain circumstances, the processes described in the flowchart above can be implemented as computer software programs. For example, a computer program product is provided, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. This computer program can be downloaded and installed from a network via communication device 509, or installed from storage device 508, or installed from read-only memory 502. When the computer program is executed by processing device 501, it performs the functions defined in the above-described methods.

[0096] It should be noted that the aforementioned computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In one case, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In another case, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.

[0097] In some cases, communication can be conducted using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), Internets (e.g., the Internet), and peer-to-peer networks (e.g., ad-hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0098] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0099] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: respond to a report generation request for an intelligent analytical agent, parse the report generation request to obtain a first task representation, the first task representation being in a structured form; based on the first task representation, obtain first data from a first data source and second data from a second data source, the first data source being used to store structured data and the second data source being used to store unstructured data; aggregate data corresponding to the same feature object in the first data and the second data to obtain a first data set; and based on the first data set, generate a first intelligent data analysis report, the first intelligent data analysis report including at least one report fragment, the at least one report fragment including a first report fragment, the first report fragment including third data and first content generated based on the third data, the first data set including the third data.

[0100] Computer program code for performing the above operations can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages, as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0101] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. Each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative cases, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0102] The modules mentioned above can be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the functionality of that module.

[0103] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Parts (ASSPs), Systems on Chips (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0104] In this context, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0105] The above description is merely illustrative and explains the technical principles employed. Those skilled in the art should understand that the scope of this document is not limited to the specific combinations of the above-described technical features, but should also cover any combination of the above-described technical features or their equivalents without departing from the above concept.

[0106] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, while some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this paper. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.

[0107] Although this document has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the aforementioned apparatus, the specific manner in which the various modules perform their operations has already been described in detail in the section concerning the method, and will not be elaborated upon here.

Claims

1. A method for generating intelligent data analysis reports based on intelligent analytical agents, comprising: In response to a report generation request from an intelligent analytical agent, the report generation request is parsed to obtain a first task representation, wherein the first task representation is in a structured form; Based on the first task representation, first data is obtained from a first data source and second data is obtained from a second data source, respectively. The first data source is used to store structured data and the second data source is used to store unstructured data. In the first data and the second data, data corresponding to the same feature object are aggregated to obtain a first data set; Based on the first data set, a first intelligent data analysis report is generated. The first intelligent data analysis report includes at least one report segment, the at least one report segment includes a first report segment, the first report segment includes third data and first content generated based on the third data, and the first data set includes the third data.

2. The method according to claim 1, wherein obtaining first data from a first data source and second data from a second data source based on the first task representation comprises: A task execution plan is generated based on the first task representation, and the task execution plan includes a first subtask and a second subtask; Based on the data object corresponding to the first subtask and the relationship between the data objects, generate a query statement and / or an analysis strategy, execute the query statement and / or the analysis strategy, and obtain the first data from the first data source; Based on the second subtask, at least one search condition is determined, and data matching the at least one search condition is obtained from the second data source to obtain the second data.

3. The method according to claim 2, further comprising: Obtain a first scope of permissions, which is determined by the requester that generated the report; The step of generating a task execution plan based on the first task representation includes: Based on the first task representation and the first permission scope, the task execution plan is generated, wherein the first permission scope is used to constrain the retrieval scope of the first subtask and / or the second subtask.

4. The method according to any one of claims 1-3, wherein aggregating data corresponding to the same feature object in the first data and the second data to obtain a first data set includes: The first data and the second data are classified according to different feature objects to obtain multiple data subsets. The data in each data subset corresponds to the same feature object, and each data has a corresponding confidence level. For each of the data subsets, filtering and / or fusion processing are performed on the data subsets; wherein, the filtering processing is used to filter data in the data subsets that do not meet a first preset condition, and the fusion processing is used to merge data in the data subsets whose similarity is greater than a preset threshold; The data within each subset of data is sorted according to the confidence level to obtain the first data set.

5. The method according to any one of claims 1-3, wherein generating a first intelligent data analysis report based on the first data set comprises: First structural information is determined based on the first data set. The first structural information is a preset report template or a dynamically generated report outline. The first structural information includes at least one segment identifier. Based on the first data set, generate content corresponding to each of the fragment identifiers, generate the at least one report fragment based on the at least one fragment identifier and the content corresponding to each of the fragment identifiers, and obtain the first intelligent data analysis report, wherein the at least one report fragment corresponds one-to-one with the at least one fragment identifier.

6. The method according to any one of claims 1-3, wherein the first intelligent data analysis report further includes second content, and the method further includes: Obtain fourth data, the fourth data including at least one of fifth data for generating the second content, recall conditions for the fifth data, and generation logic for the second content, the first data set including the fifth data; The fourth data is displayed in conjunction with the second content in the first intelligent data analysis report.

7. The method according to any one of claims 1-3, further comprising: In response to the update request for the first intelligent data analysis report, based on the first data set and the first context, the first intelligent data analysis report is updated to obtain the second intelligent data analysis report; Wherein, the first context is the intermediate data of the intelligent analysis agent during the generation process of the first intelligent data analysis report, and the update process includes modifying the content of the first intelligent data analysis report and / or adding content.

8. An intelligent data analysis report generation device based on an intelligent analytical agent, comprising: The parsing module is used to respond to a report generation request from an intelligent analysis agent, parse the report generation request, and obtain a first task representation, wherein the first task representation is in a structured form; The acquisition module is used to obtain first data from a first data source and second data from a second data source based on the first task representation. The first data source is used to store structured data, and the second data source is used to store unstructured data. The processing module is used to aggregate data corresponding to the same feature object in the first data and the second data to obtain a first data set; A generation module is configured to generate a first intelligent data analysis report based on the first data set. The first intelligent data analysis report includes at least one report segment, the at least one report segment includes a first report segment, the first report segment includes third data and first content generated based on the third data, and the first data set includes the third data.

9. A computer-readable medium having a computer program stored thereon, wherein, When executed by a processing device, the computer program implements the method described in any one of claims 1-7.

10. An electronic device, comprising: A storage device on which computer programs are stored; A processing apparatus for executing the computer program in the storage device to implement the method of any one of claims 1-7.