Municipal infrastructure data report generation method and equipment based on large language model

By using hierarchical reasoning based on a large language model, mind tree diagram construction, and multi-agent collaboration, the problems of insufficient consistency between text and graphics and insufficient analytical depth in the generation of municipal infrastructure data reports are solved, and efficient, in-depth, and reliable automated report generation is achieved.

CN121766293APending Publication Date: 2026-03-31BEIJING SCI & TECH PATENT OFFICE
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve consistency between text and graphics, depth of analysis, and stability of output quality when generating municipal infrastructure data reports. They also fail to effectively decouple the core tasks of domain knowledge, visual expression, and language organization, resulting in limited overall report usability.

Method used

Employing a large language model-based approach, this method extracts key information through hierarchical reasoning, constructs charts using mind trees, and generates reports through multi-agent collaboration. This includes structured meta-information parsing, multi-dimensional analysis, iterative key information refinement, chart configuration evaluation, and multi-agent collaboration, resulting in high-quality analysis reports.

Benefits of technology

It automates the entire process from raw data to high-quality analysis reports, improving the efficiency, depth, and reliability of data analysis and report writing, and ensuring consistency between text and graphics and depth of analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121766293A_ABST
    Figure CN121766293A_ABST
Patent Text Reader

Abstract

The invention provides a municipal infrastructure data report generation method and equipment based on a large language model. The method comprises the following steps: analyzing a statistical data table of underground municipal infrastructures to construct structured meta-information; providing the statistical data table and the structured meta-information of the statistical data table for a large language model to extract an initial key information list; analyzing the initial list to obtain an optimized key information list; generating a candidate chart type set for optimization key information objects in the optimization key information list; constructing a drawing configuration plan for each candidate chart type in the candidate chart type set, scoring the drawing configuration plan of each candidate chart type, and selecting the drawing configuration plan with the highest score as the optimal configuration plan of the corresponding optimization key information object; converting the optimal configuration plan into an executable code; generating a report text according to each optimization key information object and the optimal configuration plan thereof; and inserting a chart file generated by the executable code into the report text to form an analysis report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and device for generating municipal infrastructure data reports based on a large language model. Background Technology

[0002] In the operation and management of urban underground municipal infrastructure (mainly referring to underground pipelines, integrated utility tunnels, oil and gas pipelines, etc.), data analysis reports are the core basis for assessing the health status of infrastructure, providing early warnings of safety risks, and making decisions on upgrading and renovation. Traditional analysis processes heavily rely on manual labor, which is not only time-consuming and labor-intensive when dealing with massive amounts of multidimensional data, but also severely limits the depth and breadth of analysis to individual experience. With the advancement of artificial intelligence technology, especially the rise of Large Language Models (LLM), new technological paths have emerged for the automation of report generation. However, current explorations using this technology mostly adopt an "end-to-end" generation model. This black-box, integrated approach struggles to effectively control the consistency of text and graphics, the depth of analysis, and the quality of output, failing to guarantee continuous and stable high-quality output.

[0003] To address the aforementioned issues, a task decomposition-based method has been proposed in existing technologies. The core of this approach is to pre-divide a complete report into multiple fragmented analysis tasks, and to clearly define the data source and metrics for each task using structured tables. This aims to transform ambitious generation goals into a series of concrete steps, thereby achieving effective guidance and process control of the analysis path for large language models.

[0004] Existing technologies also propose a collaborative framework of "task library + tool library," the core idea of ​​which is to separate specialized data processing tasks from language generation tasks. Within this framework, the system uses predefined task templates to invoke external specialized tools to obtain accurate data, which is then processed by a large language model to organize and describe the objective results returned by the tools. This method aims to leverage tools to ensure the standardization and accuracy of the analysis, while simultaneously utilizing the model's advantages in language organization.

[0005] To enhance the fundamental reliability of the output results, existing technologies have proposed a process that incorporates code verification. The core idea is that the task of the large language model is to generate code for data analysis. Before execution, this code undergoes automated testing and verification by an independent code testing agent to ensure the correctness of the underlying analysis logic.

[0006] While the aforementioned existing technologies have optimized report generation methods from different perspectives, such as task decomposition, process control, and quality verification, they have not addressed the inherent limitations of the "end-to-end" model. Specifically, they fail to effectively decouple and intelligently coordinate the three distinct core tasks: data analysis requiring high domain knowledge, the generation of statistical charts demanding precise visual expression, and the writing of analytical text requiring rigorous language organization. This architecture means that a quality deficiency in any stage will affect the overall usability of the final report. Therefore, overcoming the limitations of existing technologies and providing a new method for generating high-quality reports that ensures analytical depth, consistency between text and graphics, and output stability is a pressing technical problem that needs to be solved in this field. Summary of the Invention

[0007] In view of the above problems, the present invention is proposed to provide a method and apparatus for generating municipal infrastructure data reports based on a large language model, which solves or at least partially solves the above technical problems.

[0008] One aspect of the present invention provides a method for generating municipal infrastructure data reports based on a large language model, the method comprising: The statistical data tables of underground municipal infrastructure are subjected to structural and semantic parsing, and the structured meta-information of the statistical data tables is constructed based on the parsing results; The statistical data table and its structured meta-information are provided as context to the pre-defined large language model, so that the large language model can perform multi-dimensional analysis on the input context to extract and output the initial list of key information contained in the context. The initial list of key information is analyzed and filtered to obtain an optimized list of key information; For each optimization key information object in the optimization key information list, generate a set of candidate chart types; For each candidate chart type in the candidate chart type set for each optimization key information object, a drawing configuration plan is constructed, and the drawing configuration plan corresponding to each candidate chart type is scored, so as to select the drawing configuration plan with the highest score as the optimal configuration plan for the corresponding optimization key information object. The optimal configuration plan for each key information object is transformed into executable code for generating charts; The report's text content is generated based on each optimization key information object in the optimization key information list and the optimal configuration plan for each optimization key information object; Insert chart files generated by calling the executable code corresponding to each key optimization information object into the text content of the report to form a final analysis report containing both text content and chart files.

[0009] In another aspect, the present invention provides a municipal infrastructure data report generation device based on a large language model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the municipal infrastructure data report generation method based on a large language model as described above.

[0010] A third aspect of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the municipal infrastructure data report generation method based on a large language model as described above.

[0011] The municipal infrastructure data report generation method and device based on a large language model provided in this invention introduces a key information extraction stage based on hierarchical reasoning, a chart construction stage based on mind tree, and a report generation stage based on multi-agent collaboration. This achieves full-process automation from raw data to a high-quality analytical report that combines analytical text and visual charts, significantly improving the efficiency, depth, and reliability of data analysis and report writing.

[0012] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0013] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. In the drawings: Figure 1 This is a flowchart of a municipal infrastructure data report generation method based on a large language model, according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the overall system architecture of a specific embodiment of the present invention. Detailed Implementation

[0014] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0015] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art and should not be interpreted in an idealized or overly formal sense unless specifically defined.

[0016] This invention provides a method for generating municipal infrastructure data reports based on a large language model. This method efficiently and automatically generates domain reports containing analytical text and visual charts from statistical data tables of underground municipal infrastructure data, such as structured statistical data tables of underground pipelines, thereby improving the quality and efficiency of report generation. Figure 1 As shown in the figure, the municipal infrastructure data report generation method based on a large language model proposed in this embodiment of the invention includes the following steps: S1. Perform structural and semantic parsing on the statistical data tables of underground municipal infrastructure, and construct structured meta-information of the statistical data tables based on the parsing results.

[0017] S2. Provide the statistical data table and its structured meta-information as context to the pre-defined large language model so that the large language model can perform multi-dimensional analysis on the input context to extract and output the initial list of key information contained in the context.

[0018] S3. Analyze and filter the initial key information list to obtain an optimized key information list.

[0019] S4. Generate a set of candidate chart types for each optimization key information object in the optimization key information list.

[0020] S5. For each candidate chart type in the candidate chart type set for each optimization key information object, construct a drawing configuration plan and score the drawing configuration plan corresponding to each candidate chart type, so as to select the drawing configuration plan with the highest score as the optimal configuration plan for the corresponding optimization key information object.

[0021] S6. Convert the optimal configuration plan for each key information object into executable code for generating charts.

[0022] S7. Generate the text content of the report based on each optimization key information object in the optimization key information list and the optimal configuration plan for each optimization key information object.

[0023] S8. Insert chart files generated by calling the executable code corresponding to each optimization key information object into the text content of the report to form a final analysis report containing text content and chart files.

[0024] The municipal infrastructure data report generation method based on a large language model provided in this invention introduces a key information extraction stage based on hierarchical reasoning, a chart construction stage based on mind tree, and a report generation stage based on multi-agent collaboration. This achieves full-process automation from raw data to a high-quality analytical report that combines analytical text and visual charts, significantly improving the efficiency, depth, and reliability of data analysis and report writing.

[0025] The municipal infrastructure data report generation method based on a large language model provided in this invention aims to transform structured statistical data into a domain report containing analytical text and visual charts through an automated process consisting of three core stages. The specific implementation steps of each stage will be described in detail below.

[0026] The first core stage aims to extract key information based on hierarchical reasoning. The core objective of this stage is to automatically extract a series of deeply analyzed and optimized structured key information from the raw statistical data tables of underground municipal infrastructure through a hierarchical key information extraction and reasoning process. By designing prompts at different levels, the large model is guided to perform in-depth data analysis, generating a high-quality list of key information. The first core stage involves steps S1 to S3, which are described in detail below. Specifically, this stage uses different prompts to guide a standard, untuned LLM (Limited Learning Model) to extract key information from the raw data layer by layer.

[0027] In this embodiment of the invention, step S1 is used to perform data parsing and metadata generation. Step S1 involves performing structural and semantic parsing on the statistical data table of underground municipal infrastructure, and constructing structured metadata of the statistical data table based on the parsing results. Specifically, this includes the following steps (not shown in the accompanying drawings): S11. Use a large language model to perform structural and semantic parsing on the statistical data tables of underground municipal infrastructure. Based on the parsing results, identify the data type of each data column, calculate the statistical summary of each data column, and perform natural language semantic annotation of the data columns and extract the relationships between the data columns.

[0028] S12. Construct structured metadata for the statistical data table based on the data type, statistical summary, natural language semantic annotation, and relationships between the data columns. .

[0029] This step involves retrieving the original structured statistical data of underground municipal infrastructure. As input to a large language model, the system generates structured meta-information through deep parsing. To achieve this automation and deep parsing, the system leverages the powerful data understanding capabilities of the LLM (Language Model), guiding a comprehensive analysis of the data's internal structure via prompt design. Specifically, the LLM first analyzes the data tables... A systematic analysis is performed to identify the data type of each column, such as numerical, categorical, or time series data, and statistical summaries, such as mean, standard deviation, uniqueness, and frequency, are calculated accordingly. Based on this, LLM further explores the relationships between columns, such as latent functional dependencies or correlations, and generates precise natural language semantic annotations for the columns. Finally, these analytical results are integrated and output as a structured meta-information set, denoted as... Its structure can be represented as: ; in For data columns The header of the table, For data columns Data types, For data columns Statistical summary, It is a data column Natural language semantic annotation, It is a collection of relationships between columns. By systematically encapsulating key information from data tables, it provides a comprehensive data context for subsequent stages.

[0030] In this embodiment of the invention, step S2 is used to generate initial key information. Step S2 involves providing the statistical data table and its structured meta-information as context to a preset large language model, so that the large language model can perform multi-dimensional analysis of the input context to extract and output a list of initial key information contained in the context. Specifically, this includes the following steps (not shown in the accompanying drawings): S21. Provide the statistical data table and its structured meta-information as context to the pre-defined large language model so that the large language model can perform descriptive analysis, trend analysis, correlation analysis and outlier analysis on the input context, and guide the large language model to extract significant patterns and / or anomalies in the data under each analysis dimension based on the pre-defined analysis constraints and analysis strategies. S22. Output an initial key information list in a predefined format, identifying significant patterns and / or anomalies in the data extracted from each analytical dimension. : ; Each element in the initial key information list They are all initial key information objects, and the form of the initial key information object is as follows: , For natural language descriptions of significant patterns and / or anomalies extracted under one analytical dimension. Data supporting evidence for current significant patterns and / or anomalies.

[0031] Utilizing the data context generated in the previous stage, this step systematically explores the data and generates a broad range of candidate key information. This invention employs a multi-dimensional exploration strategy, using prompt design to guide the LLM (Local Management Module) from a pre-defined analytical perspective to stimulate and constrain the analysis process, ensuring the comprehensiveness and relevance of the generated key information. First, data table T and structured metadata M are provided to the LLM as context. Then, different analytical dimensions are performed on the data, including descriptive analysis, trend analysis, correlation analysis, and outlier analysis. Within each dimension, the LLM is guided to discover significant patterns and / or anomalies in the data, such as pipeline failure trends in specific areas or the aging patterns of certain pipe materials. Each discovery is output strictly according to a predefined structure, forming a key information object containing a natural language description (Idesc) of the discovered pattern and supporting data evidence (Ievidence). Finally, the output is an initial list of key information, denoted as... This list will provide a wealth of raw material for the subsequent iterative refinement of key information.

[0032] In this embodiment of the invention, step S3 is used to implement iterative key information refinement. Step S3 involves analyzing and filtering the initial key information list to obtain an optimized key information list. Specifically, this includes: performing multi-criteria scoring on each initial key information object in the initial key information list according to preset scoring criteria; generating an evaluation vector for each initial key information object; for initial key information objects whose evaluation vectors do not meet the preset scoring criteria, further description is performed and they are merged into other initial key information objects with high content relevance to form new key information objects in the next iteration list; for initial key information objects whose evaluation vectors meet the preset scoring criteria, they are directly retained in the next iteration list; and multi-criteria scoring is performed on the key information objects in the next iteration list according to the preset scoring criteria to enter the next iteration process, until the number of iterations reaches a preset maximum iteration threshold, or all obtained key information objects meet the preset scoring criteria.

[0033] The initial key information list generated in the previous step Building upon this foundation, this step employs an iterative optimization process to refine and enhance the initial list of key information, thereby improving the analytical depth, novelty, and accuracy of each key piece of information. At its core is an iterative framework where LLM alternately acts as both "evaluator" and "refiner" in a loop, simulating the deep thinking process of an expert through self-play and correction.

[0034] In each iteration, LLM first acts as an "evaluator," evaluating the current iteration list based on scoring criteria such as analytical depth, novelty, and factual accuracy. Each key information object in Conduct multi-dimensional evaluation, including This indicates that the current number is the [number]. The LLM iterates through each criterion and assigns a numerical score between 1 and 5, thus forming an evaluation vector. Subsequently, the LLM switches to the role of the "Refiner," receiving the original key information. The LLM algorithm analyzes the evaluation vector and determines whether it meets the preset scoring criteria. Based on the specific scoring feedback, the LLM performs a corresponding optimization operation, including in-depth analysis, filtering, and merging. Key information objects that have passed the scoring are directly retained in the next generation list until the preset maximum number of iterations is reached. Or, it meets a predetermined threshold. Finally, this step outputs a high-quality list of key optimization information that has been validated through multiple rounds, denoted as... This list is the same as the input list. It has the same data structure, but the key information it contains has been significantly improved in quality, removing redundancy and shallow information and deepening the core findings.

[0035] The second core stage aims to construct a mind tree-based diagram. The core objective of this stage is to transform the abstract, textual key information from the previous stage into precise and directly executable visual code. This process employs mind tree reasoning to simulate the systematic thinking process involved in visualization design. The second core stage involves steps S4-S6, which are explained in detail below.

[0036] In this embodiment of the invention, step S4 is used to generate chart types. Step S4, which generates a candidate chart type set for each optimization key information object in the optimization key information list, specifically includes: providing each optimization key information object in the optimization key information list as context to a preset chart building language model, so that the chart building language model can perform intent analysis on the input context and call its internal knowledge base to match a set of chart types that can express the intent analysis results; and using the obtained chart type set as the candidate chart type set for the current optimization key information object.

[0037] In this step, the root node of the mind tree reasoning is the list of key information generated in the first core stage. As input, leveraging the natural language understanding and analysis capabilities of LLM, a collection of chart types containing multiple potential visualization schemes is generated. This process iterates through... For each key information object, its Idesc and Ievidence are used as context to guide the LLM in a deep analysis of its core intent. For example, a key information object describing "the failure trend of a specific pipe material in a certain area" has the core intent of trend analysis. Based on this intent analysis, the LLM will call upon its internal knowledge base to match and propose a set of suitable chart types based on the principle of being able to express the intent analysis results. Finally, a set of candidate chart types is generated for each key information object, denoted as . Here, `chart_type` is a string representing the name of a visual chart. This set forms the root node for mind tree reasoning based on key information, providing a starting point for further exploration of more detailed visualization design.

[0038] In this embodiment of the invention, step S5 is used to implement chart configuration and evaluation. The specific implementation of step S5 is as follows: After generating the chart type, the intermediate nodes of the mind tree reasoning will determine its specific drawing options, and through horizontal comparison and evaluation, transform an abstract chart selection into a clear and executable drawing configuration plan. First, the system sets a single key information object and its corresponding set of candidate chart types. As input, the LLM is guided to make visualization decisions for each chart type, developing a detailed chart configuration plan. This plan needs to clearly define the chart's components, such as data source, axis mapping, color coding, and layer division, to ensure the chart accurately reflects the core information. For example, when analyzing pipeline fault trends, it determines that the 'time' data column should be mapped to the X-axis of the chart, and the 'number of faults' to the Y-axis. The LLM then conducts a cross-sectional evaluation of the chart configuration plans within this set, scoring them based on criteria such as the effectiveness, clarity, and accuracy of information delivery, and ultimately selecting the highest-scoring plan. Ultimately, this step generates an optimal configuration plan for each critical information object.

[0039] In this embodiment of the invention, step S6 is used for executable code generation and verification. The specific implementation of step S6 is as follows: the leaf nodes of the mind tree reasoning transform the optimal configuration plan selected in the previous stage into a piece of executable code used to generate a chart. The system combines a single key information object and the optimal configuration plan... As input, an LLM with code generation capabilities is guided to translate it into standard executable code that follows the syntax of a specific visualization library. .

[0040] Furthermore, after converting the optimal configuration plan for each key information object into executable code for generating charts, the process also includes: providing the executable code as context to a pre-defined code validation language model, allowing the model to iteratively test and optimize the input executable code until the chart file generated by the optimized executable code conforms to the design intent of the corresponding optimal configuration plan in terms of data processing logic, visual element mapping, and core information delivery. Figure 1 To.

[0041] Specifically, to ensure the accuracy and robustness of the output code, this step further introduces an automated testing and verification process. This process compares the actual behavior of the code with the original visual description to ensure that the final generated chart fully matches the design intent in three aspects: data processing logic, visual element mapping, and core information delivery. Specifically, verification includes: checking whether the code accurately executes the data operations defined in the description; verifying whether the code accurately maps the specified data fields to the visual elements of the chart; and evaluating whether the chart clearly highlights its intended core message. Through continuous iteration, the LLM is guided to make targeted corrections to the generated code until it fully meets all preset standards.

[0042] The purpose of the third core stage is report generation based on multi-agent collaboration. The core objective of this stage is to simulate an efficient collaborative workflow, integrating the outputs of the first two stages—high-quality key information and visual code—into a logically coherent report through role division and pipeline operations. This third core stage involves steps S7 and S8, which will be explained in detail below.

[0043] In this embodiment of the invention, step S7, which generates the report text content based on each optimization key information object in the optimization key information list and the optimal configuration plan for each optimization key information object, specifically includes the following steps not shown in the accompanying drawings: S71. Construct a report outline for the data report based on the logical relationships between the various optimization key information objects in the optimization key information list.

[0044] The purpose of this step is to build a report outline based on a pre-defined architect agent, creating a logically clear and fluent macro framework for the final report. Specifically, this involves optimizing the list of key information to a high quality. As input, the pre-defined LLM agent, assigned the role of "architect," will perform a global analysis of all optimization-critical information objects in the list, identifying the logical relationships between them. Based on these relationships, the architect agent will group and sort the optimization-critical information objects, constructing a hierarchical report outline, denoted as... This outline defines the report's chapter divisions, the titles of each chapter, and the sequence of key information objects that each chapter should contain.

[0045] S72. Provide the report outline, the list of key optimization information, and the optimal configuration plan for each key optimization information object to the preset writing agent, so that the writing agent can expand the description of the key optimization information objects contained in each chapter of the report outline into narrative text, generate reference and interpretation text of the chart to be inserted at the preset position according to the optimal configuration plan, mark the position where the chart to be inserted should be embedded with a unique identifier, and output the initial report text containing text content and chart placeholders.

[0046] The purpose of this step is to generate a draft report based on a pre-defined writing agent. Specifically, based on the report outline obtained in step S71, structured key information and charts are filled into fluent natural language paragraphs. The input for this step is the report outline. Optimize the list of key information And the optimal configuration plan generated for each key information object. The pre-defined LLM, assigned the "writer" role, will process each chapter in the outline one by one. Within each chapter, it will optimize the sequence of key information objects contained within that chapter. It weaves the text into a narrative structure with clear transitions. At appropriate points, it inserts references and interpretations of charts and graphs, marking each chart's placement with a unique identifier. The output of this step is an initial report text containing the complete text and chart placeholders, denoted as [example text - likely a placeholder]. .

[0047] S73. Provide the initial report to the preset editing agent so that the editing agent can proofread the initial report, optimize the sentence structure and the logic between paragraphs, and output the final report text.

[0048] The purpose of this step is to review and polish the initial report text based on a pre-defined editing agent. Specifically, it involves comprehensive quality control of the draft report to ensure its accuracy, professionalism, and readability. The input for this step is the initial report text. The pre-defined "editor" role assigned to the LLM will conduct multiple rounds of review of the initial report text. This review includes: correcting grammatical errors and inappropriate wording; optimizing sentence structure and logical flow between paragraphs; and ensuring the overall language style conforms to pre-defined professional standards. The output of this step is a meticulously polished and proofread final report text, denoted as […]. .

[0049] In this embodiment of the invention, step S8, inserting a chart file generated by calling the executable code corresponding to each optimization key information object into the text content of the report, specifically includes: parsing the final report text to identify chart placeholders in the final report text; when a chart placeholder is identified, calling the executable code corresponding to the chart to be inserted at the current position to generate a chart file; and replacing the corresponding chart placeholder with the generated chart file to obtain the final analysis report.

[0050] The purpose of this step is to assemble and render the report, combining text content with visual charts into a complete document. The input for this step is the final report text. and the collection of all executable code This step involves parsing... When a chart placeholder is detected, the corresponding executable code is extracted. The executable code is run to generate a chart file (i.e., an image file) for the chart location. Then, the chart placeholders in the text are replaced with the generated chart file. The final output of this step is a complete analysis report. Furthermore, the final analysis report can be saved in various standard document formats as needed and delivered directly to the end user.

[0051] This invention provides a method for generating municipal infrastructure data reports based on a large language model. Through a structured and controllable multi-stage workflow, it systematically overcomes the core shortcomings of existing large language models in report generation tasks, such as inconsistency between text and graphics, insufficient analytical depth, and unstable quality. By introducing a key information extraction mechanism based on hierarchical reasoning, this invention can simulate the deep thinking of expert analysts, refining shallow data descriptions into more valuable deep patterns, thereby significantly improving the analytical depth of the report. Simultaneously, by decoupling key information extraction from chart construction and utilizing mind tree reasoning to tailor precise visualization blueprints for each key piece of information, this invention ensures consistency between text and graphics in the final report. Furthermore, this invention employs a transparent pipeline involving collaboration among architects, writers, and editors, ensuring that the report maintains a consistently high level of logical structure, linguistic professionalism, and factual accuracy. Ultimately, this series of designs collectively achieves a fully automated solution from raw data to a high-quality analytical report, significantly improving report writing efficiency while ensuring its depth, reliability, and professionalism.

[0052] To more clearly illustrate the technical solution of this invention, a specific implementation case will be used as an example below. This implementation case aims to analyze the fault statistics of pipelines of different materials and service lives in multiple urban areas of a city, and automatically generate a summary report on the health status of the pipeline network.

[0053] Figure 2 A schematic diagram of the overall system architecture of this specific embodiment is shown. Initial input data: The original structured underground pipeline statistics (data table) received by this embodiment. This is a tabular underground pipeline fault record form, the contents of which are as follows: Table 1 Pipeline Fault Record Table Report Month urban area Pipe Material Pipeline service life Number of failures 2025-01 Chengdong District cast iron 22 15 2025-01 Chengxi District PVC 8 2 2025-01 Chengbei District Ductile iron 15 5 2025-02 Chengdong District cast iron 22 18 2025-02 Chengxi District PVC 8 3 2025-02 Chengbei District Ductile iron 15 6 2025-03 Chengdong District cast iron 22 25 2025-03 Chengxi District PVC 8 2 2025-03 Chengbei District Ductile iron 15 5 The first core stage: extracting key information based on hierarchical reasoning.

[0054] Step a1: Data parsing and metadata generation.

[0055] Enter the above data table LLM for data tables A systematic analysis was conducted. The model first identified five columns and performed type inference and statistical summarization for each column. Specifically, "Report Month," "Urban Area," and "Pipe Material" were identified as categorical enumeration types; "Pipe Service Life" and "Number of Failures" were identified as continuous numerical types. Subsequently, the mean, maximum / minimum values ​​of the numerical columns were calculated, and the unique values ​​of the categorical columns were listed. Based on this, the model further explored the relationships between columns. Based on domain knowledge of pipeline network operation and maintenance, a potential strong correlation was marked in the relationship set F: "Number of Failures" may have a significant correlation with "Pipe Service Life" and "Pipe Material." A structured metadata set was generated. Its content structure is as follows: {“columns”:[ {"h":"Report month","type":"Categorical","S":{"unique_values":["2025-01","2025-02","2025-03"]},"Desc":"Records the month in which the failure occurred, an ordered time-categorical variable"}, {"h":"urban area","type":"Categorical","S":{"unique_values":["Eastern District","Western District","Northern District"]},"Desc":"The administrative region where the fault occurred, an unordered geographical classification variable"}, {"h":"Pipe Material","type":"Categorical","S":{"unique_values":["Cast Iron","PVC","Ductile Cast Iron"]},"Desc":"Primary Manufacturing Material of the Pipe, a Key Engineering Property Classification"}, {"h":"Pipeline service life","type":"Numerical","S":{"min":8,"max":22,"avg":15.0},"Desc":"Number of years since the pipeline was put into use (unit: years), which is the core indicator for assessing pipeline aging"} {"h":"Number of failures","type":"Numerical","S":{"min":2,"max":25,"avg":9.0},"Desc":"Total number of failure events recorded on the pipes of the corresponding month, region, and material"}], "relations":[{"description":"Based on knowledge of pipeline network operation and maintenance, 'number of failures' is likely positively correlated with 'pipeline service life' and is significantly affected by 'pipeline material'.","involved_columns":["number of failures", "pipeline service life", "pipeline material"],"type":"PotentialCorrelation"}]}.

[0056] Step a2: Initial key information generation.

[0057] Input data table and structured meta-information The model utilizes this comprehensive data context to generate preliminary, extensive candidate key information from multiple pre-defined analytical dimensions. This generates an initial list of key information. It contains several structured key information objects: { "desc": "Among all pipe materials, cast iron pipes have the highest cumulative number of failures." "evidence": "Groups and sums the 'number of failures' by 'pipe material'." "desc": "The number of monthly faults in the Chengdong District shows a clear upward trend", "evidence": The change in "number of faults" with "reporting month" after 'filtering district' = 'Chengdong District'". "desc": "The number of pipe failures involving PVC materials has remained at an extremely low level", "evidence": "Statistical summary of the number of failures after filtering 'pipe material' = 'PVC'". “desc”: “Old pipelines with a service life of more than 20 years are the main source of failures”; “evidence”: “Compare the total number of failures between the two sets of data: ‘pipeline service life’>20’ and ‘pipeline service life’<=20”}.

[0058] Step a3: Iterative refinement of key information.

[0059] Enter the initial list of key information LLM then enters an iterative closed loop of "generation-evaluation-optimization". In the first iteration, the "evaluator" role considers key information I1 and I4 to be highly related but relatively superficial, while key information I2 reveals dynamic trends with greater early warning value. Based on this feedback, the "refiner" role integrates and deepens I1, I2, and I4, and semantically enhances I3 to highlight its positive role as a "stabilizer".

[0060] After a round of refinement, a high-quality list of key information was generated. : { (Depend on (Deepening the Merger): "desc": "Although old cast iron pipes with over 20 years of service are the main and persistent source of failures throughout the city, a more alarming new development is the accelerating growth in the number of failures in the eastern urban area. This indicates that the static high-risk stockpile and the dynamic regional deterioration problem are overlapping, constituting the most critical safety risk point of the current pipeline network." "evidence": "A comprehensive analysis of the entire dataset was conducted: grouping and aggregating by 'pipeline material' and 'pipeline service life' to identify the main sources of failures; and time series trend analysis was performed on the data where 'urban area' = 'eastern urban area' to verify its growth." (Depend on (Further details): "desc": "In stark contrast to the high-risk old pipelines, new materials such as PVC pipelines have demonstrated extremely high reliability and stability in all areas. Their failure rate has remained consistently low, not only proving the superiority of the new material but also forming a 'stable foundation' for the overall health of the pipeline network, providing a buffer for risk management." "evidence": "Statistical summary (including maximum and average values) of 'failure rate' data after filtering 'pipeline material' = 'PVC', and a horizontal comparison with failure data of other materials."

[0061] The second core stage: constructing diagrams based on mind trees.

[0062] This invention uses key information objects This example illustrates the process of constructing chart specifications. The input is a high-quality optimized key information object. Mind tree reasoning: Step b1: Root node - chart type generation: Key information from LLM analysis The core narrative intent comprises two levels: 1. Composition Analysis: Shows the contribution of different regions to the total number of failures (part-whole relationship).

[0063] 2. Trend Analysis: Highlighting the dynamic growth trend of the specific area "Eastern District".

[0064] To represent both levels simultaneously, LLM matches and proposes a set of the most suitable candidate diagram types from its internal knowledge base: Candidate 1: Combination Chart (Bar Chart + Line Chart): Use a stacked bar chart to represent the total number of faults in each district, and use a line chart to highlight the growth trend in the eastern district, showing strong correlation.

[0065] Candidate 2: Stacked area map: It can well show the composition of the number of faults in each area over time, but it is not eye-catching enough to highlight the trend of a certain area.

[0066] Step b2: Intermediate Nodes - Chart Configuration and Evaluation: LLM refines the two schemes. For the combined graph scheme, the configuration plan is as follows: X-axis: Map the 'reporting month' to the X-axis as a common time dimension.

[0067] Y-axis: Maps 'number of failures' to the Y-axis.

[0068] Column layer: Chart type: Stacked column chart; Data: Data from all districts; Color coding: Map the 'District' field to colors to distinguish different stacked sections.

[0069] Line layer: Chart type: Line chart; Data: Filter only data where 'urban area' is 'Chengdong District'; Visual markers: Add data point markers to the line to enhance readability.

[0070] For the stacked area map scheme, the configuration plan is as follows:… Subsequently, LLM initiated the evaluation module to compare the generated combined map specifications with the stacked area map specifications. The evaluation conclusion was that the combined map, through independent polylines, more clearly emphasized the core key information of "accelerated growth in the eastern part of the city," resulting in higher information delivery efficiency. Therefore, this scheme was selected as the optimal visualization scheme.

[0071] Step b3: Leaf Nodes - Visual Code Generation and Verification: LLM translates the above visualization scheme into Python visualization code. Then, LLM tests and verifies the generated visualization code to ensure the correctness of the underlying analysis logic.

[0072] The third core phase: report generation based on multi-agent collaboration.

[0073] Step c1: Architect Agent: Report Structure Construction: Input is a high-quality list of key information. (Include and The "architect" agent analyzes the logical relationship between two key pieces of information: It reveals the core risks, which are the focus of the problem; It provides a stable reference point, serving as background and comparison. Therefore, it employs a narrative structure that prioritizes problem focus and supplements it with stable factors. The output generates a structured report outline. : Report Title: Quarterly Pipeline Health Status Analysis Report "chapter":[ {"Chapter Title":"Core Risk: The Combined Effect of Risks from Outdated Pipes and Key Areas","Content":[{"Type":"Describes Key Information","Key Information ID":" {"Type":"Embedded Chart","Key Information ID":"} ","Placeholder":"[CHART_FOR_ ]"}]}, {"Chapter Title":"Stable Foundation: Reliability Performance of New Material Pipelines","Content":[{"Type":"Describes Key Information","Key Information ID":"} ... "— """——————————— " " " " " " " " " " " " " " " " " " " " " " " "" " ""“ " " " "}]}, {"Chapter Title":"Conclusions and Recommendations","Content":[{"Type":"Summary","Abstract":"Summarizes the core risks and stabilizing factors, and recommends a focused investigation of the Chengdong District."}]}]} Step c2: Writing the agent - First draft generation: Input the outline, a list of key information, and chart specifications. The "Writer" agent then uses the outline to extract the key information. and Fill with fluent text paragraphs and insert chart placeholders at specified locations.

[0074] Step c3: Editing the agent - Review and polishing: The "editing" AI agent performs multiple rounds of review on the initial draft to improve its professionalism and accuracy. It performs the following optimizations: Quantitative description: The "accelerated growth trend" in the initial draft is specified as "the number of monthly failures increased from 15 in January to 25 in March, an increase of 67%".

[0075] Technical terminology: "Very reliable" is optimized to "Exhibits extremely high reliability and stability".

[0076] Logical connection: Ensure a natural and smooth transition between chapters.

[0077] Assembler - Rendering and Final Output: This is the final step in report generation, completed by an automated program: 1. Parse the final text. When a chart placeholder is detected, extract the corresponding chart visualization code based on the placeholder's ID.

[0078] 2. Execute the visualization code to generate image files.

[0079] 3. Replace the placeholders in the text with this generated image file.

[0080] The final output is a complete, visually appealing analysis report. The report is logically clear and written in professional language. Its core sections are illustrated below: {Core risk: The combined effect of outdated pipe materials and risks in key areas;} During this reporting period, the overall health of the city's pipeline network exhibited significant structural risks. Data shows that old cast iron pipes, in service for over 20 years, remain the primary source of failures throughout the city. However, a more concerning trend is the accelerated increase in pipeline failures in the Chengdong District, with monthly failures rising from 15 in January to 25 in March, an increase of 67%. This trend, coupled with the issue of aging pipe materials, strongly indicates that the health of the pipeline network in this area is rapidly deteriorating, urgently requiring priority inspection and maintenance.

[0081] [The rendered image] Stable foundation: Reliability performance of new material pipelines; In stark contrast to the high-risk aging pipelines, pipelines made of new materials, such as PVC, have demonstrated extremely high reliability and stability in all areas. The number of failures has remained consistently low, proving not only the superiority of the new materials but also forming a stable foundation for the overall health of the pipeline network, providing a buffer for risk management.

[0082] [The rendered image] }

[0083] Through all the above steps, this invention successfully transforms a raw underground pipeline network data table into a high-quality analysis report that is logically rigorous, insightful in its analysis, and completely consistent with the graphics and text.

[0084] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0085] Another embodiment of the present invention provides a municipal infrastructure data report generation device based on a large language model, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the municipal infrastructure data report generation method based on a large language model as described in the above embodiment.

[0086] For the device implementation, since its process of generating municipal infrastructure data reports based on a large language model is basically similar to that of the method implementation, it is described in a relatively simple way. For relevant details, please refer to the description of the method implementation. It also has the corresponding technical effects.

[0087] Another embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the municipal infrastructure data report generation method based on a large language model as described in the above embodiment.

[0088] The municipal infrastructure data report generation method and device based on a large language model provided in this invention realizes the full-process automation of generating high-quality analytical reports from raw data to high-quality analytical reports with both analytical text and visual charts by introducing a key information extraction stage based on hierarchical reasoning, a chart construction stage based on mind tree, and a report generation stage based on multi-agent collaboration. This significantly improves the efficiency, depth, and reliability of data analysis and report writing.

[0089] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, any of the claimed embodiments can be used in any combination.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating municipal infrastructure data reports based on a large language model, characterized in that, The method includes: The statistical data tables of underground municipal infrastructure are subjected to structural and semantic parsing, and the structured meta-information of the statistical data tables is constructed based on the parsing results; The statistical data table and its structured meta-information are provided as context to the pre-defined large language model, so that the large language model can perform multi-dimensional analysis on the input context to extract and output the initial list of key information contained in the context. The initial list of key information is analyzed and filtered to obtain an optimized list of key information; For each optimization key information object in the optimization key information list, generate a set of candidate chart types; For each candidate chart type in the candidate chart type set for each optimization key information object, a drawing configuration plan is constructed, and the drawing configuration plan corresponding to each candidate chart type is scored, so as to select the drawing configuration plan with the highest score as the optimal configuration plan for the corresponding optimization key information object. The optimal configuration plan for each key information object is transformed into executable code for generating charts; The report's text content is generated based on each optimization key information object in the optimization key information list and the optimal configuration plan for each optimization key information object; Insert chart files generated by calling the executable code corresponding to each optimization key information object into the text content of the report to form a final analysis report containing text content and chart files.

2. The method according to claim 1, characterized in that, The statistical data tables of underground municipal infrastructure are subjected to structural and semantic parsing, and the structured metadata of the statistical data tables is constructed based on the parsing results, including: A large language model is used to perform structural and semantic parsing of statistical data tables of underground municipal infrastructure. Based on the parsing results, the data type of each data column is identified, the statistical summary of each data column is calculated, and natural language semantic annotation of the data columns and the relationship between the data columns are extracted. Structured metadata for statistical data tables is constructed based on the data type, statistical summary, natural language semantic annotation, and relationships between data columns. : ; in For data columns The header of the table, For data columns Data types, For data columns Statistical summary, It is a data column Natural language semantic annotation, It is a set of relationships between columns.

3. The method according to claim 1, characterized in that, The statistical data table and its structured meta-information are provided as context to the pre-defined large language model, allowing the model to perform multi-dimensional analysis of the input context to extract and output an initial list of key information contained within the context, including: The statistical data table and its structured meta-information are provided as context to the pre-defined large language model, so that the large language model can perform descriptive analysis, trend analysis, correlation analysis and outlier analysis on the input context, and guide the large language model to extract significant patterns and / or anomalies in the data under each analysis dimension based on the pre-defined analysis constraints and analysis strategies. The significant patterns and / or anomalies extracted from the data across various analytical dimensions will be formatted according to a predefined structure to output an initial list of key information. : ; Each element in the initial key information list They are all initial key information objects, and the form of the initial key information object is as follows: , For natural language descriptions of significant patterns and / or anomalies extracted under one analytical dimension. Data supporting evidence for current significant patterns and / or anomalies.

4. The method according to claim 1, characterized in that, The initial list of key information is analyzed and filtered to obtain an optimized list of key information, including: Each initial key information object in the initial key information list is scored according to a preset scoring criterion, generating an evaluation vector for each initial key information object. For initial key information objects whose evaluation vectors do not meet the preset scoring criteria, they are further described and merged into other initial key information objects with high content relevance to form new key information objects in the next iteration list. For initial key information objects whose evaluation vectors meet the preset scoring criteria, they are directly retained in the next iteration list. The key information objects in the next iteration list are then scored according to the preset scoring criterion to enter the next iteration process, until the number of iterations reaches the preset maximum iteration threshold, or all the obtained key information objects meet the preset scoring criteria.

5. The method according to claim 1, characterized in that, For each optimization key information object in the optimization key information list, generate a set of candidate chart types, including: Each optimization key information object in the optimization key information list is provided as context to the preset chart building language model, so that the chart building language model can perform intent analysis on the input context and call its internal knowledge base to match a set of chart types that can express the intent analysis results. The resulting set of chart types is used as the candidate chart type set for the current optimization key information object.

6. The method according to claim 1, characterized in that, After converting the optimal configuration plan for each key information object into executable code for generating charts, the following is also included: The executable code is provided as context to the pre-defined code validation language model, which then iteratively tests and optimizes the input executable code until the chart file generated by the optimized executable code is consistent with the design intent of the corresponding optimal configuration plan in terms of data processing logic, visual element mapping, and core information delivery.

7. The method according to claim 1, characterized in that, The report's text content is generated based on each optimization key information object in the optimization key information list and the optimal configuration plan for each optimization key information object, including: The outline of the data report is constructed based on the logical relationships between the various key optimization information objects in the list of key optimization information. The report outline, the list of key optimization information, and the optimal configuration plan for each key optimization information object are provided to the pre-defined writing agent. The writing agent expands the description of the key optimization information objects contained in each chapter of the report outline into narrative text, generates reference and interpretation text for the chart to be inserted at the pre-defined position according to the optimal configuration plan, marks the position where the chart to be inserted should be embedded with a unique identifier, and outputs the initial report text containing text content and chart placeholders. The initial report is provided to a pre-defined editing agent, which then proofreads the initial report, optimizes sentence structure and paragraph logic, and outputs the final report text.

8. The method according to claim 7, characterized in that, Insert chart files generated by calling the executable code corresponding to each optimization key information object into the report text, including: The final report text is parsed to identify chart placeholders. When a chart placeholder is identified, the executable code corresponding to the chart to be inserted at the current position is called to generate a chart file. The generated chart file replaces the corresponding chart placeholder, resulting in the final analysis report.

9. A municipal infrastructure data report generation device based on a large language model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Personalized complex report generation method based on multi-agent system

    CN118569237A

  • Data report generation method, electronic device, storage medium and computer program product

    CN119202140A

  • Data analysis method and device, storage medium and electronic equipment

    CN119938600A

  • Large model-based data analysis report generation method, apparatus and device, and storage medium

    CN120823025A

  • Method for generating textual descriptions for a data chart and enhancing navigability on graphical interfaces

    EP4475009A1