Database information extraction and data analysis method and database intelligent analysis platform based on large language agent

Through the database information extraction method of multimodal prompts and domain knowledge fusion, the problem of semantic disassembly of large language models and insufficient data table understanding in database interaction is solved, efficient and accurate SQL generation and comprehensive analysis are achieved, and cross-domain applications are supported.

CN120277083APending Publication Date: 2025-07-08SHANGHAI PUDONG ZHONGRUAN TECH DEV CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510294296.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, database interactions based on large language models have problems such as semantic disassembly and domain knowledge, insufficient data table understanding, and lack of multimodal comprehensive analysis, resulting in the inability to accurately disassemble complex instructions, frequent SQL statement errors and fragmented analysis results.

Method used

Through multimodal prompt driving, domain knowledge fusion and dynamic feedback mechanism, the entire process automation from natural language instructions to structured analysis is realized, including semantic understanding and instruction disassembly, SQL generation and verification, adaptive data analysis and visualization, and multi-condition data analysis.

Benefits of technology

It realizes accurate disassembly of complex instructions, reduces the error rate of SQL statements, comprehensive improvement of analysis results, shortens the task time to minutes, and supports cross-domain expansion and multi-scenario decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277083A_ABST
    Figure CN120277083A_ABST
Patent Text Reader

Abstract

The invention discloses a database information extraction and data analysis method based on a large language agent, and belongs to the crossing field of database management and artificial intelligence. In order to solve the problems that traditional database query depends on manual SQL writing, and complex analysis efficiency is low, the method constructs an intelligent agent through a large language model (LLM) and a prompt engineering technology, and end-to-end conversion from a natural language instruction to structured data analysis is achieved. The method specifically comprises the following steps: 1) disassembling an instruction and associating database metadata; 2) generating a precise SQL query; 3) automatically adapting a visual chart according to a query result, and generating an analysis report; and 4) integrating data conclusions in different scenes. According to the method, in a judicial operation situation judgment scene test, the complex query generation accuracy is improved to 80%, meanwhile, the method can be applied to data analysis platforms in the fields of science, finance, medical treatment and the like, and non-technical users are supported to rapidly complete the whole process of'instruction-query-visualization-decision '.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of database management and artificial intelligence, and specifically relates to a method for database information extraction and data analysis based on large - language agents and a database intelligent analysis platform. This method realizes the intelligent conversion from natural - language instructions to structured queries, dynamic analysis of multi - modal data, and automated report generation by integrating large - language models (LLMs), domain - knowledge ontologies, and database metadata, aiming to solve the interaction barrier between non - technical users and complex database systems and improve the efficiency and accuracy of data - driven decision - making. Background Art

[0002] With the rapid development of large - language model (LLM) technology, natural - language - based database interaction has become a research hotspot. Traditional database queries rely on manually writing SQL statements, which have problems such as high technical thresholds and low efficiency in complex analysis. In the prior art, rule - or template - based NL2SQL (Natural Language to SQL) methods can simplify the generation of some queries, but their generalization ability is poor, it is difficult to adapt to dynamic business requirements, and they cannot automatically associate multiple tables to generate JOIN statements.

[0003] In recent years, agent technology based on large - language models has provided new ideas for database interaction. However, there are still significant deficiencies in directly applying existing LLMs to database scenarios: Disconnection between semantic decomposition and domain knowledge: General prompting engineering lacks guidance for domain - knowledge ontologies, resulting in complex instructions being unable to be accurately decomposed into analysis objectives, constraint conditions, and visualization requirements. Insufficient understanding of data tables: LLMs have limited perception of database metadata (such as table structures, primary and foreign key constraints), and the generated SQL statements often have syntax errors (such as missing GROUP BY clauses) or logical errors (such as incorrect table associations). Lack of multi - modal comprehensive analysis: Existing methods are difficult to automatically adapt visualization schemes based on data characteristics and lack automated support for multi - condition hypothesis analysis, resulting in fragmented analysis results. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention proposes a method for database information extraction and data analysis based on large - language agents, which realizes the full - process automation from natural - language instructions to structured analysis through multi - modal prompt driving, domain - knowledge fusion, and dynamic feedback mechanisms. The technical solution of the present invention is as follows: A method for database information extraction and data analysis based on large - language agents, characterized by including the following steps: S1. Semantic understanding and instruction decomposition: Through dynamic prompt templates and domain - knowledge ontology libraries, decompose the natural - language instructions input by the user into a three - element structure including "analysis objective, constraint condition, visualization requirement", and realize semantic mapping by associating database metadata; S2. SQL Generation and Verification: Inject database metadata based on the pattern-aware prompting mechanism, detect and correct syntax and logical errors through the syntax verification feedback mechanism, and output SQL queries that conform to the syntax specification; S3. Adaptive Data Analysis and Visualization: Automatically adapt the visualization chart type according to the data characteristics of the query results, and use the large language model to generate a semantic analysis report with key trend annotations; S4. Multi-Condition Data Analysis: Support users to append constraint conditions through natural language, dynamically update the SQL statement, combine the comparison summary technology, quantify the data differences in different scenarios, and generate comprehensive decision-making suggestions.

[0005] The core of the present invention includes the following modules:

[0006] 1. Semantic Understanding and Instruction Decomposition Module

[0007] Dynamic Prompt Template Design: Construct a multi-level prompt template to guide the LLM to parse the user's intention. Taking the analysis of the judicial trial situation as an example, given an analysis instruction, the first-level template decomposes the instruction into a JSON format of triples, such as:

[0008] {

[0009] "Analysis Target":...,

[0010] "Constraint Conditions": {"Time":..., "Elements":..., "Operations":...},

[0011] "Visualization Requirements": "Line Chart"

[0012] }

[0013] Domain Knowledge Ontology Library Integration: Define domain entities (such as in the analysis of the judicial trial situation, the entities to be defined are "indicators" and "case elements"), entity constraints, entity relationships (such as the calculation basis between indicators and elements), and synonym tables, and eliminate semantic ambiguities through entity linking technology.

[0014] 2. SQL Generation and Verification Module

[0015] Pattern-Aware Prompting Mechanism: Inject database metadata during the SQL generation stage. The following is an example of a prompt template:

[0016] [Table Metadata] - JSON Structure

[0017] [Query Statement] - Natural Language

[0018] [Condition Statement] - JSON Structure

[0018] [Output] - SQL Statement

[0019] Syntax checking - feedback iteration mechanism: Use an SQL parser (such as Apache Calcite) to detect syntax errors. If an error is found (such as a missing JOIN condition), trigger a correction prompt template to guide the LLM to regenerate.

[0020] 3. Adaptive data analysis and visualization module

[0021] Data feature perception strategy: Dynamically adjust the matching visualization rules according to the query results:

[0022] Time-series data → Line chart (X-axis is time, Y-axis is the aggregated value)

[0023] Classification comparison → Stacked bar chart (Colors distinguish categories)

[0024] Geographical association data → Heat map (Based on latitude and longitude fields)

[0025] Semantic-enhanced analysis report: Use the LLM to generate natural language conclusions.

[0026] 4. Multi-condition comprehensive decision-making module

[0027] Analysis workflow: Support users to append conditions in natural language, and the intelligent agent automatically associates the elasticity coefficient of historical data and updates the SQL statement.

[0028] Comparison summary generation: Focus on the differences in multi-scenario results and output quantitative conclusions. The method of the present invention supports cross-domain expansion, including the judicial, financial, medical, or scientific fields, where: The judicial field adapts to the analysis rules of case types, court levels, and time dimensions; The financial field adapts to the analysis of the relevance of financial indicators and the prediction of market trends; The medical field adapts to the classification of patient groups and the comparison of treatment effects. At the same time, the present invention also provides a database intelligent analysis platform, which is characterized in that it integrates the method as described above and provides the following functional modules: Natural language interaction interface, supporting text, voice, and multi-modal input; Dynamic knowledge ontology management module, supporting online editing and version control of domain ontologies; Multi-scenario comparison dashboard, which displays the difference analysis results and decision-making suggestions in real time.

[0029] The present invention realizes the following advantages through the above technical solutions: Accurate semantic mapping: Combining the domain ontology library, the efficiency of decomposing complex instructions is increased to 80%.

[0030] Efficient SQL generation: The syntax error rate is reduced to less than 20%.

[0031] Intelligent analysis closed-loop: It supports end-to-end processing from instruction parsing, data extraction, visualization to report generation, and the average task time consumption is shortened to the minute level. Description of the Drawings Figure 1 It is the implementation architecture diagram of the present invention Detailed Implementation Manner The technical solutions of the present invention will be described in detail below in conjunction with the drawings and embodiments, but the protection scope of the present invention should not be limited thereby. The implementation manner of the present invention can be deployed on a database server or a cloud platform through a software program, or integrated into a data analysis system as middleware. The implementation environment needs to include large language models (such as GPT-4, Claude, etc.), database management systems (such as MySQL, PostgreSQL), SQL parsers (such as Apache Calcite), and visualization engines (such as ECharts, Tableau). As Figure 1 shown, the implementation architecture of the present method includes the following core modules: User interaction interface: Receive natural language instructions (such as "Analyze the quarterly case ratio trend by court level"), and support text or voice input. Semantic understanding and instruction decomposition module: Based on dynamic prompt templates and domain knowledge ontologies, decompose the instructions into structured triples (analysis objectives, constraint conditions, visualization requirements). SQL generation and verification module: Generate SQL queries in combination with database metadata, and iteratively correct errors through a syntax verification feedback mechanism. Data analysis and visualization module: Execute SQL queries, match visualization rules according to data characteristics, and generate semantic analysis reports. Multi-condition comprehensive decision-making module: Support users to append conditions (such as "Only analyze data in 2023"), dynamically adjust the analysis logic and generate comparison summaries.

[0032] Taking the analysis of the trial situation of judicial cases as an example, the user inputs the instruction: "Please analyze the case ratio and change trend of each quarter by court level?". The execution process of the intelligent agent is as follows:

[0033] Instruction decomposition:

[0034] {

[0035] "Analysis objective": "Case ratio --> Number of cases / Number of pieces"

[0036] "Constraint conditions": {"Court level": ["High People's Court", "Intermediate People's Court", "Basic People's Court"]}

[0037] "Visualization requirement": "Line chart"

[0038] }

[0039] SQL Generation:

[0040] SELECT

[0041] CASE

[0042] WHEN time IN (1, 2, 3) THEN 'Q1'

[0043] WHEN time IN (4, 5, 6) THEN 'Q2'

[0044] WHEN time IN (7, 8, 9) THEN 'Q3'

[0045] WHEN time IN (10, 11, 12) THEN 'Q4'

[0046] END AS quarter,

[0047] COUNT(number) AS case count,

[0048] SUM(quantity) AS quantity,

[0049] ROUND(COUNT(number) * 1.0 / SUM(quantity), 3) AS 'case - quantity ratio'

[0050] FROM

[0051] case_ratio

[0052] WHERE

[0053] court level = 'Higher People's Court'

[0054] GROUP BY

[0055] CASE

[0056] WHEN time IN (1, 2, 3) THEN 'Q1'

[0057] WHEN time IN (4, 5, 6) THEN 'Q2'

[0058] WHEN time IN (7, 8, 9) THEN 'Q3'

[0059] WHEN time IN (10, 11, 12) THEN 'Q4'

[0060] END

[0061] ORDER BY

[0062] quarter;

[0063] Visualization and Reporting: In the quarterly case-ratio analysis at the court level, courts at different levels exhibit their respective trends and characteristics. At the Higher People's Courts, the case-ratio shows a slight downward trend, decreasing from the highest value of 0.182 in the first quarter to the lowest value of 0.175 in the fourth quarter. This trend reflects either a relative decrease in the number of cases or an improvement in processing efficiency. The overall data fluctuates within a limited range, indicating stability in case handling. Intermediate People's Courts, on the other hand, show a gradually increasing trend, rising from 0.182 in the first quarter to 0.196 in the fourth quarter, suggesting that the growth rate of the number of cases is slightly faster than that of matters, reflecting a gradual improvement in the court's handling capacity and efficiency. Compared with Higher and Intermediate Courts, the case-ratio of Basic People's Courts significantly drops to 0.164 in the third quarter and then rebounds to the highest value of 0.184 in the fourth quarter, showing significant volatility, suggesting possible periodic efficiency changes or ratio adjustments. These differential trends not only reflect the differences in case processing efficiency and rhythm faced by different courts but also provide data support for examining institutional and management improvements at different levels of the court.

[0064] This method can be deployed as database middleware or SaaS services. Through the judicial trial situation analysis test, the accuracy rate of complex analysis reaches 80%. The above embodiments are only for illustrative purposes. Those skilled in the art can adjust the module configuration, rule thresholds, or visualization schemes according to actual needs, and such variations are all within the scope of the claims of the present invention.

Claims

1. A method for database information extraction and data analysis based on large language agents, characterized in that, It includes the following steps: S1. Semantic Understanding and Instruction Decomposition: Through dynamic prompt templates and domain knowledge ontologies, decompose the natural language instructions input by the user into a three-element structure including "analysis target, constraint conditions, and visualization requirements", and associate database metadata to achieve semantic mapping; S2. SQL Generation and Verification: Inject database metadata based on the pattern-aware prompt mechanism, detect and correct syntax and logical errors through the syntax verification feedback mechanism, and output a SQL query that conforms to the syntax specification; S3. Adaptive Data Analysis and Visualization: Automatically adapt the visualization chart type according to the data characteristics of the query results, and use the large language model to generate a semantic analysis report with key trend annotations; S4. Multi-Condition Data Analysis: Support users to append constraint conditions through natural language, dynamically update the SOL statement, combine comparison summary technology, quantify data differences in different scenarios, and generate comprehensive decision-making suggestions.

2. The method according to claim 1, characterized in that, The dynamic prompt template in step S1 is a multi-level structure. The first-level template converts the user instruction into a JSON-format triple, including: Analysis target: {target type}, {aggregation method}, {associated table}; Constraint conditions: {field filtering}, {time range}, {classification limit}; Visualization requirements: {chart type}, {dimension axis}, {color coding}; The domain knowledge ontology integrates domain entities, entity relationships, synonym tables, and constraint rules, and eliminates semantic ambiguities through entity linking technology.

3. The method according to claim 2, characterized in that The domain knowledge ontology is constructed in the following ways: - Define domain entities and relationships to form an entity-attribute-value triple; - Establish a synonym mapping table to associate diverse terms expressed by users with standard database fields; - Integrate a business rule library to restrict the generation of illegal condition combinations and conflicting queries.

4. The method according to claim 1, wherein The pattern-aware prompt mechanism in step (2) is specifically as follows: In the SQL generation stage, input database metadata, natural language query statements, and structured conditional statements into the large language model to generate a SQL statement containing multi-table associations and complex conditions; the syntax verification feedback mechanism detects syntax errors through a SQL parser and triggers the large language model to iteratively optimize the SQL statement through a correction prompt template.

5. The method according to claim 4, wherein Step S2 specifically includes: S2.1 Construct a syntax verification-feedback iteration mechanism to verify the syntax correctness of the generated statement through a SQL parser; S2.2 If a syntax error is detected, trigger the correction prompt template, and combine the error information to guide the large language model to regenerate the SQL; S2.3 Support multi-table association queries, automatically identify JOIN conditions, and generate a grouped statement containing aggregation functions.

6. The method according to claim 1, characterized in that, The visualization adaptation in step S3 includes: S3.1 Match preset visualization rules according to the data distribution characteristics, generate a line chart for time-series data, and a bar chart for categorical data; S3.2 Use the large language model to generate natural language analysis conclusions, and automatically label abnormal points and key trends; S3.3 Generate a programming-based drawing scheme to support the output of multiple chart combinations.

7. The method according to claim 1, characterized in that Step S4 specifically includes: S4.1 Construct a hypothesis analysis workflow to generate predictive queries by appending constraint conditions through natural language; S4.2 Use the large prediction model component to analyze and prompt, and perform data analysis on the data chart and analysis request under a single condition; S4.3 Combine the domain knowledge base to recommend multi-condition comparative analysis, dynamically generate analysis prompts through the summary generation method, and form the final comprehensive analysis text.

8. The method according to claim 1, wherein After the visualization chart is generated, extract at least one of the following semantic analysis contents through the large language model: Data trend description, including rising, falling or fluctuating characteristics; Outlier identification and speculation on the cause; Summary of key conclusions based on statistical indicators (such as mean, variance, slope).

9. The method according to claim 1, characterized in that, The method is deployed as a database middleware, including: S10.1 Provide an API interface to receive user instructions and return analysis results; S10.2 Support the adaptation connection with databases such as MySQL and PostgreSQL; S10.3 Integrate the user feedback module to collect error cases to optimize the prompt template and ontology library.

10. The method according to claim 1, wherein The multi-condition comprehensive analysis in step (4) includes: Dynamic condition expansion: Parse the natural language constraint conditions appended by the user and update the WHERE or HAVING clause of SQL; Scenario comparison analysis: Execute multi-version queries, extract differential indicators and calculate the percentage change; Elastic coefficient adjustment: Dynamically adjust the query condition threshold according to the historical data distribution to avoid deviations caused by data sparsity.

11. The method according to claim 1, wherein When generating the comparison summary, enhance the credibility of the conclusion through the following methods: Add significance test annotations; Mark the data confidence interval; Associate domain knowledge to explain the cause of differences.

12. The method according to claim 1, wherein The method supports cross-domain expansion, including the judicial, financial, medical or scientific fields, where: The judicial field adapts to the analysis rules of case types, court levels and time dimensions; The financial field adapts to the analysis of the relevance of financial indicators and the prediction of market trends; The medical field adapts to the classification of patient groups and the comparison of treatment effects.

13. A database intelligent analysis platform, characterized in that Integrate the method described in any one of claims 1-12, and provide the following functional modules: Natural language interaction interface, supporting text, voice and multi-modal input; Dynamic knowledge ontology management module, supporting online editing and version control of domain ontologies; Multi-scenario comparison dashboard, which displays the differential analysis results and decision-making suggestions in real time.

Citation Information

Cited By

  • Interactive AI report generation method and system based on intelligent semantic driving

    CN120541091A

  • Question and answer query method and device based on large model

    CN120910222A

  • Intelligent data analysis method and system based on large model

    CN120994811A

  • Multi-scene intelligent financial analysis system based on large model and composite agent

    CN121073689A

  • Medical data intelligent query and analysis method, system and equipment and medium

    CN121256097A