Data analysis method and system based on retrieval enhancement
By introducing retrieval enhancement methods into relational database systems and utilizing neural network models to generate query statements, the problem of low data analysis efficiency in existing technologies is solved, achieving high efficiency and accuracy in natural language queries, and making it suitable for data analysis in specific fields such as finance.
Patent Information
- Application Number
- CN202511176396.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-28
AI Technical Summary
Existing relational database systems are inefficient in data analysis, especially when domain knowledge is lacking and SQL programming skills are insufficient, making it difficult to accurately locate data using natural language, leading to incorrect query results.
We employ a data analysis method based on retrieval enhancement. By acquiring natural language text query information and combining it with enhanced information from the domain knowledge base, we use a neural network model to generate query statements and execute the queries through a database query engine. This achieves accurate semantic matching and syntax adaptation, dynamically retrieves relevant knowledge, corrects field name errors, and optimizes multi-turn interactions to improve accuracy.
It improves the efficiency and accuracy of natural language queries, reduces human intervention, enhances data analysis capabilities in complex scenarios, and ensures the accuracy and consistency of query results.
Smart Images

Figure CN121029786A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial database technology, and in particular to a data analysis method and system based on retrieval enhancement. Background Technology
[0002] A Relational Database Management System (RDBMS) is a database management system based on the relational model. Relational database systems use tables to organize and store data. Each table can include rows (records) and columns (fields), and data association and querying are achieved by defining relationships between tables. Users can manage data in relational database systems in a standardized way. For example, relational database systems can be managed using Structured Query Language (SQL). Users can query, update, insert, delete, and manage data in the database by editing SQL statements.
[0003] For relational database systems, users can extract and perform data analysis by editing SQL queries. For example, by editing the SQL statement "SELECT name, age, salary FROM employees", users can query the name, age, and salary columns of the employees table in a relational database system, i.e., query specific columns. However, some data analysis processes rely on complex SQL queries. When business personnel do not fully master SQL query editing methods and lack business understanding, this can lead to inefficiency.
[0004] Furthermore, because SQL is a declarative programming language with strict programming rules and conditional constraints, it is difficult for ordinary users to query and analyze data in relational database systems according to their actual needs; that is, they cannot accurately locate data using natural language. When relational database systems are applied to specific domains, the lack of domain knowledge can also lead to the inability to obtain accurate data. For example, in the financial field, due to the strong ambiguity of concepts, the fields defined in relational databases lack correlation with the actual textual descriptions, resulting in some incorrect query results. Summary of the Invention
[0005] In view of this, embodiments of this application provide a data analysis method and system based on retrieval enhancement to solve the problem of low data analysis efficiency in relational database systems.
[0006] According to one aspect of this application, a data analysis method based on retrieval enhancement is provided, the method comprising:
[0007] Obtain query information, which includes natural language text used to query target data;
[0008] Based on the query information, enhanced information is retrieved from the current domain knowledge base. The enhanced information includes one or more combinations of data definition language annotations, query engine syntax documents, business statement examples, and custom domain metrics.
[0009] The query information and the enhancement information are input into the statement generation model, and the query statement is generated based on the retrieval results of the statement generation model. The statement generation model is a neural network model trained based on sample query data and the enhancement information.
[0010] The database query engine is invoked to execute the query statement in order to obtain the target data from the database.
[0011] In some embodiments, the method further includes:
[0012] Acquire raw domain knowledge, which includes database metadata, business metrics, and business query statements;
[0013] The original domain knowledge is enhanced according to preset knowledge enhancement items to generate domain knowledge enhanced data; the preset knowledge enhancement items include one or more combinations of data cleaning and organization, constructing semantic synthesis statements, targeted enhancement of domain terminology parsing, and query statement reverse generation of query text; the domain knowledge enhanced data includes data definition language annotations, query engine syntax documents, business statement examples, and custom domain metrics;
[0014] The current domain knowledge base is constructed based on the domain knowledge augmentation data.
[0015] In some embodiments, knowledge enhancement is performed on the original domain knowledge according to preset knowledge enhancement terms to generate domain knowledge enhanced data, including:
[0016] Retrieve data definition language statements;
[0017] Annotations are added to the data definition language statements based on the target language to construct data definition language statements with target language annotations. The type of the target language is the language type corresponding to the natural language in the current business scenario.
[0018] Retrieval information is read from the database, including table names and fields.
[0019] Establish the association between the retrieved information and the semantics of the target language to obtain the data definition language annotation.
[0020] In some embodiments, knowledge enhancement is performed on the original domain knowledge according to preset knowledge enhancement terms to generate domain knowledge enhanced data, including:
[0021] Retrieve the documentation for the functions built into the query engine;
[0022] Based on the triplet knowledge structure, the function document is split into documents to obtain triplet knowledge information, which includes function name, parameters, and purpose.
[0023] The query engine syntax document is generated based on the triple knowledge information.
[0024] In some embodiments, knowledge enhancement is performed on the original domain knowledge according to preset knowledge enhancement terms to generate domain knowledge enhanced data, including:
[0025] Identify the relevant business departments associated with the current business scenario;
[0026] Collect actual business statement examples from the business departments, which are query statements edited by the business departments when querying target data in the database;
[0027] Construct question-answer pairs that include query item ranges and domain indicators, wherein the item ranges are represented by a preset text format;
[0028] Use the question-and-answer pairs to generate custom domain metrics.
[0029] In some embodiments, retrieving enhanced information from the current domain knowledge base based on the query information includes:
[0030] Extract data description information and analysis action information from the query information;
[0031] Invoke the search enhancement engine;
[0032] Using the retrieval enhancement engine, the associated table structure is extracted from the current domain knowledge base based on the data description information;
[0033] Using the retrieval enhancement engine, analytical calculation functions are extracted from the current domain knowledge base based on the analytical action information;
[0034] The enhanced information is generated based on the association table structure and the analysis calculation function.
[0035] In some embodiments, invoking a database query engine to execute the query statement to obtain the target data from the database includes:
[0036] Obtain feedback data from the database based on the query statement;
[0037] Read field name information from the feedback data. The field name information is used to identify and describe data in the same row or column of the data table.
[0038] If the field name information is incorrect, the query statement is corrected by editing the distance to match similar field names and by using the similar field names.
[0039] The database query engine is invoked to execute the revised query statement in order to obtain the target data from the database.
[0040] In some embodiments, the method further includes:
[0041] Extract field name information from the target data;
[0042] Perform coordinate field matching and grouping logic optimization based on the field name information to determine the display method of the target data;
[0043] Generate chart generation code based on the described display method;
[0044] Execute the chart generation code to draw the display chart corresponding to the target data.
[0045] In some embodiments, coordinate field matching and grouping logic optimization are performed according to the field name information to determine the display method of the target data, including:
[0046] Obtain the number of field names contained in the field name information;
[0047] The chart type is determined based on the number of field names, and the chart type is related to the number of field names.
[0048] The associated chart template is invoked based on the chart type, and the chart template includes at least one axis matching item;
[0049] Set the field name information to the coordinate axis matching item.
[0050] According to another aspect of this application, a data analysis system based on retrieval enhancement is characterized in that the system comprises:
[0051] An information input module is used to obtain query information, which includes natural language text for querying target data.
[0052] An enhanced retrieval module is used to retrieve enhanced information from the current domain knowledge base based on the query information. The enhanced information includes one or more combinations of data definition language annotations, query engine syntax documents, business statement examples, and custom domain metrics.
[0053] The statement generation module is used to input the query information and the enhancement information into the statement generation model, and to generate a query statement based on the retrieval results of the statement generation model. The statement generation model is a neural network model trained based on sample query data and the enhancement information.
[0054] The execution module is used to call the database query engine to execute the query statement in order to obtain the target data from the database.
[0055] According to another aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described retrieval-enhanced data analysis method.
[0056] According to another aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described data analysis method based on retrieval enhancement.
[0057] By employing the above technical solutions, embodiments of this application provide a data analysis method and system based on retrieval enhancement. The method can acquire query information in natural language text form and retrieve enhanced information from the current domain knowledge base based on the query information. The enhanced information includes one or more combinations of Data Definition Language (DDL) annotations, query engine syntax documents, business statement examples, and custom domain metrics. The query information and enhanced information are then input into a statement generation model, and a query statement is generated based on the retrieval results from the statement generation model. The query statement is then executed through a database query engine to obtain target data from the database. This method can improve the efficiency and accuracy of natural language queries by constructing annotated DDL statements and query statement question-answer pairs, establishing semantic mappings for table names or fields. The method also integrates query engine syntax documents and business statement examples to train a model to generate query statements that conform to grammatical rules, achieving grammatical adaptation optimization. Furthermore, it uses a retrieval enhancement framework to dynamically retrieve relevant knowledge, combining edit distance matching to correct field name errors, reducing manual intervention. Finally, based on multi-round interaction optimization, through context history and field recall mechanisms, it gradually clarifies ambiguous query requirements, improving the generation accuracy in complex scenarios.
[0058] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0059] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0060] Figure 1 This is a schematic diagram of the relational database system architecture provided in an embodiment of this application;
[0061] Figure 2 This is a schematic diagram of the process of generating query statements based on a large language model, provided in an embodiment of this application.
[0062] Figure 3 This is a schematic diagram of the data analysis method based on retrieval enhancement provided in an embodiment of this application;
[0063] Figure 4 This is a schematic diagram of the domain data collection and processing flow provided in the embodiments of this application;
[0064] Figure 5 This is a schematic diagram of the retrieval enhancement information process provided in an embodiment of this application;
[0065] Figure 6 This is a schematic diagram illustrating the process of correcting field name errors based on edit distance, as provided in an embodiment of this application.
[0066] Figure 7 This application provides a schematic diagram illustrating the target data flow using charts as an example.
[0067] Figure 8 A schematic diagram of the structure of a data analysis system based on retrieval enhancement provided in this application embodiment;
[0068] Figure 9 This is a schematic diagram of a data analysis system architecture based on retrieval enhancement provided in an embodiment of this application. Detailed Implementation
[0069] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0070] The data analysis method based on retrieval enhancement described in this application embodiment can be applied to relational database systems. Therefore, the data analysis refers to extracting data from the relational database system and processing the extracted data to obtain analysis results.
[0071] Relational Database Management System (RDBMS) is a database management system based on the relational model. Relational database systems use tables to organize and store data. Each table can include rows (records) and columns (fields), and data association and querying are achieved by defining relationships between tables.
[0072] In relational database systems, a table is the basic unit for storing data. A table can include rows and columns, with each row representing a record containing the values of all columns in the table. Each column represents a field, defining the data type and constraints. For one or more columns in a table, a primary key can be used to uniquely identify each row. Foreign keys can be used to establish relationships between tables. A foreign key value is a valid value that references the primary key in a table.
[0073] Relational database systems can quickly locate records in a table using indexes; an index is a data structure used to improve query performance. During data operations and maintenance, users can manage data in relational database systems using standardized methods. For example... Figure 1 As shown, relational database systems can be managed using Structured Query Language (SQL). Users can query, update, insert, delete, and manage data in the database by editing SQL statements.
[0074] Based on Structured Query Language (SCL), users can define different SQL statements to suit their specific query needs. Depending on the intended operation, SQL statements can include various types. Specifically, in some embodiments, SQL statements may include Data Definition Language (DDL) statements, Data Manipulation Language (DML) statements, Data Control Language (DCL) statements, and Transaction Control Language (TCL) statements, etc.
[0075] DDL statements are used to define the database structure. For example, DDL statements can include statements such as CREATE TABLE, ALTER TABLE, and DROP TABLE to perform overall table definition operations.
[0076] DML statements are used to manipulate data. Examples of DML statements include INSERT, UPDATE, DELETE, and SELECT statements. DCL statements are used to control database access permissions. Examples of DCL statements include GRANT and REVOKE statements. TCL statements are used to manage transactions. Examples of TCL statements include COMMIT and ROLLBACK statements.
[0077] Taking DDL statements as an example, DDL is a set of SQL statements used to define and modify database structures. DDL statements can be used to create, modify, and delete objects in the database, such as tables, indexes, views, and users. When creating a table (CREATE TABLE), you can edit the following DDL statement: CREATE TABLE employees(id INT PRIMARY KEY, name VARCHAR(50) NOT NULL, age INT, salary DECIMAL(10,2)). Here, "CREATE TABLE" indicates the specific content of the current operation, i.e., creating a new table. "id INT PRIMARY KEY" is used to define a primary key field of integer type. "name VARCHAR(50) NOT NULL" is used to define a string field with a maximum length of 50, and this field cannot be empty. "age INT" is used to define an integer field. "salary DECIMAL(10,2)" is used to define a decimal field with a total length of 10 digits, of which the decimal part occupies 2 digits.
[0078] DDL statements, as a fundamental tool for managing database structures, help users efficiently create, modify, and delete database objects. For relational database systems, users can extract and analyze data by editing SQL queries. For example, by editing the SQL statement "SELECT aaaa, bbbb, cccc FROM tttt WHERE cccc>1000", users can query the "aaaa", "bbbb", and "cccc" columns of the "TTTT" table in a relational database system, i.e., query specific columns. Then, a view can be created using "CREATEVIEW high_cccc AS; SELECT * FROM TTTT WHERE cccc>1000" to display the queried data in a view format. Here, a view is a virtual table whose content is defined by the SQL query and is dynamically generated displayable data based on the query.
[0079] As we can see, users can retrieve data matching the specified content from the database using query statements with different definitions, and then perform data analysis according to the operation items defined in the query statement. Clearly, some data analysis processes rely on complex SQL query statements, which can lead to inefficiency if business personnel do not fully master SQL statement editing methods or lack business understanding.
[0080] Furthermore, because SQL is a declarative programming language with strict programming rules and conditional constraints, it is difficult for ordinary users to query and analyze data in relational database systems according to their actual needs; that is, they cannot accurately locate data using natural language. When relational database systems are applied to specific domains, the lack of domain knowledge can also lead to the inability to obtain accurate data. For example, in the financial field, due to the strong ambiguity of concepts, the fields defined in relational databases lack correlation with the actual textual descriptions, resulting in some incorrect query results.
[0081] To alleviate the problem of low data query efficiency caused by programming languages, some implementations combine relational database systems with large language models. For example... Figure 2 As shown, the system utilizes the natural language processing capabilities of large language models such as BERT and GPT to generate query statements based on the natural language text input by the user, and then performs data queries in the database based on the generated query statements.
[0082] However, combining relational database systems with large language models still presents challenges in query efficiency and accuracy. For instance, when applying relational database systems and large language models to financial scenarios, a semantic gap arises, where database table names or fields are often in English or lack explicit Chinese mappings, making it difficult for users to accurately locate data using natural language. Syntax compatibility issues also occur, as generated SQL statements must conform to the syntax of specific databases like Impala, while queries generated by large language models lack sufficient adaptability. Furthermore, there is a lack of domain knowledge; financial concepts are highly ambiguous (e.g., "position holdings") and require association with complex knowledge such as indicator definitions and product documentation. For some data, inconsistent formats can lead to recognition errors, such as time format handling issues. For example, business data storage uses the "YYYYMMDD" string format, which has poor compatibility with SQL date types, resulting in overall low analysis efficiency.
[0083] To address the problem of low data analysis efficiency in relational database systems, some embodiments of this application provide a retrieval-enhanced data analysis method. This method can be applied to electronic devices with data processing capabilities. These electronic devices may include, but are not limited to, computers, mobile terminals, servers, smart wearable devices, and industrial control hosts. In some embodiments of this application, electronic devices are used as examples to describe the retrieval-enhanced data analysis method. It should be understood that the method can also be applied to other types of electronic devices, which will not be shown in all embodiments of this application.
[0084] like Figure 3 As shown, the data analysis method based on retrieval enhancement includes:
[0085] S101. Obtain query information.
[0086] When performing data analysis, electronic devices first need to obtain query information. Query information refers to information that instructs the user to perform a data query in the database. To facilitate data queries for the user, the query information obtained by the electronic device may include natural language text used to query the target data.
[0087] The query information can be information entered by the user. In some embodiments, the electronic device can provide a user interface, which may include an input control for query information, through which the user can enter the query information.
[0088] For example, a user can enter text information containing the intent to query data, such as "Please help me query the trading volume and open interest ratio of IF and IC futures from 20200101 to 20230101", and the electronic device can then obtain query information containing the aforementioned natural language text content through the text input box.
[0089] Users can input query information via text input or other interactive methods. Specifically, in some embodiments, users can input query information via voice input, handwriting input, image recognition, etc. For query information input via different interactive methods, the electronic device can be equipped with a functional module that converts other forms of signals into text information to meet the needs of different information input interaction methods; these will not be illustrated in detail in the embodiments of this application.
[0090] The query information can also be based on business scenarios, parsed from specific business data. In some embodiments, after a user controls an electronic device to display the interactive interface corresponding to the database system, they can upload business files using a file upload control within the interface. Once uploaded, the electronic device can read the content from the business file and determine whether it contains database-related operational intents. If the business file contains database-related operational intents, query information can be extracted from the business file based on those intents.
[0091] For example, when a user needs to update the trading and open interest data of IC futures from January 1, 2020 to December 31, 2020 in the database, they can upload a business file covering the same period using the file upload control on the user interface. Since the business file contains IC futures trading and open interest data, and the database also contains data for the same period, the electronic device can determine that the uploaded file contains an intention to update the data. Therefore, it can generate query information based on this intention to retrieve the IC futures trading and open interest data for the same period, enabling subsequent data updates.
[0092] S102. Retrieve enhanced information from the current domain knowledge base based on the query information.
[0093] After obtaining the query information, the electronic device can retrieve enhanced information from the current domain knowledge base based on the query information. Enhanced information is supplementary information used to enhance the data retrieval and positioning effect. Retrieval enhanced information can be combined with the query information to jointly participate in the data query process. Retrieval enhanced information is information pre-built in the current domain knowledge base to achieve semantic accuracy matching, syntax adaptation optimization, domain knowledge enhancement, efficient retrieval enhancement, and multi-round interaction optimization.
[0094] In some embodiments, the electronic device may first acquire raw domain knowledge, which includes database metadata, business metrics, and business query statements. Then, it performs knowledge enhancement on the raw domain knowledge according to preset knowledge enhancement items to generate domain knowledge enhancement data. The preset knowledge enhancement items may include one or more combinations of data cleaning and organization, constructing semantically synthesized statements, targeted enhancement of domain terminology parsing, and the reverse generation of query text from query statements. The domain knowledge enhancement data includes data definition language annotations, query engine syntax documentation, business statement examples, and custom domain metrics. Finally, the current domain knowledge base is constructed based on the domain knowledge enhancement data.
[0095] For example, such as Figure 4As shown, to obtain knowledge-enhanced data in the financial domain, electronic devices can first acquire raw domain knowledge, such as database metadata, financial indicators, and business SQL query data. Then, the acquired data undergoes data cleaning and organization, and semantic synthesis statements are constructed, along with targeted enhancements to domain terminology parsing capabilities. Furthermore, text can be generated from SQL statements, i.e., questions in natural language text format are generated from SQL statements for knowledge enhancement. After knowledge enhancement, financial domain-enhanced data can be obtained based on the raw domain data, including DDL statements, documents, and question-and-answer pairs suitable for the financial domain.
[0096] Therefore, enhanced information includes one or more combinations of Data Definition Language (DDL) annotations, query engine syntax documentation, business statement examples, and custom domain metrics. Specifically, DDL annotations are constructed by setting annotation tags on DDL statements to achieve precise semantic matching; these are DDL statements with target language annotations and business data question-and-answer pairs.
[0097] In some embodiments, to construct Data Definition Language (DDL) annotations, the electronic device can first obtain DDL statements, and then add annotations to the DDL statements based on a target language to construct DDL statements with target language annotations. The target language type is the language type corresponding to the natural language in the current business scenario. Then, retrieval information including table names and fields is read from the database, and an association is established between the retrieval information and the semantics of the target language to obtain the DDL annotations.
[0098] For example, to build a knowledge base in the financial field, Data Definition Language (DDL) annotations can be obtained through table structure enhancement. That is, when the target language is Chinese and the SQL language is based on English, DDL statements with Chinese annotations can be constructed, such as `CREATE TABLE...COMMENT 'Institution Name'`. By adding the Chinese annotation "Institution Name", electronic devices can perform better semantic matching based on the Chinese annotation when processing queries containing Chinese content, thereby improving the accuracy of natural language queries. Simultaneously with setting Chinese annotations, a relationship can be established between table names or fields and Chinese semantics, enabling electronic devices to recognize Chinese semantics.
[0099] It is evident that by constructing DDL statements and SQL question-and-answer pairs with target language type annotations, and establishing a mapping relationship between table names and fields in the target language semantics, electronic devices can achieve higher natural language query accuracy when processing query information in the form of natural language text.
[0100] The query engine syntax document is enhanced information generated by integrating query engine documentation and actual business cases for syntax adaptation and optimization. In some embodiments, to obtain the query engine syntax document, when an electronic device performs knowledge enhancement on the original domain knowledge according to preset knowledge enhancement items to generate domain knowledge enhanced data, it can first obtain the function documentation built into the query engine, and then perform document segmentation on the function documentation according to the triple knowledge structure to obtain triple knowledge information. The triple knowledge structure includes the function name, parameters, and purpose. The query engine syntax document is then generated based on the triple knowledge information.
[0101] For example, to achieve syntax enhancement, electronic devices can access the official Impala documentation and real-world business SQL examples. Based on these examples, they can segment the Impala function documentation embedded in the official documentation to generate triples containing function names, parameters, and purposes. This integrates the official Impala documentation and the actual business SQL examples to obtain the query engine syntax document. This query engine syntax document can then be used to train a statement generation model, such as a large language model or a model built by a business operations organization based on a neural network model. After training, the statement generation model can generate SQL statements that conform to the syntax rules.
[0102] Business statement examples and custom domain metrics are enhanced information constructed from statement examples encountered during actual data querying processes for contextualized data augmentation. In some embodiments, to construct business statement examples and custom domain metrics, when an electronic device performs knowledge augmentation on the original domain knowledge according to preset knowledge augmentation items to generate domain knowledge augmented data, it can first determine the business department associated with the current business scenario, and then collect the actual business statement examples of the business department. The business statement examples are query statements edited by the business department when querying target data in the database. Then, a question-and-answer pair containing the project range and domain metrics represented by the query in a preset text format is constructed, and the question-and-answer pair is used to generate custom domain metrics.
[0103] For example, in the financial sector, electronic devices can collect actual SQL cases from the R&D and management departments, and based on the time-related content used in the actual SQL cases, construct question-and-answer pairs containing time intervals in formats such as "YYYYMMDD" and financial indicator calculations, thereby generating business statement cases and custom domain indicators for targeted data enhancement in financial scenarios.
[0104] In addition to the aforementioned enhanced information, domain knowledge can be enhanced through other means. For example, by integrating knowledge bases such as financial indicator definitions, product documents, and trading calendars, electronic devices can support polysemous word recognition and complex indicator calculations. For instance, in the polysemous word recognition process, the electronic device can equate "order execution ratio" with "transaction-to-order ratio" based on the definition of financial indicators. And for complex indicator calculations, it can support calculations of complex indicators such as the transaction-to-open-interest ratio.
[0105] Based on the enhanced information shown in the above embodiments, the electronic device can construct a knowledge base for the current business domain and store the enhanced information in the knowledge base. Upon receiving query information, the electronic device can perform an information retrieval in the knowledge base to find enhanced information that can supplement the query information.
[0106] like Figure 5 As shown, in some embodiments, to obtain enhanced information, when an electronic device retrieves enhanced information from the current domain knowledge base based on the query information, it first extracts data description information and analysis action information from the query information, then calls and uses the retrieval enhancement engine to extract the association table structure from the current domain knowledge base based on the data description information. Simultaneously, using the retrieval enhancement engine, it extracts analysis calculation functions from the current domain knowledge base based on the analysis action information. Thus, the enhanced information is generated based on the association table structure and the analysis calculation functions.
[0107] For example, to achieve efficient retrieval enhancement, electronic devices can employ the Retrieval-Augmented Generation (RAG) framework to dynamically retrieve relevant knowledge. The RAG framework is a technique combining Information Retrieval (IR) and Natural Language Generation (NLG), aiming to enhance the output of a Large Language Model (LLM) by incorporating information from external knowledge bases, thereby generating more accurate and context-appropriate responses.
[0108] After receiving a user's input stating "Please query the trading volume and open interest ratio of IF and IC futures from January 1, 2020 to January 1, 2023," the electronic device can first invoke the RAG engine and perform knowledge retrieval. Specifically, the RAG engine extracts the INTERFACE.T_DAILY_MARKETDATA table structure based on the data description information "from January 1, 2020 to January 1, 2023," "IF futures," and "IC futures." Simultaneously, based on the analysis action information "trading volume and open interest ratio," it extracts the calculation function corresponding to SUM(VOLUME) / SUM(POSI), thus obtaining the enhanced information corresponding to the extracted content.
[0109] S103. Input the query information and the enhanced information into the statement generation model, and generate a query statement based on the retrieval results of the statement generation model.
[0110] After acquiring the augmented information, the electronic device can combine the query information with the augmented information and input the combined result into a statement generation model. The statement generation model is a neural network model trained based on sample query data and the augmented information. The statement generation model is used to generate a query statement based on the input text content; that is, the input to the statement generation model is natural language text and augmented information, and the output of the statement generation model is the query statement.
[0111] The statement generation model can be obtained by reusing a large language model. In some embodiments, the database user interaction module running on the electronic device can have a data interface to the large language model. After obtaining query information and enhancement information, the electronic device can input these into the large language model through the data interface. Upon receiving the input text, the large language model performs natural language processing and information retrieval, generates a query statement based on the retrieval results, and then feeds the query statement back to the electronic device.
[0112] For example, if a user inputs the query "the ratio of trading volume and open interest of IF and IC futures during the period from 20200101 to 20230101", then, according to the enhanced information retrieval method provided in the above embodiment, after extracting the enhanced information of the INTERFACE.T_DAILY_MARKETDATA table structure and the SUM(VOLUME) / SUM(POSI) calculation function through the RAG engine, the query information and enhanced information can be input into the Llama3 series large language model. The large language model then generates a query statement with the content "SELECT PRODUCT_ID, SUM(VOLUME)*1.000 / SUM(POSI) FROM... WHERE tradingdayBETWEEN'20200101' AND'20230101' GROUP BY PRODUCT_ID" based on the query information and enhanced information.
[0113] The statement generation model can also be obtained through separate training. In some embodiments, the electronic device can first select a suitable initial model based on business needs and acquire sample query data and augmentation information. The sample query data consists of query information with query statement tags. Query statement tags are also assigned to the augmentation information. The sample query data and augmentation information are then input into the initial model to obtain its output, i.e., the output query statement. The output query statement is then compared with its tags, and the training loss is calculated based on the comparison. Backpropagation is then used to adjust the model parameters of the initial model based on the training loss. After multiple iterations, when the training loss corresponding to the model output meets the convergence requirement, a statement generation model with a certain output accuracy is obtained.
[0114] The sentence generation model can also be obtained through targeted training of a large language model. In some embodiments, targeted training data can be constructed based on actual business cases. This targeted training data may include query information collected in the current business domain and corresponding query tagging. After obtaining augmentation information such as data definition language annotations, query engine syntax documents, business sentence examples, and custom domain metrics through knowledge augmentation, this augmentation information can be input into the large language model to generate query statements. The training loss is then calculated based on the query statements and backpropagated to train the large language model in a targeted manner, thereby obtaining the sentence generation model.
[0115] S104. Call the database query engine to execute the query statement to obtain the target data from the database.
[0116] After obtaining the query statement, the electronic device can use the database query engine to execute the query and retrieve the target data from the database. For example, the electronic device can call the Impala database and execute the SQL statement for the query statement "SELECT PRODUCT_ID, SUM(VOLUME)*1.000 / SUM(POSI) FROM ... WHERE tradingday BETWEEN '20200101' AND '20230101' GROUP BY PRODUCT_ID" and return the result. That is, it queries the database for IF and IC futures data from 20200101 to 20230101 and executes the calculation function SUM(VOLUME)*1.000 / SUM(POSI) to obtain the trading volume and open interest ratio of IF and IC futures.
[0117] After electronic devices execute SQL statements through a database query engine and return data, they can also verify the results of the returned data, such as... Figure 6 As shown, in some embodiments, in order to obtain more accurate target data, when the electronic device calls the database query engine to execute the query statement to obtain the target data from the database, it can first obtain the feedback data from the database based on the query statement, and read the field name information from the feedback data. The field name information is used to identify and describe data in the same row or column of the data table.
[0118] The retrieved field name information is then evaluated based on the multi-turn dialogue content. If the field name information is correct, the feedback data can be output and displayed as the target data. If the field name information is incorrect, the query statement can be corrected by editing the distance-matched similar field names and based on the similar field names; then, the database query engine is invoked to execute the corrected query statement to obtain the target data from the database.
[0119] Edit distance, also known as Levenshtein distance, is a metric for measuring the difference between two strings. It is defined as the minimum number of single-character editing operations (insertion, deletion, or replacement) required to transform one string into another. Edit distance can be efficiently calculated using dynamic programming (DP) algorithms. The core idea of dynamic programming is to decompose the problem into subproblems and avoid redundant computation by storing the solutions to these subproblems.
[0120] For example, electronic devices can optimize queries based on multi-turn dialogue mechanisms. After initially generating an SQL statement based on a large language model, the system can check for errors in field names based on the content of the multi-turn dialogue. If duplicate or semantically similar content appears in the multi-turn dialogue, an error in the field name can be identified. Conversely, if no duplicate or semantically similar content appears, the field name is considered correct. If a field name is incorrect, it can be corrected by matching similar field names using edit distance and performing a secondary search. In other words, by leveraging contextual history and field recall mechanisms, fuzzy query requirements are gradually clarified, improving the accuracy of generation in complex scenarios.
[0121] By applying the technical solutions of the above embodiments, the data analysis method based on retrieval enhancement provided by the above embodiments can obtain query information in the form of natural language text, and retrieve enhanced information in the current domain knowledge base based on the query information. The enhanced information includes one or more combinations of data definition language annotations, query engine syntax documents, business statement examples, and custom domain indicators. The query information and enhanced information are then input into a statement generation model, and a query statement is generated based on the retrieval results of the statement generation model. The query statement is then executed by the database query engine to obtain the target data from the database. The method can improve the efficiency and accuracy of natural language queries by constructing annotated data definition language statements and query statement question-answer pairs, establishing semantic mapping of table names or fields. The method also integrates query engine syntax documents and business statement examples to train the model to generate query statements that conform to grammatical rules, achieving syntax adaptation optimization. Furthermore, it uses a retrieval enhancement framework to dynamically retrieve relevant knowledge, combines edit distance matching to correct field name errors, and reduces manual intervention. Finally, based on multi-round interaction optimization, through context history and field recall mechanisms, it gradually clarifies ambiguous query requirements, improving the generation accuracy in complex scenarios.
[0122] In some embodiments, as a refinement and extension of the specific implementation of the above embodiments, and in order to fully illustrate the specific implementation process of this embodiment, some embodiments of this application also provide a data analysis method based on retrieval enhancement, such as... Figure 7 As shown, the method, based on the above embodiments, further includes:
[0123] S201. Extract field name information from the target data;
[0124] S202. Perform coordinate field matching and grouping logic optimization according to the field name information to determine the display method of the target data;
[0125] S203. Generate chart generation code according to the described display method;
[0126] S204. Execute the chart generation code to draw the display chart corresponding to the target data.
[0127] After obtaining target data according to the above embodiments, the electronic device can first extract field name information from the target data, and then perform coordinate field matching and grouping logic optimization according to the field name information to determine the display method of the target data. Specifically, the electronic device can obtain the number of field names contained in the field name information, and determine the chart type based on the number of field names, wherein the chart type is related to the number of field names. Then, it calls the associated chart template according to the chart type, and the chart template includes at least one coordinate axis matching item. The field name information is then set to the coordinate axis matching item.
[0128] For example, after acquiring the target data, the electronic device can read the number of field names contained in the target data. When the number of field names read is 1, that is, the target data contains only one row of data, it can be determined that the chart type is a single data chart. Therefore, a single data chart template, such as a bar chart template, can be called, and the field name information can be set to the axis matching item to represent the corresponding chart coordinate content.
[0129] After determining the display method for the target data, the electronic device generates chart generation code based on the display method, and then draws the corresponding display chart for the target data by executing the chart generation code. For example, after obtaining the trading volume and open interest ratio data of IF and IC futures from January 1, 2020 to January 1, 2023, the electronic device can call a plotting function to generate a trend line chart. The horizontal axis of the line chart represents the trading day, and the vertical axis represents the trading volume and open interest ratio, so as to display the target data in a visual way.
[0130] By applying the technical solutions of the above embodiments, the data analysis method based on retrieval enhancement provided in the above embodiments can integrate intelligent drawing functions. The intelligent drawing function module extracts field name information from the target data, performs coordinate field matching and grouping logic optimization according to the field name information, and generates chart generation code after determining the display method to draw the display chart corresponding to the target data. When generating chart generation code such as line charts and bar charts, blank images are avoided and the result display effect is improved by optimizing the horizontal and vertical coordinate fields and grouping logic.
[0131] In some embodiments, as a specific implementation of the retrieval-enhanced data analysis method described in the above embodiments, some embodiments of this application also provide a retrieval-enhanced data analysis system, such as... Figure 8 As shown, the system includes:
[0132] An information input module is used to obtain query information, which includes natural language text for querying target data.
[0133] An enhanced retrieval module is used to retrieve enhanced information from the current domain knowledge base based on the query information. The enhanced information includes one or more combinations of data definition language annotations, query engine syntax documents, business statement examples, and custom domain metrics.
[0134] The statement generation module is used to input the query information and the enhanced information into the statement generation model, and to generate a query statement based on the retrieval results of the statement generation model;
[0135] The execution module is used to call the database query engine to execute the query statement in order to obtain the target data from the database.
[0136] For example, such as Figure 9 As shown, the system architecture of a retrieval-enhanced data analysis system can include an input layer, a RAG engine, a large model, and an execution layer. The input layer receives user natural language queries; the RAG engine retrieves domain knowledge bases, including DDL annotations, Impala syntax documentation, business SQL examples, and financial indicator definitions. The large model can be a Llama3 series large model, used to generate SQL statements and visualization code based on the retrieval results. The execution layer calls the Impala database to execute SQL and returns the results, generating charts through the plotting module.
[0137] It should be noted that other corresponding descriptions of the functional units involved in the data analysis system based on retrieval enhancement provided in the embodiments of this application can be found in the corresponding descriptions in the data analysis method based on retrieval enhancement provided in the above embodiments, and will not be repeated here.
[0138] This application also provides a computer device, specifically a personal computer, server, network device, etc. The computer device includes a bus, processor, memory, and communication interface, and may also include input / output interfaces and a display device. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores location information. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.
[0139] Those skilled in the art will understand that the structure of the computer device described above is only a partial structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components, or combine certain components, or have different component arrangements.
[0140] In one embodiment, a computer-readable storage medium is also provided, which may be non-volatile or volatile, having stored thereon a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0141] In one embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0142] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0143] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.
[0144] Any references to memory, database, or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc.
[0145] Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can take many forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0146] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchain. The processors involved in the embodiments provided in this application may be, but are not limited to, general-purpose processors, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc.
[0147] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0148] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A data analysis method based on retrieval enhancement, characterized in that, The method includes: Obtain query information, which includes natural language text used to query target data; Based on the query information, enhanced information is retrieved from the current domain knowledge base. The enhanced information includes one or more combinations of data definition language annotations, query engine syntax documents, business statement examples, and custom domain metrics. The query information and the enhancement information are input into the statement generation model, and the query statement is generated based on the retrieval results of the statement generation model. The statement generation model is a neural network model trained based on sample query data and the enhancement information. The database query engine is invoked to execute the query statement in order to obtain the target data from the database.
2. The method according to claim 1, characterized in that, The method further includes: Acquire raw domain knowledge, which includes database metadata, business metrics, and business query statements; The original domain knowledge is enhanced according to preset knowledge enhancement items to generate domain knowledge enhanced data; the preset knowledge enhancement items include one or more combinations of data cleaning and organization, constructing semantic synthesis statements, targeted enhancement of domain terminology parsing, and query statement reverse generation of query text; the domain knowledge enhanced data includes data definition language annotations, query engine syntax documents, business statement examples, and custom domain metrics; The current domain knowledge base is constructed based on the domain knowledge augmentation data.
3. The method according to claim 2, characterized in that, The original domain knowledge is augmented according to preset knowledge enhancement items to generate domain knowledge augmented data, including: Retrieve data definition language statements; Annotations are added to the data definition language statements based on the target language to construct data definition language statements with target language annotations. The type of the target language is the language type corresponding to the natural language in the current business scenario. Retrieval information is read from the database, including table names and fields. Establish the association between the retrieved information and the semantics of the target language to obtain the data definition language annotation.
4. The method according to claim 2, characterized in that, The original domain knowledge is augmented according to preset knowledge enhancement items to generate domain knowledge augmented data, including: Retrieve the documentation for the functions built into the query engine; Based on the triplet knowledge structure, the function document is split into documents to obtain triplet knowledge information, which includes function name, parameters, and purpose. The query engine syntax document is generated based on the triple knowledge information.
5. The method according to claim 2, characterized in that, The original domain knowledge is augmented according to preset knowledge enhancement items to generate domain knowledge augmented data, including: Identify the relevant business departments associated with the current business scenario; Collect actual business statement examples from the business departments, which are query statements edited by the business departments when querying target data in the database; Construct question-answer pairs that include query item ranges and domain indicators, wherein the item ranges are represented by a preset text format; Use the question-and-answer pairs to generate custom domain metrics.
6. The method according to claim 1, characterized in that, Based on the query information, retrieve enhanced information from the current domain knowledge base, including: Extract data description information and analysis action information from the query information; Invoke the search enhancement engine; Using the retrieval enhancement engine, the associated table structure is extracted from the current domain knowledge base based on the data description information; Using the retrieval enhancement engine, analytical calculation functions are extracted from the current domain knowledge base based on the analytical action information; The enhanced information is generated based on the association table structure and the analysis calculation function.
7. The method according to claim 1, characterized in that, Invoking the database query engine to execute the query statement to obtain the target data from the database includes: Obtain feedback data from the database based on the query statement; Read field name information from the feedback data. The field name information is used to identify and describe data in the same row or column of the data table. If the field name information is incorrect, the query statement is corrected by editing the distance to match similar field names and by using the similar field names. The database query engine is invoked to execute the revised query statement in order to obtain the target data from the database.
8. The method according to claim 1, characterized in that, The method further includes: Extract field name information from the target data; Perform coordinate field matching and grouping logic optimization based on the field name information to determine the display method of the target data; Generate chart generation code based on the described display method; Execute the chart generation code to draw the display chart corresponding to the target data.
9. The method according to claim 8, characterized in that, Based on the field name information, coordinate field matching and grouping logic optimization are performed to determine the display method of the target data, including: Obtain the number of field names contained in the field name information; The chart type is determined based on the number of field names, and the chart type is related to the number of field names. The associated chart template is invoked based on the chart type, and the chart template includes at least one axis matching item; Set the field name information to the coordinate axis matching item.
10. A data analysis system based on retrieval enhancement, characterized in that, The system includes: An information input module is used to obtain query information, which includes natural language text for querying target data. An enhanced retrieval module is used to retrieve enhanced information from the current domain knowledge base based on the query information. The enhanced information includes one or more combinations of data definition language annotations, query engine syntax documents, business statement examples, and custom domain metrics. The statement generation module is used to input the query information and the enhancement information into the statement generation model, and to generate a query statement based on the retrieval results of the statement generation model. The statement generation model is a neural network model trained based on sample query data and the enhancement information. The execution module is used to call the database query engine to execute the query statement in order to obtain the target data from the database.
Citation Information
Cited By
NL2SQL method, device and equipment based on retrieval enhancement, medium and program product
CN121705302A
Method and device for converting natural language query into GQL statement
CN121833767A