Data query model training method and device, data query method and device, medium and product

By constructing standard questioning scenarios and collecting historical data, questioning scripts and entity base datasets are generated. Randomly combining these datasets for training solves the problem of inaccurate generation of large language models in specific scenarios, and improves the accuracy and coverage of data query models.

CN121301934APending Publication Date: 2026-01-09LIAONING MOBILE COMM +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511650554.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing technologies fail to consider the adaptability to specific question scenarios when building large language models, resulting in inaccurate SQL query statements that are difficult to meet the needs of practical applications.

Method used

By constructing standard questioning scenarios, collecting historical questioning data, generating questioning scripts and entity base datasets, and randomly combining them to form a training dataset, the data query model is trained and fine-tuned to improve the accuracy of the model's generated data query analysis language.

Benefits of technology

It improves the accuracy and precision of the data query model in generating data under different standard question scenarios, ensures a balance between the amount of question wording and entity information, covers all possible expressions, and solves the problem of inaccurate model generation results in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301934A_ABST
    Figure CN121301934A_ABST
Patent Text Reader

Abstract

The invention discloses a data query model training method and device, a data query method and device, a medium and a product, and the method comprises the steps: collecting historical question data in real time, and determining a standard question scene corresponding to the historical question data; the method comprises the following steps: extracting question verbal skill and entity information according to a plurality of pieces of historical question data in the same standard question scene, and respectively generating a question verbal skill basic data set and an entity basic data set in the standard question scene; randomly combining the questioning verbal skill in the questioning verbal skill basic data set with entity information under a corresponding entity type in the entity basic data set to generate a plurality of sample questioning data; generating a training data set according to the sample question data, and training and adjusting a preset data query model according to the training data set; wherein the data query model is used for generating a data query analysis language according to the target question data. By adopting the method, the accuracy of generating the data query analysis language can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a data query model training, data query method, apparatus, medium and product. Background Technology

[0002] In the past two years, large-scale model technology has developed rapidly both domestically and internationally, with a continuous emergence of large-scale model products on both cloud and edge platforms. The commercialization of large-scale model technology across various industries is also accelerating. NL2SQL (Natural Language to Structured Query Language) is a query generation technology based on natural language and SQL (Structured Query Language). It can convert natural language questions into executable SQL query statements to retrieve data information corresponding to user queries. Using NL2SQL can improve user query efficiency and accuracy while reducing the learning curve. With the rapid development of large-scale model technology, NL2SQL solutions based on large semantic models are constantly emerging, gradually becoming an important path in the evolution of NL2SQL technology.

[0003] However, existing solutions typically use general datasets to build and train large language models. In application, these large language models are used to convert natural language questions into SQL queries. The training datasets used are relatively simple and broad, and do not consider adaptability to specific question scenarios. They also do not involve model adjustments and optimizations for practical applications, making it difficult to guarantee the accuracy of the SQL queries generated by the large language model, resulting in inaccurate data query results based on these SQL queries. Summary of the Invention

[0004] The purpose of this invention is to provide a data query model training method, device, medium, and product that can continuously update the training dataset based on historical question data under different standard question scenarios, train and fine-tune the data query model, and effectively improve the accuracy of the data query analysis language generated by the model.

[0005] To achieve the above objectives, embodiments of the present invention provide a method for training a data query model, comprising: Collect historical question data in real time and determine the standard question scenario corresponding to the historical question data; Based on several historical question data under the same standard question scenario, questioning scripts and entity information are extracted to generate a basic dataset of questioning scripts and a basic dataset of entities under the standard question scenario, respectively; wherein, the basic dataset of questioning scripts includes several questioning scripts under the standard question scenario, and the basic dataset of entities includes at least one entity type and several entity information under the entity type; The questioning scripts in the questioning script base dataset are randomly combined with the entity information under the corresponding entity type in the entity base dataset to generate several sample questioning data. A training dataset is generated based on the sample question data, and a preset data query model is trained and adjusted based on the training dataset; wherein, the data query model is used to generate a data query analysis language based on the target question data.

[0006] As an improvement to the above solution, after extracting questioning scripts and entity information from several historical questioning data under the same standard questioning scenario, and generating a basic dataset of questioning scripts and a basic dataset of entities under the standard questioning scenario respectively, the method further includes: The questioning script base dataset and the entity base dataset are respectively expanded to obtain the questioning script extended dataset and the entity extended dataset; wherein, the questioning script extended dataset includes the target number of questioning scripts under the standard questioning scenario, and the entity extended dataset includes at least one entity type and the target number of entity information under the entity type; The step of randomly combining the questioning phrases in the questioning phrase base dataset with the entity information under the corresponding entity type in the entity base dataset to generate several sample questioning data is as follows: The question phrases in the question phrase extension dataset are randomly combined with the entity information under the corresponding entity type in the entity extension dataset to generate several sample question phrases.

[0007] As an improvement to the above solution, the step of expanding the basic dataset of questioning techniques and the basic dataset of entities respectively to obtain expanded datasets of questioning techniques and entities includes: After randomly sorting all the questioning scripts in the basic questioning script dataset, they are repeatedly added to the extended questioning script dataset until the number of questioning scripts in the extended questioning script dataset reaches the target number, thus obtaining the extended questioning script dataset. After randomly sorting all entity information under each entity type in the entity base dataset, the information is repeatedly added to the corresponding entity type in the entity extended dataset until the number of entity information for each entity type in the entity extended dataset reaches the target number, thus obtaining the entity extended dataset.

[0008] As an improvement to the above scheme, the data query model includes an intent recognition sub-model, an entity extraction sub-model, and a language generation sub-model; The intent recognition sub-model is used to identify the standard question scenario corresponding to the target question data. The entity extraction sub-model is used to extract the entity type and entity information of the target question data and standardize the target question data. The language generation sub-model is used to generate the data query and analysis language corresponding to the target question data.

[0009] As an improvement to the above solution, the step of generating a training dataset based on the sample query data and training and adjusting the preset data query model based on the training dataset includes: The first training data is generated by labeling the standard questioning scenarios corresponding to the sample questioning data. Based on the sample query data, the corresponding entity type and entity information are labeled to generate the second training data; Based on the data query analysis language labeled with the sample question data, a third training data is generated; The intent recognition sub-model is trained and adjusted based on the first training data, the entity extraction sub-model is trained and adjusted based on the second training data, and the language generation sub-model is trained and adjusted based on the third training data.

[0010] This invention also provides a data query method, including: Receive the user's target question data; The target query data is input into a data query model for processing to obtain an executable data query analysis language generated by the data query model; wherein, the data query model is generated according to the training method of the data query model described above. The data query analysis language is executed to obtain the query results corresponding to the target query data.

[0011] As an improvement to the above scheme, the data query model includes an intent recognition sub-model, an entity extraction sub-model, and a language generation sub-model; The step of inputting the target query data into a data query model for processing to obtain an executable data query analysis language generated by the data query model includes: The target question data is input into the intent recognition sub-model for intent recognition, and the standard question scenario output by the intent recognition sub-model is obtained. The target question data and the standard question scenario are input into the entity extraction sub-model for entity extraction, and the target question data is standardized based on the entity extraction results to obtain standardized question data output by the entity extraction sub-model; wherein, the entity extraction results include entity type and entity information; The standardized question data and the standardized question scenario are input into the language generation sub-model for processing to obtain the data query analysis language generated by the language generation sub-model.

[0012] As an improvement to the above solution, the language generation sub-model is used to call the data information table and field definition information corresponding to the standard question scenario according to the mapping relationship between the preset standard question scenario, data information table and field definition information, and generate a data query and analysis language based on the data information table, the field definition information and the standardized question data.

[0013] This invention also provides a training apparatus for a data query model, comprising: The historical data acquisition module is used to collect historical question data in real time and determine the standard question scenario corresponding to the historical question data. The basic dataset generation module is used to extract questioning scripts and entity information from several historical questioning data under the same standard questioning scenario, and generate a basic dataset of questioning scripts and a basic dataset of entities under the standard questioning scenario respectively; wherein, the basic dataset of questioning scripts includes several questioning scripts under the standard questioning scenario, and the basic dataset of entities includes at least one entity type and several entity information under the entity type; The sample question data generation module is used to randomly combine the question scripts in the question script base dataset with the entity information under the corresponding entity type in the entity base dataset to generate several sample question data. The training dataset generation module is used to generate a training dataset based on the sample question data, and to train and adjust a preset data query model based on the training dataset; wherein, the data query model is used to generate a data query analysis language based on the target question data.

[0014] This invention also provides a data query device, comprising: The target question data receiving module is used to receive the user's target question data; The query analysis language generation module is used to input the target query data into a data query model for processing, and obtain an executable data query analysis language generated by the data query model; wherein, the data query model is generated according to the training method of the data query model as described above; The data query result generation module is used to execute the data query analysis language to obtain the query results corresponding to the target query data.

[0015] This invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements training of a data query model as described in any of the preceding claims, or a data query method as described in any of the preceding claims.

[0016] This invention also provides a computer-readable storage medium, which includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform training of a data query model as described in any of the preceding claims, or a data query method as described in any of the preceding claims.

[0017] This invention also provides a computer program product, which includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, they implement the training of the data query model as described in any of the above claims, or the data query method as described in any of the above claims.

[0018] Compared with existing technologies, the data query model training, data query method, device, medium, and product disclosed in this invention construct standard question scenarios and organize historical question data under different standard question scenarios into a basic dataset using question phrases and entity information. Based on the basic dataset, a training dataset is automatically generated by randomly arranging and combining question phrases and entities to cover all question phrases and entity information in the relevant basic dataset, ensuring a balanced quantity of each question phrase and entity information. A sufficient and balanced training dataset lays the foundation for improving the quality of model training. By combining the training datasets corresponding to each standard question scenario to train and fine-tune the data query model, it is beneficial to improve the accuracy of the data query model and the accuracy of the data query analysis language generated by the model. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating a training method for a data query model provided in an embodiment of the present invention; Figure 2This is a flowchart illustrating a data query method provided in an embodiment of the present invention; Figure 3 This is a flowchart illustrating a preferred data query method provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of the data query and analysis system provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a training device for a data query model provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a data query device provided in an embodiment of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] In the description of this application, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0022] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0023] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0024] See Figure 1This is a flowchart illustrating a data query model training method provided in an embodiment of the present invention. The embodiment of the present invention provides a data query model training method, the method comprising steps S11 to S14: S11. Collect historical question data in real time and determine the standard question scenario corresponding to the historical question data; S12. Based on several historical question data under the same standard question scenario, extract questioning scripts and entity information, and generate a basic dataset of questioning scripts and a basic dataset of entities under the standard question scenario respectively; wherein, the basic dataset of questioning scripts includes several questioning scripts under the standard question scenario, and the basic dataset of entities includes at least one entity type and several entity information under the entity type. S13. Randomly combine the questioning scripts in the questioning script base dataset with the entity information under the corresponding entity type in the entity base dataset to generate several sample questioning data. S14. Generate a training dataset based on the sample question data, and train and adjust the preset data query model based on the training dataset; wherein, the data query model is used to generate a data query analysis language based on the target question data.

[0025] It should be noted that users' questions using natural language exhibit diverse expressions, missing entity information, and inaccurate descriptions. Directly using these in generating data query and analysis languages ​​can negatively impact the quality of large-scale model generation. Existing technologies typically standardize some entities in user questions, but this approach still cannot completely resolve the issues of diverse expressions, missing entity information, and inaccurate descriptions.

[0026] In practical business applications, most users frequently use a limited number of data query scenarios, and these scenarios are highly repetitive. Therefore, this embodiment of the invention organizes frequently asked question scenarios based on users' historical question data, constructs a standard question scenario library, and builds a training dataset based on historical question data under each of these standard question scenarios. This dataset is then used to train and fine-tune a large language model, which helps improve the accuracy of data query analysis language generation.

[0027] In this embodiment of the invention, historical question data input by users is collected in real time. Based on the historical question data, scenarios that were not previously identified but have high question frequency and high business value are identified and added as standard question scenarios. Examples include "multi-time period regional indicator query", "single regional real-time indicator query", and "provincial multi-indicator comparison query".

[0028] Taking the standard question scenario of "multi-time period regional indicator query" as an example, historical question data includes: I would like to check the 5G network access rate in City A on the 4th. Show me the number of 5G users who registered yesterday; What was the number of 5G customer complaints in City B in the past 3 days? I'd like to see the number of VoNR users across the province.

[0029] For each standard questioning scenario, historical question data is used to organize relevant information in the basic dataset using a questioning script + entity approach. Taking the aforementioned standard questioning scenario as an example, a sample of the basic dataset organized from historical question data is as follows: The question format is as follows: I would like to check the [time][region][index] values; Show me the [time][indicators]; What are the [time][region][indicator] values? I'd like to take a look at the [region][indicator] situation.

[0030] Entity types include time information, geographical information, and indicator information. The entity information under each entity type is as follows: Time information: 4th, yesterday, the last 3 days; Geographic information: The entire province, City A, City B; Key metrics: 5G network access rate, number of 5G network users, number of 5G customer complaints, number of VoNR users.

[0031] This generates the basic dataset of questioning techniques and the basic dataset of entities for the standard questioning scenario of "multi-time period regional indicator query".

[0032] Understandably, the same method is used to extract questioning phrases and entity information for other standard questioning scenarios, and to generate corresponding questioning phrase base datasets and entity base datasets, which will not be elaborated here.

[0033] For each standard questioning scenario, based on the aforementioned questioning script base dataset and entity base dataset, a number of sample questioning data are automatically generated by the program using a random permutation and combination of questioning scripts and entities, thereby generating a training dataset for the large language model. The number of training data for each standard questioning scenario should be kept consistent, and the training dataset for each standard questioning scenario should cover all questioning scripts and entity information in the relevant base dataset, while ensuring that the number of questioning scripts and entity information is balanced for each scenario.

[0034] Based on the training dataset, a large language model for data querying is trained, fine-tuned, and tested to obtain the final data query model. This data query model is used to generate an executable data query analysis language, such as an SQL query statement, based on the user's input target query data. By executing the data query analysis language, the user can retrieve the data information they need.

[0035] By employing the technical means of this invention, standard questioning scenarios are constructed, and historical questioning data under different standard questioning scenarios are organized into a basic dataset using questioning phrases and entity information. Based on the basic dataset, a training dataset is automatically generated by randomly arranging and combining questioning phrases and entities to cover all questioning phrases and entity information in the relevant basic dataset, ensuring a balanced quantity of each questioning phrase and entity information. A sufficient and balanced training dataset lays the foundation for improving the quality of model training. By combining the training datasets corresponding to each standard questioning scenario to train and fine-tune the data query model, the accuracy of the data query model and the accuracy of the data query analysis language generated by the model are improved.

[0036] As a preferred embodiment, this invention further implements the above embodiments. In step S12, that is, after extracting questioning scripts and entity information from several historical questioning data under the same standard questioning scenario and generating the basic dataset of questioning scripts and the basic dataset of entities under the standard questioning scenario, the method further includes step S15: S15. Expand the questioning script base dataset and the entity base dataset respectively to obtain the questioning script extended dataset and the entity extended dataset; wherein, the questioning script extended dataset includes the target number of questioning scripts under the standard questioning scenario, and the entity extended dataset includes at least one entity type and the target number of entity information under the entity type.

[0037] Then step S13, which is to randomly combine the questioning scripts in the questioning script base dataset with the entity information under the corresponding entity type in the entity base dataset to generate several sample questioning data, specifically involves: The question phrases in the question phrase extension dataset are randomly combined with the entity information under the corresponding entity type in the entity extension dataset to generate several sample question phrases.

[0038] In this embodiment of the invention, in order to further improve the training and fine-tuning effect of the data query model and enhance its accuracy, the data in the basic dataset is expanded to generate a target number of training data for each standard question scenario.

[0039] Preferably, the step of expanding the basic dataset of questioning phrases and the basic dataset of entities to obtain expanded datasets of questioning phrases and entities includes: After randomly sorting all the questioning scripts in the basic questioning script dataset, they are repeatedly added to the extended questioning script dataset until the number of questioning scripts in the extended questioning script dataset reaches the target number, thus obtaining the extended questioning script dataset. After randomly sorting all entity information under each entity type in the entity base dataset, the information is repeatedly added to the corresponding entity type in the entity extended dataset until the number of entity information for each entity type in the entity extended dataset reaches the target number, thus obtaining the entity extended dataset.

[0040] Specifically, suppose q is a basic dataset of questioning phrases for a certain standard questioning scenario, and e1, e2...e m Let q be the basic dataset of m entity types involved, N be the number of sample question data to be generated for this standard questioning scenario, and Q be the extended dataset of questioning phrases based on q, E1, E2...E m Based on e1, e2...e m Let D be an extended entity extended dataset of m entity types, and let D be a sample question dataset for this standard question scenario. The sample question dataset D is generated through the following operations: After randomly sorting the question-and-answer script data in the basic question-and-answer script dataset q, add them sequentially to the extended question-and-answer script dataset Q. Repeat this operation until the extended question-and-answer script dataset Q has N data points. The entity base dataset e1, e2...e m After each data point is randomly sorted, it is added sequentially to the entity extended datasets E1, E2...E m In the middle, repeat this operation repeatedly until E1, E2...E m Each has N data points; Perform the following operations to generate a sample question dataset D containing N sample questions: for i=1..N do; Based on the entity information involved in the question phrasing Q[i], connect Q[i] with E1[i], E2[i]...E m [i] Combine and generate sample question data D[i], and add it to the sample question dataset D.

[0041] By employing the technical means of this invention, the data in the basic dataset is expanded to cover all possible user expressions. This helps to solve the problems of incomplete coverage of historical question scenarios and non-standard user expressions. Compared with training only using the original historical data, the expanded samples contain more expression variations, which helps to improve classification accuracy. This invention addresses the problems of data scarcity, overfitting, and weak generalization ability by expanding data diversity at a low cost.

[0042] As a preferred embodiment, the present invention is further implemented based on any of the above embodiments, and the data query model includes an intent recognition sub-model, an entity extraction sub-model, and a language generation sub-model; The intent recognition sub-model is used to identify the standard question scenario corresponding to the target question data. The entity extraction sub-model is used to extract the entity type and entity information of the target question data and standardize the target question data. The language generation sub-model is used to generate the data query and analysis language corresponding to the target question data.

[0043] Then, step S14, namely generating a training dataset based on the sample query data and training and adjusting the preset data query model based on the training dataset, includes: The first training data is generated by labeling the standard questioning scenarios corresponding to the sample questioning data. Based on the sample query data, the corresponding entity type and entity information are labeled to generate the second training data; Based on the data query analysis language labeled with the sample question data, a third training data is generated; The intent recognition sub-model is trained and adjusted based on the first training data, the entity extraction sub-model is trained and adjusted based on the second training data, and the language generation sub-model is trained and adjusted based on the third training data.

[0044] In this embodiment of the invention, the data query model processes the target query data input by the user through three stages: intent recognition, entity extraction, and data query analysis language generation, and generates accurate data query analysis language. These are implemented using three sub-models: intent recognition sub-model, entity extraction sub-model, and language generation sub-model.

[0045] Therefore, for various standard question scenarios, training datasets are generated using three stages: intent recognition, entity extraction, and data query analysis. These datasets are then used to train, fine-tune, and test the sub-models for each stage.

[0046] The intent recognition sub-model is used to identify the standard question scenario to which the target question data belongs. Its first training data in the training dataset consists of sample question data and corresponding labeled standard question scenario tags, in the format "sample question data - standard question scenario tags". The entity extraction sub-model is used to extract entity types and entity information from the target question data and standardize the question data based on the entity extraction results. Its second training data in the training dataset consists of sample question data and corresponding labeled entity annotation results, including entity types and entity information, in the format "sample question data - entity annotation results". The language generation sub-model is used to generate the data query analysis language corresponding to the target question data. Its third training data in the training dataset consists of standardized question data and corresponding labeled data query analysis language tags, in the format "standardized question data - data query analysis language".

[0047] Taking the above-mentioned standard question scenario of "multi-time period regional indicator query" as an example, the first training data is shown in Table 1: Table 1

[0048] The second training data is shown in Table 2 as an example: Table 2

[0049] The third training data is shown in Table 3 as an example: Table 3

[0050] Understandably, the same method is used to generate corresponding training data for other standard questioning scenarios, which will not be elaborated here.

[0051] The intent recognition sub-model, entity extraction sub-model, and language generation sub-model were trained and fine-tuned using the first training data, the second training data, and the third training data, respectively.

[0052] The following example illustrates the model training and fine-tuning process: The training dataset is cleaned and partitioned, including deduplication, error correction, and completion. It is then divided into a training set, a validation set, and a test set in a 7:2:1 ratio for subsequent model training and evaluation. The training set is used for learning model parameters; the validation set is used to adjust hyperparameters (such as the learning rate) during training to avoid overfitting; and the test set is used to evaluate the final performance of the model (such as the accuracy of SQL statement generation).

[0053] This invention embodiment fine-tunes the large model for each of the three core stages: intent recognition, entity extraction, and SQL statement generation. Taking the language generation sub-model as an example, the base model is selected based on business complexity and hardware resources. If complex SQL (such as multi-table joins and nested queries) is required, a model with a large number of parameters (such as LLaMA-7B or ChatGLM-6B) is selected; if only simple single-table queries are required, a lightweight model (such as BART-base or T5-small) is selected. This embodiment uses LLaMA-7B. Next, the environment is set up, including hardware, software, and environment configuration. The partitioned dataset is converted into an input format that the model can recognize. The core is to construct "Prompt-Response" pairs to adapt to the generation logic of the large model.

[0054] Prompt Construction: Includes "Scene Description + User Questions + Entity Hints + Formatting Requirements," clearly defining the boundaries of the model task. Example: plaintext Task: Generate an SQL statement based on user queries to query the network_kpi table (fields: kpi_date - date, kpi_city - region, kpi_name - metric, kpi_value - value, kpi_interval - statistical granularity).

[0055] A user asked: "I would like to check the 5G network access rate in City A on September 4, 2024." Entity information: Time = 20240904, Region = City A, Indicator = 5G network access rate, Granularity = Daily; Requirements: The SQL statement must include a WHERE condition, return only the kpi_value field, and sort in descending order of kpi_date.

[0056] Response Construction: The correct SQL statement corresponding to the Prompt. Example: SQL SELECT kpi_value FROM network_kpi WHERE kpi_date='20240904' AND kpi_city='City A' AND kpi_name='5G network access rate' AND kpi_interval='Day' ORDER BY kpi_dateDESC; Text Encoding: Encode the Prompt and Response using the corresponding LLaMA tokenizer. Convert the text into a sequence of token IDs; standardize the sequence length (e.g., set a maximum length of 512, truncate if too long, use `undefined` if too short).<pad>(Complete the missing information); generate an attention mask and mark the positions of valid tokens.

[0057] Further fine-tuning of the model is performed using a LoRA (Low-Rank Adaptation) fine-tuning strategy. Only the newly added low-rank matrix parameters (approximately 0.1% of the total parameters) of the model's attention layer are trained, avoiding the high computational cost of training all parameters. Specific steps are as follows: First, configure LoRA and set training parameters. This includes rank, learning rate, dropout (to prevent overfitting), target layer, batch size, number of training epochs, optimizer, and learning rate scheduling. After each training epoch, calculate the SQL syntax accuracy (whether the generated SQL is executable) and business matching accuracy (whether the SQL meets the user's actual needs) using the validation set. If the validation set accuracy does not improve for two consecutive epochs, trigger early stopping to avoid overfitting. Record training logs (loss value, accuracy) for subsequent analysis.

[0058] The performance of the fine-tuned model is evaluated using a test set. Key metrics include SQL syntax accuracy and business matching accuracy. SQL syntax accuracy represents the proportion of generated SQL queries that can be directly executed in the database; business matching accuracy represents the proportion of SQL query results that match user requirements. For example, in a test set of 100 samples, 96 SQL queries are grammatically correct, and 92 are business-matched, meeting the target requirements. If the syntax accuracy is low, SQL syntax error samples are added to the training set for further fine-tuning; if the business matching accuracy is low, the types of error samples are analyzed (e.g., incorrect geographic entity mapping, incorrect time format conversion), the entity database and mapping rules are optimized, and corresponding samples are added before retraining.

[0059] The fine-tuned LoRA weights are merged with the base model weights to generate the final model, which is then integrated into the model training and induction module via an API interface for use in the data query process.

[0060] It should be noted that the fine-tuning process for the intent recognition sub-model and the entity extraction sub-model is the same as that for the language generation sub-model, requiring only adjustments to the dataset and the Prompt design. The training dataset for the intent recognition sub-model is "question data – scene labels," and the Prompt focuses on identifying the intent category of the user's question; the training dataset for the entity extraction sub-model is "user question data – entity annotation results," and the Prompt focuses on extracting specific types of entities from the user's question.

[0061] Understandably, the above training and fine-tuning scenarios are only illustrative examples. In actual applications, the generated training data can be used to train and fine-tune each sub-model according to the actual situation to improve the accuracy of the model. No specific limitations are made here.

[0062] It should be noted that for non-standard question scenarios, open-source NL2SQL datasets (spider, cspider, dusql, bird, etc.) were used to fine-tune the data query analysis language generation process. Non-standard question scenarios refer to query scenarios that are infrequent, flexible in expression, and cannot be templated. The core challenge of these scenarios is the high diversity of questions and the complexity of business logic, requiring the use of open-source NL2SQL datasets to improve the model's generalization ability.

[0063] By employing the technical means of this invention, user-generated questions are transformed into standardized questions based on intent recognition and entity extraction results. Then, a fine-tuned large model is used to generate data query analysis language based on the standardized questions. This can improve the accuracy of data query analysis language generation and reduce the dataset required for fine-tuning training of the large model in the data query analysis language generation process, thereby improving the efficiency of fine-tuning training.

[0064] See Figure 2 This is a flowchart illustrating a data query method provided in an embodiment of the present invention. The embodiment of the present invention also provides a data query method, the method comprising steps S21 to S23: S21. Receive the user's target question data; S22. The target query data is input into a data query model for processing to obtain an executable data query analysis language generated by the data query model; wherein, the data query model is generated according to the training method of the data query model as described in any of the above embodiments; S23. The data query analysis language is sent to a preset database for execution to obtain the query results corresponding to the target query data.

[0065] In this embodiment of the invention, the data query model that has been constructed and trained is deployed and applied. When the target question data input by the user is received, the target question data is input into the data query model for processing to obtain an executable data query analysis language, such as an SQL statement, generated by the data query model. The SQL statement is then executed to obtain the query results required by the user.

[0066] Understandably, after the data query is completed, the target query data, as historical query data, participates in the real-time fine-tuning of the data query model to continuously improve the accuracy of the data query model.

[0067] See Figure 3 This is a flowchart illustrating a preferred data query method provided in an embodiment of the present invention. As a preferred implementation, the data query model includes an intent recognition sub-model, an entity extraction sub-model, and a language generation sub-model.

[0068] Then step S22, which is the input of the target query data into the data query model for processing to obtain the executable data query analysis language generated by the data query model, includes steps S221 to S223: S221. Input the target question data into the intent recognition sub-model to perform intent recognition, and obtain the standard question scenario output by the intent recognition sub-model. S222. Input the target question data and the standard question scenario into the entity extraction sub-model for entity extraction, and standardize the target question data according to the entity extraction result to obtain the standardized question data output by the entity extraction sub-model; wherein, the entity extraction result includes entity type and its entity information; S223. Input the standardized question data and the standardized question scenario into the language generation sub-model for processing to obtain the data query analysis language generated by the language generation sub-model.

[0069] Preferably, the language generation sub-model is used to call the data information table and field definition information corresponding to the standard question scenario according to the preset mapping relationship between the standard question scenario, the data information table, and the field definition information, and generate a data query and analysis language based on the data information table, the field definition information, and the standardized question data.

[0070] In the initial data preparation phase, it is also necessary to organize all data tables and field definition information, including field value ranges, enumerated values, and sample data. For standard question scenarios, a mapping relationship between the scenario and the data table and field definition information should be established.

[0071] Data table information is the foundation for the model to generate correct SQL statements. The model needs to clearly define which table to query, which fields to use, and what the range of field values ​​is. Therefore, it is necessary to organize all data table information in advance and establish a mapping relationship between standard scenarios and tables to avoid the model generating erroneous statements with meaningless fields or non-existent tables.

[0072] For example, for the scenario of "multi-time-period regional indicator query", a mapping relationship between the scenario and the data table is established. The associated data table is: network_kpi_data (stores network indicator data); key field definitions are: kpi_date (indicator date, format YYYYMMDD), kpi_city (indicator region, such as "City A"), kpi_name (indicator name, such as "5G network access rate"), kpi_value (indicator value), and kpi_interval (statistical granularity, such as "day").

[0073] In this embodiment of the invention, after receiving the target question data input by the user, the target question data is input into the data query model, and processed by the intent recognition sub-model, entity extraction sub-model and language generation sub-model respectively, and an executable data query analysis language is output.

[0074] Intent recognition: The user's question and prompt information are input into the big model. The big model provides the intent recognition result of the user's question and determines the standard question scenario to which the user's question belongs based on the intent recognition result. If the big model cannot recognize the intent, the user's question is marked as a non-standard question scenario. If the intent recognition result is not related to data query analysis, the user's question is marked as a non-data query scenario, and the abnormal intent recognition result is fed back to the front-end application.

[0075] The entity extraction sub-model includes three stages: entity extraction, entity standardization, and question standardization.

[0076] Entity Extraction: Based on the standard question scenario and entity extraction prompt template, generate prompt information. Input the user's question and prompt information into the large model, and the large model will provide the entity extraction results. If the entity extraction results of the user's question lack necessary key entity information, the abnormal entity extraction results of the user's question will be fed back to the front-end application. If the user's question lacks non-key entity information, the default entity information will be supplemented.

[0077] Entity standardization: Completes fuzzy matching and standardization processing of entity extraction results, converting entity information into a standardized format.

[0078] Question standardization: Based on the standardized user question template, the original user questions are transformed into standardized user questions using standardized entity information.

[0079] Data query analysis language generation: For standardized user question scenarios, relevant data tables and field definitions are selected, and prompt information is generated based on prompt word templates. The standardized user question and prompt information are input into the large model, which then provides the data query analysis language generation results. For non-standardized user question scenarios, all data tables and field definitions are selected, and prompt information is generated based on prompt word templates. The original user question and prompt information are input into the large model, which then provides the data query analysis language generation results. The data query analysis language generation results are then fed back to the front-end application, which records historical user question data.

[0080] For example, the system receives the user's target query data: "I want to see the 5G network access rate data for City A in the last 3 days." The system combines the user's target query data with preset prompts (e.g., "Please identify the user's intent and determine whether it belongs to a standard query scenario, including multi-time-period regional indicator queries, single-region real-time indicator queries, etc.") and inputs it into a finely tuned intent recognition sub-model. After model analysis, the system returns the intent recognition result. For example, if it belongs to the standard query scenario of "multi-time-period regional indicator queries," there are no anomalies. Understandably, if it is identified as a non-data query scenario, such as "inquiry about 5G packages," the system directly rejects the query from the front end. Based on the recognition result, the system calls the subsequent processing rules corresponding to the "multi-time-period regional indicator queries" scenario, such as entity extraction templates and data table association logic.

[0081] Furthermore, based on the scenario of "multi-time-period regional indicator query," the system automatically generates entity extraction prompts: "Please extract three types of entities from the user's query: time, region, and indicator. If non-key entities are missing, add default values ​​(e.g., the default region is 'the whole province'). If key entities (time, indicator) are missing, display an error message." The user's query and prompts are input into the entity extraction sub-model, and the model outputs the entity extraction results. Time frame: the last 3 days; Geographic entity: City A; Indicator entity: 5G network access rate; Result assessment: All three types of key entities are complete, no default values ​​need to be added, and there are no anomalies.

[0082] Understandably, if a user asks "I want to see the 5G network access rate" without a time entity, the default value "the last day" will be automatically added.

[0083] Because user queries may contain non-standard expressions, the system needs to standardize them. The time entity "last 3 days," combined with the query date (assumed to be September 17, 2024), is converted to the standardized date range "20240915-20240917". The region entity "City A" is matched against the region entity database and simplified to the standard format "City A", consistent with the value of the `kpi_city` field in the data table. The indicator entity "5G network access rate" is directly matched against the indicator entity database, confirming the standard name "5G network access rate", consistent with the value of the `kpi_name` field in the data table. Finally, the standardized entities are output: Time (20240915-20240917), Region (City A), and Indicator (5G network access rate).

[0084] The system calls the standardized query template for the "Multi-Time Period Regional Indicator Query" scenario (querying data at the [statistical granularity] of [time range][region][indicator]), substitutes standardized entities into the template, and generates a standardized query: "Query the data of the daily granularity indicator 5G network access rate in City A from September 15, 2024 to September 17, 2024." This step solves the problem of diverse user query expressions, provides a unified input for subsequent SQL generation, and improves the accuracy of model generation.

[0085] Furthermore, based on standardized questions, the system automatically generates SQL generation prompts: "Please generate SQL based on the data table network_kpi_data, according to the following conditions: statistical granularity kpi_interval='day', region kpi_city='City A', indicator kpi_name='5G network access rate', date range kpi_date between '20240915' and '20240917', must include indicator value kpi_value, and sorted in descending order by date." The standardized questions, prompts, and data table field definitions are input into the fine-tuned language generation sub-model, and the model outputs executable data query and analysis language, such as SQL statements. SQL SELECT kpi_date, kpi_city, kpi_name, kpi_value FROM network_kpi_data WHERE kpi_interval = 'Day' AND kpi_city = 'City A' AND kpi_name = '5G network access rate' AND kpi_date BETWEEN '20240915' AND '20240917' ORDER BY kpi_date DESC; The system sends the generated SQL statement to the database for execution and obtains the query results, such as 92% for City A on September 17, 2024, 91.5% on September 16, 2024, and 90.8% on September 15, 2024. It also records the user's original question, standardized question, and SQL statement to the historical log for subsequent training dataset updates.

[0086] The system displays the indicator data returned by the database to the user through the APP's visual interface, supporting various formats such as list viewing and trend charts. It can also provide voice broadcast functionality to complete the entire query process.

[0087] By employing the technical means in this embodiment of the invention, a data query model obtained through real-time training and fine-tuning based on historical question data is used to process the target question data input by the user to generate an executable data query analysis language. The data query analysis language is then executed to query the data information required by the user. This helps to improve the generation accuracy of the data query analysis language, thereby improving the accuracy of the queried data information and enhancing the user experience.

[0088] See Figure 4 This is a structural diagram of the data query and analysis system provided in an embodiment of the present invention. The functional structure diagram of the data query and analysis system based on a large language model is shown below. Figure 4 As shown, the system includes a system resource module, a model training and promotion module, and a capability gateway module. The main functions of each module are as follows: System resource module: includes computing resources (CPU server, GPU server), storage resources, network resources, etc. required for system deployment.

[0089] The model training and inference module includes: (1) Base Model Management: A large model platform that manages the training and inference of tasks in each stage, and provides API interfaces for users to ask questions to agents in each stage; (2) Dataset Management: Provides management functions for training datasets and evaluation datasets required for fine-tuning training of agents in various stages of application scenarios, and can update datasets based on log data; (3) Training Management: Provides fine-tuning training management functions for agent application scenarios at each stage, supports commonly used fine-tuning training methods such as ptunning and Lora, and uses updated datasets and training methods to complete the fine-tuning training of large models, thereby improving the accuracy of the generated content of large models in agent application scenarios at each stage. (4) Evaluation Management: Provides benchmark testing and evaluation management functions for the training results of agent application scenario models at each stage; (5) Agent orchestration management: Provides process orchestration, scheduling and management functions for agents at each stage, realizing end-to-end data query, analysis and processing flow; (6) Log management: Provides detailed log recording and statistical analysis functions for the data query and analysis process results.

[0090] Capability Gateway: Provides API service interface management functions for data query and analysis based on large language models, supporting the construction of data query and analysis functions for front-end APP applications or PC application systems across various channels.

[0091] See Figure 5 This is a schematic diagram of the structure of a data query model training device provided in an embodiment of the present invention. The present invention also provides a data query model training device 10, comprising: The historical data acquisition module 11 is used to collect historical question data in real time and determine the standard question scenario corresponding to the historical question data; The basic dataset generation module 12 is used to extract questioning scripts and entity information from several historical questioning data under the same standard questioning scenario, and generate a basic dataset of questioning scripts and a basic dataset of entities under the standard questioning scenario respectively; wherein, the basic dataset of questioning scripts includes several questioning scripts under the standard questioning scenario, and the basic dataset of entities includes at least one entity type and several entity information under the entity type. The sample question data generation module 13 is used to randomly combine the question scripts in the question script base dataset with the entity information under the corresponding entity type in the entity base dataset to generate several sample question data. The training dataset generation module 14 is used to generate a training dataset based on the sample question data, and to train and adjust a preset data query model based on the training dataset; wherein, the data query model is used to generate a data query analysis language based on the target question data.

[0092] In a preferred embodiment, the device further includes: An extended dataset generation module is used to extend the questioning script base dataset and the entity base dataset respectively to obtain an extended questioning script dataset and an extended entity dataset; wherein, the extended questioning script dataset includes a target number of questioning scripts under the standard questioning scenario, and the extended entity dataset includes at least one entity type and a target number of entity information under the entity type.

[0093] It should be noted that the training device for a data query model provided in this embodiment of the invention is used to execute all the process steps of the training method for a data query model in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.

[0094] See Figure 6 This is a schematic diagram of the structure of a data query device provided in an embodiment of the present invention. The embodiment of the present invention provides a data query device 20, comprising: The target question data receiving module 21 is used to receive the user's target question data; The query analysis language generation module 22 is used to input the target query data into the data query model for processing, and obtain an executable data query analysis language generated by the data query model; wherein, the data query model is generated according to the training method of the data query model as described in any of the above embodiments; The data query result generation module 23 is used to send the data query analysis language to a preset database for execution to obtain the query results corresponding to the target query data.

[0095] It should be noted that the data query device provided in this embodiment of the invention is used to execute all the process steps of the data query method in the above embodiment. The working principle and beneficial effect of the two are one-to-one, so they will not be described again.

[0096] This invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the training of a data query model as described in any of the foregoing embodiments, or the data query method as described in any of the foregoing embodiments.

[0097] This invention also provides a computer-readable storage medium, which includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform training of a data query model as described in any of the above embodiments, or a data query method as described in any of the above embodiments.

[0098] This invention also provides a computer program product, which includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, they implement the training of the data query model as described in any of the above embodiments, or the data query method as described in any of the above embodiments.

[0099] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0100] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.< / pad>

Claims

1. A training method for a data query model, characterized in that, include: Collect historical question data in real time and determine the standard question scenario corresponding to the historical question data; Based on several historical question data under the same standard question scenario, questioning scripts and entity information are extracted to generate a basic dataset of questioning scripts and a basic dataset of entities under the standard question scenario, respectively; wherein, the basic dataset of questioning scripts includes several questioning scripts under the standard question scenario, and the basic dataset of entities includes at least one entity type and several entity information under the entity type; The questioning scripts in the questioning script base dataset are randomly combined with the entity information under the corresponding entity type in the entity base dataset to generate several sample questioning data. A training dataset is generated based on the sample question data, and a preset data query model is trained and adjusted based on the training dataset; wherein, the data query model is used to generate a data query analysis language based on the target question data.

2. The training method for the data query model as described in claim 1, characterized in that, After extracting questioning phrases and entity information from several historical questioning data points under the same standard questioning scenario, and generating a basic dataset of questioning phrases and a basic dataset of entities under the standard questioning scenario, the method further includes: The questioning script base dataset and the entity base dataset are respectively expanded to obtain the questioning script extended dataset and the entity extended dataset; wherein, the questioning script extended dataset includes the target number of questioning scripts under the standard questioning scenario, and the entity extended dataset includes at least one entity type and the target number of entity information under the entity type; The step of randomly combining the questioning phrases in the questioning phrase base dataset with the entity information under the corresponding entity type in the entity base dataset to generate several sample questioning data is as follows: The question phrases in the question phrase extension dataset are randomly combined with the entity information under the corresponding entity type in the entity extension dataset to generate several sample question phrases.

3. The training method for the data query model as described in claim 2, characterized in that, The step of expanding the questioning script base dataset and the entity base dataset to obtain the questioning script expanded dataset and the entity expanded dataset includes: After randomly sorting all the questioning scripts in the basic questioning script dataset, they are repeatedly added to the extended questioning script dataset until the number of questioning scripts in the extended questioning script dataset reaches the target number, thus obtaining the extended questioning script dataset. After randomly sorting all entity information under each entity type in the entity base dataset, the information is repeatedly added to the corresponding entity type in the entity extended dataset until the number of entity information for each entity type in the entity extended dataset reaches the target number, thus obtaining the entity extended dataset.

4. The training method for the data query model as described in claim 1, characterized in that, The data query model includes an intent recognition sub-model, an entity extraction sub-model, and a language generation sub-model. The intent recognition sub-model is used to identify the standard question scenario corresponding to the target question data. The entity extraction sub-model is used to extract the entity type and entity information of the target question data and standardize the target question data. The language generation sub-model is used to generate the data query and analysis language corresponding to the target question data.

5. The training method for the data query model as described in claim 4, characterized in that, The step of generating a training dataset based on the sample query data and training and adjusting a preset data query model based on the training dataset includes: The first training data is generated by labeling the standard questioning scenarios corresponding to the sample questioning data. Based on the sample query data, the corresponding entity type and entity information are labeled to generate the second training data; Based on the data query analysis language labeled with the sample question data, a third training data is generated; The intent recognition sub-model is trained and adjusted based on the first training data, the entity extraction sub-model is trained and adjusted based on the second training data, and the language generation sub-model is trained and adjusted based on the third training data.

6. A data query method, characterized in that, include: Receive the user's target question data; The target query data is input into a data query model for processing to obtain an executable data query analysis language generated by the data query model; wherein, the data query model is generated according to the training method of the data query model as described in any one of claims 1 to 5; The data query analysis language is executed to obtain the query results corresponding to the target query data.

7. The data query method as described in claim 6, characterized in that, The data query model includes an intent recognition sub-model, an entity extraction sub-model, and a language generation sub-model. The step of inputting the target query data into a data query model for processing to obtain an executable data query analysis language generated by the data query model includes: The target question data is input into the intent recognition sub-model for intent recognition, and the standard question scenario output by the intent recognition sub-model is obtained. The target question data and the standard question scenario are input into the entity extraction sub-model for entity extraction, and the target question data is standardized based on the entity extraction results to obtain standardized question data output by the entity extraction sub-model; wherein, the entity extraction results include entity type and entity information; The standardized question data and the standardized question scenario are input into the language generation sub-model for processing to obtain the data query analysis language generated by the language generation sub-model.

8. The data query method as described in claim 7, characterized in that, The language generation sub-model is used to call the data information table and field definition information corresponding to the standard question scenario according to the mapping relationship between the preset standard question scenario, data information table and field definition information, and generate a data query and analysis language based on the data information table, the field definition information and the standardized question data.

9. A training device for a data query model, characterized in that, include: The historical data acquisition module is used to collect historical question data in real time and determine the standard question scenario corresponding to the historical question data. The basic dataset generation module is used to extract questioning scripts and entity information from several historical questioning data under the same standard questioning scenario, and generate a basic dataset of questioning scripts and a basic dataset of entities under the standard questioning scenario respectively; wherein, the basic dataset of questioning scripts includes several questioning scripts under the standard questioning scenario, and the basic dataset of entities includes at least one entity type and several entity information under the entity type; The sample question data generation module is used to randomly combine the question scripts in the question script base dataset with the entity information under the corresponding entity type in the entity base dataset to generate several sample question data. The training dataset generation module is used to generate a training dataset based on the sample question data, and to train and adjust a preset data query model based on the training dataset; wherein, the data query model is used to generate a data query analysis language based on the target question data.

10. A data query device, characterized in that, include: The target question data receiving module is used to receive the user's target question data; A query analysis language generation module is used to input the target query data into a data query model for processing, and obtain an executable data query analysis language generated by the data query model; wherein, the data query model is generated according to the training method of the data query model as described in any one of claims 1 to 5; The data query result generation module is used to execute the data query analysis language to obtain the query results corresponding to the target query data.

11. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement training of a data query model as described in any one of claims 1 to 5, or a data query method as described in any one of claims 6 to 8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform training of the data query model as described in any one of claims 1 to 5, or the data query method as described in any one of claims 6 to 8.

13. A computer program product, characterized in that, The computer program product includes a computer program or computer instructions, which, when executed by a processor, implement the training of the data query model as described in any one of claims 1 to 5, or the data query method as described in any one of claims 6 to 8.