Data operation instruction generation method and device, program product and storage medium
By extracting relevant information and configuring the database from historical data operation instructions, and generating data operation configurations in conjunction with user questions, the problem of low accuracy in the conversion of natural language to SQL statements is solved, and more efficient and accurate data operations are achieved.
Patent Information
- Application Number
- CN202510864925.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-28
AI Technical Summary
Existing natural language to SQL statement conversion technologies cannot accurately reflect the user's real needs when dealing with complex business scenarios, especially in multi-table queries and wide-table queries where accuracy is low.
By identifying historical data operation instructions that meet preset conditions from the first database, analyzing the data information generated by these instructions, and combining them with user questions to generate data operation configurations, data operation instructions are generated using various network models and configuration databases, including data transformation, keyword matching, and feedback information optimization.
It improves the accuracy and efficiency of data operation instructions, reduces communication barriers between users and the system, and enhances the system's self-learning ability and adaptability.
Smart Images

Figure CN120849447A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more specifically, to a method and apparatus for generating data operation instructions, a program product, and a storage medium. Background Technology
[0002] In traditional data manipulation request processing, users typically need a certain level of knowledge in Structured Query Language (SQL) to accurately send requests to the database and perform operations. This approach not only places high technical demands on users, but also often limits efficiency and accuracy when handling complex queries, especially those involving multi-table joins, time series analysis, or specific business logic.
[0003] To lower the barrier to entry for users and improve the convenience and intelligence of data operations, Natural Language to SQL (NL2SQL) technology has developed rapidly in recent years. NL2SQL aims to enable users to perform database queries using natural language (such as everyday speech or informal text) without needing to master professional SQL syntax. However, NL2SQL technology faces multiple challenges when handling complex queries in real-world business scenarios, especially in multi-table and wide-table queries. Accurately inferring the required SQL statement from the user's query is a problem that urgently needs to be solved. The traditional approach is to directly input the user's query into a pre-trained large model, which then attempts to generate an SQL statement, followed by simple validation or correction. While this method is intuitive, when facing complex business scenarios, the large model's lack of in-depth understanding of specific business domain terminology, field aliases, field enumeration values, metric aliases, and timeliness knowledge can lead to generated SQL statements that may not accurately reflect the user's true needs. Summary of the Invention
[0004] This application provides a method and apparatus for generating data operation instructions, a program product, and a storage medium to at least solve the problem of low accuracy in generating data operation instructions in related technologies.
[0005] According to one embodiment of this application, a method for generating data operation instructions is provided, comprising: determining N historical data operation instructions that satisfy a first preset condition from a first database based on an acquired user question, wherein the historical data operation instructions are operation instructions generated based on historical user questions, and N is a positive integer greater than or equal to 1; determining target data information from historical data information generated when executing the N historical data operation instructions; determining a data operation configuration for the target data information and the user question, wherein the data operation configuration is a configuration determined based on semantic parsing of the target data information and the user question; and generating a data operation instruction based on the data operation configuration, wherein the data operation instruction is used to perform data operations matching the user question in a second database.
[0006] In an exemplary embodiment, determining N historical data operation instructions that satisfy a first preset condition from a first database based on an acquired user question includes: performing a data transformation operation on the user question to obtain a target user question represented by a vector, wherein the target user question includes basic semantic information of the user question; determining M historical user questions that satisfy a third preset condition from the first database based on the target user question, wherein M is a positive integer greater than or equal to 1; and using the M historical user questions to determine N historical data operation instructions that satisfy the first preset condition from the first database.
[0007] In one exemplary embodiment, determining target data information from historical data information generated when executing N historical data operation instructions includes: determining P first data tables from the first database, wherein the P first data tables are data tables determined when executing N historical data operation instructions, and P is a positive integer greater than or equal to 1; determining Q target data tables from the P first data tables whose data operation frequency satisfies a first preset threshold, wherein Q is a positive integer less than or equal to P; determining R target fields from the Q target data tables, wherein the R target fields are fields determined when executing N historical data operation instructions, and all R target fields are fields whose usage frequency satisfies a second preset threshold, and R is a positive integer greater than or equal to 1; and determining the Q target data tables and R target fields as the target data information.
[0008] In one exemplary embodiment, determining the target data information and the data operation configuration of the user question includes: determining a first keyword in the user question; determining a first triggering method matching the first keyword from a third database; determining first configuration information corresponding to the first keyword from the third database using the first triggering method; and generating the data operation configuration based on the first configuration information and the user question.
[0009] In an exemplary embodiment, determining the first configuration information corresponding to the first keyword from the third database using the first triggering method includes at least one of the following: determining the keyword matching the first keyword from the third database using a regular expression to obtain the first configuration information; determining the keyword matching the first keyword based on the matching degree between the first keyword and the keywords in the configuration information stored in the third database to obtain the first configuration information; and inputting the first keyword into a pre-trained first network model to determine the keyword matching the first keyword from the third database through the first network model to obtain the first configuration information.
[0010] In one exemplary embodiment, generating a data operation instruction based on the aforementioned data operation configuration includes: inputting the aforementioned user question and the aforementioned data operation configuration into a preset prompt word template to generate a target prompt word corresponding to the aforementioned user question through the preset prompt word template; and inputting the aforementioned target prompt word into a pre-trained second network model to generate the aforementioned data operation instruction through the aforementioned second network model.
[0011] In one exemplary embodiment, after generating a data operation instruction based on the above data operation configuration, the method further includes: obtaining user feedback information; generating a target data operation instruction that satisfies a fourth preset condition based on the feedback information and using a pre-trained third network model, wherein the target data operation instruction is used to perform data operations matching the user's question in the second database.
[0012] In an exemplary embodiment, after generating a target data operation instruction that satisfies a fourth preset condition based on the feedback information and using a pre-trained third network model, the method further includes: acquiring first data information generated when executing the target data operation instruction; associating the first data information with the user question and storing it in the first database to update the first database.
[0013] In an exemplary embodiment, after generating a target data operation instruction that satisfies a fourth preset condition based on the feedback information and using a pre-trained third network model, the method further includes: inputting the target data operation instruction and the user question into the pre-trained fourth network model to determine first information through the fourth network model, wherein the first information is information not included in the user question and the data operation configuration; determining a second triggering method and second configuration information based on the first information, wherein the second configuration information is information parsed from the first information; and associating and storing the second triggering method and the second configuration information in a third database, wherein the third database stores historical triggering methods and historical configuration information corresponding to the historical user questions.
[0014] In an exemplary embodiment, storing the second triggering method and the second configuration information in a third database includes: when the third database includes target historical configuration information corresponding to the second configuration information, inputting the target historical configuration information, the second triggering method, and the second configuration information into a pre-trained fifth network model to output fused information through the fifth network model; and storing the fused information in the third database.
[0015] According to another embodiment of this application, a data operation instruction generation apparatus is provided, including a first memory, a first processor, and a first computer program stored in the first memory and executable on the first processor. When the first processor executes the first computer program, it performs the following operations: determining N historical data operation instructions that satisfy a first preset condition from a first database based on an acquired user question, wherein the historical data operation instructions are operation instructions generated based on historical user questions, and N is a positive integer greater than or equal to 1; determining target data information from historical data information generated when executing the N historical data operation instructions; determining a data operation configuration for the target data information and the user question, wherein the data operation configuration is a configuration determined based on semantic parsing of the target data information and the user question; and generating a data operation instruction based on the data operation configuration, wherein the data operation instruction is used to perform data operations matching the user question in a second database.
[0016] According to yet another embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0017] According to yet another embodiment of this application, a computer-readable storage medium is also provided, in which a computer program is stored, wherein the computer program is configured to perform the steps in any of the above method embodiments when it is run.
[0018] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0019] This application identifies N matching historical data operation instructions from a first database based on a user's question. Then, it analyzes the historical data information generated by these N instructions to determine target data information. Combining this target data information with the user's question, a data operation configuration is generated. Finally, data operation instructions are generated based on this configuration. Therefore, this approach solves the problem of low accuracy in generating data operation instructions in related technologies, thereby improving the accuracy of the generated instructions. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the hardware environment for a method of generating data operation instructions according to an embodiment of this application;
[0021] Figure 2 This is a flowchart of a method for generating data manipulation instructions according to an embodiment of this application;
[0022] Figure 3 This is a flowchart illustrating a data operation configuration for determining a user problem according to an embodiment of this application;
[0023] Figure 4 This is a flowchart illustrating a data manipulation instruction and an update of a first database according to an embodiment of this application;
[0024] Figure 5 This is a flowchart of a method for generating data operation instructions corresponding to user questions according to an embodiment of this application;
[0025] Figure 6 This is a structural block diagram of a data operation instruction generation device according to an embodiment of this application. Detailed Implementation
[0026] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0028] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a schematic diagram of the hardware environment for a method of generating data operation instructions according to an embodiment of this application. For example... Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0029] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to a data operation instruction generation method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0030] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0031] This embodiment provides a method for generating data manipulation instructions. Figure 2This is a flowchart of a method for generating data manipulation instructions according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:
[0032] Step S202: Based on the acquired user questions, determine N historical data operation instructions from the first database that meet the first preset conditions, wherein the aforementioned historical data operation instructions are operation instructions generated based on historical user questions, and the aforementioned N is a positive integer greater than or equal to 1;
[0033] Optionally, the user problem in this embodiment is a data request made by the user in natural language, including but not limited to simple data requests, such as querying sales figures, or complex business analysis needs, such as "detailed data on new energy power generation in Hebei Province during the Qingming Festival".
[0034] Optionally, the data stored in the first database of this embodiment includes, but is not limited to: historical user questions and corresponding historical data operation instructions, and keywords in the historical data operation instructions (such as the data tables used and the fields in the data tables).
[0035] Optionally, the N historical data operation instructions in this embodiment are data operation instructions extracted from the first database that have a high degree of matching with the current user's question and meet the first preset conditions.
[0036] Optionally, the first preset condition in this embodiment is a rule or standard for filtering historical data operation instructions, including but not limited to the execution efficiency, accuracy, and similarity to user questions of the data operation instructions. For example, the execution success rate of historical operation instructions is required to be higher than a certain threshold, or the semantic similarity to the current user question is required to exceed a preset value.
[0037] Optionally, the data manipulation instructions in this embodiment include, but are not limited to, SQL statements, used to perform operations such as adding, deleting, modifying, and querying data stored in the database.
[0038] Step S204: Determine the target data information from the historical data information generated when executing N of the above-mentioned historical data operation instructions;
[0039] Optionally, the historical data information in this embodiment is the actual data result generated when the historical data operation instruction is executed, including but not limited to: the data table used when the data operation instruction is executed and the information of the fields in the data table.
[0040] Optionally, the target data information in this embodiment is the set of information most relevant to the current user problem, determined through data analysis and business understanding from historical data information generated from N historical data operation instructions.
[0041] Step S206: Determine the data operation configuration of the target data information and the user question, wherein the data operation configuration is determined based on the semantic parsing of the target data information and the user question;
[0042] Optionally, the data operation configuration in this embodiment is a set of rules or parameters formed based on semantic analysis of target data information and user questions to guide the generation of data operation instructions. The data operation configuration includes, but is not limited to: the table name of the data table used, the field name of the data table used, the time range, conditional expressions, etc.
[0043] Step S208: Generate data operation instructions based on the above data operation configuration, wherein the above data operation instructions are used to perform data operations matching the above user question in the second database.
[0044] Optionally, in this embodiment, the second database is the database in which the user question ultimately needs to perform data operations, while the first database stores the data table pointed to by the user question and the specific information.
[0045] Through the above steps, N historical data operation instructions matching the user's question are determined from the first database. Then, the historical data information generated by these N instructions is analyzed to determine the target data information. A data operation configuration is generated by combining the target data information with the user question, and finally, data operation instructions are generated based on this configuration. Therefore, this method solves the problem of low accuracy in generating data operation instructions in related technologies, thereby improving the accuracy of the generated instructions.
[0046] In an exemplary embodiment, determining N historical data operation instructions that satisfy a first preset condition from a first database based on an acquired user question includes: performing a data transformation operation on the user question to obtain a target user question represented by a vector, wherein the target user question includes basic semantic information of the user question; determining M historical user questions that satisfy a third preset condition from the first database based on the target user question, wherein M is a positive integer greater than or equal to 1; and using the M historical user questions to determine N historical data operation instructions that satisfy the first preset condition from the first database.
[0047] Optionally, the data transformation operation in this embodiment refers to converting the original natural language user question into a form that a computer can understand, such as word vectors or semantic vectors. This can be achieved using natural language processing techniques, such as word embedding, Transformer models, or BERT models.
[0048] Optionally, in this embodiment, when performing data transformation operations on the user question to obtain the target user question represented as a vector, a masked matching mechanism is first used to preserve the semantic structure of the user question and reduce the weight of specific field enumeration values in the matching. Then, the user question after the above operations is transformed into the target user question represented in vector form. For example, assuming the user question is "Help me check how much new energy power generation in Hebei Province is during the Qingming Festival," a masked matching mechanism is used to reduce the weight of specific field values such as "during the Qingming Festival" and "Hebei Province" in the representation, resulting in the semantic representation of the user question, "Check the power generation of a certain province within a certain time period." Then, the semantic representation is transformed into vector form to obtain the target user question.
[0049] Optionally, the target user question in this embodiment is a vectorized representation of the question obtained through data transformation operations. It contains the basic semantic information of the original question, including but not limited to the average vector of multiple keywords, sentence vectors or paragraph vectors of the text, which serve as the basis for semantic comparison and historical user question matching in subsequent steps.
[0050] Optionally, the third preset condition in this embodiment is a standard for filtering historical user questions. The third preset condition focuses on the semantic similarity of the user questions themselves, the question category, or the query frequency.
[0051] Optionally, the N historical data operation instructions in this embodiment are selected based on M historical user questions, and are the data operation instructions that meet the first preset conditions and are most likely to meet the user's requirements.
[0052] The following is a specific example to illustrate this:
[0053] Suppose a user enters the query question "Statistics on the renewable energy power generation in Hebei Province during the Qingming Festival" into the data query system. The system first uses a masked matching mechanism to reduce the weight of specific field values such as "during the Qingming Festival" and "Hebei Province" in the representation, resulting in the semantic representation "View the power generation of a certain province within a certain time period." Then, it converts this to a vector representation using word embedding technology, converting each keyword into its corresponding word vector. Finally, it may use methods such as averaging or max pooling to obtain the vector representation of the entire sentence—the target user question. Based on the target user question, the system then starts from the first data... Five historical user questions in the database were identified as meeting the third preset condition. This third preset condition is a similarity threshold of 0.8; that is, a question is considered to meet the third preset condition only if the cosine similarity between the historical and current user questions exceeds 0.8. Then, from these five historical user questions, the system further filters for historical data operation instructions with high execution efficiency and accuracy. For example, setting the instruction success rate to over 95%, three historical data operation instructions were selected. These instructions performed well in the past and highly matched the user question "view the power generation of a certain province within a certain period." This allows the system to intelligently infer the user's current query intent using historical data and lays the foundation for generating accurate data operation instructions, thereby improving the accuracy and efficiency of data operations and reducing communication barriers between the user and the system.
[0054] Through the above steps, the user question is vectorized. After extracting basic semantic information, the user question is vectorized to obtain the target user question. Based on the target user question, M semantically similar historical user questions are found from the historical database, and N historical data operation instructions are determined. Through vectorization processing, the system can ignore superficial non-critical differences and focus on semantic matching, thereby enhancing the matching accuracy between the user question and historical user questions.
[0055] In one exemplary embodiment, determining target data information from historical data information generated when executing N historical data operation instructions includes: determining P first data tables from the first database, wherein the P first data tables are data tables determined when executing N historical data operation instructions, and P is a positive integer greater than or equal to 1; determining Q target data tables from the P first data tables whose data operation frequency satisfies a first preset threshold, wherein Q is a positive integer less than or equal to P; determining R target fields from the Q target data tables, wherein the R target fields are fields determined when executing N historical data operation instructions, and all R target fields are fields whose usage frequency satisfies a second preset threshold, and R is a positive integer greater than or equal to 1; and determining the Q target data tables and R target fields as the target data information.
[0056] Optionally, the P first data tables in this embodiment are data tables that have been accessed when executing historical data operation instructions. They are identified based on the table names mentioned in the historical operation instructions and cover all tables that are potentially related to the user's problem.
[0057] Optionally, the first preset threshold in this embodiment is used to filter tables with a high frequency of data operations. This is to ensure that the data tables recommended by the system are those that have been frequently used for data operations in the past, and are therefore more likely to contain the information required by the user's current data operations.
[0058] Optionally, the R target fields in this embodiment are a set of fields that are related to the user's problem and whose usage frequency meets the second preset threshold, determined during the execution of historical data operation instructions. This ensures that each recommended field has sufficient historical usage evidence to prove its contribution to solving the user's problem.
[0059] Optionally, the second preset threshold in this embodiment is used to filter fields with high usage frequency. The threshold is set based on the usage frequency of the fields to ensure that the recommended fields are representative.
[0060] The following is a specific example to illustrate this:
[0061] Suppose a user enters a natural language query (i.e., the user question above) into the data query system: "Total sales revenue of all departments in the company in the first quarter of 2023". In this scenario, the first database contains all historical queries and SQL operation command records related to time range, department, and sales revenue. The system identifies all data tables involved in N historical data operation commands that meet the first preset condition, such as sales (sales records), department_sales (department sales), time_range (time range), etc., assuming the total number of identified tables is P = 5. Since the minimum standard for data operation frequency (corresponding to the first preset threshold) set in the system is 50% (i.e., a table must be referenced at least half the times in past operations to be considered a frequently used table), from the above P tables, the system filters out the tables most frequently referenced by historical operation commands according to the first preset threshold, such as sales and department_sales, assuming Q = 2, i.e., the sales and department_sales tables. Within the selected sales and department_sales tables... In the `tment_sales` table, the system identifies fields whose frequency of use in historical operations meets the second preset threshold, such as `date`, `sales_amount`, and `department_id` in the `sales` table, and `total_sales` in the `department_sales` table. Assuming R=4, this translates to the four fields: `date`, `sales_amount`, `department_id`, and `total_sales`. Ultimately, the system integrates the two target data tables (`sales` and `department_sales`) and the four target fields (`date`, `sales_amount`, `department_id`, and `total_sales`) into "target data information," serving as a crucial basis for generating subsequent data operation instructions. By identifying historical patterns in the company's total sales revenue queries, the system can infer sales data from specific departments and time ranges that users might be interested in, thereby generating targeted and efficient query instructions. This avoids queries on irrelevant data tables and fields, reducing data processing complexity and improving query speed.
[0062] Through the above steps, P primary data tables involved in executing historical operation instructions are identified. Then, from the P primary data tables, Q target tables and R target fields with operation frequencies meeting the threshold are selected. Finally, the Q target tables and R target fields are integrated as target data information. Through analysis, key tables and fields are identified, reducing redundant information in the subsequent data operation instruction generation process, improving the efficiency and response speed of data operations. At the same time, the existence of the primary database can effectively improve the accuracy of subsequent data operation instructions generated in multi-table and wide-table scenarios.
[0063] In an exemplary embodiment, determining the target data information and the data operation configuration of the user question includes: determining a first keyword in the user question; determining a first triggering method matching the first keyword from a third database; determining first configuration information corresponding to the first keyword from the third database using the first triggering method; and generating the data operation configuration based on the first configuration information and the user question.
[0064] Optionally, the first keyword in this embodiment is a key entity, concept, or term identified from the user's question, used to determine the subject or direction of the query. For example, if the user's question is "Help me check how much new energy power generation in Hebei Province is during the Qingming Festival", the possible first keywords include "Qingming Festival", "Hebei Province", "new energy", and "power generation".
[0065] Optionally, the third database in this embodiment is a configuration database that stores configuration information for global configuration, field configuration, table configuration, and triggering methods. The configuration information includes various field aliases, business terms, indicator aliases, and other configuration information, which are used to address the knowledge blind spots of the pre-trained model when understanding user questions.
[0066] Optionally, the data operation configuration in this embodiment is a set of rules or parameters generated by the system based on the first configuration information and the specific requirements of the user's question, which guide how to perform data operations, including the target data table, fields, query time range, etc.
[0067] Optionally, in this embodiment, the data operation configuration is generated based on the first configuration information and the user question. In actual use, this may include, but is not limited to, using rule-based slot filling technology (using the extracted key matching information to form the elements of the configuration information) or using a general large language model (extracting the user question, performing semantic understanding and analysis on the extracted information, thereby generating more accurate configuration information).
[0068] Optionally, in this embodiment, before generating the data operation configuration based on the first configuration information and the user's question, the configuration information is sorted according to its importance based on the configuration source of the configuration information. For example, the importance of the global configuration information > the importance of the table configuration information > the importance of the field configuration information. The table configuration and field configuration information are sorted according to the importance ranking of the first data table and the target field obtained above. The global configuration information is a set of rules and settings applied to the entire system or all query scenarios.
[0069] Optionally, the global configuration in this embodiment includes, but is not limited to: time parsing rules (e.g., when a user mentions "this year", the system will by default parse the time range to all dates of the current year), unified naming conventions for tables and fields, priority and association logic for multi-table joins, and default error handling and feedback mechanisms.
[0070] The following is a specific example to illustrate this. Figure 3 This is a flowchart illustrating a data operation configuration for determining a user question according to an embodiment of this application. Assume a user raises the following question in the data query system (corresponding to the aforementioned user question): "What are the statistics on new energy power generation in Hebei Province during the 2023 Qingming Festival?" The data operation configuration for the user question is as follows:
[0071] Step S302: Identify keywords from user questions, such as "2023 Qingming Festival" (time), "Hebei Province" (geographical location), "new energy" (energy type), and "power generation" (query indicator);
[0072] Step S304: Determine the triggering method matching the keyword from the third database: For the keyword "power generation", the system queries the third database to find the triggering method that matches it, for example, by using regular expressions to determine the first configuration information;
[0073] Step S306: Determine the first configuration information from the third database using a triggering method: Extract the key information "Hebei Province: Hebei" from the user question, and the corresponding first configuration information "province field of the daily_power_detail table = ?";
[0074] Step S308: Generate data operation configuration based on the first configuration information and user questions: Use rule-based slot filling technology to generate data operation configuration. In the province field of the daily_power_detail table, "Hebei Province" should be represented as "province=Hebei".
[0075] By matching keywords in user questions with the configuration database, the system can accurately identify and convert technical terms or field aliases, and generate detailed data operation configurations. These configurations not only include the specific data tables and fields to be queried, but also specify the time range and operation logic of the query, providing key information for generating efficient and accurate data operation instructions in the future.
[0076] Through the above steps, the first keyword in the user's question is identified. Then, the first configuration information matching the first keyword is retrieved from the configuration database. This first configuration information, along with the user's question, is used to generate data operation configurations. By using keyword matching and introducing business-related configuration information, the system can effectively combine the data in the data tables to deeply understand the relationships between the objects and data tables involved in the user's statement, as well as the fields within those tables. This improves the accuracy of data operation command generation and its adaptability to business scenarios.
[0077] In an exemplary embodiment, determining the first configuration information corresponding to the first keyword from the third database using the first triggering method includes at least one of the following: determining the keyword matching the first keyword from the third database using a regular expression to obtain the first configuration information; determining the keyword matching the first keyword based on the matching degree between the first keyword and the keywords in the configuration information stored in the third database to obtain the first configuration information; and inputting the first keyword into a pre-trained first network model to determine the keyword matching the first keyword from the third database through the first network model to obtain the first configuration information.
[0078] Optionally, in this embodiment, the matching degree between the first keyword and the keywords in the configuration information stored in the third database is used to determine the keywords that match the first keyword and obtain the first configuration information. In actual use, this may include, but is not limited to, using technologies such as semantic vector libraries for matching, allowing users to select the corresponding configuration information even if the keywords in their questions do not completely match the keywords in the configuration information in the third database, but have a certain degree of similarity.
[0079] Optionally, in this embodiment, the matching degree is used to represent the semantic similarity score between the first keyword and the keywords in the configuration information stored in the third database. A higher matching degree means that the two keywords are more closely related semantically.
[0080] Optionally, the first network model in this embodiment is used to find the keyword most similar or most relevant to the first keyword in the third database, and to perform semantic understanding and matching of the user question. It not only considers the keyword, but also analyzes the fuzzy intent of the user question. The first network model includes, but is not limited to: a pre-trained deep learning model, such as Bidirectional Encoder Representations from Transformers (BERT), a Robustly Optimized BERT Pretraining Approach (RoBERTa), or other natural language processing models.
[0081] Optionally, the three methods for determining the first configuration information in this embodiment can be used individually or in combination in actual use. For example, since calling the pre-trained first network model has a certain time delay and resource cost, a hybrid matching method can be introduced. That is, before using the pre-trained first network model for matching, a brief semantic matching is performed by fuzzy matching. If the matching result meets the requirements, the pre-trained first network model is then used for matching.
[0082] Through the above steps, using one or more of the following methods—regular expressions, matching degree, or network models—configuration information matching keywords is found. This provides multiple matching strategies, enhances the system's flexibility and intelligence, and ensures that even when faced with complex or ambiguous keywords, accurate matches can be found.
[0083] In one exemplary embodiment, generating a data operation instruction based on the aforementioned data operation configuration includes: inputting the aforementioned user question and the aforementioned data operation configuration into a preset prompt word template to generate a target prompt word corresponding to the aforementioned user question through the preset prompt word template; and inputting the aforementioned target prompt word into a pre-trained second network model to generate the aforementioned data operation instruction through the aforementioned second network model.
[0084] Optionally, the preset prompt word template in this embodiment is used to transform user questions and data operation configurations into prompt information that is suitable for the second network model to understand. The prompt word template usually contains structured instructions that guide the model to generate specific types of output, including but not limited to natural language descriptions with specific markers or tags, which are used to clearly indicate the target of the query, the data source and the operation requirements.
[0085] Optionally, in this embodiment, the target prompt word is an input text generated by a preset prompt word template, which combines the user's question and data operation configuration. It is intended to clearly tell the model the context, target and rules for performing data operations, so that it can generate accurate data operation instructions.
[0086] Optionally, the second network model in this embodiment includes, but is not limited to, a pre-trained large-scale language model, such as a Generative Pre-trained Transformer (GPT) or other models with powerful text generation capabilities, used to generate specific data operation instructions, such as SQL query statements, after receiving target prompt words.
[0087] Through the above steps, user questions and configuration information are input into a preset prompt word template to generate target prompt words. Then, the target prompt words are input into a pre-trained second network model to generate data operation instructions. Through the prompt word template and the network model, it is ensured that the generated instructions are not only grammatically correct, but also semantically accurately match the user's needs, thereby improving the accuracy of instruction generation and natural language understanding capabilities.
[0088] In one exemplary embodiment, after generating a data operation instruction based on the above data operation configuration, the method further includes: obtaining user feedback information; generating a target data operation instruction that satisfies a fourth preset condition based on the feedback information and using a pre-trained third network model, wherein the target data operation instruction is used to perform data operations matching the user's question in the second database.
[0089] Optionally, the feedback information in this embodiment is the user's response to the data operation instructions (such as SQL queries) generated by the system, including whether the user is satisfied or dissatisfied, whether the instruction was executed correctly, the necessary modifications or adjustments, and the specific adjustment methods, etc.
[0090] Optionally, in this embodiment, the user feedback information may be the information collected when the user selects to adjust the target data operation instructions after obtaining the target data operation instructions generated by the model, performing a database query operation, and displaying the data operation results to the user.
[0091] Optionally, the third network model in this embodiment is used to generate or adjust data operation instructions based on user feedback information, including but not limited to a pre-trained deep learning model.
[0092] Optionally, the fourth preset condition in this embodiment is a new data operation requirement or restriction determined based on user feedback. These conditions may be specific requirements for query results (such as result sorting, filtering rules), new selection of data sources (such as obtaining data from another table), or modification of query parameters (such as modifying the query time range).
[0093] Optionally, the target data operation instruction in this embodiment is a data operation instruction generated by the third network model based on user feedback and the fourth preset condition, which is more in line with user needs. The instruction may be a modification or expansion of the original instruction to more accurately execute the data operation expected by the user.
[0094] Optionally, the second database in this embodiment is a database that actually stores the data that the user will query.
[0095] Through the above steps, user feedback on the generated instructions is obtained. This feedback information and the network model are then used to generate optimized instructions that meet preset conditions. This user feedback loop enables the system to self-adjust and optimize based on actual usage results, improving user experience and the efficiency of system iteration.
[0096] In an exemplary embodiment, after generating a target data operation instruction that satisfies a fourth preset condition based on the feedback information and using a pre-trained third network model, the method further includes: acquiring first data information generated when executing the target data operation instruction; associating the first data information with the user question and storing it in the first database to update the first database.
[0097] Optionally, the first data information in this embodiment is the result data or metadata obtained from the actual database after the system executes the data operation instruction requested by the user (such as an SQL query statement), including but not limited to the specific value returned by the query, the data format, the data table used, and the usage of the fields in the data table used.
[0098] Optionally, in this embodiment, a syntax parser can also be used to parse the target data operation instruction to obtain the table and field usage information corresponding to the target data operation instruction, so as to store the user question, the target data operation instruction, and the table and field usage information into the first database.
[0099] The following is a specific example to illustrate this. Figure 4 This is a flowchart of generating data operation instructions and updating a first database according to an embodiment of this application. Assume a user raises the following question in the data query system (corresponding to the aforementioned user question): "What is the statistical situation of new energy power generation in Hebei Province during the Qingming Festival in 2023?" The system generates and executes the corresponding SQL query instruction, returning the total data of new energy power generation in Hebei during the Qingming Festival. The first database is updated through the following steps:
[0100] Step S402: Obtain data manipulation instructions generated by the pre-trained third network model;
[0101] Step S404: Based on user feedback, generate target data operation instructions that meet user needs using a pre-trained third network model;
[0102] Step S406: Obtain the first data information generated when executing the target data operation instruction. The first data information may be the data table obtained after parsing the target data operation instruction and the usage of fields in the data table, such as "the total amount of new energy power generation in Hebei during the Qingming Festival is XX kWh". In addition to the result data, the system will also record metadata about this query, such as the query timestamp, the executed SQL statement (corresponding to the above target data operation instruction), the data table and field information used, etc.
[0103] Step S408: The user problem, the successfully executed target data operation instruction, and the first data information are associated and stored in the first database.
[0104] Through the above steps, the system can accumulate the wisdom of historical queries, providing faster and more accurate support for future queries. In particular, when dealing with ambiguity or changes in natural language expressions, referencing historical information can significantly improve the efficiency and effectiveness of NL2SQL conversion.
[0105] Through the above steps, the data information generated by the execution of optimization instructions is recorded. This data, along with user questions, is then stored in the first database to update it. By recording and updating historical data, the system can learn user preferences and common data operation patterns, providing a more accurate matching basis for subsequent data operations and enhancing the system's self-learning ability and adaptability.
[0106] In an exemplary embodiment, after generating a target data operation instruction that satisfies a fourth preset condition based on the feedback information and using a pre-trained third network model, the method further includes: inputting the target data operation instruction and the user question into the pre-trained fourth network model to determine first information through the fourth network model, wherein the first information is information not included in the user question and the data operation configuration; determining a second triggering method and second configuration information based on the first information, wherein the second configuration information is information parsed from the first information; and associating and storing the second triggering method and the second configuration information in a third database, wherein the third database stores historical triggering methods and historical configuration information corresponding to the historical user questions.
[0107] Optionally, the first information in this embodiment refers to additional information that is not directly mentioned in the user's question and data operation configuration, but is very important for constructing an accurate query. For example, the user may mention "the most recent quarter" but not specify which quarter, or indicate that they want to obtain a certain indicator but do not specify the alias or exact location information of the indicator, etc.
[0108] Optionally, in this embodiment, the information input into the pre-trained fourth network model may include target prompt words in addition to target data operation instructions and user questions.
[0109] Optionally, the fourth network model in this embodiment is used to mine and analyze missing information from target data operation instructions and user questions, including but not limited to a pre-trained deep learning model.
[0110] Optionally, the second configuration information in this embodiment is a new set of data operation configurations parsed from the first information and user questions. It serves as the basis for the system to generate more complete and accurate data operation instructions, and refines the scope, target and conditions of the query.
[0111] Through the above steps, the fourth network model is used to identify user problems and missing information in the configuration. New second configuration information is generated based on the missing information and stored in the third database. By dynamically supplementing and updating the third database, the system can continuously enrich its knowledge base, improve its adaptability to new data operation scenarios, and enhance the quality of data operation command generation.
[0112] In an exemplary embodiment, storing the second triggering method and the second configuration information in a third database includes: when the third database includes target historical configuration information corresponding to the second configuration information, inputting the target historical configuration information, the second triggering method, and the second configuration information into a pre-trained fifth network model to output fused information through the fifth network model; and storing the fused information in the third database.
[0113] Optionally, the second triggering method in this embodiment refers to a specific pattern or rule that triggers the second configuration information in the user's question. The second triggering method includes, but is not limited to, a pattern identified by keywords, regular expressions, semantic matching, or other natural language processing technologies.
[0114] Optionally, the fifth network model in this embodiment is used to generate fused information by integrating the target historical configuration information, the second triggering method, and the second configuration information. The fifth network model includes, but is not limited to, a pre-trained deep learning model.
[0115] Optionally, the target historical configuration information in this embodiment is configuration information related to the second configuration information in the third database that has been used in past user queries, such as solutions to similar problems, alias rules for specific fields, and parsing methods for time ranges.
[0116] For example, if the user's question is "Help me check how much renewable energy power generation Hebei Province will generate during the Qingming Festival," and the target data operation instructions have already been generated based on the user's needs, the large model uses the target data operation instructions, the user's question, and the target prompt to generate M pieces of information not included in the target prompt ("power generation corresponds to the power_amount field in the daily_power_detail table"). Then, the large model uses the rules of the configuration information to generate second configuration information corresponding to the M pieces of information not included. After that, the relationship between the second configuration information and other configuration information is checked to further determine whether new configuration information or configuration information needs to be added or modified. After checking, it is found that there was a previous configuration information in the third database that incorrectly matched "power generation" to the power_total (cumulative power generation) field, causing confusion in the large model. Therefore, the second configuration information is stored in the third database by modifying the original configuration information.
[0117] Optionally, in this embodiment, if there is no configuration information corresponding to the second configuration information in the third database, the second triggering method and the second configuration information can be directly associated and stored in the third database.
[0118] By following the steps above, we check whether there is already matching information in the third database. If so, we use a network model to fuse the information and store the fused information in the third database. This ensures that the updates to the configuration information reflect the latest user needs without disrupting the original knowledge structure, thus maintaining the continuity and consistency of the third database and improving the system's maintenance efficiency and data quality.
[0119] The above method will be illustrated with a specific example below. Figure 5 This is a flowchart illustrating a method for generating data operation instructions corresponding to a user question according to an embodiment of this application. Assume a user raises the following question in the data query system (corresponding to the aforementioned user question): "What are the statistics on new energy power generation in Hebei Province during the 2023 Qingming Festival?" The system generates data operation instructions corresponding to the user question through the following steps:
[0120] Step S502: Determine historical operation instructions from the first database based on user questions: First, search the first database (i.e., the historical query database) for historical user questions that are semantically similar to "Hebei Province's new energy power generation during the Qingming Festival", as well as the data tables used by their corresponding SQL query instructions;
[0121] Step S504: Based on historical operation instructions, determine N first data tables and M first fields, where N and M are both positive integers greater than or equal to 1: The system filters out N data tables and M target fields directly related to the new energy power generation in Hebei Province during the Qingming Festival from the "historical data information" obtained after the execution of the above historical operation instructions, such as the date, province, energy_type and generation_volume fields of the power_generation table;
[0122] Step S506: Determine P first data tables from N data tables that meet the preset operation frequency, and determine R target fields from M target fields that meet the preset usage frequency, where P is a positive integer less than or equal to N, and R is a positive integer less than or equal to M.
[0123] Step S508: Determine P first data tables and R target fields as target data information;
[0124] Step S510: Based on the target data information and user questions, generate the data operation configuration: {"table":"power_generation","date_range":"Qingming Festival","province":"Hebei Province","energy_type":"New Energy","fields":["generation_volume"]}.
[0125] Step S512: Generate data operation instructions using the generated data operation configuration: Using the data operation configuration generated above and combined with the structure of the current database (second database), generate a specific SQL statement: "SELECT generation_volume FROM power_generation WHERE date BETWEEN 'Qingming Festival start date' AND 'Qingming Festival end date' AND province = 'Hebei Province' AND energy_type = 'New Energy'".
[0126] Through this series of steps, the system can intelligently learn from historical queries, reduce redundant queries, and improve query efficiency and accuracy. At the same time, the data operation configuration generated based on user questions ensures the customizability and professionalism of query instructions, which can better match the user's query needs.
[0127] It should be noted that, through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0128] This embodiment also provides a data operation instruction generation apparatus, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0129] Figure 6 This is a structural block diagram of a data manipulation instruction generation apparatus according to an embodiment of this application, such as... Figure 6 As shown, the device includes a first memory 62, a first processor 64, and a first computer program 6202 stored in the first memory 62 and executable on the first processor 64. When the first processor 64 executes the first computer program 6202, it performs the following operations: determining N historical data operation instructions from a first database that satisfy a first preset condition based on the acquired user question, wherein the historical data operation instructions are operation instructions generated based on historical user questions, and N is a positive integer greater than or equal to 1; determining target data information from historical data information generated when executing the N historical data operation instructions; determining a data operation configuration for the target data information and the user question, wherein the data operation configuration is a configuration determined based on semantic parsing of the target data information and the user question; and generating data operation instructions based on the data operation configuration, wherein the data operation instructions are used to perform data operations matching the user question in a second database.
[0130] When the first processor 64 executes the first computer program 6202, it also performs the following operations: performs a data conversion operation on the user problem to obtain a target user problem represented by a vector, wherein the target user problem includes the basic semantic information of the user problem; determines M historical user problems from the first database that satisfy a third preset condition based on the target user problem, wherein M is a positive integer greater than or equal to 1; and uses the M historical user problems to determine N historical data operation instructions from the first database that satisfy the first preset condition.
[0131] When the first processor 64 executes the first computer program 6202, it also performs the following operations: determining P first data tables from the first database, wherein the P first data tables are data tables determined when executing N historical data operation instructions, and P is a positive integer greater than or equal to 1; determining Q target data tables from the P first data tables whose data operation frequency meets a first preset threshold, wherein Q is a positive integer less than or equal to P; determining R target fields from the Q target data tables, wherein the R target fields are fields determined when executing N historical data operation instructions, and all R target fields are fields whose usage frequency meets a second preset threshold, wherein R is a positive integer greater than or equal to 1; and determining the Q target data tables and R target fields as the target data information.
[0132] When the first processor 64 executes the first computer program 6202, it also performs the following operations: determining the first keyword in the user question; determining the first triggering method matching the first keyword from the third database; determining the first configuration information corresponding to the first keyword from the third database using the first triggering method; and generating the data operation configuration based on the first configuration information and the user question.
[0133] When the first processor 64 executes the first computer program 6202, it also performs the following operations: determining the first configuration information corresponding to the first keyword from the third database using the first triggering method, including at least one of the following: determining the keyword matching the first keyword from the third database using a regular expression to obtain the first configuration information; determining the keyword matching the first keyword based on the matching degree between the first keyword and the keywords in the configuration information stored in the third database to obtain the first configuration information; and inputting the first keyword into a pre-trained first network model to determine the keyword matching the first keyword from the third database through the first network model to obtain the first configuration information.
[0134] When the first processor 64 executes the first computer program 6202, it also performs the following operations: inputting the user question and the data operation configuration into a preset prompt word template to generate a target prompt word corresponding to the user question through the preset prompt word template; inputting the target prompt word into a pre-trained second network model to generate the data operation instruction through the second network model.
[0135] When the first processor 64 executes the first computer program 6202, it also performs the following operations: after generating data operation instructions based on the data operation configuration, it obtains user feedback information; based on the feedback information and using a pre-trained third network model, it generates a target data operation instruction that satisfies a fourth preset condition, wherein the target data operation instruction is used to perform data operations matching the user problem in the second database.
[0136] When the first processor 64 executes the first computer program 6202, it also performs the following operations: after generating a target data operation instruction that satisfies the fourth preset condition based on the feedback information and using a pre-trained third network model, it acquires the first data information generated when executing the target data operation instruction; and associates the first data information with the user question and stores it in the first database to update the first database.
[0137] When the first processor 64 executes the first computer program 6202, it also performs the following operations: after generating a target data operation instruction that satisfies the fourth preset condition based on the feedback information and using a pre-trained third network model, the target data operation instruction and the user question are input into the pre-trained fourth network model to determine first information through the fourth network model, wherein the first information is information not included in the user question and the data operation configuration; a second triggering method and a second configuration information are determined based on the first information, wherein the second configuration information is information parsed from the first information; the second triggering method and the second configuration information are associated and stored in a third database, wherein the third database stores historical triggering methods and historical configuration information corresponding to the historical user questions.
[0138] When the first processor 64 executes the first computer program 6202, it also performs the following operations: when the third database includes target historical configuration information corresponding to the second configuration information, the target historical configuration information, the second triggering method, and the second configuration information are all input into the pre-trained fifth network model so as to output fusion information through the fifth network model; and the fusion information is stored in the third database.
[0139] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0140] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.
[0141] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0142] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0143] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0144] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0145] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0146] The embodiments described herein also provide a computer program that includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.
[0147] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0148] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0149] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for generating data manipulation instructions, characterized in that, include: Based on the acquired user questions, N historical data operation instructions that meet the first preset conditions are determined from the first database, wherein the historical data operation instructions are operation instructions generated based on historical user questions, and N is a positive integer greater than or equal to 1; Determine the target data information from the historical data information generated when executing N historical data operation instructions; Determine the data operation configuration for the target data information and the user question, wherein the data operation configuration is determined based on the semantic parsing of the target data information and the user question; Data operation instructions are generated based on the data operation configuration, wherein the data operation instructions are used to perform data operations matching the user question in the second database.
2. The method according to claim 1, characterized in that, Based on the acquired user questions, N historical data operation instructions that meet the first preset conditions are determined from the first database, including: Perform a data transformation operation on the user question to obtain a target user question represented as a vector, wherein the target user question includes the basic semantic information of the user question; Based on the target user question, M historical user questions that satisfy the third preset condition are determined from the first database, wherein M is a positive integer greater than or equal to 1; Using the M historical user questions, N historical data operation instructions that satisfy the first preset conditions are determined from the first database.
3. The method according to claim 1, characterized in that, The target data information is determined from the historical data information generated when executing N historical data operation instructions, including: P first data tables are determined from the first database, wherein the P first data tables are data tables determined when executing N historical data operation instructions, and P is a positive integer greater than or equal to 1; From the P first data tables, determine Q target data tables whose data operation frequency satisfies a first preset threshold, where Q is a positive integer less than or equal to P; R target fields are determined from Q target data tables, wherein the R target fields are fields determined when executing N historical data operation instructions, and all R target fields are fields whose usage frequency meets a second preset threshold, and R is a positive integer greater than or equal to 1; The Q target data tables and R target fields are determined as the target data information.
4. The method according to claim 1, characterized in that, Determining the data operation configuration for the target data information and the user question includes: Identify the first keyword in the user's question; Determine the first trigger method that matches the first keyword from the third database; The first configuration information corresponding to the first keyword is determined from the third database using the first triggering method. The data operation configuration is generated based on the first configuration information and the user question.
5. The method according to claim 4, characterized in that, Determining the first configuration information corresponding to the first keyword from the third database using the first triggering method includes at least one of the following: The first configuration information is obtained by using regular expressions to determine the keywords that match the first keyword from the third database; Based on the matching degree between the first keyword and the keywords in the configuration information stored in the third database, the keywords that match the first keyword are determined, and the first configuration information is obtained; The first keyword is input into a pre-trained first network model to determine keywords matching the first keyword from the third database, thereby obtaining the first configuration information.
6. The method according to claim 1, characterized in that, Data operation instructions are generated based on the data operation configuration, including: The user question and the data operation configuration are input into a preset prompt word template to generate a target prompt word corresponding to the user question. The target prompt word is input into a pre-trained second network model to generate the data operation instructions through the second network model.
7. The method according to claim 1, characterized in that, After generating data operation instructions based on the data operation configuration, the method further includes: Obtain user feedback information; Based on the feedback information and using a pre-trained third network model, a target data operation instruction that satisfies a fourth preset condition is generated, wherein the target data operation instruction is used to perform data operations in the second database that match the user's question.
8. The method according to claim 7, characterized in that, After generating target data operation instructions that satisfy the fourth preset condition based on the feedback information and using a pre-trained third network model, the method further includes: Acquire the first data information generated when executing the target data operation instruction; The first data information and the user question are associated and stored in the first database to update the first database.
9. The method according to claim 7, characterized in that, After generating target data operation instructions that satisfy the fourth preset condition based on the feedback information and using a pre-trained third network model, the method further includes: The target data operation instruction and the user question are input into a pre-trained fourth network model to determine first information through the fourth network model, wherein the first information is information not included in the user question and the data operation configuration; The second triggering method and the second configuration information are determined based on the first information, wherein the second configuration information is information parsed from the first information; The second triggering method and the second configuration information are associated and stored in a third database, wherein the third database stores historical triggering methods and historical configuration information corresponding to the historical user issues.
10. The method according to claim 9, characterized in that, The second triggering method and the second configuration information are associated and stored in the third database, including: When the third database includes target historical configuration information corresponding to the second configuration information, the target historical configuration information, the second triggering method, and the second configuration information are all input into the pre-trained fifth network model so that the fifth network model can output fusion information. The fusion information is stored in the third database.
11. A data manipulation instruction generation apparatus, characterized in that, It includes a first memory, a first processor, and a first computer program stored in the first memory and executable on the first processor. When the first processor executes the first computer program, it performs the following operations: Based on the acquired user questions, N historical data operation instructions that meet the first preset conditions are determined from the first database, wherein the historical data operation instructions are operation instructions generated based on historical user questions, and N is a positive integer greater than or equal to 1; Determine the target data information from the historical data information generated when executing N historical data operation instructions; Determine the data operation configuration for the target data information and the user question, wherein the data operation configuration is determined based on the semantic parsing of the target data information and the user question; Data operation instructions are generated based on the data operation configuration, wherein the data operation instructions are used to perform data operations matching the user question in the second database.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 10.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 10.