Model training method, task disassembling method, data query method, equipment, medium and product
By constructing the data set and training the target model to generate task disassembly steps, the problem of inefficient SQL statement generation under complex query requests is solved, and more efficient and accurate data query is achieved.
Patent Information
- Application Number
- CN202411786548.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-06
AI Technical Summary
When query requests are more complicated, it is difficult for the computer to generate SQL statements quickly and accurately, affecting the efficiency and accuracy of data query.
By constructing a data set, including sample query requests, knowledge data and task dismantling steps, training the target model to output task dismantling steps. Then, input the target query request into the model, obtain the corresponding task disassembly steps, and convert it into SQL statements.
Improves the accuracy and generation efficiency of SQL statements, thereby improving the efficiency and accuracy of data query.
Smart Images

Figure CN119938689A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a model training, task decomposition, data query method, device, medium and product. Background Art
[0002] With the continuous development of computer technology, a large amount of data can be stored in the database. When the user needs to query the target data, he only needs to enter the query request in the form of natural language on the computer, and the computer can convert the query request into a structured query language (SQL) statement and obtain the target data from the database according to the SQL statement.
[0003] However, when the query request is more complex, the computer will not be able to quickly and accurately generate SQL statements, thus affecting the efficiency and accuracy of data query. Summary of the invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present disclosure provide a model training, task decomposition, data query method, device, medium and product, which improve the efficiency and accuracy of data query.
[0005] The present disclosure provides a model training method, which includes:
[0006] Constructing a data set, the data set includes multiple groups of data, each group of data includes a sample query request, knowledge data corresponding to the sample query request, and a task decomposition step corresponding to the sample query request, wherein the task decomposition step is described in an intermediate language between a natural language and a structured query language;
[0007] Inputting the sample query request and the knowledge data corresponding to the sample query request into the target model to be trained, so that the target model outputs the task decomposition steps;
[0008] The target model is trained according to the task decomposition steps output by the target model and the task decomposition steps corresponding to the sample query request in the data set.
[0009] The present disclosure also provides a task decomposition method, which includes:
[0010] Get the target query request;
[0011] Acquire knowledge data matching the target query request from a preset knowledge base;
[0012] The target query request and the knowledge data matching the target query request are input into the target model to obtain the task decomposition steps corresponding to the target query request. The task decomposition steps are described in an intermediate language between natural language and structured query language. The target model is trained using the method described above.
[0013] The present disclosure also provides a data query method, the method comprising:
[0014] Get the target query request;
[0015] Acquire knowledge data matching the target query request from a preset knowledge base;
[0016] Input the target query request and the knowledge data matching the target query request into the target model to obtain the task decomposition steps corresponding to the target query request, wherein the task decomposition steps are described in an intermediate language between natural language and structured query language, and the target model is trained by the method described above;
[0017] Converting the task decomposition steps corresponding to the target query request into a structured query language statement;
[0018] According to the structured query language statement, the target data is queried.
[0019] The present disclosure also provides an electronic device, the electronic device comprising:
[0020] one or more processors;
[0021] A storage device for storing one or more programs;
[0022] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.
[0023] The embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method described above is implemented.
[0024] The embodiment of the present disclosure further provides a computer program product, including computer program instructions, which implement the above method when executed by a processor.
[0025] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has at least the following advantages:
[0026] The model training, task decomposition, data query method, device, medium and product provided by the embodiments of the present disclosure construct a data set, and input the sample query request in the data set and the knowledge data corresponding to the sample query request into the target model to be trained, so that the target model outputs the task decomposition steps. Further, according to the task decomposition steps output by the target model and the task decomposition steps corresponding to the sample query request in the data set, the target model is trained so that the task decomposition steps predicted by the target model are continuously accurate. After the target model is trained, in the inference stage, it is only necessary to input the user's target query request into the target model to obtain the task decomposition steps corresponding to the target query request. Since the task decomposition steps are described in an intermediate language between natural language and SQL, when the target query request is relatively complex, by converting the target query request into a task decomposition step, and then converting the task decomposition step into an SQL statement, the accuracy and generation efficiency of the SQL statement can be improved, thereby improving the efficiency and accuracy of data query. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0028] Figure 1 is a flow chart of a model training method in an embodiment of the present disclosure;
[0029] Figure 2 A schematic diagram of a model training in an embodiment of the present disclosure;
[0030] Figure 3 A schematic diagram of task decomposition in an embodiment of the present disclosure;
[0031] Figure 4 is a schematic diagram of an application scenario in an embodiment of the present disclosure;
[0032] Figure 5 A schematic diagram of a data query in an embodiment of the present disclosure;
[0033] Figure 6 is a structural schematic diagram of a model training device in an embodiment of the present disclosure;
[0034] Figure 7 It is a structural schematic diagram of a task disassembly device in an embodiment of the present disclosure;
[0035] Figure 8 is a structural schematic diagram of a data query device in an embodiment of the present disclosure;
[0036] Fig. 9 It is a structural schematic diagram of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0037] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0038] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0039] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0040] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0041] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0042] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0043] Figure 1The flowchart of a model training method in an embodiment of the present disclosure is shown in FIG. 1 . The method can be executed by a model training device, which can be implemented in software and / or hardware, and can be configured on a server, a server cluster, or an electronic terminal. The electronic terminal specifically includes but is not limited to a smart phone, a PDA, a tablet computer, a wearable device with a display screen, a desktop computer, a laptop computer, an all-in-one machine, a smart home device, etc. The server cluster can be a cluster consisting of multiple servers.
[0044] like Figure 1 As shown, the method may specifically include the following steps:
[0045] S101. Construct a data set, wherein the data set includes multiple groups of data, each group of data includes a sample query request, knowledge data corresponding to the sample query request, and task decomposition steps corresponding to the sample query request, wherein the task decomposition steps are described in an intermediate language between natural language and structured query language.
[0046] In this embodiment, the target model is a model to be trained, which can be a large model. Before training, it is necessary to construct a data set, that is, a data set for training the target model. Specifically, the data set includes multiple groups of data, each group of data includes a sample query request, knowledge data corresponding to the sample query request, and a task decomposition step corresponding to the sample query request. Among them, the sample query request is a natural language text. The knowledge data corresponding to the sample query request can be an indicator, dimension, enumeration value, etc. in the sample query request. Specifically, the indicator is the query object, the dimension is the restriction condition on the indicator, and the enumeration value is the enumeration value of the dimension. In some embodiments, the enumeration value of the dimension can also be called a constraint. For example, the sample query request is "What is the second-hand business opportunity volume, second-hand online transaction volume, and customer source transaction volume in City A this month?" The indicators in the sample query request include "business opportunity volume", "second-hand online transaction volume", and "transaction volume". The dimensions include "city" and "business type". "City A" is the enumeration value of "city", and "second-hand" is the enumeration value of "business type". The task decomposition step corresponding to the sample query request is described in an intermediate language between natural language and structured query language. In this embodiment, the intermediate language is a query processing language (QPL). That is, the task decomposition steps corresponding to the sample query request are the processing steps for the sample query request described by QPL, which is an intermediate language between natural language and SQL. For example, the task decomposition steps corresponding to "how many second-hand business opportunities, second-hand online transactions, and customer transaction volumes are there in City A this month" are as follows:
[0047] #1 = Get data (the number of second-hand business opportunities in City A this month)
[0048] #2 = Get data (second-hand online transaction volume in City A this month)
[0049] #3 = Get data (customer transaction volume of City A this month)
[0050] That is, the first step (#1) is to query the second-hand business opportunities in City A this month. The second step (#2) is to query the second-hand online transaction volume in City A this month. The third step (#3) is to query the customer transaction volume in City A this month.
[0051] For another example, the task decomposition steps for "Year-on-year analysis of total assessment commissions from January to July in City B" are as follows:
[0052] #1 = Get data (total assessment commission from January to July in City B)
[0053] #2 = Get data (total assessment commission from January to July last year in City B)
[0054] #3=year-on-year ([#1,#2]; year-on-year)
[0055] That is, the first step (#1) is to query the total assessment commission of City B from January to July this year. The second step (#2) is to query the total assessment commission of City B from January to July last year. The third step (#3) is to calculate the year-on-year growth rate of these two data.
[0056] S102: Input the sample query request and the knowledge data corresponding to the sample query request into the target model to be trained, so that the target model outputs task decomposition steps.
[0057] Specifically, taking a group of data in a data set as an example, the sample query requests in the group of data and the knowledge data corresponding to the sample query requests are input into the target model to be trained, so that the target model outputs task decomposition steps, which can be the task decomposition steps corresponding to the sample query requests predicted by the target model.
[0058] S103: training the target model according to the task decomposition steps output by the target model and the task decomposition steps corresponding to the sample query request in the data set.
[0059] It is understandable that the task decomposition steps corresponding to each sample query request in the data set can be used as the standard answer. During the training phase, the task decomposition steps corresponding to the sample query request predicted by the target model may not be accurate. Therefore, it is necessary to train the target model according to the task decomposition steps corresponding to the sample query request predicted by the target model and the task decomposition steps corresponding to the sample query request in the data set. Specifically, the loss value is calculated according to the loss function, the task decomposition steps corresponding to the sample query request predicted by the target model, and the task decomposition steps corresponding to the sample query request in the data set. Further, according to the loss value, the parameters in the target model are updated. It is understandable that the training process described here is based on a set of data in the data set as an example, and such a training process can be recorded as one round of training. As the data selected each time is different, the number of training rounds continues to increase, and the parameters in the target model are continuously updated, so that the prediction results of the target model gradually approach the standard answer, thereby continuously improving the accuracy of the task decomposition steps predicted by the target model. For example, when the training round is greater than or equal to the preset round, or when the parameters in the target model converge, the training ends, thereby obtaining a trained target model. In addition, in some embodiments, the training process may be fine-tuning, that is, the target model is trained using a data set in a specific business field, so that the trained target model is applicable to the specific business field. In addition, in some other embodiments, each set of data in the data set can be expressed as <sample query request, knowledge data, task decomposition steps>, and 1500 sets of such data can be used when fine-tuning. In addition, the low-rank adaptation of large language models (Low-Rank Adaptation of Large Language Models, LoRA) is specifically used for fine-tuning, with a rank of 64. In addition, since the output diversity of the model is not required, in order to ensure stability during reasoning, the temperature parameter is set to 0.001 and the top_p value is set to 0.001.
[0060] The model training method provided by the embodiment of the present disclosure constructs a data set, and inputs the sample query request in the data set and the knowledge data corresponding to the sample query request into the target model to be trained, so that the target model outputs the task disassembly steps. Further, according to the task disassembly steps output by the target model and the task disassembly steps corresponding to the sample query request in the data set, the target model is trained so that the task disassembly steps predicted by the target model are continuously accurate. After the target model is trained, in the inference stage, it is only necessary to input the user's target query request into the target model to obtain the task disassembly steps corresponding to the target query request. Since the task disassembly steps are described in an intermediate language between natural language and SQL, when the target query request is more complex, by converting the target query request into a task disassembly step, and then converting the task disassembly step into an SQL statement, the accuracy and generation efficiency of the SQL statement can be improved, thereby improving the efficiency and accuracy of data query.
[0061] In addition, by breaking down the tasks into steps, the data processing flow can be described more clearly and efficiently, while reducing redundancy and improving readability.
[0062] Based on the above embodiment, constructing a data set includes: Figure 2 The following steps are shown:
[0063] S201. Generate multiple query request categories according to the data processing capability supported by the intermediate language.
[0064] Specifically, the data processing capabilities supported by the intermediate language include: data acquisition, processing, year-on-year, month-on-month, analysis, and visualization. For example, the data processing capabilities supported by QPL include: data acquisition, processing, year-on-year, month-on-month, analysis, and visualization. The task decomposition steps described by QPL can be composed of a series of lines, each line representing a processing step. Each step can be operations such as data acquisition (such as data query), processing (such as data processing), year-on-year (such as year-on-year analysis), month-on-month (such as month-on-month analysis) or visualization (such as data display). The grammatical components of QPL are as follows:
[0065] <qpl> ::= <line> +
[0066] <line> ::=# <integer> = <tool>
[0067] <tool>::=<get number>
[0068] |<Processing>
[0069] |<Year-on-year>
[0070] |<Month-on-month>
[0071] |<Analysis>
[0072] |<Visualization>
[0073] <Get number>::=Get number( <query>)
[0074] <Processing>::=Processing([<object list>]; <parameter>)
[0075] <Year-on-year>::=Year-on-year([<object list>]; <parameter>)
[0076] <Year-on-year>::=Year-on-year([<object list>]; <parameter>)
[0077] <Analysis>::=Analysis([<object list>]; <parameter>)
[0078] <Visualization>::=Visualization([<object list>]; <parameter>)
[0079] Specifically, the syntax components of QPL are introduced as follows:
[0080] <qpl>: The beginning of the entire QPL script, consisting of one or more <line>composition.
[0081] <line>: Represents a single processing step, consisting of a unique number and a <tool>composition.
[0082] <tool>:Specific data processing tools can be data acquisition, processing, year-on-year, month-on-month comparison, analysis or visualization.
[0083] <Get data>: Query the data of specific indicators, which can be subject to time and space constraints.
[0084] <Processing>: Process the data, such as calculating average, aggregation, filtering, etc.
[0085] <Year-on-year>: Calculate the year-on-year growth rate of a single indicator at the same time granularity.
[0086] <Month-on-month>: Calculates the month-on-month growth rate of a single indicator at the same time granularity.
[0087] <Analysis>: Conduct in-depth analysis of data, such as trend analysis.
[0088] <Visualization>: Display data in graphical form, such as line chart, bar chart, etc.
[0089] <Object list>: A list of one or more data objects, which can be data acquisition results or other processing results.
[0090] <parameters>: Specific parameters passed to the tool, such as time range, calculation method, etc.
[0091] In this embodiment, multiple query request categories can be generated according to the data processing capabilities supported by the QPL. The query request category is a category classified according to the solution path, for example, only data acquisition, that is, simple data acquisition, and other categories such as formula calculation and multi-table association.
[0092] Optionally, multiple query request categories are generated based on the data processing capabilities supported by the intermediate language, including: data acquisition, processing, year-on-year, quarter-on-quarter, analysis, and visualization are arranged and combined to obtain multiple query request categories.
[0093] In this embodiment, data acquisition, processing, year-on-year, quarter-on-quarter, analysis, and visualization are respectively regarded as various types of tools supported by QPL. By making different arrangements and combinations of various tools, multiple query request categories can be obtained. For example, data acquisition itself can constitute a simple data acquisition query request, data acquisition and processing can constitute a formula calculation query request, and data acquisition and processing can also constitute a multi-table association query request.
[0094] S202: For each query request category, construct one or more query request templates and a task decomposition step template corresponding to each query request template.
[0095] For example, for each query request category, one or more query request templates and a task decomposition step template corresponding to each query request template are constructed respectively. That is, for any query request category, one or more query request templates and a task decomposition step template corresponding to each query request template are constructed respectively. Specifically, each query request template constructed for any query request category can be a template designed for the query request category with indicators, dimensions, and enumeration values removed. For example, a query request template constructed for formula calculation is "What is the proportion of [Indicator 1] of each [Dimension 1] this month?". For another example, assuming that analysis itself can constitute a query request category, a query request template constructed for analysis is "Analyze [Constraint 1] this month [Indicator 1] and [Indicator 2]". In addition, each query request template corresponds to a task decomposition step template. The corresponding relationship between query request templates and task decomposition step templates is shown in Table 1 below.
[0096] Table 1
[0097]
[0098] S203: Generate multiple sample query requests according to the query request template and knowledge data in a preset knowledge base.
[0099] Specifically, in the embodiments of the present disclosure, a business knowledge base can be pre-constructed, and the business knowledge base includes an indicator knowledge base, a dimension knowledge base, and an enumeration value knowledge base. The business knowledge base can be recorded as a preset knowledge base. Among them, a large number of indicators are stored in the indicator knowledge base. A large number of dimensions are stored in the dimension knowledge base. A large number of enumeration values are stored in the enumeration value knowledge base. For example, the dimensions supported by each indicator are stored in the dimension knowledge base, and the enumeration values of each dimension are stored in the enumeration value knowledge base.
[0100] Optionally, the knowledge data includes indicators, dimensions, and enumeration values; the enumeration values are enumeration values of the dimensions.
[0101] Specifically, the indicators, dimensions and / or enumeration values in the business knowledge base are combined and brought into the query request template to obtain multiple sample query requests. For example, taking "What is the proportion of [Indicator 1] in each [Dimension 1] this month?" as an example, an indicator such as "Business Opportunity Volume" is randomly selected from the indicator knowledge base as indicator 1, and a dimension such as "Business Department" is selected from the dimension knowledge base as dimension 1. Further, "Business Department" and "Business Opportunity Volume" are brought into "What is the proportion of [Indicator 1] in each [Dimension 1] this month?" to obtain a sample query request, namely "What is the proportion of business opportunity volume in each business department this month?". If another indicator such as "Total Performance" is selected from the indicator knowledge base as indicator 1, and a dimension such as "City" is selected from the dimension knowledge base as dimension 1, and "Total Performance" and "City" are brought into "What is the proportion of [Indicator 1] in each [Dimension 1] this month?", another sample query request, namely "What is the proportion of total performance in each city this month?", can be obtained. Since a large number of indicators are stored in the indicator knowledge base and a large number of dimensions are stored in the dimension knowledge base, a large number of sample query requests can be generated by combining any indicator in the indicator knowledge base and any dimension in the dimension knowledge base and bringing them into "What is the proportion of business opportunities of each business unit this month?" That is to say, by combining the indicators, dimensions and / or enumeration values in the business knowledge base and bringing them into any query request template, multiple sample query requests can be obtained. The query request templates described in the embodiment of the present disclosure can be multiple or even a large number, so a large number of sample query requests can be obtained.
[0102] S204: Generate task decomposition steps corresponding to each sample query request according to the task decomposition step template corresponding to the query request template and the knowledge data used when generating each sample query request.
[0103] For example, the task decomposition step template for "What is the proportion of [Indicator 1] of each [Dimension 1] this month?" is "#1=Get data ([Indicator 1] of each [Dimension 1] this month); #2=Get data ([Indicator 1] of this month); #3=Process ([#1, #2]; #1 / #2, as proportion)". By bringing the "Business Division" and "Opportunity Volume" mentioned above into the task decomposition step template, the task decomposition step corresponding to "What is the proportion of business opportunities of each business division this month?" can be generated. By bringing the "Total Performance" and "City" mentioned above into the task decomposition step template, the task decomposition step corresponding to "What is the proportion of total performance of each city this month?" can be generated. That is to say, after combining the indicators, dimensions and / or enumeration values in the business knowledge base and bringing them into any query request template to obtain multiple sample query requests, the task decomposition steps corresponding to each sample query request can be obtained based on the task decomposition step template corresponding to the query request template and the combination of indicators, dimensions and / or enumeration values used to generate each sample query request in the multiple sample query requests.
[0104] S205: construct a data set according to the multiple sample query requests, the knowledge data used when generating each sample query request, and the task decomposition steps corresponding to each sample query request.
[0105] For example, a large number of sample query requests can be generated according to the above steps, and a combination of indicators, dimensions and / or enumeration values, i.e., the knowledge data used to generate each sample query request, is used when generating each sample query request. In addition, the task decomposition steps corresponding to each sample query request can also be generated according to the above steps. Therefore, a set of data can be generated according to any sample query request, the knowledge data used when generating the sample query request, and the task decomposition steps corresponding to the sample query request. Since there can be many groups of such data, many groups of such data can constitute the data set as described above. The data set can be used to train the target model. In addition, in some embodiments, some data can be selected from the data set as a test set to evaluate the performance of the trained target model.
[0106] Figure 3 The flowchart of a task decomposition method in an embodiment of the present disclosure is shown. The method can be executed by a task decomposition device, which can be implemented in software and / or hardware, and can be configured on a server, a server cluster, or an electronic terminal. The electronic terminal specifically includes but is not limited to a smart phone, a PDA, a tablet computer, a wearable device with a display screen, a desktop computer, a laptop computer, an all-in-one machine, a smart home device, etc. A server cluster can be a cluster consisting of multiple servers.
[0107] like Figure 3 As shown, the method may specifically include the following steps:
[0108] S301: Obtain a target query request.
[0109] Specifically, the task decomposition method described in this embodiment can be applied to Figure 4 The application scenario shown. The application scenario includes a server 41 and an electronic terminal 42, and the server 41 and the electronic terminal 42 can communicate. Among them, the server 41 is deployed with the target model trained as described above. The electronic terminal 42 can be a user terminal, such as a smart phone, and the electronic terminal 42 can be provided with a user interface, and the user can enter a target query request (query) on the user interface. Further, the electronic terminal 42 sends the target query request to the server 41, so that the server 41 can obtain the target query request from the electronic terminal 42.
[0110] S302: Acquire knowledge data matching the target query request from a preset knowledge base.
[0111] Specifically, the business knowledge base as described above can be stored in a database, and the database can be deployed on the server 41 or on other servers. If the database is deployed on the server 41, then when the server 41 obtains the target query request, the knowledge data matching the target query request is obtained from the business knowledge base. If the database is deployed on other servers, then when the server 41 obtains the target query request, the target query request is sent to other servers, so that the other servers obtain the knowledge data matching the target query request from the business knowledge base, and return the matched knowledge data to the server 41.
[0112] Optionally, obtaining knowledge data matching the target query request from a preset knowledge base includes: extracting text of a preset length in the target query request, where the preset length increases sequentially from 2 to the total length of the target query request; and recalling knowledge data matching the text from the preset knowledge base.
[0113] For example, the target query request is "The performance income of the operation of the Beilian Business Department of C City Tianxin District this year_South, presented by business district", and the target query request is segmented according to the preset length to obtain a text of the preset length. Specifically, the preset length is a variable, which can be increased from 2 to the total length of the target query request in sequence. For example, when the preset length is 2, the target query request is segmented, and the obtained text includes "C City", "City Bei", "Bei Lian", "Lian Shi", etc., and so on. When the preset length is 3, the target query request is segmented, and the obtained text includes "C City Bei", "City Bei Lian", "Bei Lian Shi", "Lian Shi", etc., and so on. When the preset length is the total length of the target query request, the obtained text is "The performance income of the operation of the Beilian Business Department of C City Tianxin District this year_South, presented by business district". That is to say, multiple texts can be obtained after the target query request is segmented according to the gradually increasing preset length. Further, taking each text as a search term, the knowledge data matching the text is recalled from the business knowledge base, thereby obtaining the knowledge data such as indicators, dimensions, and enumeration values related to the target query request. For example, taking "City C" as an example, searching in the indicator knowledge base, since "City C" is not an indicator, the search results may not be recalled after searching in the indicator knowledge base. Further, taking "City C" as an example, searching in the dimension knowledge base, since "City C" is not only an enumeration value of "geographic city" but also an enumeration value of "performance city", the dimension "geographic city" and the dimension "performance city" are recalled from the dimension knowledge base. Further, taking "City C" as an example, searching in the enumeration value knowledge base, the enumeration value "City C" is recalled. Therefore, the knowledge data matching "City C" includes the dimension "geographic city", the dimension "performance city", and the enumeration value "City C". It can be understood that the process of recalling knowledge data matching other texts from the business knowledge base is similar to the retrieval process taking "City C" as an example here, and will not be repeated here.
[0114] Optionally, recalling knowledge data matching the text from the preset knowledge base includes: recalling knowledge data consistent with the text expression from the preset knowledge base; and / or recalling knowledge data having a similarity with the text greater than or equal to a preset threshold from the preset knowledge base.
[0115] Specifically, the knowledge data stored in the business knowledge base may be the same as or different from the user description. For example, the target query request "Performance income_South of the Tianxin District of Beilian Business Unit in City C this year, presented by business district" includes "Performance income_South", but in the indicator knowledge base, "Performance income_South" may be stored, or "Performance income_South" may not be stored, but similar indicators such as "Total performance_South" may be stored. Or the indicator knowledge base stores both "Performance income_South" and "Total performance_South", so in this embodiment, when each text is used as a search term and the knowledge data matching the text is recalled from the business knowledge base, the knowledge data consistent with the text expression can be recalled from the business knowledge base; and / or the knowledge data whose similarity with the text is greater than or equal to a preset threshold can be recalled from the business knowledge base. For example, taking the text "Performance Income_South" in the target query request as an example, when searching in the indicator knowledge base, the indicator "Performance Income_South" that is consistent with the expression of "Performance Income_South" can be recalled, and the indicators whose similarity with the text is greater than or equal to the preset threshold can also be recalled from the indicator knowledge base based on the representation vector of "Performance Income_South". Specifically, in the retrieval process, the similarity between each indicator in the indicator knowledge base and "Performance Income_South" is calculated based on the representation vector of each indicator in the indicator knowledge base and the representation vector of "Performance Income_South", and the indicator is recalled when the similarity is greater than or equal to the preset threshold. According to the retrieval method described in this embodiment, the knowledge data recalled for "The performance income of the operation of the Beilian Business Unit in Tianxin District of City C this year_South, presented by business district" is shown in Table 2 below:
[0116] Table 2
[0117]
[0118] In the enumeration value column of Table 2, the dimension is on the left side of ":" and the enumeration value is on the right side.
[0119] S303. Input the target query request and the knowledge data matching the target query request into the target model to obtain the task decomposition steps corresponding to the target query request. The task decomposition steps are described in an intermediate language between natural language and structured query language. The target model is trained using the method described above.
[0120] For example, after recalling the knowledge data matching the target query request from the business knowledge base, the target query request and the knowledge data are input into the target model trained as described above, so that the target model outputs the task decomposition steps corresponding to the target query request. The task decomposition steps are described in an intermediate language between natural language and structured query language, and the target model is trained using the method described above. The introduction to the task decomposition steps and the training process of the target model can refer to the contents described in the above embodiment, and will not be repeated here.
[0121] The task decomposition method provided by the embodiment of the present disclosure obtains knowledge data matching the target query request from a preset knowledge base, and inputs the target query request and the knowledge data matching the target query request into a target model to obtain the task decomposition steps corresponding to the target query request. Since the task decomposition steps are described in an intermediate language between natural language and SQL, when the target query request is relatively complex, by converting the target query request into task decomposition steps, and then converting the task decomposition steps into SQL statements, the accuracy and generation efficiency of the SQL statements can be improved, thereby improving the efficiency and accuracy of data query.
[0122] Optionally, before inputting the target query request and the knowledge data matching the target query request into the target model, the method also includes: obtaining a sample query request that is most similar to the target query request from a data set, and a task decomposition step corresponding to the sample query request, the data set including multiple groups of data, each group of data including a sample query request, knowledge data corresponding to the sample query request, and a task decomposition step corresponding to the sample query request.
[0123] For example, before inputting the target query request and the knowledge data into the trained target model as described above, some embodiments may also retrieve from the data set a sample query request that is most similar to the target query request and the task decomposition steps corresponding to the sample query request, where the retrieved sample query request and the task decomposition steps corresponding to the sample query request are recorded as example samples. The data set is specifically the data set described in the above embodiment, for example, the data set includes multiple groups of data, each group of data includes a sample query request, the knowledge data corresponding to the sample query request, and the task decomposition steps corresponding to the sample query request.
[0124] Correspondingly, the target query request and the knowledge data matching the target query request are input into the target model to obtain the task decomposition steps corresponding to the target query request, including: inputting the target query request, the knowledge data matching the target query request, the sample query request, and the task decomposition steps corresponding to the sample query request into the target model to obtain the task decomposition steps corresponding to the target query request.
[0125] For example, the target query request is recorded as query1, and the sample query request retrieved from the data set that is most similar to the target query request is recorded as query2. At this time, query1, knowledge data matching query1, query2, and the task decomposition steps corresponding to query2 can be input into the target model trained as described above, so that the target model can refer to the task decomposition steps corresponding to query2 to output the task decomposition steps corresponding to query1. This improves the accuracy and generalization ability of the target model.
[0126] Figure 5 The flowchart of a data query method in an embodiment of the present disclosure is shown in FIG. 1 . The method can be executed by a data query device, which can be implemented in software and / or hardware, and can be configured on a server, a server cluster, or an electronic terminal. The electronic terminal specifically includes but is not limited to a smart phone, a PDA, a tablet computer, a wearable device with a display screen, a desktop computer, a laptop computer, an all-in-one machine, a smart home device, etc. A server cluster can be a cluster consisting of multiple servers.
[0127] like Figure 5 As shown, the method may specifically include the following steps:
[0128] S501: Obtain a target query request.
[0129] For example, the data query method described in this embodiment is applicable to Figure 4 The application scenario shown. The server 41 is deployed with the trained target model as described above. The electronic terminal 42 may be provided with a user interface, on which a user may input a target query request (query). Further, the electronic terminal 42 sends the target query request to the server 41, so that the server 41 may obtain the target query request from the electronic terminal 42.
[0130] S502: Acquire knowledge data matching the target query request from a preset knowledge base.
[0131] Specifically, the business knowledge base as described above can be stored in a database, and the database can be deployed on the server 41 or on other servers. If the database is deployed on the server 41, then when the server 41 obtains the target query request, the knowledge data matching the target query request is obtained from the business knowledge base. If the database is deployed on other servers, then when the server 41 obtains the target query request, the target query request is sent to other servers, so that the other servers obtain the knowledge data matching the target query request from the business knowledge base, and return the matched knowledge data to the server 41.
[0132] For example, the target query request is "the operating performance income of Tianxin District of Beilian Business Unit in City C this year_South, presented by business district". The process of obtaining knowledge data matching the target query request from the business database can refer to the process described in the above embodiment and will not be repeated here.
[0133] S503. Input the target query request and the knowledge data matching the target query request into the target model to obtain the task decomposition steps corresponding to the target query request. The task decomposition steps are described in an intermediate language between natural language and structured query language. The target model is trained using the method described above.
[0134] For example, after recalling the knowledge data matching the target query request from the business knowledge base, the target query request and the knowledge data are input into the target model trained as described above, so that the target model outputs the task decomposition steps corresponding to the target query request. The task decomposition steps are described in an intermediate language between natural language and structured query language, and the target model is trained using the method described above. The introduction to the task decomposition steps and the training process of the target model can refer to the contents described in the above embodiment, and will not be repeated here.
[0135] S504: Convert the task decomposition steps corresponding to the target query request into structured query language statements.
[0136] For example, the task decomposition steps for the target query request include 5 steps, and each step is converted into an SQL statement here.
[0137] S505: Query target data according to the structured query language statement.
[0138] For example, query the target data based on the SQL statements converted from the task decomposition steps.
[0139] The data query method provided by the embodiment of the present disclosure obtains knowledge data matching the target query request from a preset knowledge base, and inputs the target query request and the knowledge data matching the target query request into a target model to obtain the task decomposition steps corresponding to the target query request. Since the task decomposition steps are described in an intermediate language between natural language and SQL, when the target query request is relatively complex, by converting the target query request into task decomposition steps, and then converting the task decomposition steps into SQL statements, the accuracy and generation efficiency of the SQL statements can be improved, thereby improving the efficiency and accuracy of data query.
[0140] Figure 6 Schematic diagram of the structure of a model training device in an embodiment of the present disclosure. The device provided in the embodiment of the present disclosure can be configured in a server, a server cluster, or an electronic terminal. Figure 6 As shown, the model training device 60 specifically includes: a construction module 61, an input module 62, and a training module 63.
[0141] Among them, the construction module 61 is used to construct a data set, which includes multiple groups of data, each group of data includes a sample query request, knowledge data corresponding to the sample query request, and a task decomposition step corresponding to the sample query request, and the task decomposition step is described in an intermediate language between natural language and structured query language.
[0142] An input module 62, used for inputting the sample query request and the knowledge data corresponding to the sample query request into the target model to be trained, so that the target model outputs the task decomposition steps;
[0143] The training module 63 is used to train the target model according to the task decomposition steps output by the target model and the task decomposition steps corresponding to the sample query request in the data set.
[0144] Optionally, when constructing a data set, the construction module 61 is specifically used to:
[0145] generating a plurality of query request categories according to the data processing capability supported by the intermediate language;
[0146] For each query request category, construct one or more query request templates and a task decomposition step template corresponding to each query request template;
[0147] Generate multiple sample query requests according to the query request template and knowledge data in a preset knowledge base;
[0148] Generate task decomposition steps corresponding to each sample query request according to the task decomposition step template corresponding to the query request template and the knowledge data used when generating each sample query request;
[0149] A data set is constructed according to the multiple sample query requests, the knowledge data used when generating each sample query request, and the task decomposition steps corresponding to each sample query request.
[0150] Optionally, the knowledge data includes indicators, dimensions, and enumeration values; the enumeration values are enumeration values of the dimensions.
[0151] Optionally, the data processing capabilities supported by the intermediate language include: data acquisition, processing, year-on-year growth, quarter-on-quarter growth, analysis, and visualization; when the construction module 61 generates multiple query request categories according to the data processing capabilities supported by the intermediate language, it is specifically used to: arrange and combine data acquisition, processing, year-on-year growth, quarter-on-quarter growth, analysis, and visualization to obtain multiple query request categories.
[0152] The device provided in the embodiment of the present disclosure can execute the method steps provided in the method embodiment of the present disclosure, and the beneficial effects thereof are not described in detail here.
[0153] Figure 7 Schematic diagram of the structure of a task decomposition device in an embodiment of the present disclosure. The device provided in the embodiment of the present disclosure can be configured in a server, a server cluster, or an electronic terminal. Figure 7 As shown, the task decomposition device 70 specifically includes: a first acquisition module 71 , a second acquisition module 72 , and an input module 73 .
[0154] The first acquisition module 71 is used to acquire a target query request.
[0155] The second acquisition module 72 is used to acquire knowledge data matching the target query request from a preset knowledge base.
[0156] The input module 73 is used to input the target query request and the knowledge data matching the target query request into the target model to obtain the task decomposition steps corresponding to the target query request. The task decomposition steps are described in an intermediate language between natural language and structured query language. The target model is trained using the method described above.
[0157] Optionally, when the second acquisition module 72 acquires knowledge data matching the target query request from a preset knowledge base, it is specifically used to:
[0158] Extracting text of a preset length from the target query request, where the preset length increases sequentially from 2 to the total length of the target query request;
[0159] Recall knowledge data matching the text from the preset knowledge base.
[0160] Optionally, when the second acquisition module 72 recalls knowledge data that matches the text from the preset knowledge base, it is specifically used to: recall knowledge data that is consistent with the text expression from the preset knowledge base; and / or recall knowledge data whose similarity with the text is greater than or equal to a preset threshold from the preset knowledge base.
[0161] Optionally, the task decomposition device 70 also includes: a third acquisition module 74, the third acquisition module 74 is used to obtain a sample query request that is most similar to the target query request, and a task decomposition step corresponding to the sample query request from a data set, the data set includes multiple groups of data, each group of data includes a sample query request, knowledge data corresponding to the sample query request, and a task decomposition step corresponding to the sample query request; accordingly, the input module 73 is specifically used to: input the target query request, the knowledge data matching the target query request, the sample query request, and the task decomposition step corresponding to the sample query request into the target model to obtain the task decomposition step corresponding to the target query request.
[0162] The device provided in the embodiment of the present disclosure can execute the method steps provided in the method embodiment of the present disclosure, and the beneficial effects thereof are not described in detail here.
[0163] Figure 8 Schematic diagram of the structure of a data query device in an embodiment of the present disclosure. The device provided in the embodiment of the present disclosure can be configured in a server, a server cluster, or an electronic terminal. Figure 8 As shown, the data query device 80 specifically includes: a first acquisition module 81 , a second acquisition module 82 , an input module 83 , a conversion module 84 , and a query module 85 .
[0164] The first acquisition module 81 is used to acquire a target query request.
[0165] The second acquisition module 82 is used to acquire knowledge data matching the target query request from a preset knowledge base.
[0166] The input module 83 is used to input the target query request and the knowledge data matching the target query request into the target model to obtain the task decomposition steps corresponding to the target query request. The task decomposition steps are described in an intermediate language between natural language and structured query language. The target model is trained using the method described above.
[0167] The conversion module 84 is used to convert the task decomposition steps corresponding to the target query request into structured query language statements.
[0168] The query module 85 is used to query the target data according to the structured query language statement.
[0169] The device provided in the embodiment of the present disclosure can execute the method steps provided in the method embodiment of the present disclosure, and the beneficial effects thereof are not described in detail here.
[0170] Fig. 9 Schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Fig. 9 , which shows a schematic diagram of the structure of an electronic device 900 suitable for implementing the embodiment of the present disclosure. The electronic device 900 in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Fig. 9 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0171] like Fig. 9 As shown, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes to implement the method of the embodiment described in the present disclosure according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage device 908 to the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0172] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although Fig. 9 The electronic device 900 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0173] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart, thereby implementing the method as described above. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0174] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0175] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0176] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0177] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0178] Constructing a data set, the data set includes multiple groups of data, each group of data includes a sample query request, knowledge data corresponding to the sample query request, and a task decomposition step corresponding to the sample query request, wherein the task decomposition step is described in an intermediate language between a natural language and a structured query language;
[0179] Inputting the sample query request and the knowledge data corresponding to the sample query request into the target model to be trained, so that the target model outputs the task decomposition steps;
[0180] The target model is trained according to the task decomposition steps output by the target model and the task decomposition steps corresponding to the sample query request in the data set.
[0181] Optionally, when the above one or more programs are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.
[0182] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0183] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0184] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.
[0185] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0186] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0187] The present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, any method provided in the present disclosure is implemented.
[0188] The embodiments of the present disclosure further provide a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the method described above is implemented.
[0189] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0190] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0191] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.< / tool> < / tool> < / line> < / line> < / qpl> < / query> < / tool> < / tool> < / integer> < / line> < / line> < / qpl>
Claims
1. A model training method, characterized in that: The method comprises: Constructing a data set, the data set includes multiple groups of data, each group of data includes a sample query request, knowledge data corresponding to the sample query request, and a task decomposition step corresponding to the sample query request, wherein the task decomposition step is described in an intermediate language between a natural language and a structured query language; Inputting the sample query request and the knowledge data corresponding to the sample query request into the target model to be trained, so that the target model outputs the task decomposition steps; The target model is trained according to the task decomposition steps output by the target model and the task decomposition steps corresponding to the sample query request in the data set.
2. The method according to claim 1, characterized in that Construct a data set, including: generating a plurality of query request categories according to the data processing capability supported by the intermediate language; For each query request category, construct one or more query request templates and a task decomposition step template corresponding to each query request template; Generate multiple sample query requests according to the query request template and knowledge data in a preset knowledge base; Generate task decomposition steps corresponding to each sample query request according to the task decomposition step template corresponding to the query request template and the knowledge data used when generating each sample query request; A data set is constructed according to the multiple sample query requests, the knowledge data used when generating each sample query request, and the task decomposition steps corresponding to each sample query request.
3. The method according to claim 1 or 2, characterized in that: The knowledge data includes indicators, dimensions, and enumeration values; the enumeration values are enumeration values of the dimensions.
4. The method according to claim 2, characterized in that: The data processing capabilities supported by the intermediate language include: data acquisition, processing, year-on-year, month-on-month, analysis, and visualization; According to the data processing capability supported by the intermediate language, multiple query request categories are generated, including: Arrange and combine data acquisition, processing, year-on-year growth, month-on-month growth, analysis, and visualization to obtain multiple query request categories.
5. A task decomposition method, characterized in that: The method comprises: Get the target query request; Acquire knowledge data matching the target query request from a preset knowledge base; The target query request and the knowledge data matching the target query request are input into the target model to obtain the task decomposition steps corresponding to the target query request. The task decomposition steps are described in an intermediate language between natural language and structured query language. The target model is trained using the method described in any one of claims 1 to 4.
6. The method according to claim 5, characterized in that Acquiring knowledge data matching the target query request from a preset knowledge base includes: Extracting text of a preset length from the target query request, where the preset length increases sequentially from 2 to the total length of the target query request; Recall knowledge data matching the text from the preset knowledge base.
7. The method according to claim 6, characterized in that Recalling knowledge data matching the text from the preset knowledge base includes: Recalling knowledge data consistent with the textual representation from the preset knowledge base; and / or Recall from the preset knowledge base knowledge data whose similarity with the text is greater than or equal to a preset threshold.
8. The method according to claim 5, characterized in that Before inputting the target query request and the knowledge data matching the target query request into the target model, the method further includes: Acquire a sample query request that is most similar to the target query request and a task decomposition step corresponding to the sample query request from a data set, wherein the data set includes multiple groups of data, each group of data includes a sample query request, knowledge data corresponding to the sample query request, and a task decomposition step corresponding to the sample query request; Accordingly, the target query request and the knowledge data matching the target query request are input into the target model to obtain the task decomposition steps corresponding to the target query request, including: The target query request, the knowledge data matching the target query request, the sample query request, and the task decomposition steps corresponding to the sample query request are input into a target model to obtain the task decomposition steps corresponding to the target query request.
9. A data query method, characterized in that: The method comprises: Get the target query request; Acquire knowledge data matching the target query request from a preset knowledge base; Inputting the target query request and the knowledge data matching the target query request into a target model to obtain a task decomposition step corresponding to the target query request, wherein the task decomposition step is described in an intermediate language between a natural language and a structured query language, and the target model is trained by the method according to any one of claims 1 to 4; Converting the task decomposition steps corresponding to the target query request into a structured query language statement; According to the structured query language statement, the target data is queried.
10. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
12. A computer program product, comprising computer program instructions, which implement the method according to any one of claims 1 to 9 when executed by a processor.