Nl2SQL system capability improvement method and apparatus, device and medium
By evaluating the capability level of the NL2SQL system from a capability indicator library and a test question library, establishing an improvement module matrix, and performing efficient scheduling, the problem of low improvement efficiency in existing methods is solved, and efficient and refined improvement of the system is achieved in specific scenarios.
Patent Information
- Application Number
- PCT/CN2024/138599
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2024-12-11
- Publication Date
- 2025-10-16
AI Technical Summary
Existing NL2SQL system capacity improvement methods lack effective measurement methods and combined scheduling methods, resulting in low improvement efficiency and prone to loss spikes, which affects optimization efficiency.
By determining the capability items of the NL2SQL system to be improved in the target scenario from a pre-built capability indicator library, using a preset scenario test question bank to evaluate the capability level, establishing a capability improvement module matrix, and selecting the best tasks for collective grouping and scheduling, efficient capability perception improvement can be achieved.
It has achieved efficient and refined capability improvements for the NL2SQL system in specific scenarios, narrowed the scope of improvement, improved measurement efficiency and comprehensiveness of results, and reduced complexity.
Smart Images

Figure CN2024138599_16102025_PF_FP_ABST
Abstract
Description
NL2SQL system capability improvement method, device, equipment and medium TECHNICAL FIELD
[0001] The present application relates to the technical field of NL2SQL system capability perception, and particularly relates to an NL2SQL system capability improvement method, device, equipment and medium. BACKGROUND
[0002] Nowadays, the development of technologies such as big data, cloud computing and the Internet of Things has brought about an explosive growth in data volume. These data are usually stored in a structured form in a database and interacted using a structured query language (SQL) query language. However, for non-professional users, learning and understanding and making good use of the SQL language is a challenging task. Therefore, many researchers have begun to explore methods for converting natural language into NL2SQL (Natural Language to SQL) to reduce the learning cost of users and improve the convenience of interaction with the database.
[0003] The capabilities required in the field of NL2SQL systems are very complex. Currently, there are some problems in the general improvement method for NL2SQL systems as follows: 1) the current improvement method lacks an effective capability measurement method, lacks an effective fine-grained dataset construction method, and blindly increases data to improve the capability of the system, resulting in low efficiency of capability improvement; the current method lacks an effective model capability measurement method, and blindly increases data to improve the capability of the model, which results in poor capability improvement effect; the use of new corpus in training may cause conflicting data optimization directions, reducing optimization efficiency; 2) secondly, in the process of improving the capability of the current large model, there is a lack of an efficient combination scheduling method of data and training methods. In the existing method, the types of data and training methods are too complex, and the loss spike phenomenon often occurs in the pre-training process. The reason for this phenomenon is that the average gradient is used to optimize the direction together during pre-training, and when the number of tasks is too large, the commonality of the data is difficult to obtain, so the loss value is difficult to decrease or slowly decreases; therefore, in specific implementation, there is high complexity, and from the effect and efficiency, an NL2SQL system capability improvement method is urgently needed to solve the above technical problems. SUMMARY
[0004] Therefore, the present application provides an NL2SQL system capability improvement method, device, equipment and medium, which can comprehensively and meticulously quantify the NL2SQL system, has higher measurement efficiency, more comprehensive measurement results, and an efficient combination scheduling method, and realizes efficient improvement of the capability of the NL2SQL system with low improvement complexity.
[0005] According to an aspect of the present application, the embodiment of the present application provides a NL2SQL system capability improvement method, and the method comprises the following steps:
[0006] determining, from a pre-constructed capability index library, a capability item required by a to-be-improved NL2SQL system under a target scenario;
[0007] performing capability level evaluation on each capability item according to a preset scenario test question library, to obtain a capability level detail of the to-be-improved NL2SQL system;
[0008] determining a capability improvement module matrix corresponding to the capability item based on the capability level detail; wherein each element in the capability improvement module matrix represents a capability improvement task of the corresponding capability item; the horizontal coordinate in the capability improvement module matrix represents a capability item of the to-be-improved NL2SQL system, and the vertical coordinate represents a capability improvement task category of the to-be-improved NL2SQL system;
[0009] performing collection grouping on a best target capability improvement task based on a task improvement category to which each capability improvement task in the capability improvement module matrix belongs, to obtain a grouped target grouping combination, and scheduling the target grouping combination to improve the capability of the to-be-improved NL2SQL system.
[0010] According to another aspect of the present application, the embodiment of the present application further provides a NL2SQL system capability improvement device, and the device comprises:
[0011] a capability item determination module, configured to determine, from a pre-constructed capability index library, a capability item required by a to-be-improved NL2SQL system under a target scenario;
[0012] a level detail determination module, configured to perform capability level evaluation on each capability item according to a preset scenario test question library, to obtain a capability level detail of the to-be-improved NL2SQL system;
[0013] a improvement matrix determination module, configured to determine a capability improvement module matrix corresponding to the capability item based on the capability level detail; wherein each element in the capability improvement module matrix represents a capability improvement task of the corresponding capability item; the horizontal coordinate in the capability improvement module matrix represents a capability item of the to-be-improved NL2SQL system, and the vertical coordinate represents a capability improvement task category of the to-be-improved NL2SQL system;
[0014] The lifting module is configured to select optimal target capability lifting tasks based on a task lifting category to which each capability lifting task in the capability lifting module matrix belongs, to obtain a grouped target grouping combination, and to schedule the target grouping combination to perform capability-aware lifting on the NL2SQL system to be lifted.
[0015] According to another aspect of the present application, an electronic device is also provided, which comprises:
[0016] at least one processor; and
[0017] a memory in communication with the at least one processor; wherein
[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the NL2SQL system capability lifting method according to any one of the embodiments of the present application.
[0019] According to another aspect of the present application, a computer readable storage medium is also provided, which stores computer instructions for enabling a processor to perform the NL2SQL system capability lifting method according to any one of the embodiments of the present application when executed by the processor.
[0020] The technical solution of the embodiments of the present application can narrow the range of NL2SQL capability items to be lifted in subsequent NL2SQL capability-aware lifting and lifting by determining the capability items required by the NL2SQL system to be lifted in the target scenario from the capability index library, thereby quickly completing the adaptation of the NL2SQL system to the specific scenario. The capability grade details of the NL2SQL system to be lifted can be determined by evaluating the capability grades of each capability item through the preset scenario test question library, which can comprehensively and meticulously quantify the NL2SQL system, has higher measurement efficiency, and more comprehensive measurement results. The capability lifting module matrix corresponding to each capability item is determined through the capability grade details, each capability lifting module matrix contains one or more capability lifting tasks, and the optimal target capability lifting task is selected based on the task lifting category to which each capability lifting task belongs to perform grouping, thereby scheduling based on the grouping result, achieving an efficient combination scheduling mode, and realizing efficient and fine-grained lifting of the NL2SQL system with low lifting complexity.
[0021] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to make the technical solutions in the embodiments of the present application clearer, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0023] Fig. 1 is a flowchart of a NL2SQL system capability improvement method provided by an embodiment of the present application;
[0024] Fig. 2 is a flowchart of another NL2SQL system capability improvement method provided by an embodiment of the present application;
[0025] Fig. 3 is a construction schematic diagram of a final execution sequence provided by an embodiment of the present application;
[0026] Fig. 4 is a schematic diagram of still another NL2SQL system capability improvement method provided by an embodiment of the present application;
[0027] Fig. 5 is a structural block diagram of a NL2SQL system capability improvement device provided by an embodiment of the present application;
[0028] Fig. 6 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to make the technical solutions in the embodiments of the present application clearer, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0030] It should be noted that the terms "first", "second", and the like in the description, claims, and above drawings of the present application are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device including a series of steps or units does not necessarily have to include only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product, or device.
[0031] In an embodiment, FIG. 1 is a flowchart of a NL2SQL system capability improvement method provided by an embodiment of the present application, which can be applicable to the case of improving the capability perception of a NL2SQL system. The method can be executed by a NL2SQL system capability improvement device, which can be realized in the form of hardware and / or software and can be configured in an electronic device.
[0032] As shown in FIG. 1, the NL2SQL system capability improvement method in the embodiment is applied to a database node. The method includes the following specific steps:
[0033] S110, determining the required capability items of the NL2SQL system to be improved in the target scenario from a pre-constructed capability index library.
[0034] The NL2SQL system to be improved refers to a NL2SQL system that needs to be improved in terms of capability perception. The target scenario can be understood as one or more specific scenarios applicable to the NL2SQL system, which can be defined or selected by the user as needed.
[0035] In the embodiment, the capability items can include various types of capability items, for example, capability items of the "content" type, capability items of the "form" type, and the like, which are not limited in the embodiment. Of course, the capability items of the "content" type and the capability items of the "form" type can also be split into other capability items. For example, the capability items of the "content" type can be refined into two categories of "concept" and "operation", and the capability items of the "form" type can be refined into two categories of "explicit logic" and "implicit logic". It should be noted that each capability item corresponds to a corresponding capability index and test case, which can be understood as a one-to-one correspondence between the capability item and the capability index, and a one-to-one correspondence between each capability index and the corresponding test case. In the embodiment, each capability item corresponds to a capability index, and the capability index corresponds to an index test case.
[0036] In the embodiment, the filtering of the scenario capability index can be completed in various ways. In an embodiment, the approximate range of the required capability items of the NL2SQL system to be improved in the target scenario can be given by analyzing the log files and the like provided by the scenario, and the range of the automatically analyzed indexes can be supplemented on this basis to finally obtain the capability items of the NL2SQL system that need to be improved. In another embodiment, the approximate range of the required capability items of the NL2SQL system to be improved in the target scenario can also be filtered out by keyword matching and the like, which is not limited in the embodiment.
[0037] In an embodiment, each capability index in the capability index library corresponds to a corresponding index test case;
[0038] The establishment of the capability index library includes: using a preset keyword extraction algorithm to identify keywords in a first original document related to the NL2SQL system to be improved, and taking each keyword as a first candidate of a first capability item; wherein the preset keyword extraction algorithm includes one of a word frequency statistical method, a part-of-speech tagging method, a named entity method, a Term Frequency-Inverse Document Frequency (TF-IDF) method, or a TextRank method;
[0039] Second original documents related to the NL2SQL system to be improved are collected, and query statements and corresponding structured query statements SQL in the second original documents are extracted, logical features contained in the query statements and the corresponding structured query statements SQL are found, and the logical features are taken as second candidates of second capability items;
[0040] The first candidates and the second candidates form two categories of first target capability items and second target capability items, respectively, thereby forming the capability index library; wherein the first target capability items include the first capability items; and the second target capability items include the second capability items;
[0041] The construction of the index test case includes: for each capability item belonging to the first target capability item category, a first type test case in a first category form is constructed, for each capability item belonging to the second target capability item category, a second type test case in a second category form is constructed, and the first type test case and the second type test case are put into the capability index library.
[0042] In the embodiment, the first capability item can be understood as two categories of "concept" and "operation" capability items, i.e., belonging to "content" capability items, wherein "concept" relates to each concept-related capability item in NL2SQL, such as the definition of the "SELECT" keyword, etc.; and "operation" relates to the definition of each operation in NL2SQL, such as the "SUM" function which can calculate the sum of a specified column; the second capability item in the embodiment can be understood as two categories of "explicit logic" and "implicit logic" capability items, i.e., belonging to "form" capability items, wherein "explicit logic" is a logic knowledge that is explicitly expressed in SQL, such as the combination of logical conditions using "AND" and "OR" keywords; and "implicit logic" is a logic knowledge that is not explicitly expressed in NL2SQL, such as some industry conventions and hidden logic not expressed in text expression, such as the data duplication problem that may be caused when joining multiple tables in JOIN, which requires the introduction of distinct or group by for deduplication.
[0043] In the embodiment, the first target capability item is the capability item belonging to "content", and the second target capability item is the capability item belonging to "form", and the first capability item and the second capability item can be used as candidates to form the first target capability item and the second target capability item respectively. It can be understood that, first, original documents related to NL2SQL and containing domain knowledge are collected, such as SQL professional books, query and optimization logs of actual databases, SQL scripts, related training materials, related papers, web pages, etc.; then, methods such as word frequency statistics, part-of-speech tagging, and named entity recognition are used to identify key terms and specific entities in the domain, or key word extraction algorithms such as TF-IDF and TextRank are used to extract key words and phrases related to the domain from the text as candidates for "concept" and "operation" capability items; thirdly, related materials of NL2SQL are collected, such as training sets, database logs, and annotated SQL scripts, and the query statements and corresponding query SQLs are extracted, and then appropriate prompts are written to use large models to extract "explicit logic" and "implicit logic" contained in the query statements and query SQLs as candidates for related capability items; finally, the candidates for "concept" and "operation" capability items, and the candidates for "explicit logic" and "implicit logic" related capability items are used to form a usable capability index library, the capability candidates of "concept", "operation", "explicit logic", and "implicit logic" are screened and supplemented to form a usable capability index library, and the capability index library is updated regularly.
[0044] In the embodiment, in the case where the preset key word extraction algorithm is one of the word frequency statistics method, the part-of-speech tagging method, and the named entity method, the preset key word extraction algorithm is used to identify the key words in the first original document related to the NL2SQL system to be improved. Specifically, the word frequency statistics method, the part-of-speech tagging method, and the named entity recognition method can be used to identify key terms and specific entities in the domain. For example, first, the collected text data related to the domain is preprocessed to remove special characters, punctuation marks, etc. in the text, and then a word segmentation operation is performed to divide the text into individual words or words. The word frequency statistics method can be used to calculate the frequency of each word in the text, and the words with high frequency can be selected to identify the key terms and specific entities in the domain. In this process, the part-of-speech tagging method can be used to determine the parts of speech of the selected words, such as nouns, verbs, adjectives, etc. The named entity recognition method can be used to identify specific types of entities in the text by training a machine learning model to assist the identification of key terms and specific entities by the word frequency statistics method.
[0045] In the present embodiment, in the case where the preset keyword extraction algorithm is one of the term frequency-inverse document frequency (TF-IDF) method or the graph ranking method (TextRank), the keywords in the first original document related to the NL2SQL system to be improved are identified using the preset keyword extraction algorithm. Specifically, the TF-IDF method can be used to extract keywords and phrases related to the field. Specifically, the TF-IDF value of a word is obtained by multiplying the frequency of the word in the text (term frequency) and the inverse document frequency (IDF) in the entire text set. By sorting the TF-IDF values, keywords with high TF-IDF values are extracted. These keywords are usually the most relevant words in the text content, i.e., the keywords and phrases sought. In some embodiments, the TextRank method can also be used. TextRank is a graph-based ranking algorithm that determines the importance of phrases by analyzing the relationships between words in the text. Specifically, the words in the text can be used as nodes to construct a graph, and the co-occurrence relationship between the words can be used as the weight of the edge. Then, the TextRank algorithm is run to calculate the importance score of each word. Finally, phrases with high scores are extracted as key phrases based on the importance scores.
[0046] In the present embodiment, for each ability item belonging to the "content" category, a test case in the form of a fill-in-the-blank, selection, or judgment is constructed. For each ability item belonging to the "form" category, a test case in the form of generation, error correction, or optimization is constructed. The construction of the test case can be performed manually or automatically using template extraction, large model generation, etc. The generated test cases are finally placed in the index library for subsequent use. Specifically, when constructing the test case, a specific template can be designed, and then the relevant content can be extracted and filled into the template to generate the test case. Alternatively, a special prompt can be designed, and then a large model can be used to generate the corresponding test case. The generated test cases are finally placed in the index library for subsequent use.
[0047] In the present embodiment, in the case where the preset keyword extraction algorithm is one of the term frequency-inverse document frequency (TF-IDF) method or the graph ranking method (TextRank), the keywords in the first original document related to the NL2SQL system to be improved are identified using the preset keyword extraction algorithm. Specifically, the TF-IDF method can be used to extract keywords and phrases related to the field. Specifically, the TF-IDF value of a word is obtained by multiplying the frequency of the word in the text (term frequency) and the inverse document frequency (IDF) in the entire text set. By sorting the TF-IDF values, keywords with high TF-IDF values are extracted. These keywords are usually the most relevant words in the text content, i.e., the keywords and phrases sought. In some embodiments, the TextRank method can also be used. TextRank is a graph-based ranking algorithm that determines the importance of phrases by analyzing the relationships between words in the text. Specifically, the words in the text can be used as nodes to construct a graph, and the co-occurrence relationship between the words can be used as the weight of the edge. Then, the TextRank algorithm is run to calculate the importance score of each word. Finally, phrases with high scores are extracted as key phrases based on the importance scores.
[0048] The capability level details include a first level, a second low level, a third level, a fourth level, and a fifth level. Of course, the capability levels have a sequence, specifically, the first level is less than the second level, the second level is less than the third level, the third level is less than the fourth level, and the fourth level is less than the fifth level. The first level can be understood as that the test cases corresponding to different types of capability items do not score; the second level can be understood as that the test cases corresponding to each capability item belonging to the “content” category score to reach a preset first score requirement, but the test cases corresponding to the capability items belonging to the “form” category do not score; the third level can be understood as that the test cases of the first type score to reach a preset second score requirement, and the test cases of the generation type belonging to the “form” category reach the preset second score requirement; the fourth level can be understood as that the test cases of the generation type belonging to the “form” category reach a preset third score requirement, and the test cases of the error correction type belonging to the “form” category reach the preset third score requirement; and the fifth level can be understood as that the test cases of different types of capability items score to basically reach a preset fourth score requirement.
[0049] In the embodiment, each type of test case corresponding to each capability item in the scene test question bank can be tested to obtain a corresponding result, and the scores corresponding to various types of test cases can be determined according to the results and a preset expected result, and the capability level of the NL2SQL system to be improved can be classified by the score results to obtain the capability level details. In another embodiment, the scores of the answers to each test case can be weighted and rated, or different types of questions can be classified according to the general situation, and the classification results can be classified. The present embodiment is not limited in this regard.
[0050] S130, determining a capability improvement module matrix corresponding to the capability item based on the capability level details.
[0051] Each element in the capability improvement module matrix can also be referred to as an Improvement Matrix, and each element in the capability improvement module matrix represents an improvement task (Improvement Block) of the corresponding capability item. The horizontal coordinate in the capability improvement module matrix represents the capability item of the NL2SQL system to be improved, and the vertical coordinate represents the capability improvement task category of the NL2SQL system to be improved.
[0052] In the embodiment, the capability level details include level-compliant capability items and level-incompliant capability items. For the level-compliant capability items, since the capability items belong to compliance, a capability promotion module matrix does not need to be generated, or a capability promotion module matrix for maintaining the capability can be generated. For the level-incompliant capability items, different ways can be selected to construct the corresponding capability promotion module matrix according to the types of different incompliant capability items. Specifically, in the case where the incompliant capability item belongs to the first target capability item, if the capability level details are the first level, a data set corresponding to the first target capability item is constructed, a preset mask learning method can be used for knowledge injection as the capability promotion module matrix, or a data set for instruction fine-tuning is constructed as the capability promotion module matrix. If the capability level details are the second level, a supervised fine-tuning SFT method can be used for concept alignment, or an external knowledge base is used, or a knowledge injection method in a prompt is used to determine the capability promotion module matrix. In the case where the incompliant capability item belongs to the second target capability item, if the capability level details are the third level, a form alignment method is used to determine the capability promotion module matrix through instruction learning. If the capability level details are the fourth level, implicit logic can be converted into explicit logic through intent recognition and few-shot learning, or the capability promotion module matrix is determined through finetuning by adding an intent recognition corpus, or the capability promotion module matrix is determined through an external text or knowledge graph form knowledge base. Of course, other ways can be used to select other existing ways to determine the capability promotion module matrix in the case of different capability item levels and types, and the embodiment is not limited in this regard.
[0053] In S140, based on the task promotion categories to which the capability promotion tasks in the capability promotion module matrix belong, a best target capability promotion task is selected for grouping, a grouped target grouping combination is obtained, and the target grouping combination is dispatched to promote the capability perception of the NL2SQL system to be promoted.
[0054] Each target grouping combination includes one or more capability promotion tasks.
[0055] The task improvement categories can include, but are not limited to, mask learning, instruction learning, adaptor learning, reinforcement learning from human feedback (RLHF), and in-context learning (ICL). The target capability improvement task is the optimal capability improvement task, which can include one or more in the embodiment. The target grouping combination can include one or more grouping combinations, which can be scheduled according to the execution order of the one or more grouping combinations.
[0056] In the embodiment, the capability improvement task matrix is grouped into Group Blocks according to the task improvement categories to which the capability improvement tasks belong in the capability improvement task matrix, and a grouped Group Block is obtained; each Group Block contains one or more Improvement Blocks, all Group Blocks are scheduled, a pipeline for the final execution of the Group Block is constructed and optimized, and the Group Block is scheduled and run according to the pipeline to improve the capability of the NL2SQL system to be improved. Specifically, the capability improvement tasks belonging to the same task improvement category can be merged to obtain a first merging result, the optimal capability improvement task can be selected from the first merging result according to a preset requirement to obtain a second merging result, the optimal target capability improvement task of the second merging result can be grouped, and the execution order can be obtained according to the execution dependency relationship between the target grouping combinations to be scheduled and run to improve the capability of the NL2SQL system to be improved.
[0057] The technical scheme of the embodiment of the present application can determine the required capability items of the NL2SQL system to be improved in the target scene from the capability index library, thereby narrowing the range of the capability items for subsequent NL2SQL capability perception and improvement, and quickly completing the adaptation of the NL2SQL system to the specific scene; the capability level details of the NL2SQL system to be improved are determined through the capability items, thereby comprehensively and meticulously quantitatively measuring the NL2SQL system, the measurement efficiency is higher, the measurement result is more comprehensive, the capability improvement module matrix corresponding to the capability items is determined through the capability level details, one or more capability improvement tasks are included in each capability improvement module matrix, the best target capability improvement task is selected according to the task improvement category to which it belongs for grouping, thereby scheduling based on the grouping result, achieving an efficient combination scheduling mode, and realizing efficient and refined improvement of the NL2SQL system capability perception under the condition of low improvement complexity.
[0058] In an embodiment, after the capability perception and improvement of the NL2SQL system to be improved, the following steps are further included:
[0059] determining whether the NL2SQL system after the capability perception and improvement meets the preset requirements of the target scene;
[0060] if yes, directly applying the target NL2SQL system to the application scene;
[0061] if no, returning to the step of determining the capability improvement module matrix corresponding to the capability items based on the capability level details, until the NL2SQL system after the capability perception and improvement meets the preset requirements of the target scene.
[0062] In the present embodiment, it is determined whether the NL2SQL system after the capability perception and improvement meets the preset requirements of the target scene, in the case of meeting the preset requirements of the target scene, the target NL2SQL system is directly applied to the application scene, if not, the step of determining the capability improvement module matrix corresponding to the capability items based on the capability level details is returned, until the NL2SQL system after the capability perception and improvement meets the preset requirements of the target scene. The preset requirements refer to the pre-set scene requirements, which can be manually set according to certain requirements.
[0063] In an embodiment, FIG. 2 is a flowchart of another NL2SQL system capability improvement method provided by an embodiment of the present application. Based on the above embodiments, the embodiment determines the required capability items of the NL2SQL system to be improved in the target scenario from the pre-constructed capability index library, evaluates the capability levels of the capability items according to the preset scenario test question library, to obtain the capability level details of the NL2SQL system to be improved; determines the capability improvement module matrix corresponding to the capability items based on the capability level details; selects the best target capability improvement task for grouping based on the task improvement categories to which each capability improvement task in the capability improvement module matrix belongs, obtains the grouped target grouping combination, and schedules the target grouping combination to further refine the capability perception improvement of the NL2SQL system to be improved.
[0064] As shown in FIG. 2, the NL2SQL system capability improvement method in the embodiment can specifically include the following steps:
[0065] S210, obtaining the historical log files stored by the NL2SQL system to be improved in the target scenario.
[0066] The historical log files refer to log files generated in a specific scenario. Each scenario corresponds to a corresponding scenario log file. The log file can include the related capability items of the NL2SQL system to be improved in the specific scenario, the score of the test case corresponding to the capability item, the level details, etc.
[0067] In the embodiment, the related capability items of the NL2SQL system to be improved in the target scenario, the score of the test case corresponding to the capability item, the level details, etc. historical log file information is obtained.
[0068] S220, filtering the target range of the required capability items in the target scenario from the capability index library according to the historical log files; wherein the target range includes the capability index corresponding to the capability item, and the index test case corresponding to the capability index.
[0069] In the embodiment, the target range of the required capability items in the target scenario can be filtered from the capability index library according to the historical log files; wherein the target range includes the capability index corresponding to the capability item, and the index test case corresponding to the capability index. In the embodiment, the filtering method of the required capability items in the target scenario is not limited, and can be filtered by the existing method in the prior art.
[0070] S230, testing the first type test case in the test case corresponding to each capability item according to the preset scenario test question library to obtain a first result, and determining the first score result corresponding to each first type test case according to the first result and a preset first expected result.
[0071] The first type of test case can include, but is not limited to, fill-in-the-blank, selection, judgment, and the like. The first type of test case belongs to the test case corresponding to the ability item of the "content" category. It can be understood that each ability item belonging to the "content" category can construct test cases in the form of fill-in-the-blank, selection, and judgment. The preset first expected result refers to the standard result set by the user according to experience and the like.
[0072] In this embodiment, the scene test question bank is obtained. The first type of test case in the test case corresponding to each ability item can be tested by the scene test question bank to obtain a first result, and the first score result corresponding to each first type of test case is determined according to the first result and the preset first expected result. Specifically, the first result and the preset first expected result are compared. In the case where the first result is less than the preset first expected result, the first score result corresponding to each first type of test case is determined as unqualified, and no score is obtained. If the first result is greater than or equal to the preset first expected result, the first score result corresponding to each first type of test case is determined as qualified, and a score is obtained.
[0073] S240, the second type of test case in the test case corresponding to each ability item is tested according to the preset scene test question bank to obtain a second result, and the second score result corresponding to each second type of test case is determined according to the degree of conformity between the second result and the preset second expected result.
[0074] The second type of test case can include, but is not limited to, generation, error correction, and optimization. The second type of test case belongs to the test case corresponding to the ability item of the "form" category. It can be understood that the ability item belonging to the "form" category can construct test cases in the form of generation, error correction, and optimization. The preset second expected result refers to the standard result set by the user according to experience and the like. It should be noted that the preset first expected result and the preset second expected result can be the same or different, and are set according to requirements.
[0075] In this embodiment, the preset scene test question bank is obtained. The second type of test case in the test case corresponding to each ability item can be tested by the preset scene test question bank to obtain a second result, and the second score result corresponding to each second type of test case is determined according to the second result and the preset second expected result. Specifically, the second result and the preset second expected result are compared. In the case where the second result is less than the preset second expected result, the second score result corresponding to each second type of test case is determined as unqualified, and no score is obtained. If the second result is greater than or equal to the preset second expected result, the second score result corresponding to each second type of test case is determined as qualified, and a score is obtained.
[0076] It should be noted that the first type of test case and the second type of test case in the test case corresponding to each capability item in the preset scene test question bank can be tested in an average allocation manner or a weighted average manner. For example, in a total score of 100 points, all test cases in the index are allocated a corresponding total score according to certain rules. For test cases in the form of fill-in-the-blank, selection, and judgment, it is determined whether the result conforms to the standard result. If it conforms, a score is obtained, and if it does not conform, no score is obtained. For generation, optimization, and error correction, the degree of conformity between the answer and the standard answer is determined, and a score between 0 and the full score is given according to the degree of conformity. All judgments can be completed by artificial means, or a special model can be used to give the judgment result.
[0077] S250, grading the capability level of the NL2SQL system to be promoted according to the first score result and the second score result, to obtain a capability level detail.
[0078] The capability level detail includes: a first level, a second level, a third level, a fourth level, and a fifth level.
[0079] In this embodiment, the capability level of the NL2SQL system to be promoted is graded according to the first score result and the second score result, to obtain a corresponding capability level detail. In an embodiment, the manner of dividing the capability level detail includes: in the case where the first type of test case and the second type of test case in the NL2SQL system to be promoted have not scored, the capability level detail is determined to be the first level; in the case where the score of the first type of test case reaches a preset first score requirement, and the second type of test case has not scored, the capability level detail is determined to be the second level; in the case where the score of the first type of test case reaches a preset second score requirement, and the score of the second type of test case reaches a preset second score requirement, the capability level detail is determined to be the third level; in the case where the score of the first type of test case reaches a preset third score requirement, and the score of the second type of test case reaches a preset third score requirement, the capability level detail is determined to be the fourth level; in the case where the score of the first type of test case reaches a preset fourth score requirement, and the score of the second type of test case reaches a preset fourth score requirement, the capability level detail is determined to be the fifth level; wherein the first level is less than the second level, the second level is less than the third level, the third level is less than the fourth level, and the fourth level is less than the fifth level; the preset first score requirement is less than the preset second score requirement; the preset second score requirement is less than the preset third score requirement; and the preset third score requirement is less than the preset fourth score requirement.
[0080] In the embodiment, in the case that neither the first type test case nor the second type test case scores in the to-be-promoted NL2SQL system, the capability level details are determined to be the first level; in the case that the score of the first type test case reaches the preset first score requirement, and none of the second type test cases scores, the capability level details are determined to be the second level; in the case that the score of the first type test case reaches the preset second score requirement, and the score of the second type test case reaches the preset second score requirement, the capability level details are determined to be the third level; in the case that the score of the first type test case reaches the preset third score requirement, and the score of the second type test case reaches the preset third score requirement, the capability level details are determined to be the fourth level; in the case that the score of the first type test case reaches the preset fourth score requirement, and the score of the second type test case reaches the preset fourth score requirement, the capability level details are determined to be the fifth level; it can be understood that for each capability index, its capability is divided into five levels from low to high, i.e., "None", "understand", "unfamiliar", "available", and "master", wherein "None" means that the model does not have SQL-related knowledge at all and answers non-related SQL tasks; "understand" means that the model can master the syntax rules; "unfamiliar" means that the model can complete the corresponding SQL task with the help of few shot and prompt engineering on the basis of understanding the SQL syntax rules; "available" means that the model can complete the application of SQL query for actual query problems, including the ability to judge whether the SQL statement is correct and to generate the corresponding SQL query statement according to the natural language form of the query statement; and "master" means that the model can judge the use of the SQL statement, including correcting the SQL statement with problems or optimizing the SQL statement.
[0081] In this embodiment, in order to grade each capability item, a specific capability grading method can be designed. For example, the system can grade each capability item according to the weighted sum of the scores of the answers to the questions, and the grading is as follows: 0-20 points for "none", 20-40 points for "understand", and so on. Alternatively, the NL2SQL system to be improved can be graded according to the answers to various types of questions. If no points are scored for various types of test cases, the capability level of the capability item is "none". If points are scored for fill-in-the-blank, multiple-choice, and judgment, but not for generation, optimization, and error correction, the capability level of the capability item is "understand". If the NL2SQL system to be improved can score points for generation type test cases, the capability level of the capability item is "unfamiliar". If the NL2SQL system to be improved can score points for error correction type test cases, the capability level of the capability item is "available". If the NL2SQL system to be improved can score points for optimization type test cases, the capability level of the capability item is "master". Of course, the specific grading rules can be adjusted according to actual conditions.
[0082] S260, selecting, from the capability level details, a non-standard capability item with a non-standard capability level and a standard capability item with a standard capability level.
[0083] In this embodiment, the non-standard capability item with a non-standard capability level and the standard capability item with a standard capability level are selected from the capability level details. The non-standard capability item refers to a capability item with a capability level of the first, second, third, and fourth levels, i.e., the capability levels of "none", "understand", "unfamiliar", and "available". The standard capability item refers to a capability item with a capability level of the fifth level, i.e., the capability level of "master".
[0084] S270, for the non-standard capability item, if the non-standard capability item belongs to the first target capability item, and the capability level details are of the first level, a data set corresponding to the first target capability item is constructed, and a preset mask learning method is used for knowledge injection to construct a capability improvement module matrix.
[0085] The first target capability item is each capability item belonging to the "content" category.
[0086] In the embodiment, for the substandard ability item, in the case that the substandard ability item belongs to the first target ability item, if the ability level details are the first level, a data set corresponding to the first target ability item is constructed, a preset mask learning method is used to select full fine-tuning for knowledge injection as the ability promotion module matrix, or the instruction fine-tuning mode is used as the ability promotion module matrix. It can be understood that for the ability item belonging to "content", if the score on the ability item is very low, that is, "none", the data set corresponding to the ability item can be constructed, the mask learning method is used to select full fine-tuning for knowledge injection, or the instruction fine-tuning data is constructed, that is, the preset mask learning method is used for knowledge injection, or the instruction fine-tuning mode is used as the longitudinal coordinate of the ability promotion module matrix, and the ability item belonging to "content" is used as the horizontal coordinate. In the embodiment, the preset mask learning Mask Learning method is a learning method in the prior art, and the specific method is not described in the embodiment.
[0087] In the case that the substandard ability item belongs to the first target ability item, if the ability level details are the second level, the supervised fine-tuning SFT mode is used to construct the ability promotion module matrix.
[0088] In the embodiment, in the case that the substandard ability item belongs to the first target ability item, if the ability level details are the second level, the supervised fine-tuning SFT mode is used to construct the ability promotion module matrix.
[0089] In the case that the substandard ability item belongs to the first target ability item, if the ability level details are the second level, the supervised fine-tuning SFT mode is used to construct the ability promotion module matrix.
[0090] In the case that the substandard ability item belongs to the first target ability item, if the ability level details are the second level, the supervised fine-tuning SFT mode is used to construct the ability promotion module matrix.
[0091] In the embodiment, in the case where the substandard capability item belongs to the capability item of the second target capability item, if the capability level details are the third level, the form alignment mode is determined by the instruct learning mode. It can be understood that, for the capability item belonging to the "form", if the evaluation on the capability item is "understand", the form alignment can be performed by the instruct SFT learning mode, that is, the instruct SFT learning mode is taken as the vertical coordinate of the capability promotion module matrix, and the capability item belonging to the "form" is taken as the horizontal coordinate. In the embodiment, the instruct learning mode is the mode in the prior art, which is not specifically described herein.
[0092] S2100, in the case where the substandard capability item belongs to the capability item of the second target capability item, if the capability level details are the fourth level, the capability promotion module matrix is constructed by any one of the few-shot learning mode, the fine-tuning mode, or the external knowledge graph.
[0093] In the embodiment, in the case where the substandard capability item belongs to the capability item of the second target capability item, if the capability level details are the fourth level, the implicit logic can be converted into explicit logic by the intent recognition and the few-shot learning mode, or the fine-tuning is performed by adding the corpus of the intent recognition, or the capability promotion module matrix is determined by the method of externally hanging the knowledge base in the form of text or knowledge graph. It can be understood that, if the evaluation is "available", the implicit logic can be converted into explicit logic by the intent recognition and the few-shot learning mode, or the fine-tuning is performed by adding the corpus of the intent recognition, or the method of externally hanging the knowledge base in the form of text or KG (knowledge graph) is used to improve the performance of the model on the related capability item, that is, the few-shot learning mode, the fine-tuning mode, or any one of the modes of externally hanging the knowledge graph is taken as the vertical coordinate of the capability promotion module matrix, and the capability item belonging to the "form" is taken as the horizontal coordinate.
[0094] S2110, for the standard capability item, the capability promotion module matrix does not need to be generated, or a capability promotion module matrix maintaining the capability is generated.
[0095] In the embodiment, for the standard capability item, the capability promotion module matrix does not need to be generated, or a capability promotion module matrix maintaining the capability is generated.
[0096] In an embodiment, to facilitate better understanding of the relationship between the NL2SQL capability item and the capability improvement module matrix, Table 1 is a relationship diagram between the NL2SQL capability item and the capability improvement module matrix provided by an embodiment of the present application. In the table, Mask Learning is mask learning, Few-shot Learning is few-shot learning, instruct SFT Learing is an instructing learning method, RLHF is a language model optimized by reinforcement learning from human feedback (Reinforcement Learning from Human Feedback), and Adaptor Learning is adaptor learning.
[0097] Table 1: Relationship diagram between NL2SQL capability item and capability improvement module matrix
[0098] S2120, determine whether each capability improvement task is of the same task improvement category, if yes, execute S2130, if no, execute S2140.
[0099] In this embodiment, if the capability improvement tasks are of the same task improvement category, the capability improvement tasks of the same task improvement category are merged to obtain a merged result, and the best capability improvement task is selected in the merged result according to a preset requirement to perform secondary merging. The best target capability improvement task in the second merged result is grouped to obtain a grouped target grouping combination.
[0100] S2130, the capability improvement tasks of the same task improvement category are merged to obtain a first merged result, the best capability improvement task is selected in the first merged result according to a preset requirement to perform secondary merging to obtain a second merged result, and the best target capability improvement task in the second merged result is grouped to obtain a grouped target grouping combination.
[0101] The preset requirement includes: the capability improvement task with the largest difference in training corpus, or the capability improvement task with the most similar training corpus.
[0102] In this embodiment, if the capability improvement tasks are of the same task improvement category, the capability improvement tasks of the same task improvement category are merged to obtain a first merged result, the best capability improvement task is selected in the first merged result according to a preset requirement to perform secondary merging to obtain a second merged result, and the best target capability improvement task in the second merged result is grouped to obtain a grouped target grouping combination.
[0103] In this embodiment, the best Improvement Block is selected for merging, for example: 1. When the Improvement Block of the SFT training task is merged, the Block with the most similar training corpus can be selected, the corpus is merged to form a Group Block, which can reduce the difficulty of model training; 2. When the Improvement Block of the pre-training task is merged, the Block with the greatest difference in training corpus can be selected for merging, and the corpus of each block is scrambled according to certain rules to improve the diversity of training samples and avoid model overfitting. According to certain rules, the Improvement Matrix is merged into a series of Group Blocks, each Group Block containing one or more Improvement Blocks. The merging logic of the Group Block needs to take into account many factors. In this embodiment, the merged Improvement Block should belong to the same type of improvement method; secondly, in order to achieve the best improvement effect, the best merging granularity needs to be selected when merging, and the best Improvement Block is selected for merging according to the ability item to which the Improvement Block belongs. Of course, the merging rules are not limited to the above listed items, and the merging rules need to be designed according to the actual situation in practice.
[0104] S2140, then do not merge.
[0105] In this embodiment, when the ability improvement tasks do not belong to the same task improvement category, merging is not performed.
[0106] S2150, obtaining an execution order according to the execution dependency relationship between each target group combination.
[0107] Wherein, the dependency relationship includes: serial execution or parallel execution.
[0108] In this embodiment, the execution order is obtained according to the execution dependency relationship between each target group combination; wherein, the dependency relationship includes: serial execution or parallel execution. In this embodiment, all target group combinations Group Block are scheduled, and the execution order pipeline of the final execution of the target group combination Group Block is constructed and optimized. The execution order represents the execution dependency relationship between all target group combinations Group Block. Generally, the Group Block in the pipeline is executed in series or in parallel according to the type of the dependency relationship. When there are Group Blocks that can be executed in parallel, they are combined into a target group combination and put into the pipeline. Through the optimization of the pipeline, the overall efficiency of the model capability improvement task can be improved.
[0109] S2160, in the execution order, according to whether the resources for executing the target group combination are sufficient, one or at least two target group combinations are selected for execution until all target group combinations are scheduled.
[0110] In this embodiment, in the execution order, according to whether the resources for executing the target group combination are sufficient, one or at least two target group combinations are selected for execution until all target group combinations are scheduled. It can be understood that the Group Block is scheduled and run according to the logic of the pipeline. Specifically, the execution starts from the first Group Block or Parallel Block in the pipeline. After the execution of the Block is completed, the Group Block or Parallel Block without unexecuted dependencies is selected for execution. At the pipeline level, when there is more than one Group Block or Parallel Block that can be selected, the module will select one or more Group Blocks (Parallel Blocks) for execution according to whether the resources for supporting the execution of the Block are sufficient. At the Parallel Block level, when the resources for supporting the execution of the Group Block in the Parallel Block are sufficient, it is more likely to select more Group Blocks for parallel execution. When all Group Blocks (Parallel Blocks) in the pipeline are executed, the NL2SQL capability improvement of the system is completed, and the NL2SQL capability of the improved system needs to be evaluated again. The evaluation result can be used to determine whether the improved model meets the requirements.
[0111] In the embodiment, in order to better understand the pipeline construction, FIG. 3 is a construction diagram of a final execution sequence provided by an embodiment of the application. In the embodiment, as shown in FIG. 3, the Group Block includes: Group Block 1, Group Block 2, …, Group Block n, which respectively represent the target group combination in the above embodiment; wherein the Group Block 1 includes: Improvement Block n1, Improvement Block n2, …, Improvement Block nm, and the Improvement Block represents the target ability improvement task in the above embodiment. After scheduling, the execution dependency relationship between each target group combination obtains the execution sequence, and the execution sequence in the embodiment is as follows: Group Block 1, Group Block 2, Group Block 3 and Group Block 4 are executed in parallel, Group Block n-1, Group Block n, wherein Group Block 3 and Group Block 4 can constitute a Parallel Block.
[0112] The above technical scheme of the embodiment of the application can filter out the ability indicators corresponding to the ability items required in the target scene from the ability index library through the historical log file, and the ability indicators corresponding to the ability items and the index test cases corresponding to the ability indicators, so as to better match the requirements of the scene, improve the efficiency, improve the efficiency of scene adaptation, comprehensively and carefully quantify the NL2SQL of the NL2SQL system, the measurement efficiency is higher, and the measurement result is more comprehensive and careful, which provides data guidance for subsequent NL2SQL ability improvement; the different types of test cases in the test cases corresponding to each ability item are tested to obtain results, and the scores corresponding to each test case are determined according to the results and preset expected results, the ability level of the NL2SQL system to be improved is classified according to the score results, the ability level details are obtained, and the corresponding substandard ability items are constructed based on the substandard ability items and the standard ability items in the ability level details, the ability improvement tasks belonging to the same task improvement category are combined and scheduled at the best combination granularity, and the high-efficiency combination scheduling mode is further achieved, so that the NL2SQL system ability is perceived with high efficiency and fine improvement in the case of low complexity.
[0113] For better understanding of the NL2SQL system capability improvement method, FIG. 4 is a schematic diagram of another NL2SQL system capability improvement method provided by an embodiment of the present application. In this embodiment, it is assumed that the Indicator warehouse has completed the construction of the possible force indicator system, forming the capability indicator library. In this embodiment, the "concept" and "operation" are the first capability items in the above embodiment; the "implicit logic" and "explicit logic" are the second capability items in the above embodiment; the "content" category capability item is the first target capability item in the above embodiment; the capability item belonging to the "form" category is the second target capability item in the above embodiment; "master" is the fifth level in the capability level details; "master" is the fourth level in the capability level details; "understand" is the third level in the capability level details; "familiar" is the second level in the capability level details; "none" is the first level in the capability level details; Improvement Matrix is the capability improvement module matrix in the above embodiment; Improvement Block is the capability improvement task in the above embodiment; Group Block is the target grouping combination in the above embodiment; pipeline is the final execution order in the above embodiment;
[0114] a1, screen the required capability items of the NL2SQL system in a specific scenario from the pre-constructed capability indicator library.
[0115] In this embodiment, the Scenario indicator filter filters the required NL2SQL capability items for the new NL2SQL use scenario, which are: "SELECT", "WHERE" in the "concept" category, "COUNT operation", "MAX operation" in the "operation" category, "AND and OR logical combination" in the "explicit logic" category, and "multi-table join count needs to be removed" in the "implicit logic" category. The required capability level is "master". For these capability items, the "fill in the blank", "selection", and "judgment" questions have been constructed for the "content" category, and the "generation", "error correction", and "optimization" questions have been constructed for the "form" category.
[0116] a2, obtain the capability item classification details of the NL2SQL system.
[0117] In this embodiment, the Measurement Module uses the preset scene test question bank to test the existing NL2SQL system, and the following test results are obtained: 1) "SELECT": 2 correct fill-in-the-blank questions, 2 correct selection questions, 2 correct judgment questions, and the ability is judged as "master"; 2) "WHERE": 1 correct fill-in-the-blank question, 1 correct selection question, 1 correct judgment question, and the ability is judged as "understand"; 3) "COUNT operation": 2 correct fill-in-the-blank questions, 1 correct selection question, 2 correct judgment questions, and the ability is judged as "master"; 4) "MAX operation": 1 correct fill-in-the-blank question, 1 correct selection question, 2 correct judgment questions, and the ability is judged as "unfamiliar"; 5) "AND and OR logical combination": "generate" 2 correct questions, "correct" 1 correct question, "optimize" 1 correct question, and the ability is judged as "unfamiliar"; 6) "Multi-table join count needs to be removed": "generate" 1 correct question, "correct" 0 correct question, "optimize" 1 correct question, and the ability is judged as "understand".
[0118] a3, constructing a capability defect item for a capability item that does not meet the standard in the capability item grading details.
[0119] In this embodiment, the Planning Module specifies the improvement method Improvement Block corresponding to each capability item according to the NL2SQL capability perception result: 1) "SELECT" does not need to generate an Improvement Block because the capability judgment is "master", which meets the needs of the scene (an Improvement Block for maintaining the capability can also be generated, which is not generated in this embodiment); 2) "WHERE": because the capability judgment is "understand", an Improvement Block 1 of Instrction Learning using the full parameters containing the "WHERE" corpus is constructed; 3) "COUNT operation": because the capability judgment is "master", and according to the characteristics of the "COUNT" capability item itself, an Improvement Block 2 of Adaptor Learning containing the "COUNT" corpus is constructed; 4) "MAX operation": because the capability judgment is "unfamiliar", and according to the characteristics of the "MAX" capability item itself, an Improvement Block 3 of Adaptor Learning containing the "MAX" corpus is constructed; 5) "AND and OR logical combination": because the capability judgment is "unfamiliar", and according to the characteristics of the "AND and OR logical combination" capability item itself, an Improvement Block 4 of Adaptor Learning containing the "AND and OR logical combination" corpus is constructed; 6) "multi-table join count needs to be removed": because the capability judgment is "understand", and according to the characteristics of the "multi-table join count needs to be removed" capability item itself, an Improvement Block 5 of Mask Learning containing the "multi-table join count needs to be removed" corpus is constructed, and at the same time, an Improvement Block 6 of intent recognition is constructed, which performs intent recognition on the input text, and when the logic of "multi-table join count" is recognized, the prompt is added to the prompt.
[0120] a4, merge the Improvement Blocks in the Improvement Matrix according to the rules to form a Group Block.
[0121] In this embodiment, the Group Module groups the generated Improvement Block into Group, and in this embodiment, the method of merging Improvement Blocks belonging to the same stage into a Group Block is used to construct Group Block. 1) Improvement Block 5 belongs to the pre-training stage, and there is no other block in this stage in this embodiment, so Group Block 1 is constructed, and there is only Improvement Block 5 in the Group Block; 2) Improvement Block 1 belongs to full fine-tuning, and there is no other block in this stage in this embodiment, so Group Block 2 is constructed, and there is only Improvement Block 1 in the Group Block; 3) Improvement Block 2, Improvement Block 3, and Improvement Block 4 all belong to Adaptor Learning, so they are merged, the training corpus used in Improvement Block 2, Improvement Block 3, and Improvement Block 4 is merged to generate a total instruction fine-tuning training corpus containing "COUNT", "MAX", "AND", and "OR" logical combinations, and Group Block 3 is constructed; 4) Improvement Block 6 belongs to the prompt engineering part, and there is no other block in this stage in this embodiment, so Group Block 4 is constructed separately.
[0122] a5, construct and optimize the execution pipeline of the Group Block: the Block Scheduler uses the generated Group Block to construct the pipeline.
[0123] In this embodiment, because there is no logic that can be executed in parallel in Group Block 1-4, the final pipeline is constructed as Group Block 1->Group Block 2->Group Block 3->Group Block 4.
[0124] a6, ability improvement and evaluation.
[0125] In this embodiment, the Improvement & Eval Module runs Group Blocks 1-4 in the logical order of the pipeline. 1) Run Group Block 1 to use the "multi-table join count needs to be deduplicated" corpus to perform Mask Learning training on the model in the system to obtain an improved model; 2) Run Group Block 2 to use the Instruction Learning corpus containing "WHERE" to further fine-tune the model; 3) Run Group Block 3 to use the total instruction fine-tuning training corpus containing "COUNT", "MAX", and logical combinations of "AND" and "OR" to further fine-tune the model; 4) Run Group Block 4 to train an RNN model for intent recognition to identify whether the query statement contains the intent of "multi-table join count needs to be deduplicated", and then find a suitable prompt format through prompt engineering, add the prompt of "multi-table join count needs to be deduplicated" in the final prompt, and add the intent recognition and prompt logic to the overall flow of the system; 5) The Improvement & Eval Module calls the Measurement Module to re-evaluate the NL2SQL capability of the improved system.
[0126] a7, determine whether the improved NL2SQL capability meets the needs of the scenario, if yes, execute a8, if not, return to a3, until the improved NL2SQL system meets the preset requirements of the target scenario, and determine whether the improved system can be used according to the evaluation results of the improved system capability.
[0127] a8, directly apply the target NL2SQL system to the application scenario.
[0128] In this embodiment, the evaluation results of the improved model are as follows: 1) "SELECT": the capability is judged to be "master"; 2) "WHERE": the capability is judged to be "master"; 3) "COUNT operation": the capability is judged to be "master"; 4) "MAX operation": the capability is judged to be "master"; 5) "AND and OR logical combination": the capability is judged to be "master"; 6) "multi-table join count needs to be deduplicated": the capability is judged to be "master".
[0129] In this embodiment, the improved model capability meets the needs of the scenario and can be used in the scenario.
[0130] In an embodiment, FIG. 5 is a structural block diagram of a NL2SQL system capability improvement device provided by an embodiment of the present application, which is applicable to the case of improving the capability perception of a NL2SQL system, and can be implemented by hardware / software. The device can be configured in an electronic device to implement a NL2SQL system capability improvement processing method in an embodiment of the present application.
[0131] As shown in FIG. 5, the device applied to a database node includes a capability item determination module 510, a level detail determination module 520, an improvement matrix determination module 530, and an improvement module 540.
[0132] The capability item determination module 510 is configured to determine the required capability items of a NL2SQL system to be improved in a target scenario from a pre-constructed capability index library.
[0133] The level detail determination module 520 is configured to perform capability level evaluation on each of the capability items according to a preset scenario test question bank to obtain a capability level detail of the NL2SQL system to be improved.
[0134] The improvement matrix determination module 530 is configured to determine a capability improvement module matrix corresponding to the capability items based on the capability level detail. Each element in the capability improvement module matrix represents a capability improvement task of a corresponding capability item. The horizontal coordinate in the capability improvement module matrix represents the capability items of the NL2SQL system to be improved, and the vertical coordinate represents the capability improvement task categories of the NL2SQL system to be improved.
[0135] The improvement module 1040 is configured to select the best target capability improvement task for grouping based on the task improvement categories to which each of the capability improvement tasks in the capability improvement module matrix belongs, obtain a grouped target grouping combination, and schedule the target grouping combination to improve the capability perception of the NL2SQL system to be improved. Each target grouping combination contains at least one capability improvement task.
[0136] The embodiment of the present application, the capability item determination module, by determining the capability item required by the NL2SQL system to be promoted in the target scene from the capability index library, can narrow the range of subsequent NL2SQL capability perception and promotion of the capability index, thereby quickly completing the adaptation of the NL2SQL system to the specific scene; the promotion matrix determination module and the promotion module determine the capability level details of the NL2SQL system to be promoted through each capability item, which can comprehensively and meticulously quantify and measure the NL2SQL system, has higher measurement efficiency, and the measurement result is more comprehensive, determines the capability promotion module matrix corresponding to the capability item through the capability level details, each capability promotion module matrix contains one or more capability promotion tasks, selects the best target capability promotion task for grouping according to the task promotion category, and then schedules based on the grouping result, so as to achieve an efficient combination scheduling mode, and realize efficient and fine promotion of the NL2SQL system capability perception under the condition of low promotion complexity.
[0137] In an embodiment, the apparatus further comprises:
[0138] The judgment module is configured to judge whether the NL2SQL system after the capability perception promotion meets the preset requirement of the target scene after the capability perception promotion of the NL2SQL system to be promoted;
[0139] The application module is configured to directly apply the target NL2SQL system to the application scene if the preset requirement is met.
[0140] The cycle module is configured to return to the step of determining the capability promotion module matrix corresponding to the capability item based on the capability level details if the preset requirement is not met, until the NL2SQL system after the capability perception promotion meets the preset requirement of the target scene.
[0141] In an embodiment, each capability index in the capability index library corresponds to a corresponding index test case.
[0142] The establishment of the capability index library comprises:
[0143] The preset keyword extraction algorithm is used to identify the keywords in the first original document related to the NL2SQL system to be promoted, and each keyword is used as a first candidate of the first capability item; wherein the preset keyword extraction algorithm comprises one of the following: word frequency statistics, part-of-speech tagging, named entity, term frequency-inverse document frequency (TF-IDF), or graph ranking (TextRank).
[0144] collecting a second original document related to the NL2SQL system to be improved, extracting a query statement and a corresponding structured query language (SQL) statement from the second original document, finding logical features contained in the query statement and the corresponding structured query language (SQL) statement, and taking the logical features as second candidates of second capability items;
[0145] forming two categories of first target capability items and second target capability items from the first candidates and the second candidates, respectively, to thereby constitute the capability index library; wherein the first target capability items contain the first capability items; and the second target capability items contain the second capability items;
[0146] The construction of the index test case includes:
[0147] Each capability item belonging to the first target capability item category is constructed into a first-type test case of a first category form, each capability item belonging to the second target capability item category is constructed into a second-type test case of a second category form, and the first-type test case and the second-type test case are put into the capability index library.
[0148] In an embodiment, the capability item determination module 510 includes:
[0149] The file acquisition unit is configured to acquire a historical log file stored by the NL2SQL system to be improved in the target scenario;
[0150] The range determination unit is configured to filter a target range of capability items required in the target scenario from the capability index library according to the historical log file, wherein the target range includes capability indexes corresponding to the capability items and index test cases corresponding to the capability indexes.
[0151] In an embodiment, the grade detail determination module 520 includes:
[0152] The first result determination unit is configured to test first-type test cases in the test cases corresponding to each of the capability items according to the preset scenario test question bank to obtain a first result, and determine a first score result corresponding to each of the first-type test cases according to the first result and a preset first expected result.
[0153] The second result determination unit is configured to test second-type test cases in the test cases corresponding to each of the capability items according to the preset scenario test question bank to obtain a second result, and determine a second score result corresponding to each of the second-type test cases according to a degree of conformity between the second result and a preset second expected result.
[0154] a level determination unit configured to grade the ability level of the to-be-upgraded NL2SQL system according to the first score result and the second score result, to obtain an ability level detail; wherein the ability level detail comprises a first level, a second level, a third level, a fourth level, and a fifth level.
[0155] In an embodiment, the ability level detail is divided in the following manner: in a case where neither the first type of test case nor the second type of test case scores in the to-be-upgraded NL2SQL system, the ability level detail is determined to be the first level; in a case where the first type of test case scores reaches a preset first score requirement, and none of the second type of test case scores, the ability level detail is determined to be the second level; in a case where the first type of test case scores reaches a preset second score requirement, and the second type of test case scores reaches the preset second score requirement, the ability level detail is determined to be the third level; in a case where the first type of test case scores reaches a preset third score requirement, and the second type of test case scores reaches the preset third score requirement, the ability level detail is determined to be the fourth level; in a case where the first type of test case scores reaches a preset fourth score requirement, and the second type of test case scores reaches the preset fourth score requirement, the ability level detail is determined to be the fifth level.
[0156] The first level is less than the second level, the second level is less than the third level, the third level is less than the fourth level, and the fourth level is less than the fifth level; the preset first score requirement is less than the preset second score requirement; the preset second score requirement is less than the preset third score requirement; and the preset third score requirement is less than the preset fourth score requirement.
[0157] In an embodiment, the promotion matrix determination module 530 comprises:
[0158] a selection unit configured to select, from the ability level detail, an unqualified ability item whose ability level is unqualified, and a qualified ability item whose ability level is qualified;
[0159] a first level determination unit configured to, for the unqualified ability item, if the unqualified ability item belongs to a first target ability item, and the ability level detail is the first level, construct a data set corresponding to the first target ability item, and perform knowledge injection using a preset mask learning method to construct an ability promotion module matrix.
[0160] The second grade determining unit is configured to, in a case where the substandard ability item belongs to an ability item of the first target ability item, if the ability grade details is the second grade, adopt a supervised fine-tuning (SFT) manner to construct the ability promotion module matrix in any one of the concept alignment and the external knowledge base;
[0161] The third grade determining unit is configured to, in a case where the substandard ability item belongs to an ability item of the second target ability item, if the ability grade details is the third grade, construct the ability promotion module matrix in a formal alignment manner through an instructed learning manner;
[0162] The fourth grade determining unit is configured to, in a case where the substandard ability item belongs to an ability item of the second target ability item, if the ability grade details is the fourth grade, construct the ability promotion module matrix in any one of a few-sample learning manner, a fine-tuning manner, and an external knowledge graph.
[0163] The fifth grade determining unit is configured to, for the standard ability item, not generate the ability promotion module matrix, or generate an ability promotion module matrix for maintaining the ability.
[0164] In an embodiment, the promotion module 540 includes:
[0165] The category determining unit is configured to determine whether the ability promotion tasks are of the same task promotion category.
[0166] The grouping unit is configured to, if yes, first merge the ability promotion tasks of the same task promotion category to obtain a first merging result, second merge the best ability promotion task in the first merging result according to a preset requirement to obtain a second merging result, and group the best target ability promotion task in the second merging result to obtain a grouped target grouping combination.
[0167] The determining unit is configured to, if no, not perform the merging.
[0168] In an embodiment, the promotion module 540 includes:
[0169] The sequence determining unit is configured to obtain an execution sequence according to an execution dependency relationship between the target grouping combinations, wherein the dependency relationship includes serial execution or parallel execution.
[0170] The scheduling unit is configured to, in the execution sequence, select one or at least two target grouping combinations for execution until all the target grouping combinations are scheduled to be completed, according to whether resources for executing the target grouping combinations are sufficient.
[0171] The NL2SQL system capability improvement device provided by the embodiment of the application can execute the NL2SQL system capability improvement processing method provided by any embodiment of the application, has the function modules and beneficial effects corresponding to the execution method.
[0172] In one embodiment, FIG. 6 is a structural schematic diagram of an electronic device provided by the embodiment of the application. The electronic device 10 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the application described and / or claimed in this document.
[0173] As shown in FIG. 6, the electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0174] The plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.
[0175] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, and the like. The processor 11 performs various methods and processes described above, such as the NL2SQL system capability improvement method.
[0176] In some embodiments, the NL2SQL system capability improvement processing method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the NL2SQL system capability improvement method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the NL2SQL system capability improvement method by any other suitable means, such as by means of firmware.
[0177] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0178] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable NL2SQL system capability improvement apparatuses to produce a machine, such that the computer program, when executed, enables the machine to implement the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0179] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0180] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0181] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), blockchain network, and the Internet.
[0182] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0183] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0184] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for improving the capabilities of an NL2SQL system, characterized in that: include: Determine the capabilities required for the NL2SQL system to be improved in the target scenario from the pre-built capability indicator library; Perform capability level assessment on each capability item according to a preset scenario test question bank to obtain capability level details of the NL2SQL system to be upgraded; Determine a capability improvement module matrix corresponding to the capability item based on the capability level details; wherein each element in the capability improvement module matrix represents a capability improvement task corresponding to the capability item; the horizontal axis in the capability improvement module matrix represents the capability item of the NL2SQL system to be improved, and the vertical axis represents the capability improvement task category of the NL2SQL system to be improved; Based on the task improvement category to which each of the capability improvement tasks in the capability improvement module matrix belongs, the best target capability improvement tasks are selected for collective grouping to obtain the target group combination after grouping, and the target group combination is scheduled to perform capability perception improvement on the NL2SQL system to be improved.
2. The method according to claim 1, characterized in that After the capability awareness of the NL2SQL system to be improved is improved, the method further includes: Determine whether the NL2SQL system with enhanced capability perception meets the preset requirements of the target scenario; If satisfied, the target NL2SQL system is directly applied to the application scenario; If not, return to the step of determining the capability improvement module matrix corresponding to the capability item based on the capability level details until the NL2SQL system with improved capability perception meets the preset requirements of the target scenario.
3. The method according to claim 1, characterized in that Each capability indicator in the capability indicator library corresponds to a corresponding indicator test case; The establishment of the capability index library includes: Using a preset keyword extraction algorithm to identify keywords in a first original document related to the NL2SQL system to be improved, and using each keyword as a first candidate for the first capability item; wherein the preset keyword extraction algorithm includes: a word frequency counting method, a part-of-speech tagging method, a named entity method, a word frequency-inverse document frequency (TF-IDF) method, or a graph ranking method; Collecting a second original document related to the NL2SQL system to be improved, extracting a query statement and a corresponding structured query statement SQL from the second original document, searching for logical features contained in the query statement and the corresponding structured query statement SQL, and using the logical features as a second candidate for the second capability item; The first candidate and the second candidate form two categories, namely, a first target capability item and a second target capability item, respectively, thereby forming the capability indicator library; wherein the first target capability item includes the first capability item; and the second target capability item includes the second capability item; The construction of the indicator test case includes: For each capability item belonging to the first target capability item category, a first type of test case in the form of the first category is constructed; for each capability item belonging to the second target capability item category, a second type of test case in the form of the second category is constructed; and the first type of test case and the second type of test case are placed in the capability indicator library.
4. The method according to claim 1, wherein Determining the capability items required for the NL2SQL system to be improved in the target scenario from the pre-built capability indicator library includes: Obtain historical log files stored in the target scenario by the NL2SQL system to be upgraded; A target range of capability items required in the target scenario is filtered out from the capability indicator library according to the historical log file, wherein the target range includes capability indicators corresponding to the capability items and indicator test cases corresponding to the capability indicators.
5. The method according to claim 1, wherein The capability level assessment is performed on each capability item according to the preset scenario test question bank to obtain the capability level details of the NL2SQL system to be improved, including: Testing a first type of test case in the test cases corresponding to each of the capability items according to the preset scenario test question bank to obtain a first result, and determining a first score result corresponding to each of the first type of test cases according to the first result and a preset first expected result; Testing the second type of test cases in the test cases corresponding to each of the capability items according to the preset scenario test question bank to obtain a second result, and determining a second score result corresponding to each of the second type of test cases according to the degree of conformity between the second result and a preset second expected result; The capability level of the NL2SQL system to be improved is graded according to the first scoring result and the second scoring result to obtain capability level details; wherein the capability level details include: first level, second level, third level, fourth level, and fifth level.
6. The method according to claim 5, characterized in that The method of dividing the capability level details includes: when neither the first type test case nor the second type test case in the NL2SQL system to be improved scores, determining the capability level details as the first level; when the score of the first type test case reaches the preset first score requirement, and when neither the second type test case scores, determining the capability level details as the second level; when the score of the first type test case reaches the preset second score requirement, and when the score of the second type test case reaches the preset second score requirement, determining the capability level details as the third level; when the score of the first type test case reaches the preset third score requirement, and when the score of the second type test case reaches the preset third score requirement, determining the capability level details as the fourth level; when the score of the first type test case reaches the preset fourth score requirement, and when the score of the second type test case reaches the preset fourth score requirement, determining the capability level details as the fifth level; Among them, the first level is lower than the second level, the second level is lower than the third level, the third level is lower than the fourth level, and the fourth level is lower than the fifth level; the preset first score requirement is lower than the preset second score requirement; the preset second score requirement is lower than the preset third score requirement; and the preset third score requirement is lower than the preset fourth score requirement.
7. The method according to claim 1, characterized in that The determining of the capability improvement module matrix corresponding to the capability item based on the capability level details includes: Selecting, from the capability level details, non-compliant capability items whose capability levels do not meet the standards and compliant capability items whose capability levels meet the standards; For the unsatisfactory capability item, if the unsatisfactory capability item belongs to the first target capability item, and if the capability level detail is the first level, construct a data set corresponding to the first target capability item, and use a preset mask learning method to perform knowledge injection to construct a capability improvement module matrix; In the case where the non-compliant capability item belongs to the capability item of the first target capability item, if the capability level detail is the second level, a capability improvement module matrix is constructed by using either a supervised fine-tuning SFT method for concept alignment or an external knowledge base; In the case where the unsatisfactory competency item belongs to the competency item of the second target competency item, if the competency level detail is the third level, a competency improvement module matrix is constructed by performing formal alignment in the form of indicated learning; In the case where the unqualified capability item belongs to the second target capability item, if the capability level detail is the fourth level, a capability improvement module matrix is constructed by using any one of the methods of few-sample learning, fine-tuning, or plug-in knowledge graph; For the capability items that meet the standards, there is no need to generate a capability improvement module matrix, or to generate a capability improvement module matrix that maintains the capability.
8. The method according to claim 1, characterized in that The step of selecting the best target capability improvement tasks based on the task improvement category to which each capability improvement task in the capability improvement module matrix belongs, and grouping them together to obtain a target group combination after grouping includes: Determine whether each of the capability improvement tasks is of the same type of task improvement category; If so, the capability improvement tasks of the same task improvement category are merged for the first time to obtain a first merged result, and the best capability improvement tasks are selected from the first merged result according to preset requirements to be merged for the second time to obtain a second merged result, and the target capability improvement tasks with the best second merged result are grouped together to obtain a target group combination after grouping; If not, no merging is performed.
9. The method according to claim 1, characterized in that The scheduling of the target group combination includes: Obtaining an execution order based on the execution dependency relationship between each target group combination; wherein the dependency relationship includes: serial execution or parallel execution; In the execution sequence, one or at least two target group combinations are selected for execution according to whether the resources for executing the target group combinations are sufficient until all target group combinations are scheduled.
10. A device for improving the capabilities of an NL2SQL system, characterized in that: include: A capability item determination module is used to determine the capability items required for the NL2SQL system to be improved in the target scenario from a pre-built capability indicator library; A level detail determination module is used to evaluate the level of each capability item according to a preset scenario test question bank to obtain the capability level details of the NL2SQL system to be improved; An improvement matrix determination module is used to determine the capability improvement module matrix corresponding to the capability item based on the capability level details; wherein each element in the capability improvement module matrix represents a capability improvement task corresponding to the capability item; the horizontal axis in the capability improvement module matrix represents the capability item of the NL2SQL system to be improved, and the vertical axis represents the capability improvement task category of the NL2SQL system to be improved; The improvement module is used to select the best target capability improvement tasks based on the task improvement category to which each capability improvement task in the capability improvement module matrix belongs, group them together, obtain the target group combination after grouping, and schedule the target group combination to improve the capability perception of the NL2SQL system to be improved.
11. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the NL2SQL system capability improvement method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the NL2SQL system capability improvement method according to any one of claims 1 to 9 when executed.
Citation Information
Patent Citations
Multi-data-source NL2SQL system based on semantic rules and multi-dimensional model
CN112559550A
Method for converting natural language into SQL (Structured Query Language) statement based on deep learning
CN114880347A
SQL (Structured Query Language) generation method and device based on background knowledge enhancement, equipment and medium
CN117312372A
NL2SQL system capability improvement method and device, equipment and medium
CN118331986A
Air conditioning method and system for mobility
KR1020240030654A
Cited By
Method, device and equipment for solidifying implicit knowledge and medium
CN121434416A