TextureSql method and device based on large model and storage medium

By building text and SQL knowledge bases, data knowledge is enhanced for query conditions, and SQL tasks are generated using the LLM model, the problem of low SQL generation accuracy in the existing technology is solved, and more efficient and accurate SQL query statement generation is achieved.

CN120011548AInactive Publication Date: 2025-05-16SHENZHEN SMARTCITY TECH DEV GRP CO LTD

Patent Information

Application Number
CN202510495506.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The SQL generated by the existing construction propt method during application is low in the execution of query tasks, especially when building a query structure with complex conditions, the fields and data information extraction representations need to be invested a lot of time.

Method used

By building a text knowledge base and SQL knowledge base, we enhance the context information represented by the query problem, use the vector knowledge base to enhance the data knowledge of text query conditions and SQL query conditions, and input the enhanced results into the LLM inference model to generate SQL tasks.

Benefits of technology

The technical effect of improving the accuracy of generating SQL query statements is achieved. By enhancing the context information and data knowledge of query problems, the generated SQL tasks are more accurate, reducing the complexity and time cost of building propts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011548A_ABST
    Figure CN120011548A_ABST
Patent Text Reader

Abstract

The invention discloses a Texturesql method and device based on a large model and a storage medium. A query problem is obtained, and a text query condition and an SQL query condition are generated based on the query problem; data knowledge enhancement is conducted on the text query condition and the SQL query condition through vector knowledge bases, data knowledge enhancement results are obtained, and the vector knowledge bases comprise a text vector knowledge base and an SQL vector knowledge base; inputting the data knowledge enhancement result into an LLM inference model to generate an SQL task; and executing the SQL task. According to the method, the context information represented by the query problem is enhanced by building the text knowledge base and the SQL knowledge base, the problem data represented by the query problem is accurately positioned, the more accurate SQL task is generated through reasoning of the SQL large model, and the technical effect of improving the accuracy of generating the SQL query statement is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data, and in particular to a TextureSql method, device and storage medium based on a large model. Background Art

[0002] When natural language understanding generates SQL with various sentence types, the nested queries involved will be more complicated, which can easily lead to low accuracy of generated SQL. Therefore, it often takes a lot of training resources and time to do separate SFT training to build prompts.

[0003] However, currently, the construction of prompts is relatively complex and the accuracy is not high. Especially when it comes to constructing thought chains, the construction samples are also complex. In particular, when constructing query structures with complex conditions, it takes a lot of time to represent the fields and data information to be extracted. In addition, when setting up individual SFT training, the examples are relatively simple, which may lead to differences in query scenarios, resulting in the problem that the prompt accuracy is not high enough.

[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0005] The main purpose of this application is to provide a Texturesql method, device and storage medium based on a large model, aiming to solve the technical problem that the SQL generated by the existing prompt construction method when applied has low accuracy in executing query tasks.

[0006] To achieve the above purpose, the present application proposes a Texturesql method based on a large model, wherein the method comprises: Obtaining a query question, and generating a text query condition and an SQL query condition based on the query question; Performing data knowledge enhancement on the text query condition and the SQL query condition respectively with a vector knowledge base, and obtaining a data knowledge enhancement result, wherein the vector knowledge base includes a text vector knowledge base and an SQL vector knowledge base; Input the data knowledge enhancement result into the LLM reasoning model to generate SQL task; Execute the SQL task.

[0007] In one embodiment, the step of obtaining a query question and generating a text query condition and an SQL query condition based on the query question includes: Performing vectorized question retrieval on the query question to generate the text condition; And input the query file into the LLM big model to generate the SQL condition.

[0008] In one embodiment, the step of performing data knowledge enhancement on the text query condition and the SQL query condition respectively using a vector knowledge base and obtaining a data knowledge enhancement result includes: Converting the text query condition and the SQL query condition into corresponding query vectors, and performing similarity search in the vector knowledge base using the query vectors; Receive a result set returned by the vector knowledge base, and determine the data knowledge enhancement result according to the result set.

[0009] In one embodiment, the step of receiving a result set returned by the vector knowledge base and determining the data knowledge enhancement result according to the result set includes: Performing fusion and rearrangement processing on different query results in the result set; The data knowledge enhancement result is determined according to the processing result.

[0010] In one embodiment, before the step of fusing and rearranging different query results in the result set, the method further includes: Confirm whether the result set meets the pre-set data knowledge enhancement conditions; When it is determined that the result set meets the evaluation criteria, a step of fusing and rearranging different query results in the result set is performed.

[0011] In one embodiment, the step of inputting the data knowledge enhancement result into the LLM reasoning model to generate the SQL task includes: Extracting organization words that meet the organization information requirements of the prompt word engineering from the data knowledge enhancement results; The organization words are input into the LLM big model to generate the SQL task.

[0012] In one embodiment, the method further comprises: Acquire stock indication document material, and organize knowledge in the stock indication document material, wherein the knowledge includes text knowledge and SQL knowledge; The vector knowledge base is constructed according to the sorting results.

[0013] In one embodiment, after the step of constructing the vector knowledge base according to the sorting results, the method further includes: Performing organizational knowledge splitting on the sorting result, and converting the split organizational knowledge into a query vector; A new RAG model is created, and the data RAG model is trained with the data vector.

[0014] In addition, to achieve the above-mentioned objectives, the present application also proposes a large model-based Texturesql device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the large model-based Texturesql method as described above.

[0015] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the Texturesql method based on the large model as described above are implemented.

[0016] One or more technical solutions proposed in this application have at least the following technical effects: The technical solution of the present application obtains a query question, and generates text query conditions and SQL query conditions based on the query question; performs data knowledge enhancement on the text query condition and the SQL query condition respectively with a vector knowledge base, and obtains a data knowledge enhancement result, wherein the vector knowledge base includes a text vector knowledge base and an SQL vector knowledge base; inputs the data knowledge enhancement result into an LLM reasoning model to generate an SQL task; and executes the SQL task. The technical solution of the present application, by building a text knowledge base and an SQL knowledge base, enhances the context information represented by the query question, accurately locates the problem data represented by the query question, and then generates a more accurate SQL task through the reasoning of the SQL large model, achieving the technical effect of improving the accuracy of generating SQL query statements. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 1 A flowchart diagram of the first embodiment of the Texturesql method based on a large model of this application is provided; Figure 2 A flowchart diagram of the second embodiment of the Texturesql method based on a large model of this application is provided; Figure 3A flowchart diagram of the third embodiment of the Texturesql method based on a large model of this application is provided; Figure 4 Enhance your schematics for data knowledge; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the Texturesql method based on the large model in the embodiment of the present application.

[0020] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0022] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0023] The main solution of the embodiment of the present application is: obtain a query problem, and generate text query conditions and SQL query conditions based on the query problem; use a vector knowledge base to perform data knowledge enhancement on the text query conditions and the SQL query conditions respectively, and obtain data knowledge enhancement results, the vector knowledge base includes a text vector knowledge base and an SQL vector knowledge base; input the data knowledge enhancement results into the LLM reasoning model to generate an SQL task; and execute the SQL task.

[0024] The existing prompt construction is relatively complex and has low accuracy. Especially when it comes to constructing thought chains, the construction samples are also complex. Especially when constructing query structures with complex conditions, it takes a lot of time to represent the fields and data information to be extracted. In addition, when setting up a single SFT training, the examples are relatively simple, which may lead to differences in query scenarios, resulting in the problem of insufficient prompt accuracy.

[0025] The present application provides a solution, which enhances the context information represented by the query problem by building a text knowledge base and an SQL knowledge base, accurately locates the problem data represented by the query problem, and then generates more accurate SQL tasks through reasoning of the SQL big model, thereby achieving the technical effect of improving the accuracy of generated SQL query statements.

[0026] Based on this, the embodiment of the present application provides a Texturesql method based on a large model, referring to Figure 1 , Figure 1This is a flowchart of the first embodiment of the Texturesql method based on a large model of this application. The Texturesql method based on a large model includes steps S10 to S40: Step S10, obtaining a query question, and generating a text query condition and an SQL query condition based on the query question; Step S20, performing data knowledge enhancement on the text query condition and the SQL query condition respectively using a vector knowledge base, and obtaining a data knowledge enhancement result, wherein the vector knowledge base includes a text vector knowledge base and an SQL vector knowledge base; Step S30, inputting the data knowledge enhancement result into the LLM reasoning model to generate an SQL task; Step S40, executing the SQL task.

[0027] In this embodiment, the user inputs a query question through a set input box. After the query question is presented in the input box, the query content is parsed in the form of natural language text, so as to generate text query conditions and SQL query conditions according to the parsing results. Among them, when the query question is used to characterize the user's query needs, the user can be prompted to enter a more directional natural language text by setting a text format in the input box. For example, the text format can be indicated in the input box, and the text format includes but is not limited to related instructions such as the query result, query condition, query time, etc. of the query question, so that the user can enter the correct natural language text according to the text format. For example, the input natural language text can be a text content such as "Please query the top ten products in terms of sales volume among maternal and infant products this year".

[0028] After parsing the query content, the text query conditions and SQL query conditions obtained are actually used to indicate the text of the specific query information of the query content, and the SQL statement for executing the query task from the database for the specific query information of the query content.

[0029] That is, the step of obtaining a query question and generating a text query condition and an SQL query condition based on the query question includes: Performing vectorized question retrieval on the query question to generate the text query condition; And input the query file into the LLM big model to generate the SQL query condition.

[0030] In practical applications, the text query condition is obtained by extracting vector information from the query question, that is, extracting text vectors from the query content. Since the extracted vector information may be multiple, the text query condition generated based on the vector information is also multiple. The process of extracting vector information based on the query content can be as follows: 1) Select a data model or vector extraction method that conforms to the query language rules. In practical applications, available data models include word embedding models such as Word2Vec, GloVe, and FastText, as well as Transformer-based models such as BERT and GPT. The data model can convert text into a vector representation in a high-dimensional space by learning the semantic and grammatical information of the text.

[0031] 2) Converting the query question into text data and performing data preprocessing on the text data, wherein the data preprocessing includes removing stop words, punctuation marks and special characters, and performing operations such as word segmentation and stem extraction.

[0032] 3) Based on the training of the selected data model, in this embodiment, the data model selected for extracting the vector is a pre-trained data model that can be used normally.

[0033] 4) According to the selected data model, vector extraction is performed on the text data. The vector extraction process generally includes the following stages: converting the text data converted from the query question into a format that can be processed by the data model, such as segmenting the text into words or sub-word units; inputting the encoded text data into the data model, and the data model generates a corresponding vector representation based on the knowledge it has learned; extracting a text vector from the output of the data model, and the text vector is represented as an array of fixed dimensionality, and each element represents a dimension of the vector in space.

[0034] 5) According to the generation rule of the text query condition, the extracted text vector is normalized, cropped or padded, so that the text vector conforms to the generation rule of the text query condition. The specific processing process is related to the generation rule.

[0035] 6) Generate the text query condition through the text vector, and the text query condition may be one or more.

[0036] Specifically, the actual application of extracting vectors based on the query question and converting them into query conditions can also be characterized as generating similar query conditions based on the original query question by replacing or modifying keywords, time ranges, product categories or sorting requirements. For example, if the query question raised by the user is: Please find the top ten products in terms of sales volume among maternal and infant products this year, the similar query conditions generated based on the query question can be as follows: “Query the top ten maternal and infant products with the highest sales volume this year”; “Find out the top ten baby products sold this year”; “List the top ten baby products with the highest sales this year”; "Search for the top ten maternity and baby products sold this year"; "Showing the top ten maternity and baby products with the highest sales this year".

[0037] Specifically, the number of query conditions to be generated may be preset to avoid generating multiple repeated query conditions with the same query results. Furthermore, according to the representation method of the query condition, the query condition is defined as a text query condition in the query condition.

[0038] The SQL query condition is generated based on the query question by inputting the query question into the LLM big model, which is a big language model. The data object input into the LLM big model can be text data converted from the query question, or the query question can be converted into a data type that can be recognized by the LLM big model, which is specifically related to the content that can be recognized by the data structure of the LLM big model itself.

[0039] Alternatively, the text data converted from the query question may be used as input data for the LLM large model; alternatively, the text vector extracted from the text data converted from the query question may be used as input data for the LLM large model; as shown above, the data input to the LLM large model are obtained after data processing through the query question, and, through the input data input to the LLM large model, the SQL query conditions based on the query question are obtained through data processing by the LLM large model. As shown above, multiple text query conditions can be obtained based on the information represented by the query question, and thus the SQL query conditions generated based on the query question should also be one or more.

[0040] The LLM (Large Language Model) refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. In this embodiment, the LLM is a mature language data model that can output output data that meets the requirements after performing relevant data processing based on the input data. In this data processing, the multiple text conditions generated by the query problem are used as input data, and the corresponding SQL query conditions are obtained after data processing.

[0041] According to the generated text query conditions and SQL query conditions, the text query conditions and the SQL query conditions are subjected to data knowledge enhancement processing using a vector knowledge base. The data knowledge enhancement processing in this embodiment is essentially a data processing process that improves the breadth or accuracy of knowledge acquisition by the system by improving data collection, knowledge extraction or information fusion technology. Based on this data processing, that is, the steps of respectively performing data knowledge enhancement on the text query conditions and the SQL query conditions using a vector knowledge base and obtaining data knowledge enhancement results include: Converting the text query condition and the SQL query condition into corresponding query vectors, and performing similarity search in the vector knowledge base using the query vectors; Receive a result set returned by the vector knowledge base, and determine the data knowledge enhancement result according to the result set.

[0042] In this embodiment, a vector knowledge base is pre-built as a knowledge storage area. As shown above, the query conditions derived based on the query question include text query conditions and SQL query conditions. Therefore, the vector knowledge base is essentially a pre-set text vector knowledge base and SQL vector knowledge base, so as to perform data knowledge enhancement processing on the text query conditions and the SQL query conditions respectively. The vector knowledge base is a system that combines a vector database and related knowledge storage and retrieval technology, which is a data set stored in the form of data and is used to store, retrieve and analyze high-dimensional data. The stored vectors are represented by features extracted from the original data by a machine learning model.

[0043] The essence of the data knowledge enhancement processing is to improve the generalization ability of the model, that is, to simulate different user inputs by generating multiple types of queries, so as to perform data knowledge enhancement processing on a given query problem. Therefore, according to the data representation in the vector knowledge base, the query condition is converted into a query vector, and the knowledge vector similar to the query condition is searched from the vector knowledge base with the query vector as the condition. In this embodiment, the text query condition and the SQL query condition are respectively generated into corresponding query vectors, that is, the text query vector is used to search for vector data similar to the text query vector from the text vector knowledge base, and the SQL query vector is used to search for vector data similar to it from the SQL vector knowledge base; then, the result set fed back by the vector knowledge base is received. Among them, the result set is a data set composed of the vector data searched from the corresponding vector knowledge base with the text query vector and the SQL query vector as the condition. This is because when the text query vector and the SQL query vector are both one or more, multiple similarity searches are performed to obtain a large number of similarity vector data.

[0044] Since the search condition of the result set is one or more repeated or non-repeated text query conditions and SQL query conditions, considering the data quality of the result set, a data knowledge enhancement condition based on the result set can be set, that is, before the step of fusing and rearranging different query results in the result set, the following is further included: Confirm whether the result set meets the pre-set data knowledge enhancement conditions; When it is determined that the result set meets the evaluation criteria, a step of fusing and rearranging different query results in the result set is performed.

[0045] This data knowledge enhancement condition can include various forms: First, the similarity search of the vector knowledge base and the data processing of the LLM large model are used as a processing cycle for enhancing the knowledge parameters, and a preset number of times based on the processing cycle is set. That is, after executing the preset number of processing cycles, the output data of the last processing cycle is used as a result set for generating data knowledge enhancement results.

[0046] Secondly, the evaluation criteria of the result set are pre-set, and the accuracy of the output data of the LLM large model is evaluated according to the evaluation criteria. According to the evaluation results, it is determined whether the output data can be used as a result set of data knowledge enhancement results.

[0047] The data knowledge enhancement conditions described above can be used individually or in combination, and the specific usage can be determined based on the relevant information indicated by the query question.

[0048] As shown above, according to the data processing process of the vector knowledge base and the LLM large model, the result set is obtained, and the enhanced knowledge parameters are processed to obtain the data knowledge enhancement result, that is, the step of receiving the result set returned by the vector knowledge base and determining the data knowledge enhancement result according to the result set includes: Performing fusion and rearrangement processing on different query results in the result set; The data knowledge enhancement result is determined according to the processing result.

[0049] In the processing process based on the result set, first, the result set is preprocessed to remove duplicate items in the result set and standardize the data format. In the preprocessed result set, the corresponding query results obtained by different query conditions in the result set are determined, and the determined query results are used as fusion objects, and they are fused using the preset ReciprocalRank Fusion fusion strategy to reduce the deviation of the data set and thus improve the accuracy of the result set.

[0050] In this embodiment, the Reciprocal Rank Fusion fusion strategy is defined as a reverse sorting fusion strategy, which is used to combine multiple result sets with different relevance indicators into a single result set. The principle can be represented in the form of an algorithm, through which the multiple query results in the result set are calculated to form a corresponding search score, and a reciprocal ranking score is assigned according to the ranking position in the result set of the search score, and a fusion ranking is formed according to the reciprocal ranking score. According to the fusion result, the multiple query results represented in the fusion result are rearranged, and the rearranged result is used as the data knowledge enhancement result, which is characterized by more valuable and in-depth knowledge extracted from the query conditions after data processing. The data knowledge enhancement processing process shown above can be viewed in detail. Figure 4 , Figure 4 Enhance your diagrams for data knowledge.

[0051] According to the above, the data knowledge enhancement result obtained after processing the vector knowledge base and the LLM large model is processed by prompt organization, that is, the step of inputting the data knowledge enhancement result into the LLM reasoning model to generate the SQL task includes: Extracting organization words that meet the organization information requirements of the prompt word engineering from the data knowledge enhancement results; The organization words are input into the LLM big model to generate the SQL task.

[0052] In this embodiment, the process of prompt organization based on the data knowledge enhancement result is essentially the processing of prompt word engineering organization information based on the data knowledge enhancement result, that is, the prompt organization is also defined as the processing of prompt word engineering organization information. First, the definition of the prompt organization is clarified. The prompt organization is a prompt word organization, which is a language for communicating with the AI ​​model to indicate the characteristics of the content that the AI ​​model wants to generate. It is usually composed of multiple words, phrases or short sentences to clarify the specific requirements for generating content. To this end, the prompt organization is essentially a language organization process based on the data model, that is, the data knowledge enhancement result is processed through the prompt organization, and the processing result is obtained. The processing result is used as the input data, and the SQL task is generated after the data is processed by the LLM large model.

[0053] As shown above, the processing result is used as input data. Before the SQL task is generated after the data processing of the LLM large model, the processing result needs to be subjected to the LLM reasoning process. The essence of this process is to form the processing result into content features that can be recognized and applied by the LLM model. That is, the processing result of the prompt word engineering organization information processing of the data knowledge enhancement result is subjected to LLM reasoning, and it is converted into a data form applicable to the LLM large model, and input into the LLM large model, so as to generate the SQL task through the data structure processing of the LLM large model. The SQL task is defined as a structured query language, which is used to manage and operate data in the database. The execution of the SQL task may include but is not limited to data query, data update, data deletion, data insertion, and creation and maintenance of database structure, and the SQL task may have one or more independent or combined SQL statements. Specifically, the query language contained in the SQL task is related to the data information to be represented by the natural language text.

[0054] According to the generated SQL task, the SQL task is executed, and the execution result of the SQL task is displayed to the user. In practical applications, when executing the SQL task, a connection relationship with a corresponding database server is established in advance, that is, connected to a data source, so as to send and execute the SQL statement of the SQL task.

[0055] When the database server receives the SQL task, it first performs lexical analysis on the received SQL statement, decomposes the SQL statement into tokens (keywords, identifiers, operators, etc.), and classifies and parses these tokens to generate corresponding data structures. After that, the SQL server performs syntax analysis on the decomposed tokens, checks the correctness of the statement according to the SQL syntax rules, and generates a syntax tree. If there is a syntax error in the statement, the server returns an error message; according to the result of the syntax analysis, the database server traverses the syntax tree and performs semantic analysis. This step mainly determines the information of the table and column in the statement, including the table name, column name, column type, etc., and checks the semantic correctness of the statement. If there is a semantic error in the statement, such as referencing a non-existent table or column, the server will also return an error message. In addition, according to specific needs, the database server may also be provided with an optimizer, which is a SQL execution process component that can optimize the SQL statement, including selecting the optimal execution plan, index selection, connection method selection, etc., so as to select an efficient execution method to minimize query time and resource consumption.

[0056] Afterwards, based on the processing results of the optimizer, the execution plan generator generates a specific execution plan for the SQL statement. The execution plan describes how to access data, how to use indexes, how to connect, how to sort, etc.; the execution plan is sent to the corresponding database engine for processing. The database engine performs specific operations such as data scanning, index search, sorting, grouping, etc. according to the execution plan. These operations may involve underlying hardware resources such as disk I / O and memory access; finally, the database engine returns the execution result. Among them, for query statements, the returned result may be a result set containing all rows and columns that meet the query conditions. For update, insert, or delete statements, the returned result may be the number of rows affected by the operation or status information on whether the operation is successful.

[0057] It should be noted that different types of SQL statements (such as query statements, update statements, insert statements, etc.) included in the SQL task may be different during execution. This difference can be derived based on specific implementation details and optimization strategies set by the database management system (DBMS), and the details will not be repeated here.

[0058] The overall solution indicated by the above content can be characterized by the following calculation formula: Nsql=f((f1(t),f2(s)); Wherein, Nsql is the SQL statement parsed from natural language, t represents the original input natural language text, s represents the SQL statement converted from the original input text, f1(t) represents the data enhancement of the original input text, f2(s) represents the data enhancement of the initial SQL, and f((f1(t), f2(s)) represents the SQL generation method based on the reasoning of the enhanced text and enhanced SQL. By building an enhanced natural language parsing system, the target SQL statement can be generated more efficiently and accurately to obtain the target data. In this embodiment, by building a text knowledge base and an SQL knowledge base, the context information represented by the query question is enhanced, the problem data represented by the query question is accurately located, and then a more accurate SQL task is generated through reasoning of the SQL big model, thereby achieving the technical effect of improving the accuracy of generating SQL query statements.

[0059] Reference Figure 2 , Figure 2 This is a flow chart of the second embodiment of the Texturesql method based on a large model of this application. The Texturesql method based on a large model includes steps S50 to S80: Step S50: obtaining stock indication document material, and organizing knowledge in the stock indication document material, wherein the knowledge includes text knowledge and SQL knowledge; Step S60: construct the vector knowledge base according to the sorting results; Step S70, performing organizational knowledge splitting on the sorting result, and converting the split organizational knowledge into a query vector; Step S80: create a new RAG model, and train the data RAG model with the data vector.

[0060] In this embodiment, in order to realize data knowledge enhancement processing of query questions, a vector knowledge base based on the data knowledge enhancement processing application is set, wherein the text vector knowledge base and SQL vector knowledge base included in the vector knowledge base have corresponding construction methods respectively.

[0061] The currently stored stock knowledge document materials are obtained, the stock knowledge document materials are sorted, and document materials for text knowledge and SQL knowledge are obtained according to the sorting results. And according to the data characteristics represented by the text knowledge and SQL knowledge, a text vector knowledge base and an SQL vector knowledge base are respectively constructed.

[0062] In the process of constructing the vector knowledge base, the knowledge is further subdivided and organized to achieve a better knowledge base construction. In particular, when constructing the SQL vector knowledge base, the DDL statements of the database, the descriptive metadata of the database's own data, the database-related document descriptions, the reference sample SQL, etc. are sorted out from the stock knowledge document materials that record the SQL knowledge, and the SQL knowledge registered in the document is split and vectorized to train a RAG model suitable for database query.

[0063] That is, after the step of constructing a vector knowledge base according to the sorting results, wherein the vector knowledge base includes a text vector knowledge base and a SQL vector knowledge base, the step further includes: Splitting the organizational knowledge possessed by the sorting result, and converting the split organizational knowledge into a query vector; A new RAG model is created, and the data RAG model is trained with the data vector.

[0064] In this embodiment, based on the premise of training a RAG model suitable for database query, a new RAG model (RAG, i.e., Retrieval-Augmented Generation model, defined as a retrieval enhanced generation model) is created. The RAG model is mainly used to enhance the performance of large language models (LLMs) in specific tasks, especially tasks that require access to external knowledge bases or real-time information. It can assist the model in generating more accurate, detailed and targeted answers through an integrated retrieval mechanism, thereby overcoming the problems of limited storage capacity of LLMs, difficulty in instantly obtaining the latest information, and insufficient knowledge in specific fields.

[0065] The process of creating the RAG model can be specifically performed as follows: 1) Knowledge document preparation: Collect the stock indicator document materials related to the task, which come from various external knowledge sources, such as databases, websites, document libraries, etc. Use a special document loader or multimodal model (such as OCR technology) to convert these knowledge sources into plain text data to form the stock indicator document materials. Based on the collected plain text data converted based on the knowledge source, pre-process the plain text data, such as word segmentation, stop word removal, stem extraction, etc., to improve the directionality of the text content and thus improve the retrieval efficiency when performing retrieval.

[0066] 2) Embedding model training: select or train an embedding model to convert the text in the stock indicator document material into a vector representation, that is, convert the knowledge text collected in the stock indicator document material into a vector unit. Commonly used embedding models include Word2Vec, BERT, GPT, etc. The vector representation can capture the semantic information of the document, making similar documents closer in the vector space.

[0067] 3) Vector database construction: vectors generated from the knowledge text of the stock indicating document materials are stored in the vector database. Through the knowledge vectors stored in the vector database, users can be provided with fast retrieval of query operations. The vector database optimizes the efficiency of processing and storing large-scale vector data, and can quickly retrieve the most relevant information to the user's query.

[0068] 4) Query retrieval: After the user enters a query, the system converts it into a vector representation and retrieves knowledge text or historical conversation records that are semantically similar to the user's query vector in the vector database.

[0069] 5) Information integration and generation: Integrate the retrieved documents with the user query to form a rich context. Use large language models (such as the GPT series, BERT series, etc.) as generators to use this context to generate the final answer or text output.

[0070] As shown above, the created RAG model is trained through the knowledge content contained in the existing indication document material. Through the training process of the RAG model, the context information of the knowledge content in the existing indication document material is improved. The context information includes database structure, table description, column description, sample query, question SQL pair, business document, historical query question, etc., that is, the training process of the RAG model can also be defined as the process of arranging and improving the knowledge text in the existing indication document material. It can be understood that the training process not only improves the data structure of the RAG model, but also realizes the knowledge improvement of the vector database.

[0071] According to the training process of the RAG model, the relevant information provided in the stock indication document material is cleaned, processed, vectorized, and stored in the vector knowledge base. The vector knowledge base indicated here is a text vector knowledge base.

[0072] The above shows the process of creating a text vector knowledge base based on the text knowledge possessed by the stock indication document material. When creating the SQL vector knowledge base based on the SQL knowledge recorded in the stock indication document material, the SQL knowledge of the SQL vector knowledge base is also improved through the training of the SQL knowledge of the RAG model. Specifically, an important factor that can improve the accuracy of system output is the quality of training data. Therefore, the training data of the RAG model that can be provided in the stock indication document material is SQL knowledge such as known correct problem SQL pairs, detailed DDL statements and rich business documents. The RAG model is trained through the SQL knowledge to improve the context information of the SQL knowledge stored in the SQL knowledge database.

[0073] That is, in the training process, different training effects can be achieved based on the different knowledge contents in the stock indication document material. The knowledge contents are the text knowledge for generating a text vector knowledge base and the SQL knowledge for generating an SQL knowledge base. The knowledge contents are obtained by extracting the contents recorded in the stock indication document material. The information that can be represented by the knowledge contents can be shown as follows: Database DDL statement: The DDL statement can be used to obtain information such as data tables, data columns, and data types in the database. The DDL statement is indicated as: ddl="CREATE TABLE table1 (id INT, name TEXT, age INT, descTEXT)".

[0074] Document information: includes any documents about database, business or industry, which can help LLM understand the context of user problems. The document information is represented by documentation="Our work is operating the platform to defines xxx as yyy to do sth".

[0075] SQL statements: Multiple commonly used SQL query statements provide data assistance for the LLM large model to understand the SQL query mode, so as to understand the context information of the query problem. The SQL query statement can be indicated as: sql="SELECT id, name, desc FROM table1", etc.

[0076] Question-SQL pair: The question-SQL pair contains a large amount of mapping information, which can be used to understand the context information of the question and to help the LLM model to improve the understanding of ambiguous and vague questions.

[0077] Through the training of the RAG model with the above SQL knowledge, the context information of the SQL knowledge in the SQL vector knowledge base is improved, and the trained knowledge vectors are stored in the SQL vector knowledge base as the initial storage objects, thus completing the creation of the SQL vector knowledge base.

[0078] In this embodiment, when constructing the vector knowledge base, the relevant knowledge content in the stock instruction document materials is disassembled and differentiated to form the organizational knowledge of the vector knowledge base. By constructing the RAG model and training the RAG model with the organizational information, further network knowledge retrieval can be performed on the organizational knowledge, so as to improve the context information of the organizational knowledge in the vector knowledge base, thereby improving the accuracy of the data knowledge enhancement result.

[0079] In addition, reference can also be made to Figure 3 , Figure 3 which is a schematic flowchart of the third embodiment of the Texturesql method based on the large model of the present application. The Texturesql method based on the large model includes the following content: In this embodiment, during the process of inferring and generating SQL tasks based on the LLM large model, it can be understood as the parsing between text and SQL. An appropriate prompt P is used to guide a large language model M to obtain the probability distribution on the SQL query Y, and SQL query tokens are generated one by one. The generation process of the SQL query Y can be expressed as follows:

[0080] In the above formula, the Y<i is the prefix of the SQL query Y, and the PM(Yi|·) is the conditional probability of the i-th token in the SQL query Y given the prefix Y<i, the prompt P, the mode S, and the question Q. By providing a richer background corresponding to the question and SQL query knowledge, the accuracy of SQL generation can be improved.

[0081] The implementation steps of text-to-SQL parsing are as follows: 1) User input. The user submits a query request in the form of natural language.

[0082] 2) Preliminary retrieval and recognition. The user's text input is parsed, and the NLP large model engine analyzes the user's query, identifies its intention and key information, and retrieves relevant knowledge information in the text vector knowledge base.

[0083] 3) Preliminary SQL generation. Using the large model, the user input is identified to generate a preliminary SQL statement. Retrieve relevant knowledge information in the SQL vector knowledge base.

[0084] 4) Enhanced retrieval: Based on the content and SQL identified in the above two steps, query question data is enhanced, and related data in the database is retrieved, integrated, and rearranged.

[0085] 5) Prompt word engineering: Based on the retrieval knowledge information integrated above, prompt word engineering is used to organize information.

[0086] 6) SQL generation: Generate corresponding SQL statements based on the prompt word engineering information through the large model reasoning capability.

[0087] 7) Execution and feedback: the generated SQL statements are executed on the database and the results are fed back to the user.

[0088] In this implementation, when a user asks a question, the RAG model of the multi-dimensional knowledge base retrieves relevant information, generates a prompt, and hands it over to the LLM to generate an SQL query. The generated SQL query will be executed in the database and the result will be returned to the user. The technical effect of improving the accuracy of SQL queries can be achieved without maintaining the prompt.

[0089] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the Texturesql method based on a large model of the present application. More forms of simple transformations based on this technical concept are all within the scope of protection of the present application.

[0090] The present application provides a Texturesql device based on a large model, and the Texturesql device based on a large model includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the Texturesql method based on the large model in the above-mentioned embodiment one.

[0091] Reference below Figure 5 , which shows a schematic diagram of the structure of a Texturesql device based on a large model suitable for implementing the embodiment of the present application. The Texturesql device based on a large model in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The large model-based Texturesql device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present application.

[0092] like Figure 5 As shown, the Texturesql device based on the large model may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. Various programs and data required for the operation of the Texturesql device based on the large model are also stored in RAM1004. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the Texturesql device based on the large model to communicate wirelessly or wired with other devices to exchange data. Although the diagram shows a Texturesql device based on the large model with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.

[0093] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0094] The large-model-based Texturesql device provided by the present application adopts the large-model-based Texturesql method in the above-mentioned embodiment, which can solve the technical problems of the existing complex prompt construction and low accuracy. Compared with the prior art, the beneficial effects of the large-model-based Texturesql device provided by the present application are the same as the beneficial effects of the large-model-based Texturesql method provided in the above-mentioned embodiment, and the other technical features in the large-model-based Texturesql device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0095] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0096] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0097] The present application provides a storage medium, which is a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the large model-based Texturesql method in the above-mentioned embodiment.

[0098] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.

[0099] The computer-readable storage medium may be included in a Texturesql device based on a large model; or may exist independently without being assembled into a Texturesql device based on a large model.

[0100] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the Texturesql device based on the large model, the Texturesql device based on the large model implements the technical content of the Texturesql method embodiment based on the large model as shown above.

[0101] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0102] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0103] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0104] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned Texturesql method based on a large model, and can solve the technical problem that the existing prompt construction is relatively complex and has low accuracy. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the Texturesql method based on a large model provided in the above-mentioned embodiment, and will not be repeated here.

Claims

1. A TextureSql method based on a large model, characterized in that: The method comprises the following steps: Obtaining a query question, and generating a text query condition and an SQL query condition based on the query question; Performing data knowledge enhancement on the text query condition and the SQL query condition respectively with a vector knowledge base, and obtaining a data knowledge enhancement result, wherein the vector knowledge base includes a text vector knowledge base and an SQL vector knowledge base; Input the data knowledge enhancement result into the LLM reasoning model to generate SQL task; Execute the SQL task.

2. The large model-based TextureSql method according to claim 1, characterized in that: The step of obtaining a query question and generating a text query condition and an SQL query condition based on the query question includes: Performing vectorized question retrieval on the query question to generate the text condition; And input the query file into the LLM big model to generate the SQL condition.

3. The large model-based TextureSql method according to claim 1, characterized in that: The step of performing data knowledge enhancement on the text query condition and the SQL query condition respectively using the vector knowledge base and obtaining the data knowledge enhancement result comprises: Converting the text query condition and the SQL query condition into corresponding query vectors, and performing similarity search in the vector knowledge base using the query vectors; Receive a result set returned by the vector knowledge base, and determine the data knowledge enhancement result according to the result set.

4. The large model-based TextureSql method according to claim 3, characterized in that: The step of receiving the result set returned by the vector knowledge base and determining the data knowledge enhancement result according to the result set comprises: Performing fusion and rearrangement processing on different query results in the result set; The data knowledge enhancement result is determined according to the processing result.

5. The large model-based TextureSql method according to claim 4, characterized in that: Before the step of fusing and rearranging different query results in the result set, the method further includes: Confirm whether the result set meets the pre-set data knowledge enhancement conditions; When it is determined that the result set meets the evaluation criteria, a step of fusing and rearranging different query results in the result set is performed.

6. The large model-based TextureSql method according to claim 1, characterized in that: The step of inputting the data knowledge enhancement result into the LLM reasoning model to generate the SQL task includes: Extracting organization words that meet the organization information requirements of the prompt word engineering from the data knowledge enhancement results; The organization words are input into the LLM big model to generate the SQL task.

7. The large model-based TextureSql method according to any one of claims 1 to 5, characterized in that: The method further comprises: Acquire stock indication document material, and organize knowledge in the stock indication document material, wherein the knowledge includes text knowledge and SQL knowledge; The vector knowledge base is constructed according to the sorting results.

8. The large model-based TextureSql method according to claim 7, characterized in that: After the step of constructing the vector knowledge base according to the sorting results, the method further includes: Performing organizational knowledge splitting on the sorting result, and converting the split organizational knowledge into a query vector; A new RAG model is created, and the data RAG model is trained with the data vector.

9. A Texturesql device based on a large model, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the large model-based TextureSql method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the TextureSql method based on a large model as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Method, system and equipment for generating SQL (Structured Query Language) statement based on large model

    CN119127913A

  • Method for automatically generating sql statement in natural Chinese language

    CN119149571A

  • Data query method and device, electronic equipment and computer program product

    CN119149579A

  • SQL (Structured Query Language) generation and error correction method based on multi-stage feedback

    CN119474129A

Cited By

  • Text-to-SQL (Structured Query Language) conversion method based on semantic modeling

    CN121350124A

  • Self-adaptive context engineering method, device, equipment, medium and product

    CN121706959A