Data query method and device based on text conversion, medium and equipment

By annotating and optimizing the text content, text participle with labels is generated, and query logic and objects are determined using preset query statements to transform relationships and lexicons, the problem of inefficient SQL statement query in the existing technology is solved, and efficient SQL query is realized.

CN120336508APending Publication Date: 2025-07-18PING AN PAY ELECTRONIC PAYMENT CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510504700.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing SQL statement queries need to be written manually, consume a lot of human resources, and cannot meet the complex query requirements of multi-table join and aggregation operations, and the query efficiency is inefficient.

Method used

The text content is annotated by a large language model trained based on annotation samples, and text participles with intent, conditions, tables and/or field labels are generated. The preset query statements are used to transform relationships and lexicons to determine the query logic and objects, combine them into query statements, and optimize the query through association indexes.

Benefits of technology

It improves the conversion accuracy of SQL queries, reduces the cumbersomeness and labor consumption of manual compilation, and improves query efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336508A_ABST
    Figure CN120336508A_ABST
Patent Text Reader

Abstract

The invention discloses a data query method and device based on text conversion, a medium and equipment, relates to the technical field of big data and financial science and technology, and mainly aims at solving the problem that existing SQL statement query is poor in effectiveness. Comprising the steps of obtaining to-be-queried text content; the text content is labeled on the basis of a labeling model after model training is completed, text segmented words with labels are obtained, the labeling model is obtained by training a large language model on the basis of labeling samples, and the labeling samples comprise sample data with intention labels, condition labels, table and / or field labels and function labels; determining query logic matched with the tag according to a preset query statement conversion relationship, and determining a query object matched with the text segmented word based on a preset query statement word bank; and combining the query logic with the query object to obtain a query statement, and querying a query result matched with the text content based on the query statement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of big data and fintech, and particularly to a data query method, device, medium, and equipment based on text conversion. Background Art

[0002] With the rapid development of big data, in order to quickly and accurately obtain effective data, data query capabilities are essential in different applications and databases. In particular, SQL statement queries can meet the diverse query needs of users.

[0003] Currently, existing SQL statement data queries usually require writing corresponding SQL query statements. However, compiling SQL query statements requires manual writing, consuming a large amount of human resources, and users need to understand the meaning of the statements in advance when using them, resulting in low query efficiency and inability to meet the complex query needs of multi-table joins and aggregation operations. Summary of the Invention

[0004] In view of this, the present invention provides a data query method, device, medium, and equipment based on text conversion, mainly aiming to solve the problem of poor effectiveness of existing SQL statement queries.

[0005] According to one aspect of the present invention, a data query method based on text conversion is provided, including:

[0006] Obtain the text content to be queried;

[0007] Annotate the text content based on a trained annotation model to obtain text tokens with labels. The annotation model is trained based on a large language model using annotation samples, and the annotation samples include sample data with intent labels, condition labels, table and / or field labels, and function labels;

[0008] Determine the query logic that matches the label according to a preset query statement conversion relationship, and determine the query object that matches the text token based on a preset query statement vocabulary;

[0009] Combine the query logic and the query object to obtain a query statement, and optimize the query statement based on the selected association index to perform a query based on the optimized query statement.

[0010] Further, before annotating the text content based on the trained annotation model to obtain text tokens with labels, the method further includes:

[0011] Obtain text samples and perform label annotation on the text samples to obtain the annotation samples;

[0012] Load a large language model and perform model training on the large language model based on the labeled samples;

[0013] When the model loss of the large language model is less than a preset loss value, the model training of the large language model is completed, and the model loss is determined based on the loss function of the large language model and the label features of the labeled samples.

[0014] Further, the labeling the text samples to obtain the labeled samples includes:

[0015] Output at least one text instance in the text sample and receive the label content obtained by labeling the text instance;

[0016] Perform word segmentation labeling on the text sample according to the word segmentation results, intention recognition results corresponding to the text instances, and the label content, and generate word segmentation libraries corresponding to different labels. The intention recognition result is obtained by performing intention enhancement recognition on the word segmentation results based on an adversarial network, and the adversarial network is composed of a generator and a discriminator;

[0017] When the word segmentation library passes the verification, generate the label sample based on the word segmentation library.

[0018] Further, the method further includes;

[0019] After outputting the word segmentation library, receive evaluation information for evaluating the word segmentation library, and verify the word segmentation library based on the evaluation features obtained by parsing the evaluation information; or,

[0020] Obtain the label types of the word segmentation library, and perform recognition verification on the word segmentation library according to the label types based on a trained natural language processing model.

[0021] Further, before combining the query logic with the query object to obtain a query statement and optimizing the query statement based on the selected association index for querying based on the optimized query statement, the method further includes:

[0022] Obtain historical query results, and construct an index relationship according to the query statements and query results in the historical query results. The index relationship includes the mapping relationships between different intents, conditions, and functions and different fields and tables;

[0023] The optimizing the query statement based on the selected association index for querying based on the optimized query statement includes:

[0024] Query at least one associated index from the index relationship based on the condition field corresponding to the condition tag;

[0025] Add the associated index to the query statement to obtain the optimized query statement.

[0026] Further, after optimizing the query statement based on the selected associated index and querying based on the optimized query statement, the method further includes:

[0027] Output the query result;

[0028] When a confirmation instruction for the query result is received, update the query result and the query statement to the index relationship.

[0029] Further, the method further includes:

[0030] Output a query text entry instruction to instruct the user to enter the text content;

[0031] When the text content lacks query logic keywords, output the filling instruction to instruct the user to fill in keywords for the text content.

[0032] According to another aspect of the present invention, there is provided a data query device based on text conversion, including:

[0033] An acquisition module for acquiring the text content to be queried;

[0034] A marking module for marking the text content based on a marked model that has completed model training to obtain text word segmentation with tags. The marked model is obtained by training a large language model based on marked samples, and the marked samples include sample data with intent tags, condition tags, table and / or field tags, and function tags;

[0035] A determination module for determining the query logic matching the tag according to a preset query statement conversion relationship and determining the query object matching the text word segmentation based on a preset query statement thesaurus;

[0036] A query module for combining the query logic and the query object to obtain a query statement, and optimizing the query statement based on the selected associated index for querying based on the optimized query statement.

[0037] Further, the device further includes:

[0038] A marking module for acquiring a text sample and performing label marking on the text sample to obtain the marked sample;

[0039] A loading module, configured to load a large language model and perform model training on the large language model based on the labeled samples;

[0040] A training module, configured to complete the model training of the large language model when the model loss of the large language model is less than a preset loss value, where the model loss is determined based on the loss function of the large language model and the label features of the labeled samples.

[0041] Further, the annotation module is specifically configured to output at least one text instance in the text sample and receive the label content obtained by annotating the text instance; perform word segmentation annotation on the text sample according to the word segmentation result, intent recognition result, and the label content corresponding to the text instance, and generate a word segmentation vocabulary corresponding to different labels, where the intent recognition result is obtained by performing intent enhancement recognition on the word segmentation result based on an adversarial network, and the adversarial network is composed of a generator and a discriminator; when the word segmentation vocabulary passes the verification, generate the labeled samples based on the word segmentation vocabulary.

[0042] Further, the apparatus further includes:

[0043] A verification module, configured to, after outputting the word segmentation vocabulary, receive evaluation information for evaluating the word segmentation vocabulary, and verify the word segmentation vocabulary based on evaluation features obtained by parsing the evaluation information; or, obtain the label types of the word segmentation vocabulary, and perform recognition verification on the word segmentation vocabulary according to the label types based on a trained natural language processing model.

[0044] Further, the apparatus further includes: a construction module,

[0045] The construction module is configured to obtain historical query results and construct an index relationship according to the query statements and query results in the historical query results, where the index relationship includes the mapping relationships between different intents, conditions, functions and different fields and tables;

[0046] The query module is specifically configured to query at least one associated index from the index relationship based on the conditional fields corresponding to the conditional labels; add the associated index to the query statement to obtain the optimized query statement.

[0047] Further, the apparatus further includes:

[0048] An output module, configured to output query results;

[0049] An update module, configured to update the query results and the query statements to the index relationship when receiving a confirmation instruction for the query results.

[0050] Further, the output module is further configured to output a query text entry instruction to instruct the user to enter the text content; when the text content lacks query logic keywords, the filling instruction is output to instruct the user to fill in keywords for the text content.

[0051] According to another aspect of the present invention, there is provided a storage medium storing at least one executable instruction, and the executable instruction causes a processor to perform operations corresponding to the above-described data query method based on text conversion.

[0052] According to still another aspect of the present invention, there is provided a computer device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0053] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above-described data query method based on text conversion.

[0054] By means of the above technical solutions, the technical solutions provided by the embodiments of the present invention have at least the following advantages:

[0055] The present invention provides a data query method, device, medium, and device based on text conversion. Compared with the prior art, the embodiments of the present invention obtain the text content to be queried; annotate the text content based on a labeled model that has completed model training to obtain text word segments with labels, and the labeled model is obtained by training a large language model based on labeled samples, and the labeled samples include sample data with intent labels, condition labels, table and / or field labels, and function labels; determine the query logic that matches the label according to a preset query statement conversion relationship, and determine the query object that matches the text word segment based on a preset query statement thesaurus; combine the query logic and the query object to obtain a query statement, and optimize the query statement based on the selected association index to perform a query based on the optimized query statement, which greatly improves the conversion accuracy of SQL queries based on text content, reduces the tediousness of manual compilation and human consumption, and thus improves the query efficiency of SQL content.

[0056] The above description is only an overview of the technical solutions of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. In order to make the above and other objects, features, and advantages of the present invention more obvious and understandable, the following specific embodiments of the present invention are specifically described. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the following detailed description of the preferred embodiments. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. Also, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0058] Figure 1 The flowchart of a data query method based on text conversion provided by an embodiment of the present invention is shown;

[0059] Figure 2 The block diagram of a data query device based on text conversion provided by an embodiment of the present invention is shown;

[0060] Figure 3 The schematic structural diagram of a computer device provided by an embodiment of the present invention is shown. Detailed implementation manners

[0061] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0062] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of the present invention and the drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such objects can be interchanged under appropriate circumstances so that the embodiments of the embodiments of the present invention can be implemented in an order other than the following illustration or description. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, or products.

[0063] In one embodiment, the embodiments of the present invention provide a data query method based on text conversion. Taking the application of this method to a computer device such as a server as an example, where the server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, such as financial trading platforms, financial information management systems, etc.

[0064] An embodiment of the present invention provides a data query method based on text conversion, as Figure 1 shown, the method includes:

[0065] 101. Obtain the text content to be queried.

[0066] In the embodiment of the present application, the text content to be queried is input by the user through the front-end interface, and the front-end interface can be the front-end interface of a system supporting different query services. The query services include but are not limited to banking services, insurance services, etc., and the embodiment of the present application does not make specific limitations. Among them, the text content is text content applicable to different languages, including but not limited to Chinese, English, German, etc., so that the user can input the text content for querying SQL table data through the front-end interface. For example, the user enters "Query the number of flight ticket orders of Zhang San in the past month" in the system front-end interface. The embodiment of the present application does not make specific limitations.

[0067] It should be noted that the current execution end, as the execution end for data query, can be a client device, such as a mobile phone, a tablet, etc., or a cloud server, a cloud server, etc. The embodiment of the present application does not make specific limitations.

[0068] 102. Label the text content based on the labeled model that has completed model training to obtain text segmentation with labels.

[0069] In the embodiment of the present application, after the current execution end obtains the text content input by the user, it uses the labeled model that has completed model training to perform labeling processing on the text content to obtain text segmentation with labels. Among them, the labeled model is trained based on labeled samples for a large language model, that is, the large language model learns from the labeled samples with labels. At this time, the labeled samples include sample data with intent labels, condition labels, table and / or field labels, and function labels. The large language model is an artificial intelligence processing algorithm with the ability to simulate human language processing and generation. It is widely used in the field of natural language processing (NLP), such as text generation, machine translation, sentiment analysis, question answering systems, etc. By training with data from a large amount of text, it can understand, generate, and reason about natural language. In the embodiment of the present application, the LLM large language model (Large Language Model) can be selected, or the DeepSeek large language model can also be used, and no specific limitations are made.

[0070] It should be noted that in the embodiments of the present application, in order to automatically match the SQL objects to be queried from the text content, the labeled samples adopted include sample data with intent tags, condition tags, table and / or field tags, and function tags, so that the large language model can learn from the numerous sample data how to label the intent, query conditions, tables or fields, functions, etc. in the text content, so as to perform SQL object matching.

[0071] 103. Determine the query logic matching the label according to the preset query statement conversion relationship, and determine the query object matching the text tokenization based on the preset query statement vocabulary.

[0072] In the embodiments of the present application, after determining the label corresponding to the text content, the current execution end retrieves the preset query statement conversion relationship and matches the corresponding query logic based on the labeled label. At this time, the query logic is the logic for retrieving SQL content, including but not limited to logics such as AND, OR, and NOT, which are not specifically limited in the embodiments of the present application. Among them, the preset query statement conversion relationship is the mapping relationship between different labels and different query logics configured in advance. At the same time, the preset query statement vocabulary is the mapping relationship between different text tokenizations and different query objects generated in advance. At this time, the query object can be used to represent the specific scripting language that makes up the SQL query statement. For example, conditions represent the condition description in the SQL query statement, which is not specifically limited in the embodiments of the present application.

[0073] It should be noted that in the embodiments of the present application, the mapping relationships in the preset query statement conversion relationship and the preset query statement vocabulary can be entered in advance or obtained through learning based on artificial intelligence algorithms, which are not specifically limited in the embodiments of the present application.

[0074] 104. Combine the query logic and the query object to obtain a query statement, and optimize the query statement based on the selected association index, so as to perform a query based on the optimized query statement.

[0075] In the embodiments of the present application, after matching the query logic and the query object, the two are combined to obtain a query statement for performing an SQL query, that is, an SQL query statement that can be recognized by a computer. Furthermore, an optimized query is performed based on this query statement to obtain a query result. Among them, the association index is used to represent the query relationship for optimizing the query statement, and can include the index relationship after classifying fields such as conditions and intents, so that the query statement performs a query based on the index relationship, improving the query efficiency.

[0076] It should be noted that in the embodiments of the present application, when querying based on the optimized SQL query statement, the query can be performed from the SQL database, and the query result obtained is output to the front-end interface to achieve the purpose of realizing SQL query after transformation based on the text.

[0077] In another embodiment of the present invention, for further illustration and limitation, before the step of annotating the text content based on the annotated model that has completed model training to obtain text word segments with labels, the method further includes:

[0078] Obtain a text sample, and perform label annotation on the text sample to obtain the annotated sample;

[0079] Load a large language model, and perform model training on the large language model based on the annotated sample;

[0080] When the model loss of the large language model is less than a preset loss value, the model training of the large language model is completed.

[0081] In order to perform multi-dimensional annotation on the text content to construct an accurate conversion relationship between the text and the SQL statement, the current execution end pre-trains the large language model. First, load the text sample. At this time, the text sample is the content composed of text words, including but not limited to Chinese text, English text, etc., and then perform label annotation on the text sample to obtain the annotated sample. Second, since the large language model includes but not limited to natural language processing models such as DeepSeek, it can be loaded through a cloud server or an application server to obtain the large language model to be trained, and use the annotated sample to perform model training on the large language model. In order to make the large language model more adaptable to the annotation scenario between the text content and the query statement, during the training process, for each training iteration round, it is necessary to calculate the model loss of the large language model. At this time, the model loss is determined based on the loss function of the large language model and the label features of the annotated sample, so as to compare with the preset loss value until the model loss is less than the preset loss value, and the model training of the large language model is completed. Among them, the label feature is the feature for classifying the label, including intent feature, condition feature, table and / or field feature, and function feature. Therefore, calculate the model loss in intent annotation learning, the model loss in condition annotation learning, the model loss in table and / or field annotation learning, and the model loss in function annotation learning respectively, so as to compare with the corresponding preset loss values respectively. Among them, the loss function for calculating each model loss is not specifically limited, and may include but not limited to BPR loss function, weighted loss function, exponential loss function, etc., which are not specifically limited in the embodiments of the present application.

[0082] In some embodiments, the text meanings corresponding to the fields in the labeled samples include: text: natural language query text; label: a dictionary containing labeled information; intent: query intent, such as query, insert, update, delete, etc.; conditions: query conditions, each condition containing a natural language description and the corresponding SQL condition; tables: table names involved in the query, each table containing a natural language description and the corresponding table name; fields: field names involved in the query, each field containing a natural language description and the corresponding SQL field; aggregation: aggregation functions involved in the query.

[0083] In some embodiments, after labeling the text sample "Help me query the order quantities of air tickets, hotels, and taxi rides for Zhang San and Li Si in the most recent 1 month", the obtained labeled sample includes: intent: query; condition 1: Zhang San and Li Si -> customer_name IN('Zhang San','Li Si'); condition 2: the most recent 1 month -> order_date >= DATE_SUB(CURDATE(), INTERVAL 1 MONTH); table name 1: air ticket order -> t_order_air_ticket; table name 2: hotel order -> t_order_hotel_order; table name 3: taxi ride order -> t_order_car_order; field name 1: air ticket order quantity -> COUNT(*) AS air_ticket_count; field name 2: hotel order quantity -> COUNT(*) AS hotel_order_count; field name 3: taxi ride order quantity -> COUNT(*) AS car_order_count; aggregation function: COUNT(*).

[0084] In some embodiments, the labeling method in the embodiments of the present application is executed by dynamic labeling. The specific steps include:

[0085] 1. Input natural language query text: For example, "Query the product with the highest sales in 2024";

[0086] 2. The natural language processing steps include:

[0087] ① Word segmentation: Decompose the text into "query", "2024", "sales", "highest", "product";

[0088] ② Syntactic analysis: Extract the sentence structure and identify the subject, predicate, and object;

[0089] ③ Semantic understanding: Extract key information, such as, time "2024", field "sales", product "product";

[0090] 3. Intent Recognition: Recognize that the query intent is "query".

[0091] 4. Condition Extraction: Extract the conditions "2024" and "highest sales".

[0092] 5. Table and Field Matching:

[0093] ① Map "product" to the "product" table in the database.

[0094] ② Map "sales amount" to the "sales" field.

[0095] 6. Aggregation Function Recognition: Recognize "highest" as an aggregation operation, e.g., MAX.

[0096] 7. Generate labeled data by annotating the above word segmentation results:

[0097]

[0098]

[0099] 8. Compose an SQL statement according to the labeled data, for example:

[0100] SELECT product.name, MAX(sales) AS max_sales

[0101] FROM product

[0102] WHERE year = '2024'

[0103] GROUP BY product.name;

[0104] 9. Finally, perform a query according to the above SQL statement.

[0105] In another embodiment of the present invention, for further illustration and limitation, the steps of label annotating the text sample to obtain the labeled sample include:

[0106] Output at least one text instance in the text sample and receive the label content obtained by annotating the text instance.

[0107] Perform word segmentation annotation on the text sample according to the word segmentation result, intent recognition result corresponding to the text instance, and the label content, and generate a word segmentation dictionary corresponding to different labels.

[0108] When the word segmentation dictionary passes the verification, generate the labeled sample based on the word segmentation dictionary.

[0109] In order to meet the query requirements of different query users and improve the query efficiency of converting text content into query statements, when performing label annotation, at least one text instance of the text sample is first output to the query front-end interface, which can be screened randomly from the text sample or in a polling manner. The embodiments of the present application do not make specific limitations. After outputting the text instance, the user can perform detailed annotation on the text instance based on the query requirements, including but not limited to labels such as conditions and intentions, so that the current execution end can learn based on artificial intelligence from the text instance with label content. Furthermore, the current execution end uses the word segmentation result, intention recognition result, and label content of the text instance to perform word segmentation annotation on the text sample to generate a word segmentation library corresponding to different labels. Among them, the word segmentation result is obtained by performing grammatical divisions such as subject, predicate, and adverbial on the text instance based on natural language processing technology, and the intention recognition result is obtained by enhancing the intention recognition of the word segmentation result based on an adversarial network. The adversarial network is composed of a generator and a discriminator, that is, the intention of the text entity is recognized through the adversarial network that has completed intention learning training to obtain the intention recognition result. Furthermore, the text sample is annotated with word segmentation using the word segmentation result, intention recognition result, and label content. At this time, it is to train the annotation model using the word segmentation result, intention recognition result, and label content, and use the trained annotation model to perform word segmentation annotation on the text sample to generate a word segmentation library corresponding to different labels. Among them, the annotation model can be a convolutional neural network or a deep learning model, and the embodiments of the present application do not make specific limitations.

[0110] In some embodiments, during the intent recognition process, a semantic enhancement module based on a knowledge graph is introduced. By constructing a knowledge graph of the database (including information such as tables, fields, relationships, etc.), the model can more deeply understand the semantic intent of the user's query. At this time, for the annotation process, high-quality annotation data can be generated in real time according to the user's natural language annotation input through a dynamic annotation method at the stage of generating annotation data. At this time, syntax analysis, semantic analysis, and database context information can be combined to automatically identify query intent, conditions, table relationships, and fields, thereby reducing the workload of manual annotation and improving the accuracy of the annotation data. The algorithm process of dynamic annotation is as follows: In the natural language processing module, the input natural language text is tokenized, syntactically analyzed, and semantically understood. That is, after tokenizing the input text through a pre-trained tokenization model (such as BERT, RoBERTa), a syntax analysis tool (such as Stanford Parser or dependency syntax analyzer) is used to analyze the sentence structure, extract components such as the subject, predicate, and object. Finally, semantic analysis is performed in combination with a pre-trained language model (such as BERT) to extract key information (such as time, location, product name, etc.). In addition, during intent recognition, learning is carried out through an adversarial network, enabling the tokenization model to deeply understand the language characteristics of different domains, accurately segment words, and improve the quality of the generated SQL statements. Specifically, first, an adversarial network composed of a generator and a discriminator is constructed. The generator tokenizes the input natural language query, and the discriminator determines whether the tokenization result conforms to the language pattern of the corresponding business domain. Through training with a large number of natural language queries in different domains and their corresponding SQL statement data, the generator continuously adjusts the tokenization strategy to make it difficult for the discriminator to distinguish, and the discriminator continuously improves its discrimination ability. For example, in the financial field, when the user inputs "Query the average deposit interest rates of each bank in the past year", the initial generator may mis-segment "deposit interest rate". The discriminator recognizes that the tokenization does not conform to the language habits of the financial field and feeds it back to the generator. The generator then combines the weight priorities of the graph database and adjusts again to accurately segment the words, thereby generating a correct intent recognition result.

[0111] It should be noted that to ensure the effective use of the thesaurus, the thesaurus for tokenization can also be verified. When the thesaurus for tokenization passes the verification, the label samples are generated based on the thesaurus for tokenization. At this time, the verification method can be manual verification or thesaurus comparison verification, and the embodiments of the present application do not make specific limitations.

[0112] In some embodiments, when condition labels are marked, query conditions are first extracted from the text content of natural language, such as time range, numerical range, etc., and then marked. The form of the marked labels is not specifically limited. Among them, when querying conditions, the keyword matching method can also be used, that is, through predefined keywords, such as "greater than", "less than", "equal to", to determine the extracted conditions. It is also possible to based on the dependency analysis method, that is, through the syntactic analysis results, extract the subject and object as conditions. Finally, multiple conditions can also be combined to obtain the final label that needs to be marked. For example, the sales amount is greater than 1000 and less than 5000, which is not specifically limited in the embodiments of the present application.

[0113] In some embodiments, for the marking of tables or fields, the tables and fields in the text content can be mapped to specific table names and field names in the database. Specifically, it can be determined by means of rule-based matching, that is, through predefined mapping rules, such as mapping "product" to the "product" table for matching. It is also possible to perform matching based on machine learning, that is, automatic matching of tables and fields based on the trained BERT model, which is not specifically limited in the embodiments of the present invention.

[0114] In some embodiments, the identification and marking of functions can be determined by identifying specific function operations, such as SUM, COUNT, AVG, etc. Specifically, it can be determined by the keyword matching method, that is, after identifying aggregation operations through predefined keywords, such as "total", "average", "quantity", etc., function marking is performed. It is also possible to based on semantic analysis, that is, combined with the semantic understanding results, to judge whether an aggregation operation is required, which is not specifically limited in the embodiments of the present application.

[0115] In another embodiment of the present invention, for further illustration and limitation, the steps further include;

[0116] After outputting the word segmentation thesaurus, receive the evaluation information for evaluating the word segmentation thesaurus, and verify the word segmentation thesaurus based on the evaluation features obtained by parsing the evaluation information; or,

[0117] Obtain the label type of the word segmentation thesaurus, and identify and verify the word segmentation thesaurus based on the trained natural language processing model according to the label type.

[0118] To ensure the effective use of the word segmentation dictionary and improve the matching accuracy between the text and the query statement, in a specific implementation scenario during the current execution verification, after the word segmentation dictionary is first output, the user can make a judgment based on the word segmentation dictionary and give corresponding evaluation information. After the current execution end receives the evaluation information on the evaluation of this word segmentation dictionary, it verifies the word segmentation dictionary based on the evaluation features obtained by parsing this evaluation information. At this time, the evaluation information includes multiple evaluation features, including but not limited to accurate word segmentation, effective word segmentation, usable word segmentation, etc., so that the current execution end can determine that this word segmentation dictionary passes the verification. In another specific implementation scenario, the current execution end can also obtain the label type of the word segmentation dictionary and perform recognition verification on the word segmentation dictionary according to the label type based on the trained natural language processing model. At this time, the label type includes conditional type, table type, etc., and the natural language processing model can be BERT, etc., so as to identify whether the word segmentation in the word segmentation dictionary is correct. The embodiments of the present application do not make specific limitations.

[0119] In another embodiment of the present invention, for further illustration and limitation, before combining the query logic with the query object to obtain a query statement and optimizing the query statement based on the selected association index for querying based on the optimized query statement, the method further includes:

[0120] Obtain historical query results and construct an index relationship according to the query statements and query results in the historical query results.

[0121] To improve the query efficiency based on the query statement, the current execution end pre - establishes an index relationship. Specifically, first obtain historical query results and construct an index relationship according to the query statements and corresponding query results in this historical query result. At this time, the index relationship includes the mapping relationships between different intents, conditions, functions and different fields and tables. Among them, when constructing the index relationship, directly bind the intent, condition, function in the query statement to the field or table in the query result. The embodiments of the present application do not make specific limitations.

[0122] Correspondingly, the step of optimizing the query statement based on the selected association index for querying based on the optimized query statement includes:

[0123] Query at least one association index from the index relationship based on the conditional field corresponding to the condition label;

[0124] Add the association index to the query statement to obtain the optimized query statement.

[0125] In order to optimize the initially constructed query statement to ensure the query efficiency of the query statement, when the current execution end performs optimization based on the associated index, it first determines the associated index from the index relationship based on the conditional fields corresponding to the conditional tags marked by word segmentation. That is, the associated index is used to represent the query relationship for optimizing the query statement, including the index relationship between the fields or tables in the query results, conditions, intents, or functions in historical query statements, so that the query statement can be queried based on the associated index to improve the query efficiency. After determining the associated index, the current execution end adds the associated index to the query statement to obtain the optimized query statement. For example, it identifies the conditional fields of the query SQL, that is, which specific field of which table is queried. Through all the conditional fields, it determines which conditional fields have indexes. According to the conditional fields of the query SQL, it selects the associated indexes with relatively strong relevance. At this time, there can be multiple associated indexes, and each associated index is composed of one or more conditional fields. After selecting the indexes with relatively strong relevance, they are added to the initially composed query SQL, such as USEINDEX(idx_order_date). After adding the indexes to the query SQL, the optimized query SQL can be obtained.

[0126] In another embodiment of the present invention, for further illustration and limitation, after the step of optimizing the query statement based on the selected associated index and querying based on the optimized query statement, the method further includes:

[0127] Output the query result;

[0128] When receiving the confirmation instruction for the query result, update the query result and the query statement to the index relationship.

[0129] In order to meet the visualization requirements of different users for the query results, after the current execution end queries according to the query statement, it outputs the query results for the users to view. At the same time, the current execution end provides a confirmation trigger button for the users to confirm the validity of the query results after viewing. After the current execution end receives the confirmation instruction for the query results, it updates the query results and the query statement as historical query statements to the index relationship to improve the effectiveness of using the index relationship to select associated indexes to optimize the query statement.

[0130] In some instances, a graph database is used to store SQL data. Specifically, metadata is stored in a graphical manner, with tables, fields, primary keys, foreign keys, etc. as nodes, and their relationships, such as belonging, association, constraints, etc. as edges. For example, an "order" table node is connected to the "order number", "customer ID", etc. field nodes it contains through an "includes" edge, and the "customer ID" field node is connected to the primary key field node of the "customer" table through a "foreign key association" edge, which can visually display complex relationships and facilitate complex queries. Among them, there are weight sizes and shape features between each node. If it is the main order table, it is marked as a circular red node; if it is the order detail table, it is marked as a triangular yellow node; if it is a sub-table of the detail table, it is marked as a square gray node. The weights of red, yellow, and gray decrease in sequence, indicating that the search results can be queried based on the weight priority by word segmentation, thereby improving the query speed.

[0131] In another embodiment of the present invention, for further illustration and limitation, the steps further include:

[0132] Output a query text input instruction to instruct the user to input the text content;

[0133] When the text content lacks query logic keywords, output the filling instruction to instruct the user to fill in keywords for the text content.

[0134] To meet the flexible needs of users for querying SQL content from text, the current execution end outputs a query text input instruction in the front-end interface to instruct the user to input text content. At this time, the text input instruction can be configured based on the user's commonly used language. For example, if the user's commonly used language is Chinese, the output text input instruction is "Please input the query Chinese text" in Chinese. The embodiments of the present application do not make specific limitations. To avoid the situation of unrecognizable text content, improve the effectiveness of annotating the text content, and thus improve the accuracy of determining the query statement, the current execution end determines whether the text content lacks query logic keywords. Among them, query logic keywords are used to represent words that compose the effectiveness of the query statement, including but not limited to units, numbers, locations, etc., and can be screened through a preset keyword library to determine whether there are missing query logic keywords. The embodiments of the present application do not make specific limitations. When there are missing query logic keywords, output a filling instruction so that the user can re-enter the text content and fill in the keywords. For example, if the text content is "23 accommodation location" and lacks a time unit, an instruction to fill in the unit is output, and the user fills it in again as "Location of accommodation on the 23rd".

[0135] An embodiment of the present invention provides a data query method based on text conversion. Compared with the prior art, the embodiment of the present invention obtains the text content to be queried; annotates the text content based on an annotated model that has completed model training to obtain text word segments with labels. The annotated model is obtained by training a large language model based on annotated samples, and the annotated samples include sample data with intent labels, condition labels, table and / or field labels, and function labels; determines the query logic that matches the label according to a preset query statement conversion relationship, and determines the query object that matches the text word segment based on a preset query statement thesaurus; combines the query logic with the query object to obtain a query statement, and optimizes the query statement based on the selected association index, so as to perform a query based on the optimized query statement, greatly improving the conversion accuracy of SQL queries based on text content, reducing the tediousness and manpower consumption of manual compilation, and thus improving the query efficiency of SQL content.

[0136] Further, as an implementation of the method described above Figure 1 An embodiment of the present invention provides a data query device based on text conversion, as Figure 3 shown. The device includes:

[0137] An acquisition module 21, configured to acquire the text content to be queried;

[0138] An annotation module 22, configured to annotate the text content based on an annotated model that has completed model training to obtain text word segments with labels. The annotated model is obtained by training a large language model based on annotated samples, and the annotated samples include sample data with intent labels, condition labels, table and / or field labels, and function labels;

[0139] A determination module 23, configured to determine the query logic that matches the label according to a preset query statement conversion relationship, and determine the query object that matches the text word segment based on a preset query statement thesaurus;

[0140] A query module 24, configured to combine the query logic with the query object to obtain a query statement, and optimize the query statement based on the selected association index, so as to perform a query based on the optimized query statement.

[0141] Further, the device further includes:

[0142] An annotation module, configured to acquire a text sample and perform label annotation on the text sample to obtain the annotated sample;

[0143] A loading module, configured to load a large language model and perform model training on the large language model based on the annotated sample;

[0144] A training module, configured to complete the model training of the large language model when the model loss of the large language model is less than a preset loss value, where the model loss is determined based on the loss function of the large language model and the label features of the labeled samples.

[0145] Further, the annotation module is specifically configured to output at least one text instance in the text sample, and receive the label content obtained by annotating the text instance; perform word segmentation annotation on the text sample according to the word segmentation result, intention recognition result, and the label content corresponding to the text instance, and generate a word segmentation dictionary corresponding to different labels. The intention recognition result is obtained by enhancing the intention recognition of the word segmentation result based on an adversarial network, and the adversarial network is composed of a generator and a discriminator; when the word segmentation dictionary passes the verification, generate the label sample based on the word segmentation dictionary.

[0146] Further, the device further includes:

[0147] A verification module, configured to, after outputting the word segmentation dictionary, receive the evaluation information for evaluating the word segmentation dictionary, and verify the word segmentation dictionary based on the evaluation features obtained by parsing the evaluation information; or, obtain the label type of the word segmentation dictionary, and perform recognition verification on the word segmentation dictionary according to the label type based on a trained natural language processing model.

[0148] Further, the device further includes: a construction module

[0149] The construction module is configured to obtain historical query results, and construct an index relationship according to the query statements and query results in the historical query results. The index relationship includes the mapping relationships between different intents, conditions, and functions and different fields and tables;

[0150] The query module is specifically configured to query at least one associated index from the index relationship based on the conditional field corresponding to the conditional label; add the associated index to the query statement to obtain the optimized query statement.

[0151] Further, the device further includes:

[0152] An output module, configured to output query results;

[0153] An update module, configured to update the query result and the query statement to the index relationship when receiving a confirmation instruction for the query result.

[0154] Further, the output module is further configured to output a query text input instruction to instruct the user to input the text content; when the text content lacks query logic keywords, the filling instruction is output to instruct the user to fill in keywords for the text content.

[0155] An embodiment of the present invention provides a data query device based on text conversion. Compared with the prior art, in the embodiment of the present invention, the text content to be queried is obtained; the text content is annotated based on an annotated model that has completed model training to obtain text word segmentation with labels. The annotated model is obtained by training a large language model based on annotated samples, and the annotated samples include sample data with intent labels, condition labels, table and / or field labels, and function labels; the query logic matching the label is determined according to a preset query statement conversion relationship, and a query object matching the text word segmentation is determined based on a preset query statement thesaurus; the query logic and the query object are combined to obtain a query statement, and the query statement is optimized based on the selected association index, so as to perform a query based on the optimized query statement, greatly improving the conversion accuracy of SQL queries based on text content, reducing the tediousness and manpower consumption of manual compilation, and thus improving the query efficiency of SQL content.

[0156] According to an embodiment of the present invention, a storage medium is provided. The storage medium stores at least one executable instruction, and the computer executable instruction can execute the data query method based on text conversion in any of the above method embodiments.

[0157] Figure 3 FIG. shows a schematic structural diagram of a computer device according to an embodiment of the present invention. The specific implementation of the computer device in the specific embodiment of the present invention is not limited.

[0158] As Figure 3 shown, the computer device may include: a processor 302, a communication interface 304, a memory 306, and a communication bus 308.

[0159] Among them: the processor 302, the communication interface 304, and the memory 306 complete mutual communication through the communication bus 308.

[0160] The communication interface 304 is used to communicate with network elements of other devices such as clients or other servers.

[0161] The processor 302 is configured to execute the program 310, and specifically may execute the relevant steps in the above-mentioned data query method embodiment based on text conversion.

[0162] Specifically, the program 310 may include program code, which includes computer operation instructions.

[0163] The processor 302 may be a central processing unit (CPU), or a specific application integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the computer device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0164] The memory 306 is used to store the program 310. The memory 306 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0165] The program 310 is specifically configured to cause the processor 302 to perform the following operations:

[0166] Obtain the text content to be queried;

[0167] Annotate the text content based on an annotation model that has completed model training, to obtain text word segmentation with labels. The annotation model is obtained by training a large language model based on annotation samples, and the annotation samples include sample data with intent labels, condition labels, table and / or field labels, and function labels;

[0168] Determine the query logic that matches the label according to a preset query statement conversion relationship, and determine the query object that matches the text word segmentation based on a preset query statement thesaurus;

[0169] Combine the query logic and the query object to obtain a query statement, and optimize the query statement based on the selected association index, so as to perform a query based on the optimized query statement.

[0170] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be centralized on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0171] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data query method based on text conversion, characterized in that, Including: Obtain the text content to be queried; Annotate the text content based on the annotated model that has completed model training, and obtain text word segments with labels. The annotated model is obtained by training a large language model based on annotation samples, and the annotation samples include sample data with intent labels, condition labels, table and / or field labels, and function labels; Determine the query logic that matches the label according to the preset query statement conversion relationship, and determine the query object that matches the text word segment based on the preset query statement thesaurus; Combine the query logic and the query object to obtain a query statement, and optimize the query statement based on the selected association index, so as to perform a query based on the optimized query statement.

2. The method according to claim 1, wherein Before the step of annotating the text content based on the annotated model that has completed model training to obtain text word segments with labels, the method further includes: Obtain text samples, and perform label annotation on the text samples to obtain the annotation samples; Load a large language model, and perform model training on the large language model based on the annotation samples; When the model loss of the large language model is less than the preset loss value, complete the model training of the large language model. The model loss is determined based on the loss function of the large language model and the label features of the annotation samples.

3. The method according to claim 2, wherein The step of performing label annotation on the text samples to obtain the annotation samples includes: Output at least one text instance in the text sample, and receive the label content obtained by annotating the text instance; Perform word segment annotation on the text sample according to the word segment result, intent recognition result, and label content corresponding to the text instance, and generate a word segment thesaurus corresponding to different labels. The intent recognition result is obtained by enhancing the intent recognition of the word segment result based on an adversarial network, and the adversarial network is composed of a generator and a discriminator; When the word segment thesaurus passes the verification, generate the label sample based on the word segment thesaurus.

4. The method according to claim 3, wherein The method further includes; After outputting the word segment thesaurus, receive the evaluation information for evaluating the word segment thesaurus, and verify the word segment thesaurus based on the evaluation features obtained by parsing the evaluation information; Or, Obtain the label type of the word segment thesaurus, and perform recognition verification on the word segment thesaurus according to the label type based on the trained natural language processing model.

5. The method according to claim 1, characterized in that, Before the step of combining the query logic and the query object to obtain a query statement, and optimizing the query statement based on the selected association index to perform a query based on the optimized query statement, the method further includes: Obtain historical query results, and construct an index relationship according to the query statements and query results in the historical query results. The index relationship includes the mapping relationship between different intents, conditions, and functions and different fields and tables; The step of optimizing the query statement based on the selected association index to perform a query based on the optimized query statement includes: Query at least one association index from the index relationship based on the condition field corresponding to the condition label; Add the associated index to the query statement to obtain the optimized query statement.

6. The method according to claim 5, wherein After optimizing the query statement based on the selected associated index and performing a query based on the optimized query statement, the method further includes: Output the query result; When a confirmation instruction for the query result is received, update the query result and the query statement to the index relationship.

7. The method according to any one of claims 1 to 6, characterized in that The method further includes: Output a query text entry instruction to instruct the user to enter the text content; When the text content lacks query logic keywords, output the filling instruction to instruct the user to fill in keywords for the text content.

8. A data query device based on text conversion, characterized in that, Includes: An acquisition module for acquiring the text content to be queried; A labeling module for labeling the text content based on a labeled model that has completed model training, obtaining text segmentations with labels. The labeled model is obtained by training a large language model based on labeled samples, and the labeled samples include sample data with intent labels, condition labels, table and / or field labels, and function labels; A determination module for determining the query logic that matches the label according to a preset query statement conversion relationship and determining the query object that matches the text segmentation based on a preset query statement thesaurus; A query module for combining the query logic and the query object to obtain a query statement, and optimizing the query statement based on the selected associated index to perform a query based on the optimized query statement.

9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to claim 1.

Citation Information

Cited By

  • OLED material information retrieval method, system and equipment

    CN121009111A

  • Oled material information retrieval method, system and device

    CN121009111B