A method and device for improving the recognition of valid data in an artificial intelligence model query
Standardized problems are generated through context recognition and rewriting processing, and the NER model and preset keyword list are used for entity encoding, which solves the misunderstanding and misjudgment problems of the artificial intelligence model when identifying user problems, and improves the accuracy of question-and-answer and the efficiency of structured data query.
Patent Information
- Application Number
- CN202411866433.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Artificial intelligence models are prone to misunderstandings, misjudgments and incomprehensions when identifying users' natural language problems, resulting in low accuracy of question-and-answer.
By obtaining the current natural language problems and historical problems entered by the user, conducting context recognition, selecting target historical problems in the same context for rewriting and processing, generating standardized problems, and using the NER model and preset keyword list for entity annotation and encoding, and generating structured query statements to improve accuracy.
It reduces the risk of misunderstanding and misjudgment in the identification of user problems and improves the accuracy of question-and-answer questions, especially in the fields of finance, research and enterprise information data query.
Smart Images

Figure CN119336875B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and device for improving the recognition and query of valid data by an artificial intelligence model. Background Art
[0002] An artificial intelligence model is generally a model trained using a large amount of vectorized data, which can generate natural language text or understand the meaning of natural language text. This enables users to interact with data in a more intuitive way and simplifies the information acquisition process. However, the natural language questions input by users contain a large number of expressions with complex business logics, strong professional attributes, and colloquialisms. These expressions can cause misunderstandings, misjudgments, and incomprehensions when the artificial intelligence model recognizes user questions, and further result in a low accuracy rate of the answers output by the artificial intelligence model during question and answer. Summary of the Invention
[0003] The present invention provides a method and device for improving the recognition and query of valid data by an artificial intelligence model, so as to solve the defects of misunderstandings, misjudgments, and incomprehensions that occur when the artificial intelligence model in the prior art recognizes user questions, and improve the accuracy rate of the answers output by the artificial intelligence model during question and answer.
[0004] The present invention provides a method for improving the recognition and query of valid data by an artificial intelligence model, including:
[0005] Obtaining a current natural language question input by a user;
[0006] When it is determined that there is a historical natural language question input by the user, performing context recognition on the current natural language question and the historical natural language question, and obtaining a target historical natural language question in the same context as the current natural language question from the historical natural language question;
[0007] Performing rewriting processing on the current natural language question based on the target historical natural language question to obtain a standardized question;
[0008] Inputting the standardized question into the artificial intelligence model to obtain an answer output by the artificial intelligence model.
[0009] According to the method for improving the recognition and query of valid data by an artificial intelligence model provided by the present invention, obtaining a target historical natural language question in the same context as the current natural language question from the historical natural language question specifically includes:
[0010] Obtain multiple historical natural language questions in the same context as the current natural language question from the historical natural language questions, obtain the time length ranking between the multiple historical natural language questions and the current natural language question in the time dimension, and select the top k historical natural language questions with the top k time length rankings as the target historical natural language questions.
[0011] According to a method for improving the recognition of valid data in an artificial intelligence model provided by the present invention, perform a rewriting process on the current natural language question based on the target historical natural language question, specifically including:
[0012] Determine the reference relationship between the target historical natural language question and the current natural language question, and perform a reference substitution and rewriting process on the current natural language question based on the reference relationship; and / or
[0013] Determine the ellipsis relationship between the target historical natural language question and the current natural language question, and perform an ellipsis completion and rewriting process on the current natural language question based on the ellipsis relationship.
[0014] According to a method for improving the recognition of valid data in an artificial intelligence model provided by the present invention, after obtaining the current natural language question input by the user, the method further includes:
[0015] In the case where the historical natural language question input by the user is not determined, perform a reference substitution and rewriting process on the current natural language question based on a preset rule to obtain a standardized question.
[0016] According to a method for improving the recognition of valid data in an artificial intelligence model provided by the present invention, after obtaining the standardized question, the method further includes:
[0017] Perform part-of-speech analysis on the standardized question to obtain a part-of-speech analysis result, and label the nouns in the standardized question based on the part-of-speech analysis result;
[0018] Input the labeled standardized question into the NER model to obtain the entity labels of the nouns output by the NER model; the NER model is trained based on the sample standardized questions with noun entity labels;
[0019] Perform element recognition processing on the standardized question based on the entity labels of the nouns to obtain an element extraction question;
[0020] Obtain a preset keyword list, perform a matching process on the element extraction question based on the preset keyword list to generate an element parameter pair of the standardized question; the element parameter pair includes an entity, an entity name, and an entity code;
[0021] Generate a coded question by replacing the entities in the standardized question with the said entity codes.
[0022] According to a method for improving the recognition of valid data in an artificial intelligence model provided by the present invention, after obtaining the current natural language question input by the user, the method further includes:
[0023] Input the current natural language question into a question classification model to obtain the question category output by the question classification model; the question classification model is trained based on sample natural language questions with question category labels.
[0024] Inputting the standardized question into the artificial intelligence model specifically includes:
[0025] Input the question category and the standardized question into the artificial intelligence model.
[0026] According to a method for improving the recognition of valid data in an artificial intelligence model provided by the present invention, after annotating the nouns in the standardized question based on the part-of-speech analysis result, the method further includes:
[0027] When the entities in the standardized question include a name entity and a company entity, obtain the name corresponding to the name entity, match the name in the enterprise name library corresponding to the company entity, and obtain the natural person code of the name entity.
[0028] When the entities in the standardized question include the name entity and do not include the company entity, obtain the name corresponding to the name entity, match the name in a preset well-known person name library, and obtain the natural person code of the name entity.
[0029] According to a method for improving the recognition of valid data in an artificial intelligence model provided by the present invention, after generating a coded question, the method further includes:
[0030] Input the coded question into a table column parameter acquisition model to obtain the table column parameters of the data table output by the table column parameter acquisition model; the table column parameter acquisition model is trained based on sample coded questions and sample table column parameters of the data table.
[0031] Generate a structured query statement based on the table column parameters of the data table and the coded question.
[0032] Input the structured query statement, the coded question, and the element parameter pair into the artificial intelligence model to obtain the answer output by the artificial intelligence model.
[0033] A method for improving the recognition and query of valid data by an artificial intelligence model provided by the present invention, wherein the table column parameter acquisition model includes a data table selection unit and a field selection unit; the data table column parameters include a data table name parameter and corresponding field parameters;
[0034] Inputting the encoded problem into the table column parameter acquisition model to obtain the data table column parameters output by the table column parameter acquisition model, specifically including:
[0035] Inputting the encoded problem into the data table selection unit to obtain the data table name parameter output by the data table selection unit based on a preset multi-dimensional matrix; the preset multi-dimensional matrix includes multiple matrices, each matrix includes multiple business domains, each business domain includes multiple data tables, and each data table corresponds to a data table name parameter and multiple field parameters;
[0036] Inputting the encoded problem and the data table name parameter into the field selection unit to obtain the data table name parameter output by the field selection unit and corresponding field parameters.
[0037] The present invention also provides a device for improving the recognition and query of valid data by an artificial intelligence model, including:
[0038] A current problem acquisition module, configured to acquire a current natural language problem input by a user;
[0039] A historical problem acquisition module, configured to, when determining a historical natural language problem input by the user, perform context recognition on the current natural language problem and the historical natural language problem, and acquire a target historical natural language problem in the same context as the current natural language problem from the historical natural language problem;
[0040] A standardized problem acquisition module, configured to perform rewriting processing on the current natural language problem based on the target historical natural language problem to obtain a standardized problem;
[0041] An answer acquisition module, configured to input the standardized problem into the artificial intelligence model to obtain an answer output by the artificial intelligence model.
[0042] The method and device for improving the recognition of valid data in an artificial intelligence model provided by the present invention obtain the current natural language question and historical natural language questions input by the user, select target historical natural language questions from the historical natural language questions according to the same context, and rewrite the current natural language question based on the target historical natural language questions to obtain a standardized question. It can use the relatively high - relevance context information provided by the target historical natural language questions to rewrite the current natural language question into a standardized question with clearer meaning, reduce the risks of misunderstanding, misjudgment, and non - understanding when the artificial intelligence model recognizes the user's question, and improve the accuracy of the answer output when the artificial intelligence model conducts question - answering. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following - described drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0044] Figure 1 It is a schematic flowchart of the method for improving the recognition of valid data in an artificial intelligence model provided by the present invention.
[0045] Figure 2 It is one of the schematic framework diagrams of the method for improving the recognition of valid data in an artificial intelligence model provided by the present invention.
[0046] Figure 3 It is the second schematic framework diagram of the method for improving the recognition of valid data in an artificial intelligence model provided by the present invention.
[0047] Figure 4 It is the third schematic framework diagram of the method for improving the recognition of valid data in an artificial intelligence model provided by the present invention.
[0048] Figure 5 It is the fourth schematic framework diagram of the method for improving the recognition of valid data in an artificial intelligence model provided by the present invention.
[0049] Figure 6 It is the schematic structural diagram of the device for improving the recognition of valid data in an artificial intelligence model provided by the present invention.
[0050] Figure 7 It is the schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To make the objectives, technical solutions and advantages of the present invention more clear, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0052] The following is based on Figures 1-6 Describe the method and device for improving the recognition and query of valid data by an artificial intelligence model of the present invention.
[0053] Figure 1 is a schematic flowchart of the method for improving the recognition and query of valid data by an artificial intelligence model provided by the present invention. As Figure 1 shown, the method includes:
[0054] Step 101, obtain the current natural language question input by the user.
[0055] The current natural language question refers to the latest specific query request put forward by the user during the interaction with a dialogue system such as an artificial intelligence model. The content of the current natural language question may include ambiguous information such as unclear reference or incomplete expression of information. For example, the current natural language question may be "What are the recently established companies nearby?" or "Which enterprises has he invested in?", etc.
[0056] Exemplarily, the problem text input by the user to the artificial intelligence model at the current moment can be obtained through an input device such as a keyboard or a touch screen.
[0057] Step 102, in the case of determining the historical natural language question input by the user, perform context recognition on the current natural language question and the historical natural language question, and obtain the target historical natural language question in the same context as the current natural language question from the historical natural language question.
[0058] The historical natural language question refers to the specific query request put forward before the current natural language question during the interaction with a dialogue system such as an artificial intelligence model.
[0059] The context refers to the surrounding environment or background of the natural language question input by the user. Context recognition refers to recognizing the context of the natural language question input by the user, such as the context of the current natural language question and the context of a certain historical natural language question, to determine whether the context of the current natural language question and the context of a certain historical natural language question are the same.
[0060] Exemplarily, in the case where the current natural language question input by the user is not the first question, it is determined that there is a historical natural language question input by the user and the historical natural language question input by the user can be obtained according to a preset method. For example, a retrieval process can be preset. When the current natural language question input by the user is received and it is determined that there is a historical natural language question, the historical natural language question is retrieved from a specified area.
[0061] After obtaining the historical natural language question input by the user, context recognition can be performed on the current natural language question and the obtained historical natural language question to select a target historical natural language question in the same context as the current natural language question from the historical natural language questions.
[0062] Step 103: Rewrite the current natural language question based on the target historical natural language question to obtain a standardized question.
[0063] A standardized question refers to a question after reducing the possible number of understandings of the question text or eliminating the ambiguity in the understanding of the question text. The possible number refers to the number of possible understanding meanings.
[0064] The target historical natural language question in the same context can provide relatively high - relevance context information related to the current natural language question. Based on the relatively high - relevance context information, the meaning that the current natural language question wants to express can be understood more clearly, improving the accuracy and coherence of question understanding. On this basis, rewriting the current natural language question can obtain a standardized question with a clearer meaning.
[0065] Exemplarily, when the current natural language question can be: "Is it raining heavily?", the current natural language question itself has multiple possible meanings. Based on this, when answering, the probability of getting a wrong answer is relatively high. When the target historical natural language question in the same context can be: "Will it rain in Haidian today?", the current natural language question can be rewritten to obtain a standardized question: "Is it raining heavily in Haidian today?" This can make the meaning of the standardized question clearer and reduce the probability of getting a wrong answer.
[0066] Step 104: Input the standardized question into the artificial intelligence model to obtain the answer output by the artificial intelligence model.
[0067] The artificial intelligence model is any trained question - answering model. In this embodiment, the type of the artificial intelligence model is not specifically limited. It can be a general - purpose artificial intelligence question - answering model or an artificial intelligence question - answering model in a specific business aspect after fine - tuning, etc.
[0068] The method for improving the recognition of valid data by an artificial intelligence model provided in the embodiments of the present invention obtains the current natural language question and historical natural language questions input by the user, selects a target historical natural language question from the historical natural language questions according to the same context, and performs rewriting processing on the current natural language question based on the target historical natural language question to obtain a standardized question. It can use the relatively high relevance context information provided by the target historical natural language question to rewrite the current natural language question into a standardized question with clearer meaning, reduce the risks of misunderstanding, misjudgment, and incomprehension when the artificial intelligence model recognizes the user's question, and improve the accuracy of the answer output by the artificial intelligence model when answering questions.
[0069] Based on the above embodiments, obtaining a target historical natural language question from the historical natural language questions that is in the same context as the current natural language question specifically includes:
[0070] Obtain multiple historical natural language questions from the historical natural language questions that are in the same context as the current natural language question, obtain the time length ranking between the multiple historical natural language questions and the current natural language question in the time dimension, and select the top k historical natural language questions with the top k time length rankings as the target historical natural language questions.
[0071] The time length refers to the time difference between two moments. For example, the time length between 12:05 and 12:10 is 300 seconds. The time length ranking refers to the time difference ranking between the corresponding moments of each historical natural language question in the same context and the corresponding moment of the current natural language question.
[0072] When selecting the top k historical natural language questions with the top k time length rankings, it can be the case of ranking in ascending order, and select the k historical natural language questions including the first one; or it can be the case of ranking in descending order, and select the k historical natural language questions including the last one. The top k historical natural language questions with the top k time length rankings are the k historical natural language questions including the historical natural language question with the smallest time difference between the corresponding moment and the corresponding moment of the current natural language question.
[0073] Among them, k in the top k time length rankings can be determined according to the actual application scenario, and the specific value of k in this embodiment is not specifically limited. For example, the specific value of k can be 9 to obtain more context information; the specific value of k can also be 5 to reduce the computing resources required in subsequent processing, etc.
[0074] In this embodiment, when rewriting the current natural language question, not all historical natural language questions in the same context are used. Instead, k target historical natural language questions are selected from them to reduce the risk of reasoning confusion caused by the excessive number of target historical natural language questions and enhance the clarity of the obtained standardized question meaning.
[0075] To obtain a standardized question that better conforms to natural semantics, based on any of the above embodiments, the current natural language question is rewritten based on the target historical natural language question, specifically including: determining the reference relationship between the target historical natural language question and the current natural language question, and performing a reference substitution and rewriting process on the current natural language question based on the reference relationship.
[0076] The reference relationship refers to that a certain word or phrase in the current natural language question refers to another word or phrase in a certain target historical natural language question. In other words, the relationship between the words in the current natural language question and another corresponding word in the target historical natural language question is the reference relationship. Therefore, identifying the reference relationship between the target historical natural language question and the current natural language question and replacing the pronouns in the current natural language question can generate a standardized question that better conforms to natural semantics.
[0077] Exemplarily, the current natural language question is "Which enterprises has he invested in?" The target historical natural language questions include "Who is the legal representative of Vision XX Company?" The current natural language question and the target historical natural language questions can be input into the pre-trained qwen-max model. After obtaining the specific reference relationship through the qwen-max model combined with Few-shot-Cot reasoning, a reference substitution and rewriting process is performed based on the specific reference relationship, and the processed standardized question is output. The standardized question can be "Which enterprises has the legal representative of Vision XX Company invested in?"
[0078] To obtain a standardized question that better conforms to natural semantics, in another embodiment, the current natural language question is rewritten based on the target historical natural language question, and specifically further includes: determining the ellipsis relationship between the target historical natural language question and the current natural language question, and performing an ellipsis completion and rewriting process on the current natural language question based on the ellipsis relationship.
[0079] The ellipsis relationship refers to the information that is omitted or not directly expressed in the current natural language question. This omitted or not directly expressed information can be obtained from the target historical natural language question to facilitate a complete understanding of the meaning expressed by the text of the current natural language question. Therefore, identifying the ellipsis relationship between the target historical natural language question and the current natural language question and filling in the omitted part in the current natural language question can generate a more natural semantic standardized question.
[0080] Exemplarily, the current natural language question is "What are there in Beijing". The target historical natural language questions include "Who is the legal representative of Vision XX Company" and "Which enterprises has he invested in", etc. The current natural language question and the target historical natural language question can be input into a pre-trained large language model. After obtaining the specific ellipsis relationship through the inference of the large language model, based on the specific ellipsis relationship, an ellipsis completion and rewriting process is performed, and the processed standardized question is output. The standardized question can be "Which enterprises in Beijing are invested in by the legal representative of Vision XX Company".
[0081] Further, in an embodiment, after obtaining the current natural language question, the current natural language question is sent to the specified area for associated storage with the historical language question according to a preset method. In this way, when determining the historical natural language question input by the user to obtain the target historical natural language question, it is convenient to use the rewritten question as the target historical natural language question, thereby reducing the number of target historical natural language questions while being able to determine the richness of the scenario information provided by the target historical natural language question.
[0082] Exemplarily, the current natural language questions are "Which ones are engaged in real estate" and "Which enterprises has he invested in". The target historical natural language questions include "Who is the legal representative of Vision XX Company", "Which enterprises has he invested in", and "What are there in Beijing", etc.
[0083] In this way, the current natural language question and the target historical natural language question can be input into the pre-trained qwen-max model. After obtaining the specific reference relationship through the inference of the qwen-max model combined with Few-shot-Cot, based on the specific reference relationship, a reference replacement and rewriting process is performed, and the processed standardized question is output. The standardized question can be "Which real estate enterprises in Beijing are invested in by the legal representative of Vision XX Company". Compared with "Which real estate enterprises has he invested in", the standardized question in this embodiment can provide richer scenario information.
[0084] Exemplarily, a prompt word can be constructed based on a pre-constructed prompt word template, reference relationship and / or ellipsis relationship, the current natural language question, and the target historical natural language question. The prompt word is input into a general artificial intelligence model to obtain the standardized question output by the general artificial intelligence model.
[0085] The following shows a prompt template:
[0086] You are a language analysis expert, good at dealing with anaphora resolution and ellipsis completion in human conversations. Infer the anaphoric relationships and ellipsis content in the user's new question from the historical records provided by the user and complete them. Do not answer the user's questions and do not add fabricated content. Only rewrite the questions;
[0087] Task description:
[0088] Given a series of historical conversation records, your task is to analyze the user's question and, if necessary, perform anaphora resolution and ellipsis completion. Analyze the historical records step by step;
[0089] It is necessary to complete the anaphoric relationships in the user's question to ensure that the new question can express more complete semantic information;
[0090] Historical conversation records: %s;
[0091] User's question: %s;
[0092] Task requirements:
[0093] The returned result must not contain explanatory information, and the most accurate inference result must be extracted from it. The data returned is preceded by # and followed by #;
[0094] Do not analyze non-interrogative sentences in the historical records. If the user's question is a non-interrogative sentence, the user's question must be returned. Otherwise, the returned result must be an interrogative sentence;
[0095] If there is no context relationship between the user's question and the historical records, no processing is required, and the user's question must be returned directly;
[0096] If conjunctions such as "at the same time" or "and" appear in the user's question, the user's question needs to be modified and completed according to the historical conversation records;
[0097] If the user's question is an interrogative sentence and contains anaphora or ellipsis, replace it with specific information to make the sentence more complete;
[0098] If the user's question is not an interrogative sentence or does not contain anaphora or ellipsis, the user's question must be returned;
[0099] Ensure that the modified question can be understood independently, even without reading the historical conversation;
[0100] The following are sample examples:
[0101] Example 1: xxxxx? Return: #xxxxx#;
[0102] Example 2: xxxxx? Return: #xxxxx#.
[0103] The %s in the above prompt template is specific relevant text in actual applications.
[0104] During the interaction with a dialogue system such as an artificial intelligence model, the questions input by users are generally short, and it is relatively difficult for the artificial intelligence model to guess the user's intention based on the questions input by the user. For example, if the question input by the user is "What are the recently established companies nearby?", in the absence of historical questions, the specific meanings of "nearby" and "recently" are difficult to determine.
[0105] To solve the above technical problems, based on any of the above embodiments, after obtaining the current natural language question input by the user, the method further includes:
[0106] In the case where the historical natural language question input by the user is not determined, based on a preset rule, perform a substitution rewriting process on the current natural language question to obtain a standardized question.
[0107] Exemplarily, the location-related information in the user's question can be replaced according to the geographical location coordinates of the user device. For example, "nearby" in the above example can be replaced with the geographical location coordinates of the user device. Some words can also be mapped and replaced with specified meanings according to the business definition. For example, "recently" can be replaced with "in the last three months". In this way, for the question text input by the user located in Haidian District, Beijing to the large language model on October 1, 2024: "What are the recently established companies nearby?", the processed standardized question is: "What are the companies established after July 1, 2024 in Haidian District, Beijing?".
[0108] Exemplarily, a prompt can be constructed based on the current natural language question and a prompt template pre-constructed according to a preset rule, and the prompt is input into a general artificial intelligence model to obtain a standardized question output by the general artificial intelligence model.
[0109] The following shows a prompt template:
[0110] Please extract all dates from the user's question, including words representing dates;
[0111] For example, related words and phrases such as "this year", "the year before last", "in the last three years", etc. must be converted into the standard date format YYYY-MM-DD according to the current time, and the date should include the first day to the last day within that date, rather than randomly taking a time point. Note that the current date is: 【2024-09-25】;
[0112] The following are examples:
[0113] Example 1: What is the production volume of civilian drones in Guangdong Province in the past two years? Return 2022-01-01,2023-12-31;
[0114] Example 2: Which enterprises were established in 2023? Return 2023-01-01,2023-12-31;
[0115] Example 3: What is the latest stock price of Company Wan XX? Return 2024-09-25;
[0116] Example 4: What is the revenue of Company Wan XX in the past three years? Return 2021,2022,2023;
[0117] Requirements:
[0118] (1) If there is no relevant information about date or time in the user's question, return "None";
[0119] (2) Only return the extracted standard dates without any other unnecessary explanations;
[0120] (3) "Recently" by default means three months before the current date;
[0121] (3) If the date in the user's question is up to today, it should include the latest date up to the present;
[0122] (4) If there is only year and month in the user's question, return it in the YYYY-MM format;
[0123] (5) If there is only year, return it in the YYYY format;
[0124] (6) If there is a quarter in the user's question, return the end month of the quarter;
[0125] The current natural language question is: Which companies were recently established in Haidian District, Beijing?
[0126] During the interaction process with a dialogue system such as an artificial intelligence model, since the company name, industry name, etc. may be relatively long or subject to changes, when the user inputs a question, they generally use its alias, abbreviation, former name, or APP name, etc. If the artificial intelligence model directly performs subsequent reasoning and query processing based on the question input by the user, it may only obtain information that matches the alias or abbreviation, etc., increasing the risk of incorrect reply answers.
[0127] To solve the above technical problems, based on any of the above embodiments, after obtaining the standardized question, the method further includes:
[0128] Perform part-of-speech analysis on the standardized question to obtain a part-of-speech analysis result, and label the nouns in the standardized question based on the part-of-speech analysis result;
[0129] Input the labeled standardized question into the NER model to obtain the entity labels of nouns output by the NER model; the NER model is trained based on sample standardized questions with noun entity labels.
[0130] Perform element recognition processing on the standardized question based on the entity labels of the nouns to obtain an element extraction question.
[0131] Obtain a preset keyword list, and perform matching processing on the element extraction question based on the preset keyword list to generate an element parameter pair for the standardized question; the element parameter pair includes an entity, an entity name, and an entity code.
[0132] Replace the entity in the standardized question with the entity code to generate a coded question.
[0133] Among them, part-of-speech analysis refers to performing word segmentation on the standardized question to obtain multiple words for analysis to determine the part-of-speech analysis result of each word. The part-of-speech analysis result refers to whether this word is a noun, a verb, an adjective, etc. Exemplarily, after obtaining the part-of-speech analysis result of each word, position tags can be added to the nouns in the standardized question to label the nouns in the standardized question, so as to determine the nouns in the standardized question through the position tags.
[0134] A sample standardized question refers to a question including keywords in the preset keyword list. According to the context of the sample standardized question, some nouns in the sample standardized question are entities, some nouns in the sample standardized question are not entities, or at least some nouns are not entities.
[0135] The following table is an analysis example of whether the noun online game is an industry entity.
[0136]
[0137] It can be understood that it is essential to include sample standardized questions with some nouns not being entities. The sample element extraction question is the question after deleting the entity from the sample standardized question.
[0138] The entity label can include NLB and YLB; NLB can indicate that the noun in the standardized question is not an entity, and YLB can indicate that the noun in the standardized question is an entity, and vice versa. In this embodiment, no specific limitation is made in this regard.
[0139] In one embodiment, the preset keyword list refers to a prefabricated list including keywords, keyword names, and keyword codes. The preset keyword list can be a list prefabricated by industry experts. Among them, the keywords are all nouns.
[0140] Exemplarily, all data in a preset keyword list can be loaded and obtained, and the obtained data can be regularized to obtain a legal regular expression. Based on the legal regular expression, the element extraction problem is matched, and the keywords, keyword names, and keyword codes that match successfully are obtained, so as to generate element parameter pairs of a standardized problem based on these keywords, keyword names, and keyword codes that match successfully. The entity, entity name, and entity code in an element parameter pair correspond to a group of keywords, keyword names, and keyword codes.
[0141] It can be understood that the nouns marked in the standardized problem include entities and non-entities. Exemplarily, such as Figure 2 As shown, entities can include enterprises, industries, product services, natural persons, regions, and time, etc. The NER model can be used to identify whether a noun in a standardized problem is an entity in the context of the standardized problem. For example, which companies in Beijing have obtained online game licenses? Among them, Beijing is related to the region in the context, and Beijing is an entity; online games are not related to the industry in the context, and online games are non-entities. On this basis, the Beijing code can be used to replace Beijing in the standardized problem, and the standardized problem obtained is: which companies with province code 11 have obtained online game licenses.
[0142] In this embodiment, the entity in the standardized problem is identified through the NER model and the preset keyword list, and the entity code is used to replace the entity in the standardized problem to generate a coded problem. Since the entity code is unique, it is convenient for the artificial intelligence model to identify the correct entity information and improve the efficiency of subsequent database queries.
[0143] The NER model can be used to determine the nouns in the standardized problem according to the position tags in the standardized problem, perform identification processing on the nouns to obtain the entity tags of the nouns, and output the entities or standardized problems carrying the position tags and entity tags.
[0144] Element identification processing means that when it is determined that a noun in the standardized problem after marking is an entity, the noun is not processed; when it is determined that a noun in the standardized problem after marking is not an entity, the noun is deleted.
[0145] Exemplarily, after obtaining the sample standardization problem, it can be imported into the doccano platform for data annotation (also known as data tagging), and the annotated JSON file can be exported. Based on the annotated JSON file, the UIE model is trained to obtain the NER model. Exemplarily, by inputting the annotated JSON file into the UIE model, the predicted positions and predicted labels of the nouns output by the UIE model in the sample standardization problem can be obtained, and the predicted positions and predicted labels in the sample standardization problem can be compared with the actual positions and actual labels in the sample standardization problem to adjust the parameters of the UIE model and obtain the trained NER model.
[0146] In one embodiment, the preset keyword list includes multiple ones, such as the enterprise keyword list, the industry keyword list, the industrial chain keyword list, and the product service keyword list, etc.
[0147] The following table is an example of a partial industry keyword list.
[0148]
[0149] The following table is an example of a partial industrial chain keyword list.
[0150]
[0151] The following table is an example of a partial product service keyword list.
[0152]
[0153] In another embodiment, the preset keyword list refers to a prefabricated list including entities and entity names; based on the preset keyword list, matching processing is performed on the element extraction problem to generate the element parameter pairs of the standardization problem, specifically including:
[0154] Based on the preset keyword list, matching processing is performed on the element extraction problem to obtain the entity names corresponding to the entities in the element extraction problem;
[0155] According to the entity name, a pull-through request is executed to obtain the entity code corresponding to the entity name;
[0156] Based on the entity, the entity name, and the entity code, the element parameter pairs of the standardization problem are generated.
[0157] Through data preprocessing, unique entity names and corresponding entity codes can be established for entities such as enterprises, industries, product services, natural persons, regions, and time in the historical database, and the entity names and corresponding entity codes are associated and stored in the specified database. Exemplarily, the entity name corresponding to Douyin is Douyin Co., Ltd.
[0158] The pull-through request refers to sending an entity name to a specified database to obtain the entity code corresponding to the entity name returned by the specified database. In this embodiment, a set of entity, entity name, and entity code forms a parameter pair.
[0159] In one embodiment, when the query result of the specified database is empty, the entity code corresponding to the entity name is newly created and stored. Exemplarily, the current maximum entity code (Max(Id)) can be stored in Redis. When creating the entity code corresponding to the entity name, the maximum entity code is obtained from Redis and incremented. For example, the new entity code is Max(Id)+1.
[0160] Compared with loading and obtaining all the data in the preset keyword list, performing regularization processing on the obtained data to obtain a legal regular expression, and then matching the element extraction problem based on the legal regular expression to obtain the successfully matched keywords, keyword names, and keyword codes, in this embodiment, the entity name corresponding to the entity in the element extraction problem is obtained through the preset keyword list, and the pull-through request is executed according to the entity name to obtain the entity code corresponding to the entity in the element extraction problem, which can reduce the components of the legal regular expression and reduce the risk of errors in matching the element extraction problem based on the legal regular expression.
[0161] As Figure 3 shown, further, in one embodiment, executing the pull-through request according to the entity name to obtain the entity code corresponding to the entity name specifically includes: synchronizing the entity name to Open Search, and performing text retrieval based on the entity name to obtain the entity code corresponding to the entity name.
[0162] In another embodiment, executing the pull-through request according to the entity name to obtain the entity code corresponding to the entity name specifically includes: synchronizing the entity name to the MySQL database, and obtaining the entity code corresponding to the entity name based on the preset pull-through function in the MySQL database.
[0163] The preset pull-through function is used to query the corresponding entity code in the MySQL database according to the entity name, and encapsulate the query result into JSON format data for return, so that the interface layer can parse the JSON format data to obtain the entity code corresponding to the entity name.
[0164] As Figure 4As shown, based on any of the above embodiments, after annotating the nouns in the standardized question based on the part-of-speech analysis result, the method further includes: when the entities in the standardized question include a name entity and a company entity, obtaining the name corresponding to the name entity, matching the name in the enterprise name library corresponding to the company entity, and obtaining the natural person code of the name entity. Exemplarily, when the standardized question involves personnel information with a company as an attributive, such as "Help me check the profile of Yu X from Wan XX Company", it can be determined that the entities in the standardized question include a name entity and a company entity. Otherwise, it is considered that the entities in the standardized question include a name entity.
[0165] In this embodiment, the name entity is different from the name. The name entity can be a nickname, alias, etc. For example, in the standardized question "What is the theme of Luo A's New Year's speech", Luo A is the name entity, but Luo A is not the name. Exemplarily, after obtaining that the name corresponding to the name entity Luo A is Luo Yi, the natural person code of Luo A can be obtained by querying in the corresponding enterprise name library according to the company entity corresponding to Luo Yi. The corresponding element parameter pair can be: Luo A, Luo Yi, 1786788XXXX.
[0166] In another embodiment, after annotating the nouns in the standardized question based on the part-of-speech analysis result, the method further includes: when the entities in the standardized question include the name entity and do not include the company entity, obtaining the name corresponding to the name entity, matching the name in the preset well-known person name library, and obtaining the natural person code of the name entity.
[0167] Exemplarily, in the standardized question "Which dynasty was Li Si from", the name Li Si corresponding to the name entity can be obtained, and the natural person code of Li Si can be obtained by matching Li Si in the preset well-known person name library. The corresponding element parameter pair can be: Li Si, Li Si, 1786782XXXX.
[0168] Based on any of the above embodiments, after obtaining the current natural language question input by the user, the method further includes:
[0169] Inputting the current natural language question into a question classification model to obtain the question category output by the question classification model; the question classification model is trained based on sample natural language questions with question category labels;
[0170] Inputting the standardized question into the artificial intelligence model specifically includes:
[0171] Inputting the question category and the standardized question into the artificial intelligence model.
[0172] A sample natural language question with a problem category label refers to a situation where the problem category label indicates the true problem category of the sample natural language question. In this way, during the training process, after obtaining the predicted problem category output by the problem classification model, the predicted problem category can be compared with the true problem category to adjust the parameters of the problem classification model and obtain a trained problem classification model.
[0173] The sample natural language questions with problem category labels can be obtained by experts' writing or by using an artificial intelligence model to fill in the sample templates written by experts.
[0174] Exemplarily, the sample templates written by experts may include sample natural language questions with slots and their problem category labels. Inputting the sample templates into a pre-trained artificial intelligence model, sample natural language questions with problem category labels output by the pre-trained artificial intelligence model can be obtained; the pre-trained artificial intelligence model can be used to obtain the corresponding information of the slots from a specified database and fill the corresponding information as slot values into the slots.
[0175] In this embodiment, by determining the problem category through the problem classification model and inputting the problem category and the standardized question into the artificial intelligence model, more information related to the standardized question can be provided, thereby improving the accuracy of the answer output by the artificial intelligence model.
[0176] In fields such as financial, research, and enterprise information data query, a large amount of data with direct value and high business value is organized and stored in a fixed format as structured data. The large language model cannot directly perform operations such as querying, retrieving, sorting, and filtering on the structured data, and it is unrealistic to fully vectorize the large-scale structured data for all scenarios, resulting in a relatively low accuracy of the answers output by the large language model during the question and answer process.
[0177] To solve the above technical problems, based on any of the above embodiments, after generating the encoded question, the method further includes:
[0178] Inputting the encoded question into a table column parameter acquisition model to obtain the table column parameters of the data table output by the table column parameter acquisition model; the table column parameter acquisition model is trained based on sample encoded questions and sample data table column parameters;
[0179] Generating a structured query statement based on the data table column parameters and the encoded question;
[0180] Inputting the structured query statement, the encoded question, and the element parameter pair into the artificial intelligence model to obtain the answer output by the artificial intelligence model.
[0181] The data table column parameters are data table information including the data table name and data table fields, etc. Among them, the data table name can include the table English name and table Chinese name, etc., and the data table fields can include the column English name, column Chinese name, constant mapping, and field type, etc.
[0182] A structured query statement refers to a query statement applicable to structured data. For example, the structured query statement can be an SQL query statement that conforms to the MySQL specification.
[0183] Exemplarily, a prompt can be generated based on the data table column parameters and the encoded problem in combination with a Prompt template, and the prompt is sent to the artificial intelligence model to obtain the generated structured query statement.
[0184] The following is a prompt template:
[0185] You are a database expert who can generate high-performance SQL statements that conform to the MySQL specification according to the user's needs;
[0186] Table information:
[0187] - Table name: ${Table English name} (${Table Chinese name});
[0188] - Fields:
[0189] - ${Column English name} (${Column Chinese name} ${Constant mapping}, type: ${Field type});
[0190] - Table name: ${Table English name} (${Table Chinese name});
[0191] - Fields:
[0192] - ${Column English name} (${Column Chinese name} ${Constant mapping}, type: ${Field type});
[0193] Table association:
[0194] - ${Table association method};
[0195] User requirements:
[0196] - ${User requirements};
[0197] Requirements:
[0198] - The output SQL statement should only include the tables related to the user requirements;
[0199] - The SQL statement needs to start and end with #;
[0200] - For all tables, add dataStatus != 3;
[0201] - No SQL analysis process is required. Directly provide the query statement;
[0202] Please write an SQL query that meets the above requirements.
[0203] The following is an example of a prompt generated according to the prompt template:
[0204] You are a database expert who can generate high-performance SQL statements that comply with MySQL specifications according to the user's requirements;
[0205] Table information:
[0206] - Table name: sy_cs_ms_base_entity (entity wide table);
[0207] - Fields:
[0208] - compCode (company code);
[0209] - compName (company name);
[0210] - creditCode (uniform social credit code);
[0211] - industryName (industry name);
[0212] - regStatusNm (registration status, in operation / not in operation);
[0213] - estiblishDate (establishment date, type: date);
[0214] - provinceName (province name);
[0215] - Table name: sy_cd_ms_cn_comp_innocom_zsyx (high-tech enterprise table);
[0216] - Fields:
[0217] - compCode (company code);
[0218] - compName (company name);
[0219] - provinceCode (province code);
[0220] Table association:
[0221] - The table sy_cs_ms_base_entity and the table sy_cd_ms_cn_comp_innocom_zsyx are associated through the compCode field;
[0222] User requirements:
[0223] - Query the list of high-tech enterprises established in 2020 with a provincial code of 11;
[0224] Requirements:
[0225] - The output SQL statement should only include the tables related to the user requirements;
[0226] - The SQL statement needs to start and end with #;
[0227] - All tables need to add dataStatus != 3;
[0228] - No SQL analysis process is required, directly provide the query statement;
[0229] Please write an SQL query that meets the above requirements.
[0230] Furthermore, after generating a structured query statement based on the data table column parameters and the encoded problem, the structured query statement can be parsed to obtain the first table column information, and the first table column information can be compared with the data table name parameters retrieved by RAG; the comparison result is fed back to the table column parameter acquisition model to fine-tune the table column parameter acquisition model. It is also possible to perform a retrieval in a preset database based on the structured query statement, and feed the retrieval result back to the table column parameter acquisition model to fine-tune the table column parameter acquisition model.
[0231] In this embodiment, by using the table column parameter acquisition model to obtain the data table column parameters corresponding to the encoded problem and generating a structured query statement based on the data table column parameters and the encoded problem, the accuracy of querying, retrieving, sorting, and filtering in structured data can be increased.
[0232] In addition, in this embodiment, by inputting the encoded problem and the element parameters into the artificial intelligence model, it is convenient for the artificial intelligence model to correctly understand the meaning of the encoded problem.
[0233] As Figure 5 shown, based on any of the above embodiments, the table column parameter acquisition model includes a data table selection unit and a field selection unit; the data table column parameters include data table name parameters and corresponding field parameters;
[0234] Input the encoded problem into the tabular parameter acquisition model to obtain the tabular parameter of the data table output by the tabular parameter acquisition model, specifically including:
[0235] Input the encoded problem into the data table selection unit to obtain the data table name parameter output by the data table selection unit based on the preset multi-dimensional matrix; the preset multi-dimensional matrix includes multiple matrices, each matrix includes multiple business domains, each business domain includes multiple data tables, and each data table corresponds to a data table name parameter and multiple field parameters;
[0236] Input the encoded problem and the data table name parameter into the field selection unit to obtain the data table name parameter and the corresponding field parameter output by the field selection unit.
[0237] Exemplarily, the business domain can be the field of patent rights and the field of copyright, etc., the matrix can be intellectual property and civil litigation, etc., and the multi-dimensional matrix can be legal affairs.
[0238] The sample encoded problem comes from the data (pattern knowledge) manually annotated in a single business domain. In this way, after inputting the encoded problem into the data table selection unit, the trained data table selection unit can perform RAG (Retrieval Augmented Generation) in the preset vector library to sequentially determine the multi-dimensional matrix, matrix, and business domain to which the sample encoded problem belongs, and then determine the data table name parameter corresponding to the sample encoded problem in the business domain to which the sample encoded problem belongs. The data table name parameter can be a data table name or multiple data table names. Exemplarily, the preset vector library can be the lindorm library, and the preset vector library can store data table information and data column information.
[0239] The following table is an example of some data table information.
[0240]
[0241] The following table is an example of some data column information.
[0242]
[0243] After determining the data table name parameter, the corresponding field parameter of the sample encoded problem can be retrieved and determined in the preset vector library based on the corresponding data column information.
[0244] In this embodiment, when selecting the data table name parameter of the encoded problem, it is possible to determine the multi-dimensional matrix, matrix, and business domain to which the data table name parameter of the encoded problem belongs, narrow the data query range when determining the data table name parameter, further narrow the data query range when determining the field parameter, and improve the accuracy of structured data query.
[0245] To specifically illustrate the function of the method for improving the recognition and query of valid data by an artificial intelligence model provided in this embodiment, a specific example is provided below.
[0246] A method for improving the recognition and query of valid data by an artificial intelligence model includes: obtaining a current natural language question input by a user, and performing context recognition on the current natural language question and the historical natural language question when determining the historical natural language question input by the user; obtaining multiple historical natural language questions in the same context as the current natural language question from the historical natural language question, obtaining the time length ranking between the multiple historical natural language questions and the current natural language question in the time dimension, and selecting the top k historical natural language questions with the time length ranking as the target historical natural language questions;
[0247] determining the reference relationship between the target historical natural language question and the current natural language question, and performing a reference substitution and rewriting process on the current natural language question based on the reference relationship to obtain a standardized question; and / or determining the ellipsis relationship between the target historical natural language question and the current natural language question, and performing an ellipsis completion and rewriting process on the current natural language question based on the ellipsis relationship to obtain a standardized question;
[0248] performing part-of-speech analysis on the standardized question to obtain a part-of-speech analysis result, and annotating the nouns in the standardized question based on the part-of-speech analysis result; inputting the annotated standardized question into a NER model to obtain the entity labels of the nouns output by the NER model; the NER model is trained based on the sample standardized questions carrying the noun entity labels; performing element recognition processing on the standardized question based on the entity labels of the nouns to obtain an element extraction question; obtaining a preset keyword list, and performing a matching process on the element extraction question based on the preset keyword list to generate an element parameter pair of the standardized question; the element parameter pair includes an entity, an entity name, and an entity code; replacing the entity in the standardized question with the entity code to generate an encoded question;
[0249] Input the encoded problem into the tabular parameter acquisition model to obtain the tabular parameters of the data table output by the tabular parameter acquisition model; the tabular parameter acquisition model is trained based on sample encoded problems and sample tabular parameters of the data table; generate a structured query statement based on the tabular parameters of the data table and the encoded problem; input the structured query statement, the encoded problem, and the element parameter pair into the artificial intelligence model to obtain the answer output by the artificial intelligence model.
[0250] The method for improving the recognition and query of valid data by the artificial intelligence model provided by the embodiments of the present invention provides high-quality preprocessing for generating SQL query statements with relatively high execution accuracy for the artificial intelligence large model and NL2SQL, etc. through the optimization, processing, and data dimension refinement of natural language business problems. It preprocesses a large number of colloquial or natural language business problems that are not easy for the large model or machine and system to understand, and refines and processes more data that is helpful for more accurate and precise data output for subsequent business problem answering and structured data query. And by combining with the large model and the database, many natural language problems related to data query can be accurately answered and queried, and according to the technology and method of the present invention, the accuracy rate can reach more than 85%, reaching a trustworthy and usable level.
[0251] The device for improving the recognition and query of valid data by the artificial intelligence model provided by the present invention will be described below. The device for improving the recognition and query of valid data described below can be mutually corresponding and referenced with the method for improving the recognition and query of valid data by the artificial intelligence model described above.
[0252] Figure 6 is a schematic structural diagram of the device for improving the recognition and query of valid data by the artificial intelligence model provided by the present invention, as Figure 6 shown. The device includes:
[0253] The current problem acquisition module 601 is used to acquire the current natural language problem input by the user;
[0254] The historical problem acquisition module 602 is used to, when determining the historical natural language problem input by the user, perform context recognition on the current natural language problem and the historical natural language problem, and acquire the target historical natural language problem in the same context as the current natural language problem from the historical natural language problem;
[0255] The standardized problem acquisition module 603 is used to rewrite the current natural language problem based on the target historical natural language problem to obtain a standardized problem;
[0256] An answer acquisition module 604, configured to input the standardized question into the artificial intelligence model to obtain an answer output by the artificial intelligence model.
[0257] Figure 7 An example is a schematic structural diagram of an electronic device, as Figure 7 shown. The electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 may call logic instructions in the memory 730 to execute a method for improving the recognition of valid data by the artificial intelligence model. The method includes: obtaining a currently input natural language question of a user; in the case of determining a historically input natural language question of the user, performing context recognition on the currently input natural language question and the historically input natural language question, and obtaining a target historically input natural language question in the same context as the currently input natural language question from the historically input natural language question; performing rewriting processing on the currently input natural language question based on the target historically input natural language question to obtain a standardized question; inputting the standardized question into the artificial intelligence model to obtain an answer output by the artificial intelligence model.
[0258] In addition, when the logic instructions in the foregoing memory 730 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0259] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for improving the identification and query of valid data by an artificial intelligence model provided by the above-mentioned various methods. The method includes: obtaining a current natural language question input by a user; when it is determined that there is a historical natural language question input by the user, performing context recognition on the current natural language question and the historical natural language question, and obtaining a target historical natural language question in the same context as the current natural language question from the historical natural language question; performing rewriting processing on the current natural language question based on the target historical natural language question to obtain a standardized question; inputting the standardized question into the artificial intelligence model to obtain an answer output by the artificial intelligence model.
[0260] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for improving the identification and query of valid data by an artificial intelligence model provided by the above-mentioned various methods. The method includes: obtaining a current natural language question input by a user; when it is determined that there is a historical natural language question input by the user, performing context recognition on the current natural language question and the historical natural language question, and obtaining a target historical natural language question in the same context as the current natural language question from the historical natural language question; performing rewriting processing on the current natural language question based on the target historical natural language question to obtain a standardized question; inputting the standardized question into the artificial intelligence model to obtain an answer output by the artificial intelligence model.
[0261] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0262] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0263] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for improving the recognition of valid data in an artificial intelligence model, characterized in that Including: Obtain the current natural language question input by the user; When determining the historical natural language question input by the user, perform context recognition on the current natural language question and the historical natural language question, and obtain the target historical natural language question in the same context as the current natural language question from the historical natural language question; Rewrite the current natural language question based on the target historical natural language question to obtain a standardized question; Input the standardized question into the artificial intelligence model to obtain the answer output by the artificial intelligence model; After obtaining the standardized question, the method further includes: Identify the entities in the standardized question through the NER model and the preset keyword list. Among them, according to the context of the sample standardized question, some nouns in some sample standardized questions are entities, some nouns in some sample standardized questions are not entities, or at least some nouns are not entities; Replace the entities in the standardized question with entity codes to generate a coded question; Among them, entities include enterprises, industries, product services, natural persons, regions, and time; Identify the entities in the standardized question through the NER model and the preset keyword list, and replace the entities in the standardized question with entity codes to generate a coded question, including: Perform part-of-speech analysis on the standardized question to obtain a part-of-speech analysis result, and label the nouns in the standardized question based on the part-of-speech analysis result; Input the labeled standardized question into the NER model to obtain the entity labels of the nouns output by the NER model; the NER model is trained based on the sample standardized questions with noun entity labels; Perform element recognition processing on the standardized question based on the entity labels of the nouns to obtain an element extraction question; among them, the element recognition processing refers to not processing the nouns when it is determined that the nouns in the labeled standardized question are entities; when it is determined that the nouns in the labeled standardized question are not entities, delete the nouns; Obtain the preset keyword list, and perform matching processing on the element extraction question based on the preset keyword list to generate the element parameter pair of the standardized question; the element parameter pair includes an entity, an entity name, and an entity code; After generating the coded question, the method further includes: Obtain the data table column parameters corresponding to the coded question through the table column parameter acquisition model; Generate a structured query statement based on the data table column parameters and the coded question; The data table column parameters include a data table name parameter and the corresponding field parameters; Determine the data table name parameter, and retrieve and obtain the field parameters corresponding to the determined sample coded question based on the corresponding data column information in the preset vector library; when selecting the data table name parameter of the coded question, the multi-dimensional matrix, matrix, and business domain to which the data table name parameter of the coded question belongs can be determined, narrowing the data query range when determining the data table name parameter; After generating the structured query statement, the method further includes: Execute a retrieval in a preset database based on the structured query statement, and feedback the retrieval result to the table column parameter acquisition model to fine-tune the table column parameter acquisition model; Among them, the preset keyword list includes keywords, keyword names, and keyword encodings, and the entity, entity name, and entity encoding in the affiliated element parameter pair correspond to a set of keywords, keyword names, and keyword encodings; The table column parameter acquisition model includes a data table selection unit and a field selection unit; Input the encoded problem into the table column parameter acquisition model to obtain the data table column parameters output by the table column parameter acquisition model, specifically including: Input the encoded problem into the data table selection unit to obtain the data table name parameter output by the data table selection unit based on a preset multi-dimensional matrix; the preset multi-dimensional matrix includes multiple matrices, each matrix includes multiple business domains, each business domain includes multiple data tables, and each data table corresponds to a data table name parameter and multiple field parameters; Input the encoded problem and the data table name parameter into the field selection unit to obtain the data table name parameter and the corresponding field parameters output by the field selection unit.
2. The method for improving the recognition of query valid data by an artificial intelligence model according to claim 1, wherein Obtain the target historical natural language problem in the same context as the current natural language problem from the historical natural language problems, specifically including: Obtain multiple historical natural language problems in the same context as the current natural language problem from the historical natural language problems, obtain the time length ranking between the multiple historical natural language problems and the current natural language problem in the time dimension, and select the top k historical natural language problems with the time length ranking as the target historical natural language problems.
3. The method for improving the recognition of valid data in an artificial intelligence model for queries according to claim 1, characterized in that, Perform a rewriting process on the current natural language problem based on the target historical natural language problem, specifically including: Determine the reference relationship between the target historical natural language problem and the current natural language problem, and perform a reference substitution rewriting process on the current natural language problem based on the reference relationship; and / or Determine the ellipsis relationship between the target historical natural language problem and the current natural language problem, and perform an ellipsis completion rewriting process on the current natural language problem based on the ellipsis relationship.
4. The method for improving the recognition of valid data in a query by an artificial intelligence model according to claim 1 or 3, characterized in that, After obtaining the current natural language problem input by the user, the method further includes: In the case where the historical natural language problem input by the user is not determined, perform a reference substitution rewriting process on the current natural language problem based on a preset rule to obtain a standardized problem.
5. The method for improving the recognition of valid data in a query by an artificial intelligence model according to claim 1, characterized in that, After obtaining the current natural language problem input by the user, the method further includes: Input the current natural language problem into a problem classification model to obtain the problem category output by the problem classification model; the problem classification model is trained based on sample natural language problems with problem category labels; Input the standardized problem into the artificial intelligence model, specifically including: Input the problem category and the standardized problem into the artificial intelligence model.
6. The method for improving the recognition of valid data in an artificial intelligence model according to claim 1, wherein After annotating the nouns in the standardized problem based on the part-of-speech analysis result, the method further includes: When the entities in the standardization problem include name entities and company entities, obtain the names corresponding to the name entities, match the names in the enterprise name library corresponding to the company entities, and obtain the natural person codes of the name entities; When the entities in the standardization problem include the name entities and do not include the company entities, obtain the names corresponding to the name entities, match the names in the preset well-known person name library, and obtain the natural person codes of the name entities.
7. The method for improving the recognition of valid data in a query by an artificial intelligence model according to claim 1 or 6, characterized in that, After generating the encoded problem, the method further includes: Input the encoded problem into a table column parameter acquisition model to obtain the table column parameters of the data table output by the table column parameter acquisition model; the table column parameter acquisition model is trained based on sample encoded problems and sample data table column parameters; Generate a structured query statement based on the data table column parameters and the encoded problem; Input the structured query statement, the encoded problem, and the element parameter pair into an artificial intelligence model to obtain the answer output by the artificial intelligence model.
8. A device for improving the recognition of valid data in an artificial intelligence model, characterized in that, including: A current problem acquisition module for acquiring the current natural language problem input by the user; A historical problem acquisition module for, when determining the historical natural language problem input by the user, performing context recognition on the current natural language problem and the historical natural language problem, and obtaining the target historical natural language problem in the same context as the current natural language problem from the historical natural language problem; A standardization problem acquisition module for rewriting the current natural language problem based on the target historical natural language problem to obtain a standardization problem; An answer acquisition module for inputting the standardization problem into the artificial intelligence model to obtain the answer output by the artificial intelligence model; The standardization problem acquisition module is further used for: Identifying the entities in the standardization problem through an NER model and a preset keyword list. Among them, according to the context of the sample standardization problem, some nouns in some sample standardization problems are entities, some nouns in some sample standardization problems are not entities, or at least some nouns are not entities; Using entity encoding to replace the entities in the standardization problem to generate an encoded problem; where entities include enterprises, industries, product services, natural persons, regions, and time; The standardization problem acquisition module is specifically used for: Performing part-of-speech analysis on the standardization problem to obtain a part-of-speech analysis result, and annotating the nouns in the standardization problem based on the part-of-speech analysis result; Inputting the annotated standardization problem into an NER model to obtain the entity labels of the nouns output by the NER model; the NER model is trained based on sample standardization problems with noun entity labels; Performing element recognition processing on the standardization problem based on the entity labels of the nouns to obtain an element extraction problem; where the element recognition processing means that when determining that the nouns in the annotated standardization problem are entities, no processing is performed on the nouns; when determining that the nouns in the annotated standardization problem are not entities, the nouns are deleted. Obtain a preset keyword list, and perform matching processing on the element extraction problem based on the preset keyword list to generate an element parameter pair of the standardized problem; the element parameter pair includes an entity, an entity name, and an entity code; It further includes: A structured query statement generation module, which is used to obtain the data table column parameters corresponding to the encoded problem through a table column parameter acquisition model; generate a structured query statement based on the data table column parameters and the encoded problem; the data table column parameters include a data table name parameter and corresponding field parameters; determine the data table name parameter, and retrieve based on the corresponding data column information in the preset vector library to obtain the field parameters corresponding to the determined sample encoded problem; among them, when selecting the data table name parameter of the encoded problem, the multi-dimensional matrix, matrix, and business domain to which the data table name parameter of the encoded problem belongs can be determined, narrowing the data query range when determining the data table name parameter; A fine-tuning module, which is used to perform a retrieval in a preset database based on the structured query statement, and feedback the retrieval result to the table column parameter acquisition model to fine-tune the table column parameter acquisition model; Among them, the preset keyword list includes keywords, keyword names, and keyword codes, and the entities, entity names, and entity codes in the element parameter pair correspond to a set of keywords, keyword names, and keyword codes; The table column parameter acquisition model includes a data table selection unit and a field selection unit; Input the encoded problem into the table column parameter acquisition model to obtain the data table column parameters output by the table column parameter acquisition model, specifically including: Input the encoded problem into the data table selection unit to obtain the data table name parameter output by the data table selection unit based on the preset multi-dimensional matrix; the preset multi-dimensional matrix includes multiple matrices, each matrix includes multiple business domains, each business domain includes multiple data tables, and each data table corresponds to a data table name parameter and multiple field parameters; Input the encoded problem and the data table name parameter into the field selection unit to obtain the data table name parameter output by the field selection unit and the corresponding field parameters.
Citation Information
Patent Citations
Interactive intelligent question and answer method based on power grid practical training question and answer knowledge base
CN114417880A
Data query method and device, equipment and storage medium
CN116737756A