Data processing method and device, computer equipment, storage medium and program product
By filtering text paragraphs related to user questions and evaluating the accuracy and practicality of their answers, the questions that caused inaccurate answers are solved by recalling irrelevant information, and the effect of improving the accuracy of answers is achieved.
Patent Information
- Application Number
- CN202411731589.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, when answering user questions, the recalled document paragraphs contain irrelevant information, resulting in inaccurate answers. How to improve the accuracy of the answers becomes a problem that needs to be solved.
By obtaining problem data, multiple target text paragraphs are identified that have a degree of correlation to the problem exceeding the preset threshold, and input them into a pre-trained search text correlation scoring model to filter out reliable text paragraphs. Then, use the answer factual scoring model and the answer usability scoring model to evaluate the answers corresponding to reliable text paragraphs to ensure the accuracy and practicality of the answers.
By expanding the search scope, the possibility of recalling relevant document paragraphs is increased, and the factual accuracy and practicality of the answers are improved, ensuring that the final target answer data meets both the facts and the user needs.
Smart Images

Figure CN119938815A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a data processing method, device, computer equipment, storage medium and program product. Background Art
[0002] Currently, before answering user questions, information retrieval technology is first used to retrieve a set of documents or paragraphs that are most relevant to the question from large-scale document repositories (such as web pages, academic papers, professional databases, etc.). These documents serve as additional knowledge sources and provide more specific and accurate information background for large language models. Subsequently, the original question and the retrieved relevant documents are input into the large language model, and the model generates an answer based on this information.
[0003] However, when the retrieval model recalls document paragraphs containing answers to questions, it also recalls some document paragraphs that are not related to the question. This irrelevant information will cause the large model to give inaccurate answers. Therefore, how to improve the accuracy of answers to questions becomes a problem that needs to be solved. Summary of the invention
[0004] In view of this, the present invention provides a data processing method, apparatus, computer equipment, storage medium and program product.
[0005] In a first aspect, the present invention provides a data processing method, which includes: obtaining question data; determining multiple target text paragraphs based on the question data; wherein the target text paragraph is a text paragraph whose relevance to the question data exceeds a preset relevance threshold; inputting the multiple target text paragraphs and the question data into a pre-trained retrieval text relevance scoring model to obtain reliable text paragraphs; determining the reliability answer corresponding to each reliable text paragraph based on the reliable text paragraphs and the question data; inputting the multiple reliable text paragraphs, the question data and the reliability answers corresponding to the reliable text paragraphs into a pre-trained answer factual scoring model to obtain answer factual scoring results; inputting the multiple reliable text paragraphs, the question data and the reliability answers corresponding to the reliable text paragraphs into a pre-trained answer availability scoring model to obtain answer availability scoring results; determining the target answer data corresponding to the question data based on the answer factual scoring results and the answer availability scoring results.
[0006] The data processing method provided in this embodiment determines multiple target text paragraphs through question data, and the relevance of these paragraphs to the question data exceeds a preset relevance threshold. By expanding the retrieval scope, the possibility of recalling relevant document paragraphs is increased, thereby compensating for the problem of insufficient recall rate.
[0007] At the same time, by inputting multiple target text paragraphs and question data into the pre-trained retrieval text relevance scoring model, reliable text paragraphs that are highly relevant to the question data can be further screened out. In addition, the answer factuality scoring model is introduced, which can factually score the reliability answers corresponding to reliable text paragraphs, ensuring that the generated answers have high factual accuracy and reducing incorrect answers caused by insufficient precision.
[0008] At the same time, an answer usability scoring model is also introduced. The answer usability scoring model scores the reliability answers corresponding to reliable text paragraphs, and further selects answers that are both accurate and usable. This step improves the practicality and quality of the answers, ensuring that the final target answer data is both in line with the facts and meets user needs.
[0009] In one possible implementation, multiple target text paragraphs are determined based on question data, including: inputting the question data into a pre-trained vector embedding model to obtain a question data vector corresponding to the question data; and determining, based on the question data vector, multiple target text paragraphs whose similarity exceeds a preset similarity from a target vector library.
[0010] The data processing method provided in this embodiment can quickly convert question data into question data vectors by using a pre-trained vector embedding model, and can quickly find multiple target text paragraphs that are most relevant to the question data by calculating the similarity between the question data vector and the vectors of each text paragraph in the target vector library.
[0011] In one possible implementation, target answer data corresponding to the question data is determined based on the answer factuality scoring results and the answer availability scoring results, including: determining the mean of the answer factuality scoring results and the answer availability scoring results corresponding to each reliable text paragraph; comparing each mean, and determining a target mean with the largest mean from multiple means; and using the reliable text paragraph corresponding to the target mean as the target answer data corresponding to the question data.
[0012] The data processing method provided in this embodiment can further reduce the errors that may be caused by a single score and improve the accuracy and stability of the evaluation by calculating the average of the answer factuality score result and the answer availability score result corresponding to each reliable text paragraph.
[0013] At the same time, after obtaining multiple means, by comparing the sizes of each mean, the target mean with the largest mean can be determined. The reliable text paragraph corresponding to this target mean is the answer that performs best under the double standard.
[0014] In one possible implementation, the process of constructing a pre-trained retrieval text relevance scoring model includes: obtaining a text relevance training data set; wherein the training data set includes: a first historical question data, a first historical text paragraph related to the first historical question data, and a second historical text paragraph unrelated to the first historical question data; using the training data set to perform model training on a first model to obtain a pre-trained retrieval text relevance scoring model; wherein the first model is trained to provide a first coefficient result exceeding a first coefficient threshold for the first historical text paragraph, and to provide a second coefficient result below the first coefficient threshold for the second historical text paragraph.
[0015] In the data processing method provided in this embodiment, the training data set includes a first historical text paragraph related to the first historical question data and an unrelated second historical text paragraph, which can fully cover various possibilities of the question data, so that the model can learn more features of relevance and irrelevance during the training process. The first model is trained to provide a first coefficient result that exceeds the first coefficient threshold for the first historical text paragraph, and to provide a second coefficient result that is lower than the first coefficient threshold for the second historical text paragraph, so as to improve the first model's ability to quickly learn the criteria for distinguishing relevance and irrelevance during the training process.
[0016] In one possible implementation, the step of constructing a target vector library includes: obtaining factual text data; dividing the factual text data into a plurality of target paragraphs; wherein the target paragraph indicates a paragraph whose length is within a preset range; inputting the plurality of target paragraphs into a pre-trained vector embedding model respectively to obtain target paragraph vectors corresponding to the target paragraphs; and caching the target paragraph vectors into an initial vector library to obtain a target vector library.
[0017] The data processing method provided in this embodiment divides factual text data into multiple target paragraphs, which can make the information more structured and facilitate subsequent processing and management. At the same time, by setting the indicated length of the target paragraph within a preset range, it can ensure that the content of each paragraph is relatively concentrated and easy to understand.
[0018] At the same time, caching the target paragraph vector into the initial vector library to form the target vector library can speed up the subsequent retrieval and processing of text information. When you need to query a text paragraph or related text, you only need to match the vector in the target vector library to quickly find the required information.
[0019] In one possible implementation, the process of constructing a pre-trained answer factual scoring model includes: obtaining a factual scoring training data set; wherein the factual scoring training data set includes: a second historical question data, a historical text paragraph containing the second historical question data, a first reliability answer that is consistent with the facts of the text paragraph content, and a second reliability answer that is inconsistent with the facts of the text paragraph content; using the factual scoring training data set to perform model training on a second model to obtain a pre-trained answer factual scoring model; wherein the second model is trained to provide a third coefficient result that exceeds a second coefficient threshold for the first reliability answer, and to provide a fourth coefficient result that is lower than the second coefficient threshold for the second reliability answer.
[0020] The data processing method provided in this embodiment has a training data set that covers various types of questions and answers, and can fully reflect various situations that may be encountered in practical applications. This comprehensive design helps the model learn more features and rules during the training process and improves the generalization ability of the model.
[0021] At the same time, the second model is trained to provide a third coefficient result that exceeds the second coefficient threshold for the first reliability answer, and to provide a fourth coefficient result that is lower than the second coefficient threshold for the second reliability answer. This clear training goal helps the model quickly learn how to give corresponding factual scores based on the content of the answer during the training process.
[0022] In a second aspect, the present invention provides a data processing device, which includes: a question data acquisition module for acquiring question data; a first determination module for determining multiple target text paragraphs based on the question data; wherein the target text paragraph is a text paragraph whose relevance to the question data exceeds a preset relevance threshold; a first processing module for inputting the multiple target text paragraphs and the question data into a pre-trained retrieval text relevance scoring model to obtain reliable text paragraphs; a second determination module for determining the reliability answer corresponding to each reliable text paragraph based on the reliable text paragraphs and the question data; the second processing module for inputting the multiple reliable text paragraphs, the question data and the reliability answers corresponding to the reliable text paragraphs into a pre-trained answer factual scoring model to obtain answer factual scoring results; a third processing module for inputting the multiple reliable text paragraphs, the question data and the reliability answers corresponding to the reliable text paragraphs into a pre-trained answer availability scoring model to obtain answer availability scoring results; and a third determination module for determining the target answer data corresponding to the question data based on the answer factual scoring results and the answer availability scoring results.
[0023] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the data processing method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0024] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the data processing method of the first aspect or any corresponding embodiment thereof.
[0025] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the data processing method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0027] Figure 1 is a flow chart of a data processing method according to an embodiment of the present invention;
[0028] Figure 2 is a schematic diagram of constructing a pre-trained retrieval text relevance scoring model according to an embodiment of the present invention;
[0029] Figure 3 is a schematic diagram of constructing a pre-trained answer factuality scoring model / or constructing a pre-trained answer availability scoring model according to an embodiment of the present invention;
[0030] Figure 4 is a flow chart of a data processing method provided according to an embodiment of the present invention;
[0031] Figure 5 is a structural block diagram of a data processing device according to an embodiment of the present invention;
[0032] Figure 6 A schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0034] According to an embodiment of the present invention, a data processing method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0035] In this embodiment, a data processing method is provided, which can be used in computer equipment, such as computers, servers, etc. Figure 1 is a flow chart of a data processing method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0036] Step S101, obtaining question data.
[0037] Question data is used to represent the questions that users ask and need to be answered, usually in the form of natural language text. For example, "How long does it take for the earth to revolve around the sun?" Another example is "What is artificial intelligence, etc."
[0038] As an example, a question input by a user may be received through a user interface.
[0039] As an example, natural language processing techniques can be used to extract questions from documents or web pages.
[0040] Step S102, based on the question data, a plurality of target text paragraphs are determined; wherein the target text paragraphs are text paragraphs whose relevance to the question data exceeds a preset relevance threshold.
[0041] The target text paragraph is used to represent the text paragraph whose relevance to the question data exceeds the preset relevance threshold. The preset relevance threshold can be M1 or M2, etc., which is not specifically limited here. Specifically, keyword matching, semantic similarity calculation, etc. can be used to determine multiple target text paragraphs whose relevance to the question data exceeds the preset relevance threshold.
[0042] As an example, text paragraph G: "Learning programming requires patience and perseverance. You need to choose a suitable programming language and keep learning. At the same time, practice is very important, and you can consolidate what you have learned by writing code." Text paragraph H: "Cooking is an art that requires mastering the combination of various ingredients and cooking techniques. Different dishes have different methods and need to be adjusted according to personal taste."
[0043] Using the correlation calculation method, the following correlation scores can be obtained:
[0044] The relevance score of text paragraph G is 0.85 (assuming the full score is 1);
[0045] The relevance score of text paragraph H is 0.20 (assuming the full score is 1); if the preset relevance threshold is 0.5, then the relevance score of text paragraph G exceeds the threshold and is therefore regarded as a target text paragraph; while the relevance score of text paragraph H is lower than the threshold and is therefore not regarded as a target text paragraph.
[0046] Step S103 , inputting the plurality of target text paragraphs and question data into a pre-trained retrieval text relevance scoring model respectively to obtain reliable text paragraphs.
[0047] The pre-trained retrieval text relevance scoring model can be used to evaluate the relevance between the target text paragraph and the question data, and output a score or coefficient to represent the relevance. The pre-trained retrieval text relevance scoring model can be a deep learning model, such as BERT, RoBERTa, etc., or a model based on a traditional machine learning algorithm (such as SVM, logistic regression, etc.). The reliable text paragraph can represent a text paragraph extracted or generated from the reliable text paragraph for the question data.
[0048] Specifically, a threshold method may be used, that is, a text paragraph is considered reliable only when the score exceeds a preset threshold. A sorting method may also be used to sort all target text paragraphs according to the score, and select the top ranked ones as reliable text paragraphs, etc., which is not specifically limited here and can be implemented by those skilled in the art.
[0049] Each target text paragraph and question data are taken as a pair of input data and fed into the model for relevance scoring. Based on the input target text paragraph and question data, a score representing the relevance between them is calculated. This score is usually a floating point number between 0 and 1, indicating the degree of relevance between the two. According to actual needs, a threshold for the relevance score is set. This threshold is used to determine which target text paragraphs are highly relevant to the question data and can be considered as reliable text paragraphs. The relevance score of each target text paragraph is compared with the threshold, and the text paragraphs with a score exceeding the threshold are considered reliable text paragraphs. The screened out reliable text paragraphs are used as the output of the model.
[0050] Step S104, determining the reliability answer corresponding to each reliable text paragraph according to the reliable text paragraph and question data.
[0051] A reliable answer can represent an answer to a question data that is extracted or generated from a reliable text paragraph. This answer may be a direct answer or a text paragraph containing the answer. Specifically, machine learning or deep learning models can be used to extract answers or generate reliable answers. Natural language processing technology and domain knowledge bases can also be combined to reason about answers and generate reliable answers.
[0052] Step S105 , inputting the multiple reliable text paragraphs, question data and reliability answers corresponding to the reliable text paragraphs into a pre-trained answer factuality scoring model to obtain an answer factuality scoring result.
[0053] The answer factuality scoring model can be used to evaluate the factuality of the answer, that is, whether the answer truly and accurately reflects the facts involved in the question data. The answer factuality scoring model can be a model built based on deep learning (such as BERT, Transformer, etc.) or traditional machine learning algorithms.
[0054] Specifically, each reliable text paragraph, question data, and corresponding reliability answer are used as a set of input data and fed into the answer factuality scoring model, and an answer factuality scoring result is calculated based on the input data. This answer factuality scoring result is a floating point number between 0 and 1, where a higher score indicates that the answer factuality scoring result is more likely to be accurate and consistent with the facts.
[0055] It should be noted that the question data used in the model training process is the same question data, but the corresponding reliable text paragraphs are different and the reliability answers are different. For example: question G, reliable text paragraph R1, and reliability answer T1 are a group; question G, reliable text paragraph R2, and reliability answer T2 are a group; question G, reliable text paragraph R3, and reliability answer T3 are a group.
[0056] In a possible implementation, each answer factuality score result may be checked, and a threshold value (the threshold value may be N1 or N2, etc.) may be set according to actual needs. Answers whose factuality score results exceed the threshold value are regarded as more accurate and reliable answers.
[0057] Analyze factual scoring results for low-scoring answers to determine if there is potentially false or misleading information and consider whether further verification or correction is needed.
[0058] Step S106, inputting the multiple reliable text paragraphs, question data and reliability answers corresponding to the reliable text paragraphs into a pre-trained answer availability scoring model to obtain an answer availability scoring result.
[0059] The pre-trained answer usability scoring model can be a model built based on deep learning or traditional machine learning algorithms. Specifically, each reliable text paragraph, question data, and corresponding reliability answer are used as a set of input data and sent to the answer usability scoring model. Based on the input data, a score representing the usability of the answer is calculated. This score is usually a floating point number between 0 and 1, where a higher score indicates that the answer better meets user needs and has higher usability.
[0060] As an example, the question data is how to make pizza at home. There are two reliable text paragraphs, which provide two methods of making pizza and generate corresponding reliability answers:
[0061] Reliable text paragraph A: "Making pizza requires preparing dough, sauce, cheese, and various toppings. First, roll out the dough and spread it on a baking sheet, then spread it with sauce, sprinkle it with cheese and toppings, and finally bake it in the oven until golden."
[0062] Reliability answer A: "The steps to making pizza include preparing ingredients, rolling out dough, applying sauce, sprinkling toppings, and baking."
[0063] Reliable text paragraph B: "Making pizza at home is easy. First, you need a pizza crust, which can be homemade or purchased. Then, spread tomato sauce on the crust, sprinkle with mozzarella cheese and various toppings of your choice, such as ham, green peppers, onions, etc. Finally, bake the pizza in an oven preheated to 200 degrees for 15-20 minutes."
[0064] Reliability answer B: "Making pizza requires crust, tomato sauce, cheese and toppings, and the baking temperature is 200 degrees and the baking time is 15-20 minutes."
[0065] After inputting these two reliable text paragraphs, question data, and corresponding reliability answers into the answer usability scoring model, the following usability scores can be obtained:
[0066] The usability score of answer A is 0.8 (assuming the full score is 1); the usability score of answer B is 0.9 (assuming the full score is 1).
[0067] Step S107, determining the target answer data corresponding to the question data according to the answer factuality scoring result and the answer availability scoring result.
[0068] Determine the answer factuality scoring result and the answer factuality scoring result of each answer. Then, the answer factuality scoring result and the answer factuality scoring result of each answer can be averaged, weighted summed, etc. to determine the total score, and then determine the target answer data based on the total score.
[0069] In a possible implementation, after determining the total score, it is necessary to determine whether the total score meets the requirements (whether the total score of the target answer data exceeds a preset threshold), etc. Only when the total score exceeds the preset threshold, the target answer data can be returned to the user.
[0070] As an example, two answers K and L have the following factuality and usability scores:
[0071] Answer K: factuality score = 0.9, usability score = 0.7;
[0072] Answer L: Factualness score = 0.8, Usability score = 0.9.
[0073] If you value the accuracy (factuality) of the answer more, you can set a higher weight for the factuality score, such as 0.7, and a lower weight for the usability score, such as 0.3. Then, the comprehensive score of answer K is 0.9*0.7+0.7*0.3=0.84, and the comprehensive score of answer L is 0.8*0.7+0.9*0.3=0.83. In this case, answer K will be selected as the target answer data.
[0074] However, if we place more emphasis on the degree to which the answer satisfies user needs (usability), we can set a higher weight, such as 0.7, for the usability score and a lower weight, such as 0.3, for the factual score. Then, the comprehensive score of answer K is 0.9*0.3+0.7*0.7=0.7, and the comprehensive score of answer L is 0.8*0.3+0.9*0.7=0.77. In this case, answer L is selected as the target answer data.
[0075] The data processing method provided in this embodiment determines multiple target text paragraphs through question data, and the relevance of these paragraphs to the question data exceeds a preset relevance threshold. By expanding the retrieval scope, the possibility of recalling relevant document paragraphs is increased, thereby compensating for the problem of insufficient recall rate.
[0076] At the same time, by inputting multiple target text paragraphs and question data into the pre-trained retrieval text relevance scoring model, reliable text paragraphs that are highly relevant to the question data can be further screened out. In addition, the answer factuality scoring model is introduced, which can factually score the reliability answers corresponding to reliable text paragraphs, ensuring that the generated answers have high factual accuracy and reducing incorrect answers caused by insufficient precision.
[0077] At the same time, an answer usability scoring model is also introduced. The answer usability scoring model scores the reliability answers corresponding to reliable text paragraphs, and further selects answers that are both accurate and usable. This step improves the practicality and quality of the answers, ensuring that the final target answer data is both in line with the facts and meets user needs.
[0078] In a possible implementation, the above step S102 includes:
[0079] Step a1: input the question data into a pre-trained vector embedding model to obtain a question data vector corresponding to the question data.
[0080] The pre-trained vector embedding model is a machine learning model that converts text data into a fixed-length vector representation to capture the semantic similarities and relationships between texts. Specifically, load the pre-trained vector embedding model. Take the question data as input and process it through the model to obtain its vector representation.
[0081] As an example, you can use a pre-trained word vector model such as Word2Vec, GloVe, FastText, etc., and convert each word in the question data into a corresponding vector, and then average or weighted sum these vectors to obtain the question data vector.
[0082] As an example, you can use a Transformer-based pre-trained language model such as BERT, RoBERTa, GPT, etc., take the question data as input, and extract its last hidden state or pooling layer output as the question data vector.
[0083] Step a2: determining, based on the question data vector, a plurality of target text paragraphs whose similarities exceed a preset similarity from the target vector library.
[0084] The preset similarity can be used to determine which target text paragraphs are similar enough to the question data to be considered as potential sources of answers. The preset similarity can be q1, q2, etc., which are not specifically limited here. Specifically, each target text paragraph vector in the target vector library is traversed. The similarity between the question data vector and the target text paragraph vector is calculated. The target text paragraphs whose similarity exceeds the preset similarity are screened out.
[0085] As an example, cosine similarity can be used as a similarity measure to calculate the cosine value between the question data vector and the target text paragraph vector.
[0086] As an example, the Euclidean distance can be used as a variant of the similarity measurement method, by calculating the Euclidean distance between two vectors and then taking its inverse or performing other transformations to obtain the similarity value (note that this method requires appropriate adjustment of the preset similarity setting).
[0087] As an example, a K-nearest neighbor (KNN) algorithm or a similar method may be used to find K target text paragraphs in the target vector library that are most similar to the question data vector.
[0088] The data processing method provided in this embodiment can quickly convert question data into question data vectors by using a pre-trained vector embedding model, and can quickly find multiple target text paragraphs that are most relevant to the question data by calculating the similarity between the question data vector and the vectors of each text paragraph in the target vector library.
[0089] In a possible implementation, the above step S107 includes:
[0090] Step b1, determining the mean of the answer factuality score result and the answer availability score result corresponding to each reliable text paragraph.
[0091] For each reliable text paragraph, obtain its answer factuality score and answer availability score respectively, and calculate the average of these two scores as the comprehensive score of the reliable text paragraph.
[0092] As an example, the answer factuality scoring result and the answer availability scoring result may be added together and then divided by 2 to obtain the mean. Different weights may also be set for the answer factuality scoring result and the answer availability scoring result, and then weighted summed and divided by the weighted sum to obtain the mean.
[0093] Step b2, compare each mean and determine the target mean with the largest mean from multiple means.
[0094] Sort the comprehensive scores (means) of all reliable text paragraphs. Find the largest comprehensive score after sorting, which is the target mean. For example, if there are three reliable text paragraphs with comprehensive scores of 0.85 (paragraph A), 0.78 (paragraph B), and 0.92 (paragraph C), then the target mean can be determined to be 0.92 through sorting or comparison, corresponding to reliable text paragraph C.
[0095] Step b3, taking the reliable text paragraph corresponding to the target mean as the target answer data corresponding to the question data.
[0096] The data processing method provided in this embodiment can further reduce the errors that may be caused by a single score and improve the accuracy and stability of the evaluation by calculating the average of the answer factuality score result and the answer availability score result corresponding to each reliable text paragraph.
[0097] At the same time, after obtaining multiple means, by comparing the sizes of each mean, the target mean with the largest mean can be determined. The reliable text paragraph corresponding to this target mean is the answer that performs best under the double standard.
[0098] In a possible implementation, the method further includes:
[0099] Step c1, obtaining a text relevance training data set; wherein the training data set includes: a first historical question data, a first historical text paragraph related to the first historical question data, and a second historical text paragraph unrelated to the first historical question data.
[0100] Step c2, using the training data set to train the first model to obtain a pre-trained retrieval text relevance scoring model; wherein the first model is trained to provide a first coefficient result exceeding a first coefficient threshold for the first historical text paragraph, and to provide a second coefficient result below the first coefficient threshold for the second historical text paragraph.
[0101] In order to train a pre-trained retrieval text relevance scoring model, we first need to construct a corresponding data set. The format of each data sample in the data set is as follows: (x, y1, y2). For the training data set of the retrieval text relevance scoring model, x is used to represent the first historical question data, y1 is used to represent the first historical text paragraph related to the first historical question data, and y2 is used to represent the second historical text paragraph unrelated to the first historical question data.
[0102] The ultimate goal of the pre-trained retrieval text is to be able to provide a first coefficient result exceeding a first coefficient threshold for a first historical text paragraph, and to provide a second coefficient result below the first coefficient threshold for a second historical text paragraph.
[0103] Please refer to Figure 2, Figure 2 It is a schematic diagram of constructing a pre-trained retrieval text relevance scoring model provided according to an embodiment of the present invention.
[0104] Figure 2 In the example, the first historical question data is to introduce a freshwater lake in China. Four texts are obtained: A: The Yangtze River is...; B: Qinghai Lake is...; C: Huangshan is located at...; D: Poyang Lake is located at... Among them, D>B>A=C. Then, the first model (text relevance scoring model) is trained according to the sorting results to obtain a pre-trained retrieval text relevance scoring model.
[0105] In the data processing method provided in this embodiment, the training data set includes a first historical text paragraph related to the first historical question data and an unrelated second historical text paragraph, which can fully cover various possibilities of the question data, so that the model can learn more features of relevance and irrelevance during the training process. The first model is trained to provide a first coefficient result that exceeds the first coefficient threshold for the first historical text paragraph, and to provide a second coefficient result that is lower than the first coefficient threshold for the second historical text paragraph, so as to improve the first model's ability to quickly learn the criteria for distinguishing relevance and irrelevance during the training process.
[0106] In a possible implementation, the method further includes:
[0107] Step d1, obtaining factual textual information.
[0108] Factual text materials may be texts containing objective facts, data, events and other information, such as news reports, scientific papers, historical records, etc. Specifically, web crawler technology may be used to crawl relevant text materials from the Internet.
[0109] Step d2, dividing the factual text data into a plurality of target paragraphs; wherein the target paragraphs indicate paragraphs whose length is within a preset range.
[0110] The target paragraph is used to represent the paragraphs that are segmented from factual text materials and whose length is within a preset range. These paragraphs are usually used for subsequent text processing and analysis. The preset range can be (O1-O2) or (O3-O4), etc., which is not specifically limited here. Specifically, text processing software (such as Python's TextBlob, NLTK and other libraries) can be used for automatic segmentation.
[0111] In step d3, multiple target paragraphs are respectively input into a pre-trained vector embedding model to obtain target paragraph vectors corresponding to the target paragraphs.
[0112] Each target paragraph is taken as input and vectorized through a pre-trained vector embedding model to obtain the target paragraph vector.
[0113] As an example, an open source vector embedding model (such as Word2Vec, BERT, GPT, etc.) can be used for vector conversion. Preferably, the target paragraph is preprocessed (such as removing stop words, stemming, word form restoration, etc.) to improve the effect of vector conversion.
[0114] Step d4, caching the target paragraph vector into the initial vector library to obtain the target vector library.
[0115] The initial vector library is used to store and retrieve vector information. In this embodiment, the initial vector library may be an empty database or a database containing a small amount of basic vectors. Specifically, each target paragraph vector is stored in the initial vector library to form a target vector library containing multiple target paragraph vectors.
[0116] As an example, find this text about climate change on the Internet. Divide the text into two target paragraphs: "Climate change refers to long-term changes caused by natural or human factors (such as...)." and "In recent years, with the continuous increase in human activities...it has had a profound impact". Input the two target paragraphs into the pre-trained BERT model respectively to obtain two target paragraph vectors. Store the two target paragraph vectors in the MySQL database to form a target vector library.
[0117] The data processing method provided in this embodiment divides factual text data into multiple target paragraphs, which can make the information more structured and facilitate subsequent processing and management. At the same time, by setting the indicated length of the target paragraph within a preset range, it can ensure that the content of each paragraph is relatively concentrated and easy to understand.
[0118] At the same time, caching the target paragraph vector into the initial vector library to form the target vector library can speed up the subsequent retrieval and processing of text information. When you need to query a text paragraph or related text, you only need to match the vector in the target vector library to quickly find the required information.
[0119] In a possible implementation, the method further includes:
[0120] Step e1, obtaining a factual scoring training data set; wherein the factual scoring training data set includes: the second historical question data, the historical text paragraph containing the second historical question data, the first reliability answer consistent with the facts of the text paragraph content, and the second reliability answer inconsistent with the facts of the text paragraph content.
[0121] Step e2, using the factual scoring training data set to train the second model to obtain a pre-trained answer factual scoring model; wherein the second model is trained to provide a third coefficient result exceeding the second coefficient threshold for the first reliability answer, and to provide a fourth coefficient result below the second coefficient threshold for the second reliability answer.
[0122] The first reliability answer is used to characterize the answer that is consistent with the facts of the text paragraph content, that is, this answer is factually correct. The second reliability answer is used to characterize the answer that is inconsistent with the facts of the text paragraph content, that is, this answer is factually wrong. Specifically, the second model is trained using a factual scoring training data set. The second model is trained to provide a third coefficient result that exceeds the second coefficient threshold for the first reliability answer (i.e., the correct answer), and to provide a fourth coefficient result that is lower than the second coefficient threshold for the second reliability answer (i.e., the wrong answer). The coefficient here can be understood as the model's score or confidence in the factuality of the answer.
[0123] It should be noted that the process of building a pre-trained answer factuality scoring model is the same as the process of building a pre-trained answer usability scoring model.
[0124] Please refer to Figure 3 , Figure 3 It is a schematic diagram of constructing a pre-trained answer factuality scoring model / or constructing a pre-trained answer availability scoring model provided according to an embodiment of the present invention.
[0125] Figure 3 In the data, the second historical question data and the historical text paragraph containing the second historical question data may include: Introducing a freshwater lake in China + Poyang Lake is located in the northern part of Jiangxi Province, with an area of 2,933 square kilometers, and is the largest freshwater lake in China. The scenery of Duyang Lake varies throughout the year, and it is beautiful in spring...., and the corresponding answers are A: Fanyang Lake is located in the west of China; B: Poyang Lake is located in the north of Jiangbei, China, and is the second largest freshwater lake in China. C: Poyang Lake is located in Hangzhou; D: Poyang Lake is the largest freshwater lake in China. Among them, the answer ranking is: B>A=C=D. Then, the second model (factual scoring model or reliability scoring model) is trained according to the answer ranking to obtain the pre-trained answer corresponding to the factual scoring model and the pre-trained answer availability scoring model corresponding to the factual scoring model or the reliability scoring model.
[0126] The data processing method provided in this embodiment has a training data set that covers various types of questions and answers, and can fully reflect various situations that may be encountered in practical applications. This comprehensive design helps the model learn more features and rules during the training process and improves the generalization ability of the model.
[0127] At the same time, the second model is trained to provide a third coefficient result that exceeds the second coefficient threshold for the first reliability answer, and to provide a fourth coefficient result that is lower than the second coefficient threshold for the second reliability answer. This clear training goal helps the model quickly learn how to give corresponding factual scores based on the content of the answer during the training process.
[0128] Please refer to Figure 4 , Figure 4 is a flow chart of a data processing method provided according to an embodiment of the present invention.
[0129] First, the format of each data sample of the dataset for scoring model training is as follows: (x, y1, y2). For the first model (retrieval text relevance scoring model), x represents the question, y1 represents the content related to the question, and y2 represents the content unrelated to the question; for a second model (answer factuality scoring model), x represents the question + the text paragraph containing the question content, y1 represents the answer that is consistent with the facts of the text paragraph content, and y2 represents the answer that is inconsistent with the facts of the text paragraph content; for another second model (answer availability scoring model), x represents the question + the text paragraph containing the question content, y1 represents the answer with higher readability, and y2 represents the answer with lower readability.
[0130] Based on the training mechanism of the reward model, the above datasets are used to train the text relevance scoring model, the answer factuality scoring model, and the answer availability scoring model respectively.
[0131] For user questions, we first use the vector embedding model to encode them into vectors, and then use the vectors to match and sort all the vectors in the FAISS vector library to get the n texts with the highest similarity (y1, y2, ..., y n ).
[0132] Combine the user's question with the n related text paragraphs {(x,y1),(x,y2),…,(x,y n )} Input the search text relevance scoring model one by one to obtain the relevance score between the question and each text paragraph. Exclude the text with relevance score lower than the threshold to obtain reliable text paragraphs
[0133] Combine the question with multiple reliable text passages Input them one by one to the large language model, and let the large language model make corresponding answers (a1, a2, ..., a m ).
[0134] Question + reliable text and corresponding answer Input the answer factual scoring model and the answer availability scoring model one by one to obtain the corresponding scores Filter out the ones with the highest overall score If the comprehensive score is higher than the threshold, indicating that the answer is more factual and reliable, then h The answer is returned to the user as an answer. If the comprehensive score is lower than the threshold, it means that the factuality and reliability of the answer are low. In order to avoid misleading the user, a signal of rejection will be returned to the user.
[0135] In one possible implementation, a large language model is used as a reward model. The output of the reward model is similar to a regression task. The output vector of the large model is passed through a linear layer to obtain a scalar output. The loss function during training is as follows:
[0136] Among them, r θ (x,y1) represents the score of the reward model for question x+text content y1. Depending on the training set, it can be a relevance score, a factual score, or an availability score. D represents the corresponding training data set. The training goal of the reward model is to increase the gap between the model's scores for question x+text content y1 and question x+text content y2, so as to train an effective scoring model.
[0137] In the present embodiment, a data processing device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0138] This embodiment provides a data processing device, such as Figure 5 As shown, including:
[0139] The problem data acquisition module 501 is used to acquire problem data;
[0140] A first determination module 502 is used to determine a plurality of target text paragraphs based on the question data; wherein the target text paragraphs are text paragraphs whose relevance to the question data exceeds a preset relevance threshold;
[0141] The first processing module 503 is used to input the multiple target text paragraphs and question data into the pre-trained retrieval text relevance scoring model to obtain reliable text paragraphs;
[0142] The second determination module 504 is used to determine the reliability answer corresponding to each reliable text paragraph according to the reliable text paragraph and the question data;
[0143] The second processing module 505 is used to input the multiple reliable text paragraphs, question data and reliability answers corresponding to the reliable text paragraphs into the pre-trained answer factuality scoring model to obtain the answer factuality scoring result;
[0144] The third processing module 506 is used to input the plurality of reliable text paragraphs, the question data and the reliability answers corresponding to the reliable text paragraphs into the pre-trained answer availability scoring model to obtain the answer availability scoring result;
[0145] The third determination module 507 is used to determine the target answer data corresponding to the question data according to the answer factuality scoring result and the answer availability scoring result.
[0146] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0147] The data processing device in this embodiment is presented in the form of a functional unit, where the functional unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0148] The embodiment of the present invention also provides a computer device having the above Figure 5 The data processing device shown.
[0149] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.
[0150] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0151] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0152] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0153] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0154] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0155] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0156] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0157] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A data processing method, characterized in that: The method comprises: Get the problem data; Based on the question data, a plurality of target text paragraphs are determined; wherein the target text paragraphs are text paragraphs whose relevance to the question data exceeds a preset relevance threshold; Inputting the plurality of target text paragraphs and question data into a pre-trained retrieval text relevance scoring model to obtain reliable text paragraphs; Determining a reliability answer corresponding to each reliable text paragraph according to the reliable text paragraph and the question data; Inputting the plurality of reliable text paragraphs, the question data and the reliability answers corresponding to the reliable text paragraphs into a pre-trained answer factuality scoring model to obtain an answer factuality scoring result; Inputting the plurality of reliable text paragraphs, the question data and the reliability answers corresponding to the reliable text paragraphs into a pre-trained answer availability scoring model to obtain an answer availability scoring result; According to the answer factuality scoring result and the answer availability scoring result, target answer data corresponding to the question data is determined.
2. The data processing method according to claim 1, characterized in that: Based on the question data, multiple target text paragraphs are determined, including: Inputting the question data into a pre-trained vector embedding model to obtain a question data vector corresponding to the question data; According to the problem data vector, a plurality of target text paragraphs having similarities exceeding a preset similarity are determined from a target vector library.
3. The data processing method according to claim 1, characterized in that: Determining target answer data corresponding to the question data according to the answer factuality scoring result and the answer availability scoring result includes: Determine the mean of the answer factuality score and the answer usability score for each reliable text paragraph; Compare the various means and determine the target mean with the largest mean from multiple means; The reliable text paragraph corresponding to the target mean is used as the target answer data corresponding to the question data.
4. The data processing method according to claim 1, characterized in that: The process of building a pre-trained retrieval text relevance scoring model includes: Acquire a text relevance training data set; wherein the training data set includes: first historical question data, a first historical text paragraph related to the first historical question data, and a second historical text paragraph unrelated to the first historical question data; The first model is trained using the training data set to obtain a pre-trained retrieval text relevance scoring model; wherein the first model is trained to provide a first coefficient result exceeding a first coefficient threshold for a first historical text paragraph, and to provide a second coefficient result below the first coefficient threshold for a second historical text paragraph.
5. The data processing method according to claim 2, characterized in that: The steps to build the target vector library include: Obtain factual textual information; Dividing the factual text material into a plurality of target paragraphs; wherein the target paragraphs indicate paragraphs whose lengths are within a preset range; Inputting the plurality of target paragraphs into a pre-trained vector embedding model respectively to obtain target paragraph vectors corresponding to the target paragraphs; The target paragraph vector is cached in an initial vector library to obtain a target vector library.
6. The data processing method according to claim 1, characterized in that: The process of building a pre-trained answer factuality scoring model includes: Obtain a factual scoring training data set; wherein the factual scoring training data set includes: the second historical question data, the historical text paragraph containing the second historical question data, the first reliability answer consistent with the facts of the text paragraph content, and the second reliability answer inconsistent with the facts of the text paragraph content; The second model is trained using the factual scoring training data set to obtain a pre-trained answer factual scoring model; wherein the second model is trained to provide a third coefficient result exceeding the second coefficient threshold for the first reliability answer, and to provide a fourth coefficient result below the second coefficient threshold for the second reliability answer.
7. A data processing device, characterized in that: The device comprises: A problem data acquisition module, used for acquiring problem data; A first determination module is used to determine a plurality of target text paragraphs based on the question data; wherein the target text paragraphs are text paragraphs whose relevance to the question data exceeds a preset relevance threshold; A first processing module is used to input the plurality of target text paragraphs and question data into a pre-trained retrieval text relevance scoring model to obtain reliable text paragraphs; A second determination module is used to determine the reliability answer corresponding to each reliable text paragraph according to the reliable text paragraph and the question data; A second processing module is used to input the plurality of reliable text paragraphs, the question data and the reliability answers corresponding to the reliable text paragraphs into a pre-trained answer factuality scoring model to obtain an answer factuality scoring result; A third processing module is used to input the plurality of reliable text paragraphs, the question data and the reliability answers corresponding to the reliable text paragraphs into a pre-trained answer availability scoring model to obtain an answer availability scoring result; The third determination module is used to determine the target answer data corresponding to the question data according to the answer factuality scoring result and the answer availability scoring result.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data processing method according to any one of claims 1 to 6 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data processing method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the data processing method according to any one of claims 1 to 6.