Context-based question and answer processing method and device
Through the improved question-and-answer model, the efficiency and accuracy problems of the existing question-and-answer model in the case of more contexts are solved, and efficient and accurate question-and-answer processing is achieved.
Patent Information
- Application Number
- CN202311695051.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-17
AI Technical Summary
The existing question and answer model has low processing efficiency and cannot skip sentences when there are many processing contexts, resulting in reduced processing accuracy.
A modified question and answer model is adopted, which consists of multiple modules, including an embedded encoding module, an encoder, a correlation processing module, an answer generator, a correct answer probability prediction model and an answer sorting module. The model uses one-time encoding of context and question text, perform similarity correlation and clustering, generate answers, and evaluate the correctness of the answers.
The coding cycle is shortened, the Q&A efficiency is improved, and multiple sentences in the context can be used as references to generate answers, which improves the accuracy of the Q&A, and further improves the accuracy by evaluating the correctness of the answers.
Smart Images

Figure CN120162398A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a context-based question and answer processing method and device. Background Art
[0002] A question and answer model is a type of artificial intelligence model implemented based on a language model (LM). This type of model is typically composed of an encoding module, a retrieval module, and an answer generation module (also known as a language generator or text generator). The working principle of this type of model is that the encoding module encodes each sentence of the input question text to output the corresponding encoded feature tensor, the retrieval module performs approximate text retrieval on the text information library based on each encoded feature tensor to obtain one or more corresponding retrieval texts, and the answer generation module generates the answer text based on the set language / grammar logic according to the question text and all retrieval texts. This type of model is commonly used in application scenarios such as robot question and answer and knowledge retrieval.
[0003] With the gradual enrichment of the application scenarios of the question and answer model, recently some developers have also tried to use it to solve reading comprehension problems. The so-called reading comprehension problem is to answer or search for the answer or associated clues of a given question with the given context as the retrieval object. However, during the actual operation process, developers found some problems: 1) The encoding module of the conventional question and answer model only performs one-sentence encoding each time, that is, the processing process of the conventional question and answer model requires multiple loops of encoding + retrieval + answer generation processes, which will lead to a decrease in the processing efficiency of the question and answer model in the case of a large amount of context content; 2) The conventional question and answer model only calculates the similarity between each sentence of the context and the question text one by one, and uses the sentence with the highest similarity as the reference text to generate the corresponding answer text, and does not perform skip-sentence understanding of the full text like the human reading habit. Skip-sentence understanding means taking multiple adjacent or non-adjacent sentences as a whole to be associated with the question text, which will reduce the processing accuracy of the question and answer model; 3) The conventional question and answer model does not discriminate whether the generated answer text fits the context and the question text, which will also reduce the processing accuracy of the question and answer model. Summary of the Invention
[0004] The object of the present invention is to provide a context-based question and answer processing method, device, electronic device and computer-readable storage medium in view of the defects of the prior art. The present invention performs question and answer processing based on an improved question and answer model, which is composed of a first embedding encoding module, a second embedding encoding module, a first encoder, a second encoder, a relevance processing module, an answer generator, a third embedding encoding module, a third encoder, a correct answer probability prediction model and an answer sorting module; among them, the first embedding encoding module + the first encoder are responsible for splitting the context text into sentences and performing one-time encoding on each sentence; the second embedding encoding module + the second encoder are responsible for performing one-time encoding on the question text; the relevance processing module not only associates the similarity between each sentence in the context and the question text based on the context encoding tensor and the question encoding tensor, but also clusters multiple sentences in the context with approximate similarity to the question text and further associates the similarity between each clustering vector and the question text; the answer generator generates corresponding answer texts according to various association combinations output by the relevance processing module; the third embedding encoding module + the third encoder are responsible for performing one-time encoding on each answer text; the correct answer probability prediction model evaluates the correctness of each answer text output by the answer generator according to various association combinations output by the relevance processing module and each answer feature encoding output by the third encoder; the answer sorting module sorts the correct answer texts according to the evaluation results of the correct answer probability prediction model. Through the question and answer model provided by the present invention, on the one hand, the encoding cycle can be shortened and the model processing efficiency can be improved, on the other hand, multiple sentences in the context can be used as a reference to improve the processing accuracy of the model, and on the other hand, the correctness of the generated answer texts can be evaluated and used to improve the processing accuracy of the model.
[0005] To achieve the above object, a first aspect of an embodiment of the present invention provides a context-based question and answer processing method, the method comprising:
[0006] Receiving a context text, a question text and a preferred number of answers as a corresponding first context, first question and first preferred number; the first preferred number is an integer greater than or equal to 0;
[0007] Inputting the first context and the first question into a question and answer model, and predicting an answer to the first question by the question and answer model according to the first context to obtain a corresponding first predicted answer sequence;
[0008] Performing question and answer result screening according to the first preferred number and the first predicted answer sequence to generate a corresponding first screening result.
[0009] Preferably, the Q&A model includes a first embedding and encoding module, a second embedding and encoding module, a first encoder, a second encoder, a relevance processing module, an answer generator, a third embedding and encoding module, a third encoder, a correct answer probability prediction model, and an answer ranking module;
[0010] The Q&A model has two model input ends, namely the first and second model input ends; the model output end of the Q&A model is the first model input end; the first input end is used to receive the first context, the second model input end is used to receive the first question, and the first model input end is used to output the first predicted answer sequence;
[0011] The input end of the first embedding and encoding module is connected to the first model input end, and the output end is connected to the input end of the first encoder; the input end of the second embedding and encoding module is connected to the second model input end, and the output end is connected to the input end of the second encoder; the output end of the first encoder is connected to the first input end of the relevance processing module; the output end of the second encoder is connected to the second input end of the relevance processing module; the first output end of the relevance processing module is connected to the input end of the answer generator, and the second output end is connected to the second input end of the correct answer probability prediction model; the first output end of the answer generator is connected to the third embedding and encoding module, and the second output end is connected to the second input end of the answer ranking module; the output end of the third embedding and encoding module is connected to the input end of the third encoder; the output end of the third encoder is connected to the first input end of the correct answer probability prediction model; the output end of the correct answer probability prediction model is connected to the first input end of the answer ranking module.
[0012] Further, the first, second, and third encoders are implemented based on the same encoder model; the encoder model includes at least the encoder of the Transformer model;
[0013] The embedding and encoding mechanisms of the first, second, and third embedding and encoding modules are implemented based on the same text embedding and encoding mechanism; the text embedding and encoding mechanism includes at least one-hot encoding mechanism and Word2Vec encoding mechanism;
[0014] The answer generator is implemented based on a type of LM model for text generation; the LM model for text generation includes RNN model, LSTM model, SeqGAN model, Transformer model, GPT model, BERT model, GPT-2 model, and BART model;
[0015] The correct answer probability prediction model consists of a feature fusion unit and a binary classification model; the binary classification model includes a binary classifier implemented based on the MLP structure and a binary classifier implemented based on the CNN structure.
[0016] Preferably, the answer model predicts the answer to the first question according to the first context to obtain a corresponding first predicted answer sequence, specifically including:
[0017] The first embedding and encoding module performs text segmentation on the first context to obtain a corresponding first sentence sequence TC{tc i}; and performs embedding encoding on each first sentence tc i to obtain a corresponding first text vector c i ; and sorts all the obtained first text vectors c i to form a corresponding first text vector sequence C and sends it to the first encoder; the first sentence sequence TC includes a plurality of the first sentences tc i , and the sentence index i is an integer greater than 0;
[0018] And the second embedding and encoding module performs embedding encoding on the first question to obtain a corresponding second text vector q and sends it to the second encoder;
[0019] And the first encoder performs feature encoding on the first text vector sequence C to obtain a corresponding first encoded tensor sequence X C {x c,i} and sends it to the correlation processing module; the first encoded tensor sequence X C includes a plurality of first encoded tensors x c,i ;
[0020] And the second encoder performs feature encoding on the second text vector q to obtain a corresponding second encoded tensor x q and sends it to the correlation processing module;
[0021] And the correlation processing module performs correlation prediction according to the first encoded tensor sequence X C and the second encoded tensor x q to obtain a corresponding first correlation tensor sequence R; and sends the first correlation tensor sequence R to the answer generator and the correct answer probability prediction model respectively; the first correlation tensor sequence R includes a plurality of first correlation tensors r j , and the correlation tensor index j is an integer greater than 0; the first correlation tensor r j includes a first tensor y j , the second encoded tensor x q and a first correlation coefficient s j; The first tensor y j is composed of one or more of the first encoded tensors x c,i ;
[0022] And the answer generator generates an answer text based on each of the first correlation tensors r in the first correlation tensor sequence R j of the first tensor y j and the second encoded tensor x q to obtain a corresponding first answer w j ; And all the obtained first answers w j form a corresponding first answer sequence W; and the first answer sequence W is sent to the third embedding encoding module and the answer sorting module respectively;
[0023] And the third embedding encoding module performs embedding encoding on each of the first answers w in the first answer sequence W j to obtain a corresponding third text vector a j ; And all the obtained third text vectors a j are sorted to form a corresponding third text vector sequence A and sent to the third encoder;
[0024] And the third encoder performs feature encoding on the third text vector sequence A to obtain a corresponding third encoded tensor sequence X A {x a,j} and send it to the correct answer probability prediction model; The third encoded tensor sequence X A includes multiple third encoded tensors x a,j ;
[0025] And the feature fusion unit of the correct answer probability prediction model fuses the features of the third encoded tensor sequence X A and the first correlation tensor sequence R to obtain a corresponding first fusion feature tensor M; and the binary classification model of the correct answer probability prediction model performs binary classification prediction of correct / wrong answers based on each first feature tensor m of the first fusion feature tensor M j to obtain a corresponding first prediction vector; and identify whether the first correct answer probability of each first prediction vector is greater than the first wrong answer probability, if so, set the corresponding first prediction probability z j as the first correct answer probability, if not, set the corresponding first prediction probability z j to 0; and all the obtained first prediction probabilities z j form a corresponding first prediction probability vector Z and send it to the answer sorting module; The first fusion feature tensor M consists of multiple first feature tensors m jComposition; the first feature tensor m j Composed of the corresponding third encoding tensor x a,j And the first correlation tensor r j Composition; the first prediction vector includes the first correct answer probability and the first wrong answer probability;
[0026] And the answer sorting module marks the first prediction probability z with a probability value of 0 in the first prediction probability vector Z j All as screening probabilities; and the first answer w corresponding to each of the screening probabilities in the first answer sequence W j Delete; and each of the remaining first answers w in the first answer sequence W j As the corresponding first predicted answer; and for all the first predicted answers according to the corresponding first prediction probability z j Sort in descending order to form the corresponding first predicted answer sequence; and output the first predicted answer sequence as the prediction result of the question-answering model.
[0027] Furthermore, the correlation processing module performs correlation prediction according to the first encoding tensor sequence X C And the second encoding tensor x q To obtain the corresponding first correlation tensor sequence R, specifically including:
[0028] The correlation processing module calculates the similarity between each first encoding tensor x C In the first encoding tensor sequence X c,i And the second encoding tensor x q To obtain the corresponding first similarity d i ;
[0029] And each first encoding tensor x c,i As a corresponding first tensor y j ; And each first encoding tensor x c,i The corresponding first similarity d i As a corresponding first correlation coefficient s j ; And by each first encoding tensor x c,i The corresponding first tensor y j And the first correlation coefficient s j And the second encoding tensor x q To form a corresponding first correlation tensor r j ;
[0030] And based on the preset piecewise similarity clustering principle for all the first similarities d iPerform clustering to obtain multiple first clustering sets; each of the first clustering sets includes multiple of the first similarity degrees d i ; each of the first clustering sets corresponds to a preset similarity range, and all of the first similarity degrees d i in each of the first clustering sets satisfy the corresponding similarity range;
[0031] And calculate the mean of all of the first similarity degrees d i in each of the first clustering sets to obtain the corresponding first clustering similarity degree; and from all of the first similarity degrees d i corresponding to each of the first clustering sets, all of the first encoded tensors x c,i form a corresponding first tensor y j ; and use the first clustering similarity degree corresponding to each of the first clustering sets as a corresponding first correlation coefficient s j ; and from the first tensor y j corresponding to each of the first clustering sets and the first correlation coefficient s j and the second encoded tensor x q form a corresponding first correlation tensor r j ;
[0032] And from all of the first correlation tensors r j obtained, form the corresponding first correlation tensor sequence R.
[0033] Preferably, the screening of the question-answering result according to the first preferred quantity and the first predicted answer sequence to generate the corresponding first screening result specifically includes:
[0034] Count the number of the first predicted answers in the first predicted answer sequence to obtain the corresponding first quantity;
[0035] Identify the first preferred quantity and the first quantity;
[0036] If the first preferred quantity is equal to 0 or the first quantity is less than the first preferred quantity, then use the first predicted answer sequence as the corresponding first screening result;
[0037] If the first quantity is greater than or equal to the first preferred quantity, then extract the first preferred quantity of the first predicted answers with the earliest ranking in the first predicted answer sequence to form the corresponding first screening result.
[0038] In the second aspect of the embodiments of the present invention, there is provided an apparatus for implementing the context-based question-answering processing method described in the first aspect above. The apparatus includes: a data receiving module, a question-answering model processing module, and a data output module;
[0039] The data receiving module is configured to receive context text, question text, and the preferred number of answers as the corresponding first context, first question, and first preferred number; the first preferred number is an integer greater than or equal to 0.
[0040] The question and answer model processing module is configured to input the first context and the first question into a question and answer model, and the question and answer model predicts the answer to the first question according to the first context to obtain a corresponding first predicted answer sequence.
[0041] The data output module is configured to perform question and answer result screening according to the first preferred number and the first predicted answer sequence to generate a corresponding first screening result.
[0042] A third aspect of the embodiments of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0043] The processor is configured to be coupled to the memory, read and execute instructions in the memory to implement the method steps described in the first aspect above;
[0044] The transceiver is coupled to the processor, and the processor controls the transceiver to perform message sending and receiving.
[0045] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium storing computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.
[0046] An embodiment of the present invention provides a context-based question-answering processing method, apparatus, electronic device, and computer-readable storage medium. As can be seen from the above, the embodiment of the present invention performs question-answering processing based on an improved question-answering model, which is composed of a first embedding encoding module, a second embedding encoding module, a first encoder, a second encoder, a relevance processing module, an answer generator, a third embedding encoding module, a third encoder, a correct answer probability prediction model, and an answer sorting module; among them, the first embedding encoding module + the first encoder are responsible for splitting the context text into sentences and performing one-time encoding on each sentence; the second embedding encoding module + the second encoder are responsible for performing one-time encoding on the question text; the relevance processing module not only associates the similarity between each sentence in the context and the question text based on the context encoding tensor and the question encoding tensor, but also clusters multiple sentences in the context with approximate similarity to the question text and further associates the similarity between each clustering vector and the question text; the answer generator generates corresponding answer texts according to various association combinations output by the relevance processing module; the third embedding encoding module + the third encoder are responsible for performing one-time encoding on each answer text; the correct answer probability prediction model evaluates the correctness of each answer text output by the answer generator according to various association combinations output by the relevance processing module and each answer feature encoding output by the third encoder; the answer sorting module sorts the correct answer texts according to the evaluation results of the correct answer probability prediction model. Through the improvement of the question-answering model in the embodiment of the present invention, on the one hand, the encoding cycle is shortened and the question-answering efficiency is improved. On the other hand, multiple sentences in the context can be used as references to generate answer texts, improving the question-answering accuracy. On the other hand, the correctness of the generated answer texts can be evaluated and the wrong answers can be deleted, further improving the question-answering accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of a context-based question-answering processing method provided in Embodiment 1 of the present invention;
[0048] Figure 2 It is a module structure diagram of the question-answering model provided in Embodiment 1 of the present invention;
[0049] Figure 3 It is a module structure diagram of a context-based question-answering processing apparatus provided in Embodiment 2 of the present invention;
[0050] Figure 4 It is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0052] Embodiment 1 of the present invention provides a context-based question-and-answer processing method, as Figure 1 shown in the schematic diagram of a context-based question-and-answer processing method provided in Embodiment 1 of the present invention, this method mainly includes the following steps:
[0053] Step 1, receive the context text, the question text, and the preferred number of answers as the corresponding first context, first question, and first preferred number;
[0054] Among them, the first preferred number is an integer greater than or equal to 0.
[0055] Here, the context text of the embodiment of the present invention is a text object with punctuation marks, which can be one or more articles, one or more paragraphs, or one or more sentences. The question text is usually one or more sentences under normal circumstances.
[0056] Step 2, input the first context and the first question into the question-and-answer model, and the question-and-answer model predicts the answer to the first question based on the first context to obtain the corresponding first predicted answer sequence;
[0057] Here, the question-and-answer model of the embodiment of the present invention is as Figure 2 shown in the module structure diagram of the question-and-answer model provided in Embodiment 1 of the present invention, and includes: a first embedding and encoding module, a second embedding and encoding module, a first encoder, a second encoder, a relevance processing module, an answer generator, a third embedding and encoding module, a third encoder, a correct answer probability prediction model, and an answer sorting module;
[0058] The model input end of the question-and-answer model in the embodiment of the present invention has two, namely the first and second model input ends; the model output end of the question-and-answer model is the first model input end; the first input end is used to receive the first context, the second model input end is used to receive the first question, and the first model input end is used to output the first predicted answer sequence;
[0059] The connection relationship of each module of the question-and-answer model in the embodiment of the present invention is: as Figure 2As shown, the input end of the first embedding and encoding module is connected to the input end of the first model, and the output end is connected to the input end of the first encoder; the input end of the second embedding and encoding module is connected to the input end of the second model, and the output end is connected to the input end of the second encoder; the output end of the first encoder is connected to the first input end of the correlation processing module; the output end of the second encoder is connected to the second input end of the correlation processing module; the first output end of the correlation processing module is connected to the input end of the answer generator, and the second output end is connected to the second input end of the correct answer probability prediction model; the first output end of the answer generator is connected to the third embedding and encoding module, and the second output end is connected to the second input end of the answer sorting module; the output end of the third embedding and encoding module is connected to the input end of the third encoder; the output end of the third encoder is connected to the first input end of the correct answer probability prediction model; the output end of the correct answer probability prediction model is connected to the first input end of the answer sorting module;
[0060] It should be noted that the first, second, and third encoders in the embodiments of the present invention are implemented based on the same encoder model; the encoder model at least includes the encoder of the Transformer model; the embedding and encoding (Embedding) mechanisms of the first, second, and third embedding and encoding modules in the embodiments of the present invention are implemented based on the same text embedding and encoding mechanism; the text embedding and encoding mechanism at least includes the one-hot encoding mechanism and the Word2Vec encoding mechanism; the answer generator in the embodiments of the present invention is implemented based on a type of LM model for text generation; the type of LM model for text generation at least includes the RNN model, the LSTM model, the SeqGAN model, the Transformer model, the GPT model, the BERT model, the GPT-2 model, and the BART model; the correct answer probability prediction model in the embodiments of the present invention is composed of a feature fusion unit and a binary classification model; the binary classification model includes a binary classifier implemented based on the MLP structure and a binary classifier implemented based on the CNN structure;
[0061] The main execution steps of step 2 specifically include:
[0062] Step 21, the first embedding and encoding module performs text sentence segmentation on the first context to obtain the corresponding first sentence sequence TC{tc i}; and performs embedding and encoding on each first sentence tc i to obtain the corresponding first text vector c i ; and sorts all the obtained first text vectors c i to form the corresponding first text vector sequence C and send it to the first encoder;
[0063] Among them, the first sentence sequence TC includes multiple first sentences tc i , and the sentence index i is an integer greater than 0;
[0064] Step 22, and the second embedding encoding module performs embedding encoding on the first question to obtain the corresponding second text vector q and sends it to the second encoder;
[0065] Step 23, and the first encoder performs feature encoding on the first text vector sequence C to obtain the corresponding first encoded tensor sequence X C {x c,i} and sends it to the correlation processing module;
[0066] Among them, the first encoded tensor sequence X C includes multiple first encoded tensors x c,i ;
[0067] Step 24, and the second encoder performs feature encoding on the second text vector q to obtain the corresponding second encoded tensor x q and sends it to the correlation processing module;
[0068] Step 25, and the correlation processing module performs correlation prediction according to the first encoded tensor sequence X C and the second encoded tensor x q to obtain the corresponding first correlation tensor sequence R; and sends the first correlation tensor sequence R to the answer generator and the correct answer probability prediction model respectively;
[0069] Among them, the first correlation tensor sequence R includes multiple first correlation tensors r j , the correlation tensor index j is an integer greater than 0; the first correlation tensor r j includes the first tensor y j , the second encoded tensor x q and the first correlation coefficient s j ; the first tensor y j is composed of one or more first encoded tensors x c,i ;
[0070] The correlation processing module performs correlation prediction according to the first encoded tensor sequence X C and the second encoded tensor x q to obtain the corresponding first correlation tensor sequence R, specifically including:
[0071] Step A1, the correlation processing module calculates the similarity between each first encoded tensor x C of the first encoded tensor sequence X c,i and the second encoded tensor x q to obtain the corresponding first similarity d i ;
[0072] Step A2, and each first encoded tensor x c,iAs a corresponding first tensor y j ; and for each first encoded tensor x c,i the corresponding first similarity d i is used as a corresponding first correlation coefficient s j ; and from each first encoded tensor x c,i the corresponding first tensor y j and the first correlation coefficient s j and the second encoded tensor x q to form a corresponding first correlation tensor r j ;
[0073] Step A3, and based on a preset piecewise similarity clustering principle, cluster all the first similarities d i to obtain multiple first clustering sets;
[0074] Among them, the first clustering set includes multiple first similarities d i ; each first clustering set corresponds to a preset similarity range, and all the first similarities d i in each first clustering set satisfy the corresponding similarity range;
[0075] Here, the piecewise similarity clustering principle of the embodiments of the present invention is to preset multiple similarity ranges, and use each similarity range as a clustering basis to cluster all the first similarities d i , that is, cluster multiple first similarities d i that satisfy each similarity range into one category; it should be noted that the number of first similarities d i in the first clustering set must be greater than 1, that is, if there is only one first similarity d i that satisfies a certain similarity range, the corresponding first clustering set is not generated;
[0076] Step A4, and calculate the mean of all the first similarities d i in each first clustering set to obtain the corresponding first clustering similarity; and from all the first similarities d i in each first clustering set, the corresponding all first encoded tensors x c,i form a corresponding first tensor y j ; and use the first clustering similarity corresponding to each first clustering set as a corresponding first correlation coefficient s j ; and from the first tensor y j corresponding to each first clustering set and the first correlation coefficient s j and the second encoded tensor x q to form a corresponding first correlation tensor r j ;
[0077] Step A5, and all the obtained first correlation tensors r j form the corresponding first correlation tensor sequence R;
[0078] Step 26, and the answer generator generates the answer text based on each first correlation tensor r in the first correlation tensor sequence R j of the first tensor y j and the second encoded tensor x q to obtain the corresponding first answer w j ; and all the obtained first answers w j form the corresponding first answer sequence W; and the first answer sequence W is sent to the third embedding encoding module and the answer sorting module respectively;
[0079] Step 27, and the third embedding encoding module performs embedding encoding on each first answer w in the first answer sequence W j to obtain the corresponding third text vector a j ; and all the obtained third text vectors a j are sorted to form the corresponding third text vector sequence A and sent to the third encoder;
[0080] Step 28, and the third encoder performs feature encoding on the third text vector sequence A to obtain the corresponding third encoded tensor sequence X A {x a,j} and send it to the correct answer probability prediction model;
[0081] Among them, the third encoded tensor sequence X A includes multiple third encoded tensors x a,j ;
[0082] Step 29, and the feature fusion unit of the correct answer probability prediction model fuses the features of the third encoded tensor sequence X A and the first correlation tensor sequence R to obtain the corresponding first fusion feature tensor M; and the binary classification model of the correct answer probability prediction model performs binary classification prediction of correct / wrong answers based on each first feature tensor m of the first fusion feature tensor M j to obtain the corresponding first prediction vector; and identify whether the first correct answer probability of each first prediction vector is greater than the first wrong answer probability. If so, set the corresponding first prediction probability z j as the first correct answer probability, and if not, set the corresponding first prediction probability z j as 0; and all the obtained first prediction probabilities z j form the corresponding first prediction probability vector Z and send it to the answer sorting module;
[0083] Among them, the first fusion feature tensor M consists of multiple first feature tensors m jComposition; the first feature tensor m j Composed of the corresponding third encoding tensor x a,j And the first correlation tensor r j Composition; the first prediction vector includes the first correct answer probability and the first wrong answer probability;
[0084] Step 30, and the answer sorting module marks the first prediction probability z with a probability value of 0 in the first prediction probability vector Z j All are marked as screening probabilities; and the first answer w corresponding to each screening probability in the first answer sequence W j Delete; and each remaining first answer w in the first answer sequence W j As the corresponding first predicted answer; and all the first predicted answers are sorted in descending order according to the corresponding first prediction probability z j Sorted from high to low to form the corresponding first predicted answer sequence; and the first predicted answer sequence is output as the prediction result of the Q&A model.
[0085] Step 3, perform Q&A result screening according to the first preferred quantity and the first predicted answer sequence to generate the corresponding first screening result;
[0086] Specifically include: Step 31, count the number of the first predicted answers in the first predicted answer sequence to obtain the corresponding first quantity;
[0087] Step 32, identify the first preferred quantity and the first quantity;
[0088] Step 33, if the first preferred quantity is equal to 0 or the first quantity is less than the first preferred quantity, then use the first predicted answer sequence as the corresponding first screening result;
[0089] Step 34, if the first quantity is greater than or equal to the first preferred quantity, then extract the first preferred quantity of the first predicted answers with the highest ranking in the first predicted answer sequence to form the corresponding first screening result.
[0090] Figure 3 This is the module structure diagram of a context-based Q&A processing device provided in the second embodiment of the present invention. The device is a terminal device or a server for implementing the foregoing method embodiment, or can also be a device that enables the foregoing terminal device or server to implement the foregoing method embodiment. For example, the device can be a device or chip system of the foregoing terminal device or server. As Figure 3 Shown, the device includes: a data receiving module 201, a Q&A model processing module 202, and a data output module 203.
[0091] The data receiving module 201 is configured to receive the context text, the question text, and the preferred number of answers as the corresponding first context, first question, and first preferred number; the first preferred number is an integer greater than or equal to 0.
[0092] The question and answer model processing module 202 is configured to input the first context and the first question into the question and answer model, and the question and answer model predicts the answer to the first question based on the first context to obtain the corresponding first predicted answer sequence.
[0093] The data output module 203 is configured to perform question and answer result screening according to the first preferred number and the first predicted answer sequence to generate the corresponding first screening result.
[0094] A context-based question and answer processing device provided by an embodiment of the present invention can execute the method steps in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0095] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; they can also be partially implemented in the form of software called by a processing element and partially implemented in the form of hardware. For example, the data receiving module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above determined modules. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.
[0096] For example, the above modules may be one or more integrated circuits configured to implement the above methods. For example: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain above module is implemented in the form of a processing element scheduler code, the processing element may be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a System-on-a-chip (SOC).
[0097] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The above available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0098] Figure 4 A schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be a terminal device or a server that implements the method of the foregoing embodiments, or may be a terminal device or a server that is connected to the foregoing terminal device or server and implements the method of the foregoing embodiments. As Figure 4As shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the method of the foregoing embodiments. Preferably, the electronic device according to the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.
[0099] As mentioned in Figure 4 the system bus 305 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0100] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0101] It should be noted that the embodiment of the present invention further provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is enabled to execute the methods and processing procedures provided in the foregoing embodiments.
[0102] The embodiment of the present invention further provides a chip for running instructions, and the chip is used to execute the processing steps described in the foregoing method embodiments.
[0103] An embodiment of the present invention provides a context-based question-answering processing method, apparatus, electronic device, and computer-readable storage medium; as can be seen from the above, the embodiment of the present invention performs question-answering processing based on an improved question-answering model, which is composed of a first embedding encoding module, a second embedding encoding module, a first encoder, a second encoder, a relevance processing module, an answer generator, a third embedding encoding module, a third encoder, a correct answer probability prediction model, and an answer sorting module; among them, the first embedding encoding module + the first encoder are responsible for splitting the context text into sentences and performing a one-time encoding on each sentence; the second embedding encoding module + the second encoder are responsible for performing a one-time encoding on the question text; the relevance processing module not only associates the similarity between each sentence in the context and the question text based on the context encoding tensor and the question encoding tensor, but also clusters multiple sentences in the context with approximate similarity to the question text and further associates the similarity between each cluster vector and the question text; the answer generator then generates corresponding answer texts according to various association combinations output by the relevance processing module; the third embedding encoding module + the third encoder are responsible for performing a one-time encoding on each answer text; the correct answer probability prediction model evaluates the correctness of each answer text output by the answer generator according to various association combinations output by the relevance processing module and each answer feature encoding output by the third encoder; the answer sorting module then sorts the correct answer texts according to the evaluation results of the correct answer probability prediction model. Through the improvement of the question-answering model in the embodiment of the present invention, on the one hand, the encoding cycle is shortened and the question-answering efficiency is improved, on the other hand, multiple sentences in the context can be used as a reference to generate answer texts, improving the question-answering accuracy, and on the other hand, the correctness of the generated answer texts can be evaluated and the wrong answers can be deleted, further improving the question-answering accuracy.
[0104] Those skilled in the art should also be able to further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0105] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well known in the art.
[0106] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A context-based question and answer processing method, characterized in that, The method includes: Receiving context text, question text, and the preferred number of answers as the corresponding first context, first question, and first preferred number; the first preferred number is an integer greater than or equal to 0; Inputting the first context and the first question into a question-answering model, and the question-answering model predicts the answer to the first question based on the first context to obtain a corresponding first predicted answer sequence; Performing question-answering result screening according to the first preferred number and the first predicted answer sequence to generate a corresponding first screening result.
2. The context-based question and answer processing method according to claim 1, characterized in that, The question-answering model includes a first embedding encoding module, a second embedding encoding module, a first encoder, a second encoder, a relevance processing module, an answer generator, a third embedding encoding module, a third encoder, a correct answer probability prediction model, and an answer ranking module; The model has two input ends, namely the first and second model input ends; the model output end of the question-answering model is the first model input end; the first input end is used to receive the first context, the second model input end is used to receive the first question, and the first model input end is used to output the first predicted answer sequence; The input end of the first embedding encoding module is connected to the first model input end, and the output end is connected to the input end of the first encoder; the input end of the second embedding encoding module is connected to the second model input end, and the output end is connected to the input end of the second encoder; the output end of the first encoder is connected to the first input end of the relevance processing module; the output end of the second encoder is connected to the second input end of the relevance processing module; the first output end of the relevance processing module is connected to the input end of the answer generator, and the second output end is connected to the second input end of the correct answer probability prediction model; the first output end of the answer generator is connected to the third embedding encoding module, and the second output end is connected to the second input end of the answer ranking module; the output end of the third embedding encoding module is connected to the input end of the third encoder; the output end of the third encoder is connected to the first input end of the correct answer probability prediction model; the output end of the correct answer probability prediction model is connected to the first input end of the answer ranking module.
3. The context-based question and answer processing method according to claim 2, characterized in that, The first, second, and third encoders are implemented based on the same encoder model; the encoder model at least includes the encoder of the Transformer model; The embedding encoding mechanisms of the first, second, and third embedding encoding modules are implemented based on the same text embedding encoding mechanism; the text embedding encoding mechanism at least includes one-hot encoding mechanism and Word2Vec encoding mechanism; The answer generator is implemented based on a type of LM model for text generation; the LM model for text generation includes RNN model, LSTM model, SeqGAN model, Transformer model, GPT model, BERT model, GPT-2 model, and BART model; The correct answer probability prediction model consists of a feature fusion unit and a binary classification model; the binary classification model includes a binary classifier implemented based on the MLP structure and a binary classifier implemented based on the CNN structure.
4. The context-based question and answer processing method according to claim 2, characterized in that, The answer prediction of the first question by the question-answering model according to the first context to obtain a corresponding first predicted answer sequence specifically includes: The first embedding encoding module performs text segmentation on the first context to obtain a corresponding first sentence sequence TC{tc i}; and performs embedding encoding on each first sentence tc i to obtain a corresponding first text vector c i ; and all the obtained first text vectors c i are sorted to form a corresponding first text vector sequence C and sent to the first encoder; the first sentence sequence TC includes a plurality of the first sentences tc i , and the sentence index i is an integer greater than 0; And the second embedding encoding module performs embedding encoding on the first question to obtain a corresponding second text vector q and sends it to the second encoder; And the first encoder performs feature encoding on the first text vector sequence C to obtain a corresponding first encoded tensor sequence X C {x c,i} is sent to the correlation processing module; the first encoded tensor sequence X C includes a plurality of first encoded tensors x c,i ; And the second encoder performs feature encoding on the second text vector q to obtain a corresponding second encoded tensor x q Send it to the correlation processing module; and the correlation processing module predicts the correlation based on the first encoded tensor sequence X C and the second encoded tensor x q to obtain the corresponding first correlation tensor sequence R; and sends the first correlation tensor sequence R to the answer generator and the correct answer probability prediction model respectively; the first correlation tensor sequence R includes a plurality of first correlation tensors r j , and the correlation tensor index j is an integer greater than 0; the first correlation tensor r j includes the first tensor y j , the second encoded tensor x q and the first correlation coefficient s j ; the first tensor y j is composed of one or more of the first encoded tensors x c,i ; and the answer generator generates a corresponding first answer w j from the first tensor y j of each of the first association tensors r in the first association tensor sequence R q and the second encoded tensor x j ; and all the obtained first answers w j form a corresponding first answer sequence W; and the first answer sequence W is sent to the third embedding encoding module and the answer sorting module respectively; And each of the first answers w in the first answer sequence W is embedded and encoded by the third embedding and encoding module j to obtain the corresponding third text vector a j ; and all the obtained third text vectors a j are sorted to form the corresponding third text vector sequence A and sent to the third encoder; And the third encoder performs feature encoding on the third text vector sequence A to obtain a corresponding third encoded tensor sequence X A {x a,j} is sent to the correct answer probability prediction model; the third encoded tensor sequence X A includes a plurality of third encoded tensors x a,j ; and the feature fusion unit of the correct answer probability prediction model fuses the features of the third encoded tensor sequence X A and the first correlation tensor sequence R to obtain a corresponding first fused feature tensor M; and the binary classification model of the correct answer probability prediction model is based on each first feature tensor m of the first fused feature tensor M j to perform binary classification prediction of correct / wrong answers to obtain a corresponding first prediction vector; and identify whether the first correct answer probability of each first prediction vector is greater than the first wrong answer probability, and if so, set the corresponding first prediction probability z j as the first correct answer probability, and if not, set the corresponding first prediction probability z j to 0; and all the obtained first prediction probabilities z j form a corresponding first prediction probability vector Z and send it to the answer sorting module; the first fused feature tensor M is composed of multiple first feature tensors m j ; the first feature tensor m j is composed of the corresponding third encoded tensor x a,j and the first correlation tensor r j ; the first prediction vector includes the first correct answer probability and the first wrong answer probability; And the answer sorting module marks the first prediction probabilities \(z\) with a probability value of 0 in the first prediction probability vector \(Z\) as screening probabilities; and deletes the first answers \(w\) in the first answer sequence \(W\) corresponding to the respective screening probabilities; and takes the remaining first answers \(w\) in the first answer sequence \(W\) as the corresponding first predicted answers; and sorts all the first predicted answers according to the corresponding first prediction probabilities \(z\) in descending order to form the corresponding first predicted answer sequence; and outputs the first predicted answer sequence as the prediction result of the question-answering model. j And deletes the first answer \(w\) in the first answer sequence \(W\) corresponding to each of the screening probabilities; j And takes the remaining first answers \(w\) in the first answer sequence \(W\) j As the corresponding first predicted answers; and sorts all the first predicted answers according to the corresponding first prediction probabilities \(z\) j In descending order to form the corresponding first predicted answer sequence; and outputs the first predicted answer sequence as the prediction result of the question-answering model.
5. The context-based question and answer processing method according to claim 4, characterized in that, The correlation processing module performs correlation prediction according to the first encoded tensor sequence X C and the second encoded tensor x q to obtain a corresponding first correlation tensor sequence R, which specifically includes: The correlation processing module calculates the similarity between each of the first encoded tensors x C in the first encoded tensor sequence X c,i and the second encoded tensor x q to obtain the corresponding first similarity d i ; and take each of the first encoded tensors x c,i as a corresponding first tensor y j ; and take each of the first encoded tensors x c,i corresponding first similarity d i as a corresponding first correlation coefficient s j ; and from each of the first encoded tensors x c,i corresponding first tensor y j and the first correlation coefficient s j and the second encoded tensor x q form a corresponding first correlation tensor r j ; and perform clustering on all the first similarities d based on a preset principle of segment similarity clustering i to obtain multiple first clustering sets; the first clustering sets include multiple of the first similarities d i ; each of the first clustering sets corresponds to a preset similarity range, and all of the first similarities d in each of the first clustering sets i satisfy the corresponding similarity range; And calculate the mean of all the first similarities d of each of the first clustering sets to obtain the corresponding first clustering similarity; and from all the first similarities d of each of the first clustering sets i calculate the corresponding first clustering similarity; and from all the first similarities d of each of the first clustering sets i all the corresponding first encoding tensors x c,i form a corresponding first tensor y j ; and use the first clustering similarity corresponding to each of the first clustering sets as a corresponding first correlation coefficient s j ; and from the first tensor y corresponding to each of the first clustering sets j and the first correlation coefficient s j and the second encoding tensor x q form a corresponding first association tensor r j ; and all of the obtained first correlation tensors r j constitute the corresponding first correlation tensor sequence R.
6. The context-based question and answer processing method according to claim 1, characterized in that, The screening of the question-answering result according to the first preferred quantity and the first predicted answer sequence to generate a corresponding first screening result specifically includes: Count the number of the first predicted answers in the first predicted answer sequence to obtain a corresponding first quantity; Identify the first preferred quantity and the first quantity; If the first preferred quantity is equal to 0 or the first quantity is less than the first preferred quantity, then use the first predicted answer sequence as the corresponding first screening result; If the first quantity is greater than or equal to the first preferred quantity, then extract the first preferred quantity of the first predicted answers with the earliest ranking in the first predicted answer sequence to form the corresponding first screening result.
7. An apparatus for performing the context-based question-answering processing method according to any one of claims 1-6, characterized in that, The device includes: a data receiving module, a question-answering model processing module, and a data output module; The data receiving module is used to receive the context text, the question text, and the answer preferred quantity as the corresponding first context, first question, and first preferred quantity; the first preferred quantity is an integer greater than or equal to 0; The question-answering model processing module is used to input the first context and the first question into the question-answering model, and the question-answering model predicts the answer to the first question according to the first context to obtain a corresponding first predicted answer sequence; The data output module is used to screen the question-answering result according to the first preferred quantity and the first predicted answer sequence to generate a corresponding first screening result.
8. An electronic device, characterized in that, Including: A memory, a processor, and a transceiver; The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method according to any one of claims 1-6; The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by the computer, the computer is caused to execute the method according to any one of claims 1-6.