Answer generation methods, devices, equipment, and storage media

By using an answer recognition model for encoding, perception, and decoding in an automated question-answering system, and combining the question type and historical answers to generate input text, the problem of existing systems being unable to generate reasonable answers is solved, thus achieving accuracy and completeness of the answers.

CN116628161BActive Publication Date: 2025-11-14PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310593894.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-11-14
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

Existing automated question-answering systems are unable to generate accurate and reasonable answers to questions, especially in the medical field when it comes to counting questions, they cannot generate reasonable answer formats.

Method used

The input text is generated by acquiring the question to be tested, the question type, the evidence text, and the historical answers. The pre-trained answer recognition model is used for encoding, perception, and decoding. The answer is generated by combining the first output probability and the second output probability. The model includes multiple encoders, perceptual network layers, and prediction output layers.

Benefits of technology

It improves the accuracy and rationality of the generated answers, ensuring that the answers match the question type and avoiding the generation of incorrect or redundant information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116628161B_ABST
    Figure CN116628161B_ABST
Patent Text Reader

Abstract

This invention relates to the fields of artificial intelligence and digital healthcare, providing a method, apparatus, device, and storage medium for generating answers. The method generates input text based on the question to be tested, the question type, evidence text, and historical answers. The input text is encoded to obtain a text code. The text code undergoes perceptual processing to obtain a perceptual vector, which is then normalized to obtain a first output probability. Based on the text code, the perceptual vector and the input text are decoded to obtain decoded information. The decoded information is then predicted to obtain a second output probability. Finally, the answer to the question is generated based on the first and second output probabilities. This method utilizes a neural network to accurately output the answer to the question. Furthermore, this invention also relates to blockchain technology, allowing the answer to the question to be stored in the blockchain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and digital medical technology, and in particular to an answer generation method, apparatus, device and storage medium. Background Technology

[0002] With the development of artificial intelligence, automated question-answering systems have also evolved, and current automated question-answering systems can be applied to the medical field. In existing automated question-answering systems, the corresponding answer is usually retrieved directly from the evidence text based on the question. This method cannot provide an answer format that meets the requirements of the question. For example, for a counting question, the evidence text may only contain enumeration descriptions; therefore, current automated question-answering systems cannot generate reasonable answer information.

[0003] Therefore, how to accurately and reasonably generate the answer text corresponding to the question has become a technical problem that urgently needs to be solved. Summary of the Invention

[0004] In view of the above, it is necessary to provide an answer generation method, apparatus, device, and storage medium that can solve the technical problem of how to accurately and reasonably generate the answer text corresponding to the question.

[0005] On the one hand, the present invention proposes an answer generation method, the answer generation method comprising:

[0006] The input text is generated based on the obtained question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to the historical questions associated with the question to be tested;

[0007] Obtain a pre-trained answer recognition model, which includes multiple encoders, a perceptual network layer, multiple decoders, and a prediction output layer;

[0008] The input text is encoded using the multiple encoders to obtain the text encoding.

[0009] Based on the perceptual network layer, the text encoding is perceptually processed to obtain a perceptual vector, and the perceptual vector is normalized to obtain the first output probability of each text word in the input text.

[0010] Based on the multiple decoders and the text encoding, the perceptual vector and the input text are decoded to obtain decoded information;

[0011] Based on the prediction output layer, the decoded information is predicted to obtain the second output probability of each template word in the plurality of decoders;

[0012] The answer to the question to be tested is generated based on the first output probability and the second output probability.

[0013] According to a preferred embodiment of the present invention, the step of generating input text based on the obtained question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to historical questions associated with the question to be tested includes:

[0014] Extract interrogative words from the question to be tested;

[0015] The question type is identified by matching the question words with preset words;

[0016] Identify the generation time and issue session of the problem to be tested;

[0017] The requested question is obtained based on the generation time and the question session, and the request time of the requested question is also obtained.

[0018] The historical questions are selected from the requested questions based on the requested time.

[0019] Obtain the text of the answer corresponding to the historical question as the historical answer;

[0020] The input text is a combination of the question type, the question to be tested, the preset identifier, the evidence text, and the historical answers.

[0021] According to a preferred embodiment of the present invention, the step of encoding the input text based on the plurality of encoders to obtain text encoding includes:

[0022] Each text word in the input text is represented to obtain a text vector;

[0023] For any encoder, a focus vector is generated based on each set of weight matrices for the text vector;

[0024] Multiple attention vectors are concatenated to obtain a concatenated vector;

[0025] The attention information is obtained by calculating the product of the configuration matrix and the concatenated vector;

[0026] Based on the preset matrix and preset bias value, a fully connected operation is performed on the attention information to obtain an initial vector. The formula for generating the initial vector is: y = max(a, xW1+1)2+2, where y represents the initial vector, a represents a preset constant, x represents the attention information, W1 and W2 represent the preset matrix, and b1 and b2 represent the preset bias value.

[0027] The initial vector is used as the text vector for the next encoder for encoding, until all the encoders participate in the encoding to obtain the text encoding.

[0028] According to a preferred embodiment of the present invention, the perceptual processing of the text encoding based on the perceptual network layer to obtain the perceptual vector includes:

[0029] For each hidden layer of the perceptual network layer, a fully connected operation is performed on the text encoding based on the network parameters of the neurons in that hidden layer to obtain the output vector of that hidden layer.

[0030] The output vector is used as the text encoding for the next hidden layer for perceptual processing until each hidden layer of the perceptual network layer participates in the processing, thus obtaining the perceptual vector.

[0031] According to a preferred embodiment of the present invention, the step of decoding the perceptual vector and the input text based on the plurality of decoders and the text encoding to obtain decoded information includes:

[0032] The input text is masked to obtain a mask vector;

[0033] Generate a target vector based on the perception vector and the mask vector;

[0034] For any decoder, attention analysis is performed on the target vector based on the decoder and the text encoding to obtain attention information;

[0035] The attention information is fully connected based on the decoding weight matrix and decoding bias of any of the decoders to obtain the initial decoding. The initial decoding is then used as the target vector for the next decoder to perform decoding, until all the decoders participate in the decoding to obtain the decoding information.

[0036] According to a preferred embodiment of the present invention, the step of predicting the decoded information based on the prediction output layer to obtain the second output probability of each template word in the plurality of decoders includes:

[0037] The decoded information is activated based on the activation function of the prediction output layer to obtain activation information;

[0038] The activation information is normalized to obtain the second output probability.

[0039] According to a preferred embodiment of the present invention, generating the answer to the test question based on the first output probability and the second output probability includes:

[0040] The first output probability and the second output probability are weighted and summed to obtain the target probabilities of the text words and the template words.

[0041] Words whose target probability is greater than a preset probability threshold are identified as target words;

[0042] Generate the answer to the question based on the target vocabulary.

[0043] On the other hand, the present invention also proposes an answer generation device, the answer generation device comprising:

[0044] The generation unit is used to generate input text based on the obtained question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to historical questions associated with the question to be tested;

[0045] The acquisition unit is used to acquire a pre-trained answer recognition model, which includes multiple encoders, a perceptual network layer, multiple decoders, and a prediction output layer.

[0046] An encoding unit is used to encode the input text based on the plurality of encoders to obtain text encoding;

[0047] A perception unit is used to perform perception processing on the text encoding based on the perception network layer to obtain a perception vector, and to normalize the perception vector to obtain the first output probability of each text word in the input text.

[0048] A decoding unit is used to decode the perceptual vector and the input text based on the plurality of decoders and the text encoding to obtain decoded information;

[0049] A prediction unit is used to predict the decoded information based on the prediction output layer to obtain a second output probability for each template word in the plurality of decoders;

[0050] The generation unit is further configured to generate the answer to the question to be tested based on the first output probability and the second output probability.

[0051] On the other hand, the present invention also proposes an electronic device, the electronic device comprising:

[0052] Memory, which stores computer-readable instructions; and

[0053] The processor executes computer-readable instructions stored in the memory to implement the answer generation method.

[0054] On the other hand, the present invention also proposes a computer-readable storage medium storing computer-readable instructions, which are executed by a processor in an electronic device to implement the answer generation method.

[0055] As can be seen from the above technical solutions, this application generates input text by combining the question to be tested, the question type, the evidence text, and the historical answers. Since different types of questions correspond to different answer formats, adding the question type to the input text can improve the accuracy and rationality of the generated answer. Adding the historical answers to the input text not only makes the subsequently generated answer more complete but also guides the answer recognition model to conform to the answer corresponding to the question type. By setting the multiple encoders in the answer recognition model, the text encoding can focus on semantic information of different granularities from shallow to deep layers. Thus, the first output probability and the second output probability can accurately generate the question answer, thereby avoiding errors in the generation of the question answer or the generation of redundant information. Attached Figure Description

[0056] Figure 1 This is a flowchart of a preferred embodiment of the answer generation method of the present invention.

[0057] Figure 2 This is a network structure diagram of the answer recognition model in the answer generation method of this invention.

[0058] Figure 3 This is a functional block diagram of a preferred embodiment of the answer generation device of the present invention.

[0059] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the answer generation method of the present invention. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the answer generation method of the present invention. Depending on different needs, the order of the steps in this flowchart can be changed, and some steps can be omitted.

[0062] The answer generation method described above can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0063] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0064] The answer generation method is applied to one or more electronic devices. The electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored computer-readable instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0065] The electronic device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.

[0066] The electronic devices may include network devices and / or user devices. The network devices include, but are not limited to, single network electronic devices, groups of multiple network electronic devices, or cloud computing-based systems consisting of a large number of hosts or network electronic devices.

[0067] The network in which the electronic device is located includes, but is not limited to: the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.

[0068] 101. Generate input text based on the obtained question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to the historical questions associated with the question to be tested.

[0069] In at least one embodiment of the present invention, the question to be tested can be in English or Chinese, and this application makes no limitation thereto. The question to be tested can be a user question in any round of an automated question-and-answer system. For example, the question to be tested can be a counting problem. The question to be tested can be a medical-related question.

[0070] The question type refers to the category corresponding to the question to be tested. For example, the question type can be a yes / no type, etc.

[0071] The evidence text refers to the source text of the answer to the question to be tested. For example, if the question to be tested is a counting problem, the evidence text may include specific enumeration values, etc. The evidence text may be medical text, which may be an electronic healthcare record, an electronic personal health record, including medical records, electrocardiograms, medical images, and other electronic records with archival value.

[0072] The historical question refers to a question that is in the same question-and-answer session as the question to be tested, and the request time of the historical question is less than the generation time of the question to be tested.

[0073] The historical answers refer to the answer information corresponding to the historical questions.

[0074] The input text refers to the text generated by concatenating the question to be tested, the question type, the evidence text, and the historical answers.

[0075] In at least one embodiment of the present invention, the electronic device generates input text based on the acquired question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to historical questions associated with the question to be tested, including:

[0076] Extract interrogative words from the question to be tested;

[0077] The question type is identified by matching the question words with preset words;

[0078] Identify the generation time and issue session of the problem to be tested;

[0079] The requested question is obtained based on the generation time and the question session, and the request time of the requested question is also obtained.

[0080] The historical questions are selected from the requested questions based on the requested time.

[0081] Obtain the text of the answer corresponding to the historical question as the historical answer;

[0082] The input text is a combination of the question type, the question to be tested, the preset identifier, the evidence text, and the historical answers.

[0083] The interrogative words refer to words that express a question in the question to be tested, such as what, how, is, are, etc.

[0084] The preset vocabulary includes, but is not limited to, words corresponding to multiple question-related categories. For example, the preset vocabulary may be what, how, is, are, etc.

[0085] The question type refers to the category corresponding to the preset words that successfully match the question words.

[0086] The generation time refers to the specific point in time when the question to be tested is generated, and the question session can be a specific question and answer session in the automatic question and answer system, for example, the question session can be 1001, etc.

[0087] The question-and-answer session for the requested question is the same as the question session, and the request time for the requested question is less than the generation time.

[0088] The historical question refers to a request question where the requested time falls within the target time period. The target time period can be generated based on the difference between the generated time and the preset time period. For example, if the generated time is 9:00 and the preset time period is [1 min, 5 min], then the target time period is 8:55-8:59.

[0089] The preset identifier can be any identifier that can separate the problem to be tested and the evidence text. For example, the preset identifier can be [SEP], etc.

[0090] The question type can be accurately identified by the interrogative words. By combining the generation time and the question session, the historical questions can be reasonably filtered out. The preset identifier can separate the question to be tested from the evidence text, thereby reducing the recognition difficulty of the model. By combining the question type, the question to be tested, the preset identifier, the evidence text, and the historical answers to generate the input text, not only can the generated question answer conform to the answer format of the question type, but the generated question answer can also be more complete.

[0091] 102. Obtain a pre-trained answer recognition model, which includes multiple encoders, a perceptual network layer, multiple decoders, and a prediction output layer.

[0092] In at least one embodiment of the present invention, the answer recognition model is used to identify the answer to the question to be tested. The answer recognition model can be applied to an intelligent question-answering system, which can be applied to intelligent diagnosis and treatment, and remote consultation.

[0093] The plurality of encoders are used to generate text encodings of the input text.

[0094] The perceptual network layer is used to generate the perceptual vector of the text encoding.

[0095] The plurality of decoders are used to generate decoding information for the text encoding, and the number of the plurality of decoders is equal to the number of the plurality of encoders.

[0096] The prediction output layer is used to generate the probability of the decoded information for each template word.

[0097] In at least one embodiment of the present invention, before obtaining the pre-trained answer recognition model, the method further includes:

[0098] Construct an answer recognition network;

[0099] Obtain training samples;

[0100] The loss value of the training samples on the answer recognition network is calculated based on a preset cross-entropy loss function;

[0101] The preset parameters of the answer recognition network are adjusted based on the loss value until the loss value no longer decreases, thus obtaining the answer recognition model.

[0102] like Figure 2 The diagram shown is a network structure diagram of the answer recognition model in the answer generation method of this invention. Figure 2 The answer recognition model includes encoder 1, encoder 2, perceptual network layer, decoder 1, decoder 2 and prediction output layer.

[0103] 103. The input text is encoded based on the multiple encoders to obtain the text encoding.

[0104] In at least one embodiment of the present invention, the text encoding refers to the vector information output by the encoder with the largest concatenation order among the plurality of encoders based on the input text.

[0105] In at least one embodiment of the present invention, the electronic device encodes the input text based on the plurality of encoders to obtain text encoding including:

[0106] Each text word in the input text is represented to obtain a text vector;

[0107] For any encoder, a focus vector is generated based on each set of weight matrices for the text vector;

[0108] Multiple attention vectors are concatenated to obtain a concatenated vector;

[0109] The attention information is obtained by calculating the product of the configuration matrix and the concatenated vector;

[0110] Based on the preset matrix and preset bias value, a fully connected operation is performed on the attention information to obtain an initial vector. The formula for generating the initial vector is: y = max(a, xW1+1)2+2, where y represents the initial vector, a represents a preset constant, x represents the attention information, W1 and W2 represent the preset matrix, and b1 and b2 represent the preset bias value.

[0111] The initial vector is used as the text vector for the next encoder for encoding, until all the encoders participate in the encoding to obtain the text encoding.

[0112] Among them, multiple text words can be generated by segmenting the input text based on a preset dictionary.

[0113] Each set of weight matrices may include multiple matrices. The attention vector refers to the vector generated by performing attention analysis on the text vector based on the multiple matrices.

[0114] Multiple sets of the weight matrices, configuration matrices, preset matrices, and preset bias values ​​can be obtained from the model parameters of the answer recognition model.

[0115] Through the above implementation, in each encoder, attention analysis is performed on the text vector simultaneously using multiple sets of weight matrices, which can improve the accuracy of the attention information in representing the input text. By combining the preset matrix and the preset bias value to analyze the attention information, the representation ability of the initial vector is further improved, thereby improving the representation ability of the text encoding.

[0116] 104. Based on the perceptual network layer, perform perceptual processing on the text encoding to obtain a perceptual vector, and normalize the perceptual vector to obtain the first output probability of each text word in the input text.

[0117] In at least one embodiment of the present invention, the perception vector includes a representation vector for each text word in the input text.

[0118] The first output probability refers to the probability generated by combining the multiple encoders and the perceptual network layer for the multiple text words.

[0119] In at least one embodiment of the present invention, the electronic device performs perceptual processing on the text encoding based on the perceptual network layer to obtain a perceptual vector, including:

[0120] For each hidden layer of the perceptual network layer, a fully connected operation is performed on the text encoding based on the network parameters of the neurons in that hidden layer to obtain the output vector of that hidden layer.

[0121] The output vector is used as the text encoding for the next hidden layer for perceptual processing until each hidden layer of the perceptual network layer participates in the processing, thus obtaining the perceptual vector.

[0122] The network parameters include a parameter matrix and parameter bias values.

[0123] By processing the text encoding through neurons in all hidden layers of the perceptual network, the representational power of the perceptual vector can be improved.

[0124] 105. Based on the multiple decoders and the text encoding, the perceptual vector and the input text are decoded to obtain decoded information.

[0125] In at least one embodiment of the present invention, the decoding information refers to the information output by the decoder with the largest splicing order among the plurality of decoders.

[0126] In at least one embodiment of the present invention, the electronic device decodes the perceptual vector and the input text based on the plurality of decoders and the text encoding to obtain decoded information including:

[0127] The input text is masked to obtain a mask vector;

[0128] Generate a target vector based on the perception vector and the mask vector;

[0129] For any decoder, attention analysis is performed on the target vector based on the decoder and the text encoding to obtain attention information;

[0130] The attention information is fully connected based on the decoding weight matrix and decoding bias of any of the decoders to obtain the initial decoding. The initial decoding is then used as the target vector for the next decoder to perform decoding, until all the decoders participate in the decoding to obtain the decoding information.

[0131] The target vector refers to the vector generated based on the average value of each element in the perception vector and each element in the mask vector.

[0132] The method for generating the initial decoder is similar to the method for generating the initial vector, and will not be described in detail here.

[0133] By masking the input text, the multiple decoders can avoid relying on future information to decode the text encoding. By combining the perceptual vector and the mask vector, the accuracy of the attention analysis of the target vector by the text encoding can be improved, thereby improving the accuracy of the generated decoding information.

[0134] Specifically, the electronic device performs masking processing on the input text to obtain a mask vector including:

[0135] Extract the word vector for each word in the input text from the text vector;

[0136] Count the number of elements in each word vector;

[0137] Based on the number of elements with the largest value, the multiple word vectors are padded to obtain padded vectors.

[0138] Generate a vector matrix based on the multiple completion vectors;

[0139] Identify the diagonal of the vector matrix, and identify the lower triangular elements of the vector matrix based on the diagonal;

[0140] The lower triangular elements are masked, and the resulting matrix is ​​determined as the mask vector.

[0141] The above implementation method not only ensures that the vector length of each text word is the same, but also avoids the multiple decoders from relying on future information to decode the text encoding.

[0142] 106. Based on the prediction output layer, the decoded information is predicted to obtain the second output probability of each template word in the plurality of decoders.

[0143] In at least one embodiment of the present invention, the plurality of template words refer to words pre-configured in the answer recognition model.

[0144] The second output probability refers to the probability generated for each template word based on the prediction output layer and the decoding information.

[0145] In at least one embodiment of the present invention, the electronic device predicts the decoded information based on the prediction output layer to obtain a second output probability for each template word in the plurality of decoders, including:

[0146] The decoded information is activated based on the activation function of the prediction output layer to obtain activation information;

[0147] The activation information is normalized to obtain the second output probability.

[0148] Through the above implementation method, the decoding information can be accurately mapped to multiple template words, thereby improving the accuracy of the second output probability.

[0149] 107. Generate the answer to the question to be tested based on the first output probability and the second output probability.

[0150] It should be emphasized that, to further ensure the privacy and security of the answers to the above questions, the answers can also be stored in a node of a blockchain.

[0151] In at least one embodiment of the present invention, the question answer refers to the answer text corresponding to the question to be tested. When the question to be tested is a question entered by a user in an automatic question-and-answer system, the question answer may be the answer output by the automatic question-and-answer system.

[0152] In at least one embodiment of the present invention, the electronic device generates the answer to the test question based on the first output probability and the second output probability, including:

[0153] The first output probability and the second output probability are weighted and summed to obtain the target probabilities of the text words and the template words.

[0154] Words whose target probability is greater than a preset probability threshold are identified as target words;

[0155] Generate the answer to the question based on the target vocabulary.

[0156] The preset probability threshold can be set according to actual needs.

[0157] By combining the first output probability and the second output probability, the target probability can be accurately generated, thereby improving the accuracy of the answer to the question.

[0158] As can be seen from the above technical solutions, this application generates input text by combining the question to be tested, the question type, the evidence text, and the historical answers. Since different types of questions correspond to different answer formats, adding the question type to the input text can improve the accuracy and rationality of the generated answer. Adding the historical answers to the input text not only makes the subsequently generated answer more complete but also guides the answer recognition model to conform to the answer corresponding to the question type. By setting the multiple encoders in the answer recognition model, the text encoding can focus on semantic information of different granularities from shallow to deep layers. Thus, the first output probability and the second output probability can accurately generate the question answer, thereby avoiding errors in the generation of the question answer or the generation of redundant information.

[0159] like Figure 3 The diagram shown is a functional block diagram of a preferred embodiment of the answer generation device of the present invention. The answer generation device 11 includes a generation unit 110, an acquisition unit 111, an encoding unit 112, a sensing unit 113, a decoding unit 114, a prediction unit 115, a construction unit 116, a calculation unit 117, and an adjustment unit 118. The module / unit referred to in this invention is a series of computer-readable instruction segments that can be acquired by the processor 13 and perform a fixed function, stored in the memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0160] The generation unit 110 is used to generate input text based on the acquired question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to historical questions associated with the question to be tested;

[0161] The acquisition unit 111 is used to acquire a pre-trained answer recognition model, which includes multiple encoders, a perceptual network layer, multiple decoders, and a prediction output layer.

[0162] Encoding unit 112 is used to encode the input text based on the plurality of encoders to obtain text encoding;

[0163] The perception unit 113 is used to perform perception processing on the text encoding based on the perception network layer to obtain a perception vector, and to normalize the perception vector to obtain the first output probability of each text word in the input text.

[0164] Decoding unit 114 is used to decode the perceptual vector and the input text based on the plurality of decoders and the text encoding to obtain decoded information;

[0165] Prediction unit 115 is used to predict the decoded information based on the prediction output layer to obtain the second output probability of each template word in the plurality of decoders;

[0166] The generation unit 110 is further configured to generate the answer to the question to be tested based on the first output probability and the second output probability.

[0167] In at least one embodiment of the present invention, the generation unit 110 is further configured to extract interrogative words from the question to be tested;

[0168] The question type is identified by matching the question words with preset words;

[0169] Identify the generation time and issue session of the problem to be tested;

[0170] The requested question is obtained based on the generation time and the question session, and the request time of the requested question is also obtained.

[0171] The historical questions are selected from the requested questions based on the requested time.

[0172] Obtain the text of the answer corresponding to the historical question as the historical answer;

[0173] The input text is a combination of the question type, the question to be tested, the preset identifier, the evidence text, and the historical answers.

[0174] In at least one embodiment of the present invention, the encoding unit 112 is further configured to represent each text word in the input text to obtain a text vector;

[0175] For any encoder, a focus vector is generated based on each set of weight matrices for the text vector;

[0176] Multiple attention vectors are concatenated to obtain a concatenated vector;

[0177] The attention information is obtained by calculating the product of the configuration matrix and the concatenated vector;

[0178] Based on the preset matrix and preset bias value, a fully connected operation is performed on the attention information to obtain an initial vector. The formula for generating the initial vector is: y = max(a, xW1+1)2+2, where y represents the initial vector, a represents a preset constant, x represents the attention information, W1 and W2 represent the preset matrix, and b1 and b2 represent the preset bias value.

[0179] The initial vector is used as the text vector for the next encoder for encoding, until all the encoders participate in the encoding to obtain the text encoding.

[0180] In at least one embodiment of the present invention, the perception unit 113 is further configured to perform a fully connected operation on the text encoding based on the network parameters of the neurons in each hidden layer of the perception network layer to obtain the output vector of the hidden layer.

[0181] The output vector is used as the text encoding for the next hidden layer for perceptual processing until each hidden layer of the perceptual network layer participates in the processing, thus obtaining the perceptual vector.

[0182] In at least one embodiment of the present invention, the decoding unit 114 is further configured to perform masking processing on the input text to obtain a mask vector;

[0183] Generate a target vector based on the perception vector and the mask vector;

[0184] For any decoder, attention analysis is performed on the target vector based on the decoder and the text encoding to obtain attention information;

[0185] The attention information is fully connected based on the decoding weight matrix and decoding bias of any of the decoders to obtain the initial decoding. The initial decoding is then used as the target vector for the next decoder to perform decoding, until all the decoders participate in the decoding to obtain the decoding information.

[0186] In at least one embodiment of the present invention, the prediction unit 115 is further configured to perform activation processing on the decoded information based on the activation function of the prediction output layer to obtain activation information;

[0187] The activation information is normalized to obtain the second output probability.

[0188] In at least one embodiment of the present invention, the generation unit 110 is further configured to perform a weighted sum operation on the first output probability and the second output probability to obtain the target probabilities of the text words and the template words;

[0189] Words whose target probability is greater than a preset probability threshold are identified as target words;

[0190] Generate the answer to the question based on the target vocabulary.

[0191] In at least one embodiment of the present invention, before obtaining the pre-trained answer recognition model, a construction unit 116 is used to construct an answer recognition network;

[0192] The acquisition unit 111 is also used to acquire training samples;

[0193] The calculation unit 117 is used to calculate the loss value of the training sample on the answer recognition network based on a preset cross-entropy loss function;

[0194] The adjustment unit 118 is used to adjust the preset parameters of the answer recognition network based on the loss value until the loss value no longer decreases, thereby obtaining the answer recognition model.

[0195] As can be seen from the above technical solutions, this application generates input text by combining the question to be tested, the question type, the evidence text, and the historical answers. Since different types of questions correspond to different answer formats, adding the question type to the input text can improve the accuracy and rationality of the generated answer. Adding the historical answers to the input text not only makes the subsequently generated answer more complete but also guides the answer recognition model to conform to the answer corresponding to the question type. By setting the multiple encoders in the answer recognition model, the text encoding can focus on semantic information of different granularities from shallow to deep layers. Thus, the first output probability and the second output probability can accurately generate the question answer, thereby avoiding errors in the generation of the question answer or the generation of redundant information.

[0196] like Figure 4 The diagram shown is a schematic diagram of the structure of an electronic device that implements the answer generation method of the present invention.

[0197] In one embodiment of the present invention, the electronic device 1 includes, but is not limited to, a memory 12, a processor 13, and computer-readable instructions, such as an answer generation program, stored in the memory 12 and executable on the processor 13.

[0198] Those skilled in the art will understand that the schematic diagram is merely an example of electronic device 1 and does not constitute a limitation on electronic device 1. It may include more or fewer components than shown in the diagram, or combine certain components, or different components. For example, electronic device 1 may also include input / output devices, network access devices, buses, etc.

[0199] The processor 13 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor 13 is the computing core and control center of the electronic device 1, connecting various parts of the electronic device 1 through various interfaces and lines, and executing the operating system of the electronic device 1, as well as various installed application programs and program code.

[0200] For example, the computer-readable instructions can be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units can be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer-readable instructions in the electronic device 1. For example, the computer-readable instructions can be divided into a generation unit 110, an acquisition unit 111, an encoding unit 112, a sensing unit 113, a decoding unit 114, a prediction unit 115, a construction unit 116, a calculation unit 117, and an adjustment unit 118.

[0201] The memory 12 can be used to store the computer-readable instructions and / or modules. The processor 13 implements various functions of the electronic device 1 by running or executing the computer-readable instructions and / or modules stored in the memory 12 and calling the data stored in the memory 12. The memory 12 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. The memory 12 may include non-volatile and volatile memory, such as: hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other storage devices.

[0202] The memory 12 can be the external memory and / or internal memory of the electronic device 1. Furthermore, the memory 12 can be a physical memory, such as a memory module, a TF card (Trans-flash Card), etc.

[0203] If the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by instructing related hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when executed by a processor, the computer-readable instructions can implement the steps of the various method embodiments described above.

[0204] The computer-readable instructions include computer-readable instruction code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer-readable instruction code, recording medium, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), and random access memory (RAM).

[0205] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed answer generation, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0206] Combination Figure 1 The memory 12 in the electronic device 1 stores computer-readable instructions to implement an answer generation method, and the processor 13 can execute the computer-readable instructions to achieve the following:

[0207] The input text is generated based on the obtained question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to the historical questions associated with the question to be tested;

[0208] Obtain a pre-trained answer recognition model, which includes multiple encoders, a perceptual network layer, multiple decoders, and a prediction output layer;

[0209] The input text is encoded using the multiple encoders to obtain the text encoding.

[0210] Based on the perceptual network layer, the text encoding is perceptually processed to obtain a perceptual vector, and the perceptual vector is normalized to obtain the first output probability of each text word in the input text.

[0211] Based on the multiple decoders and the text encoding, the perceptual vector and the input text are decoded to obtain decoded information;

[0212] Based on the prediction output layer, the decoded information is predicted to obtain the second output probability of each template word in the plurality of decoders;

[0213] The answer to the question to be tested is generated based on the first output probability and the second output probability.

[0214] Specifically, the specific implementation method of the processor 13 of the above-mentioned computer-readable instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.

[0215] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0216] The computer-readable storage medium stores computer-readable instructions, which, when executed by the processor 13, are used to perform the following steps:

[0217] The input text is generated based on the obtained question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to the historical questions associated with the question to be tested;

[0218] Obtain a pre-trained answer recognition model, which includes multiple encoders, a perceptual network layer, multiple decoders, and a prediction output layer;

[0219] The input text is encoded using the multiple encoders to obtain the text encoding.

[0220] Based on the perceptual network layer, the text encoding is perceptually processed to obtain a perceptual vector, and the perceptual vector is normalized to obtain the first output probability of each text word in the input text.

[0221] Based on the multiple decoders and the text encoding, the perceptual vector and the input text are decoded to obtain decoded information;

[0222] Based on the prediction output layer, the decoded information is predicted to obtain the second output probability of each template word in the plurality of decoders;

[0223] The answer to the question to be tested is generated based on the first output probability and the second output probability.

[0224] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0225] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0226] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.

[0227] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described may also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.

[0228] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for generating answers, characterized in that, The answer generation method includes: The input text is generated based on the obtained question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to the historical questions associated with the question to be tested; Obtain a pre-trained answer recognition model, which includes multiple encoders, a perceptual network layer, multiple decoders, and a prediction output layer; The input text is encoded using the multiple encoders to obtain the text encoding. Based on the perceptual network layer, the text encoding is perceptually processed to obtain a perceptual vector, and the perceptual vector is normalized to obtain the first output probability of each text word in the input text. Based on the multiple decoders and the text encoding, the perceptual vector and the input text are decoded to obtain decoded information, including: masking the input text to obtain a mask vector; generating a target vector based on the perceptual vector and the mask vector; for any decoder, performing attention analysis on the target vector based on the any decoder and the text encoding to obtain attention information; performing fully connected processing on the attention information based on the decoding weight matrix and decoding bias of any decoder to obtain an initial decoder, and using the initial decoder as the target vector for the next decoder to perform decoding processing, until all multiple decoders participate in decoding to obtain the decoded information; Based on the prediction output layer, the decoded information is predicted to obtain the second output probability of each template word in the plurality of decoders; The answer to the question to be tested is generated based on the first output probability and the second output probability.

2. The answer generation method as described in claim 1, characterized in that, The step of generating input text based on the obtained question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to historical questions associated with the question to be tested includes: Extract interrogative words from the question to be tested; The question type is identified by matching the question words with preset words; Identify the generation time and issue session of the problem to be tested; The requested question is obtained based on the generation time and the question session, and the request time of the requested question is also obtained. The historical questions are selected from the requested questions based on the requested time. Obtain the text of the answer corresponding to the historical question as the historical answer; The input text is a combination of the question type, the question to be tested, the preset identifier, the evidence text, and the historical answers.

3. The answer generation method as described in claim 1, characterized in that, The process of encoding the input text based on the multiple encoders to obtain the text encoding includes: Each text word in the input text is represented to obtain a text vector; For any encoder, a focus vector is generated based on each set of weight matrices for the text vector; Multiple attention vectors are concatenated to obtain a concatenated vector; The attention information is obtained by calculating the product of the configuration matrix and the concatenated vector; Based on a preset matrix and a preset bias value, a fully connected operation is performed on the attention information to obtain an initial vector. The formula for generating the initial vector is as follows: ,in, This represents the initial vector. This represents a preset constant. This refers to the attention information. and Represents the preset matrix, and This represents the preset bias value; The initial vector is used as the text vector for the next encoder for encoding, until all the encoders participate in the encoding to obtain the text encoding.

4. The answer generation method as described in claim 1, characterized in that, The perceptual processing of the text encoding based on the perceptual network layer to obtain the perceptual vector includes: For each hidden layer of the perceptual network layer, a fully connected operation is performed on the text encoding based on the network parameters of the neurons in that hidden layer to obtain the output vector of that hidden layer. The output vector is used as the text encoding for the next hidden layer for perceptual processing until each hidden layer of the perceptual network layer participates in the processing, thus obtaining the perceptual vector.

5. The answer generation method as described in claim 1, characterized in that, The step of predicting the decoded information based on the prediction output layer to obtain the second output probability of each template word in the plurality of decoders includes: The decoded information is activated based on the activation function of the prediction output layer to obtain activation information; The activation information is normalized to obtain the second output probability.

6. The answer generation method as described in claim 1, characterized in that, The method of generating the answer to the test question based on the first output probability and the second output probability includes: The first output probability and the second output probability are weighted and summed to obtain the target probabilities of the text words and the template words. Words whose target probability is greater than a preset probability threshold are identified as target words; Generate the answer to the question based on the target vocabulary.

7. An answer generation device, characterized in that, The answer generation device includes: The generation unit is used to generate input text based on the obtained question to be tested, the question type of the question to be tested, the evidence text corresponding to the question to be tested, and the historical answers corresponding to historical questions associated with the question to be tested; The acquisition unit is used to acquire a pre-trained answer recognition model, which includes multiple encoders, a perceptual network layer, multiple decoders, and a prediction output layer. An encoding unit is used to encode the input text based on the plurality of encoders to obtain text encoding; A perception unit is used to perform perception processing on the text encoding based on the perception network layer to obtain a perception vector, and to normalize the perception vector to obtain the first output probability of each text word in the input text. A decoding unit is configured to decode the perceptual vector and the input text based on the plurality of decoders and the text encoding to obtain decoding information, including: performing masking processing on the input text to obtain a mask vector; generating a target vector based on the perceptual vector and the mask vector; for any decoder, performing attention analysis on the target vector based on any decoder and the text encoding to obtain attention information; performing fully connected processing on the attention information based on the decoding weight matrix and decoding bias of any decoder to obtain an initial decoder, and using the initial decoder as the target vector for decoding processing of the next decoder, until all the plurality of decoders participate in decoding to obtain the decoding information; A prediction unit is used to predict the decoded information based on the prediction output layer to obtain a second output probability for each template word in the plurality of decoders; The generation unit is further configured to generate the answer to the question to be tested based on the first output probability and the second output probability.

8. An electronic device, characterized in that, The electronic device includes: Memory, which stores computer-readable instructions; and The processor executes computer-readable instructions stored in the memory to implement the answer generation method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions that are executed by a processor in an electronic device to implement the answer generation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Text processing method, text processing device, electronic equipment and storage medium

    CN115130432A

  • Method for generating chapter-level complex problem based on dual programming

    CN115510814A