Question Answering Matching Method, Device, Equipment and Storage Medium Based on Attention Mechanism

Through the Q&A matching method based on attention mechanism, the user problem vector is converted using the BERT model and attention mechanism, which solves the problem of difficulty in taking into account both the effect and efficiency in the traditional Q&A matching method, and achieves more efficient Q&A matching and user experience.

CN113886550BActive Publication Date: 2025-07-22PINGAN INT SMART CITY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111182254.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-11
Publication Date
2025-07-22
Estimated Expiration
2041-10-11

AI Technical Summary

Technical Problem

Traditional Q&A matching methods cannot take into account the matching effect and matching efficiency. The dual encoder loses user problem information, resulting in poor results, while the interactive encoder takes a long time and has a large data processing capacity.

Method used

The question-answer matching method based on attention mechanism is adopted, and the word vector sequence of user questions is obtained, and the hidden state problem vector is output using the BERT model, and converted into problem feature vectors based on the attention mechanism, and weighted summation is combined with the answer vector to determine the correct answer.

Benefits of technology

It improves the matching effect of candidate answers and user questions, reduces the amount of data processing, and thus improves the efficiency and user experience of question-and-answer matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113886550B_ABST
    Figure CN113886550B_ABST
Patent Text Reader

Abstract

The present invention is used in the field of artificial intelligence and relates to the field of blockchain. It discloses a question-answer matching method, device, equipment and storage medium based on an attention mechanism. Among them, the method part includes: obtaining a user question input by a user and obtaining answer vectors corresponding to a plurality of candidate answers; inputting a character vector sequence of the user question into a BERT model to obtain a plurality of hidden state question vectors output by the BERT model; converting the plurality of hidden state question vectors based on the attention mechanism to obtain m question feature vectors for characterizing the user question; converting the m question feature vectors according to the answer vectors to obtain a user question vector corresponding to the answer vector; determining the correct answer to the user question among the plurality of candidate answers according to the matching value between the corresponding user question vector and the answer vector; the present invention improves the matching effect between the candidate answer and the user question, reduces the data processing amount in the matching process, and improves the efficiency of question-answer matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a question-answer matching method, device, equipment and storage medium based on an attention mechanism. Background Art

[0002] A question-answer system is an advanced form of an information retrieval system, which can answer questions raised by users in accurate and concise natural language. The question-answer system generally performs question-answer matching through a matching model based on the BERT network. The matching model pre-trains the BERT model through a large-scale corpus in the general domain, and then fine-tunes the encoder built on top of the BERT model to improve the matching effect of the matching model, thereby ensuring the retrieval accuracy of the question-answer system.

[0003] In the traditional question-answer matching method, the encoder built on top of the BERT model is generally a dual encoder or a cross encoder. Both the dual encoder and the cross encoder have defects, resulting in the traditional question-answer matching method being unable to balance the matching effect and the matching efficiency. Among them, the core of the matching model of the dual encoder is to encode the user question and the candidate answer into vectors respectively, and finally calculate the similarity between the two vectors through a correlation discrimination function. This question-answer matching method has a fast matching speed, but the dual encoder will lose some user question information, resulting in a poor matching effect between the user question and the answer. The interactive encoder can achieve a finer-grained matching between the question and the candidate answer, making the matching model of the interactive encoder have a better matching effect. However, this question-answer matching method needs to traverse all combinations of user questions and candidate answers and calculate the correlation of each question-answer combination, with a large amount of data processing and long time consumption, reducing the question-answer matching efficiency. Summary of the Invention

[0004] The present invention provides a question-answer matching method, device, equipment and storage medium based on an attention mechanism to solve the problem that the traditional question-answer matching method cannot balance the matching effect and the matching efficiency.

[0005] Provided is a question-answer matching method based on an attention mechanism, including:

[0006] Obtain a user question input by a user, and determine a plurality of candidate answers for the user question and the answer vectors corresponding to the candidate answers;

[0007] Convert the user question into a sequence of word vectors, and input the sequence of word vectors of the user question into the BERT model to obtain a plurality of hidden state question vectors output by the BERT model;

[0008] Based on the attention mechanism, convert the plurality of hidden state question vectors to obtain m question feature vectors for characterizing the user question, where m is an integer greater than 1;

[0009] Convert the m question feature vectors according to the answer vector to obtain the user question vector corresponding to the answer vector;

[0010] Determine the correct answer to the user question from multiple candidate answers according to the matching value between the corresponding user question vector and the answer vector.

[0011] Provide a question-answer matching device based on the attention mechanism, comprising:

[0012] An acquisition module, configured to acquire the user question input by the user, and determine multiple candidate answers to the user question and the answer vectors corresponding to the candidate answers;

[0013] An encoding module, configured to convert the user question into a sequence of word vectors, and input the sequence of word vectors of the user question into the BERT model to obtain multiple hidden state question vectors output by the BERT model;

[0014] A first conversion module, configured to convert the multiple hidden state question vectors based on the attention mechanism to obtain m question feature vectors for characterizing the user question, where m is an integer greater than 1;

[0015] A second conversion module, configured to convert the m question feature vectors according to the answer vector to obtain the user question vector corresponding to the answer vector;

[0016] A determination module, configured to determine the correct answer to the user question from multiple candidate answers according to the matching value between the corresponding user question vector and the answer vector.

[0017] Provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned question-answer matching method based on the attention mechanism are implemented.

[0018] Provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned question-answer matching method based on the attention mechanism are implemented.

[0019] In one solution provided by the above question-answering matching method, device, equipment and storage medium based on the attention mechanism, a user question input by a user is obtained, and multiple candidate answers and answer vectors corresponding to the candidate answers are determined; then the user question is converted into a sequence of word vectors, and the sequence of word vectors of the user question is input into a BERT model to obtain multiple hidden state question vectors output by the BERT model; and based on the attention mechanism, the multiple hidden state question vectors are transformed to obtain m question feature vectors for characterizing the user question, where m is an integer greater than 1; then the m question feature vectors are transformed according to the answer vectors to obtain a user question vector corresponding to the answer vectors; finally, according to the matching value between the corresponding user question vector and the answer vector, the correct answer to the user question is determined among the multiple candidate answers; in the present invention, the user question vector is improved based on the attention mechanism to obtain m question feature vectors, which can represent more global features, and then the m question feature vectors are weighted into the corresponding user question vector according to the answer vectors, so that the candidate answers and the user question can be mutually fused, improving the accuracy of the matching value between the corresponding user question vector and the answer vector, and further improving the matching effect between the candidate answers and the user question. On this basis, the data processing amount in the matching process is reduced, thereby improving the efficiency of question-answering matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0021] Figure 1 is a schematic diagram of an application environment of the question-answering matching method based on the attention mechanism in an embodiment of the present invention;

[0022] Figure 2 is a schematic flowchart of the question-answering matching method based on the attention mechanism in an embodiment of the present invention;

[0023] Figure 3 is Figure 2 a schematic implementation flowchart of step S30 in

[0024] Figure 4 is Figure 2 a schematic implementation flowchart of step S40 in

[0025] Figure 5 is Figure 2 a schematic implementation flowchart of step S50 in

[0026] Figure 6Yes Figure 2 It is a schematic flowchart of an implementation of step S10 in

[0027] Figure 7 Yes Figure 2 It is another schematic flowchart of an implementation of step S10 in

[0028] Figure 8 Yes Figure 7 It is a schematic flowchart of an implementation of step S03 in

[0029] Figure 9 It is a schematic structural diagram of a question - answering matching device based on an attention mechanism in an embodiment of the present invention;

[0030] Figure 10 It is a schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0032] The question - answering matching method based on an attention mechanism provided by the embodiments of the present invention can be applied in, for example, Figure 1In the application environment, the terminal device communicates with the server through the network. The server obtains the user question input by the user through the terminal device, and determines multiple candidate answers to the user question and the answer vectors corresponding to the candidate answers; then converts the user question into a sequence of word vectors, and inputs the sequence of word vectors of the user question into the BERT model to obtain multiple hidden state question vectors output by the BERT model; and based on the attention mechanism, converts the multiple hidden state question vectors to obtain m question feature vectors for characterizing the user question, where m is an integer greater than 1; then converts the m question feature vectors according to the answer vectors to obtain the user question vectors corresponding to the answer vectors; finally, determines the correct answer to the user question among the multiple candidate answers according to the matching value between the corresponding user question vectors and the answer vectors; in the present invention, the user question vectors are improved based on the attention mechanism to obtain m question feature vectors, which can represent more global features, and then the m question feature vectors are weighted into the corresponding user question vectors according to the answer vectors, enabling the candidate answers and the user questions to be mutually integrated, improving the accuracy of the matching value between the corresponding user question vectors and the answer vectors, and further improving the matching effect between the candidate answers and the user questions. On this basis, the data processing amount in the matching process is reduced, thereby improving the efficiency of question-answer matching. Finally, the artificial intelligence of the question-answer system is further improved, and the user experience is improved.

[0033] Among them, relevant data such as multiple candidate answers and the answer vectors corresponding to the candidate answers are stored in the database of the server. When it is necessary to perform answer matching on the user question, the relevant data is directly obtained from the database of the server, improving the efficiency of question-answer matching.

[0034] The database in this embodiment is stored in the blockchain network and is used to store the data used and generated in the question-answer matching method based on the attention mechanism, such as relevant data such as candidate answers and the answer vectors corresponding to the candidate answers. The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc. Deploying the database on the blockchain can improve the security of data storage.

[0035] Among them, the terminal device can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers.

[0036] In one embodiment, as Figure 2 shown, a question-and-answer matching method based on an attention mechanism is provided. Taking the server in Figure 1 as an example, the method includes the following steps:

[0037] S10: Obtain the user question input by the user, and determine multiple candidate answers for the user question and the answer vectors corresponding to the candidate answers.

[0038] Obtain the user question input by the user through the terminal device, then obtain multiple standard answers from the database, perform keyword matching between the user question and the standard answers to obtain multiple candidate answers, and further obtain the answer vectors corresponding to the multiple candidate answers. Among them, the answer vector is the answer vector obtained after extracting the feature vector of the candidate answer.

[0039] Among them, the answer vector can be obtained in the following ways:

[0040] The first way: Obtain a candidate answer from the database, then input the candidate answer into the BERT model to obtain multiple hidden state answer vectors output by the BERT model, and then aggregate the multiple hidden state answer vectors to obtain the answer vector of the candidate answer. So that after obtaining the user question, match the answer vector and the user question vector of the user question. If the answer vector and the user question vector do not match, continue to obtain a candidate answer from the database, convert the candidate answer into an answer vector to match the user question vector.

[0041] The second way, after clarifying multiple standard answers, perform offline conversion on the multiple standard answers to obtain the answer vectors corresponding to the multiple standard answers, store the multiple standard answers and the answer vectors corresponding to the standard answers in a one-to-one manner in the database. After obtaining the user question input by the user, determine the standard answer corresponding to the user question in the database as the candidate answer for the user question, and then directly retrieve the answer vector corresponding to the candidate answer in the database, without performing online conversion on the candidate answer, reducing the calculation amount and improving the question-and-answer matching efficiency. In the offline conversion, the standard answer is input into the BERT model offline in advance to obtain multiple hidden state answer vectors output by the BERT model, and then the multiple hidden state answer vectors are aggregated to obtain the answer vector of the standard answer, so as to obtain the answer vectors of all standard answers offline.

[0042] S20: Convert the user question into a sequence of word vectors, and input the sequence of word vectors of the user question into the BERT model to obtain multiple hidden state question vectors output by the BERT model.

[0043] After obtaining the user question input by the user, the user question is segmented into single characters to obtain the character vector sequence of the user question. Then, the character vector sequence of the user question is input into the BERT model for encoding to obtain the hidden states of the BERT model, which are used as multiple hidden state question vectors output after the BERT model encodes the user question. The hidden states of the BERT model are used as the representation of the user question, making the vector have a higher correlation with the user question information, facilitating subsequent feature extraction based on the hidden state question vectors to ensure the accuracy of the extracted feature vectors.

[0044] For example, if the user question is q, the character vector sequence of the user question is represented as: The character vector sequence of the user question is input into the BERT model for encoding to obtain the hidden states of the last layer of the transformer of the BERT model, which are used as the representation vector h of the user question j , that is, multiple hidden state question vectors output by the BERT model are obtained. The hidden state question vectors are represented by the following formula:

[0045]

[0046] where N x is the length of the character vector sequence of the user question, is the Nth x character vector in the user question, h j is the hidden state question vector of the jth character in the user question, and j ∈ [0, N x .

[0047] S30: Based on the attention mechanism, multiple hidden state question vectors are transformed to obtain m question feature vectors for representing the user question.

[0048] After obtaining multiple hidden state question vectors output by the BERT model, based on the attention mechanism, multiple hidden state question vectors are transformed to filter and obtain m question feature vectors for representing the user question. Among them, m is an integer greater than 1, and m can be valued according to actual needs.

[0049] Convert multiple hidden state problem vectors based on the attention mechanism to obtain m problem feature vectors for characterizing the user's problem. The specific process is as follows: First, determine the attention weight matrix according to the multiple hidden state problem vectors, and perform weighted summation on the multiple hidden state problem vectors according to the multiple attention weights in the attention weight matrix, and screen to obtain a problem feature vector for characterizing the user's problem; then update the attention weight matrix, and perform weighted summation on the multiple hidden state problem vectors according to the multiple attention weights in the updated attention weight matrix, and screen to obtain the next problem feature vector. Repeat the screening process multiple times in sequence until m problem feature vectors for characterizing the user's problem are obtained. Obtaining m problem feature vectors based on the attention mechanism can make the problem feature vectors have better semantic relevance. On the basis of ensuring the accuracy of the problem feature vectors, capturing m problem feature vectors also reduces the computational complexity.

[0050] S40: Convert the m problem feature vectors according to the answer vector to obtain the user problem vector corresponding to the answer vector.

[0051] After obtaining m problem feature vectors for characterizing the user's problem, perform weighted summation on the m problem feature vectors according to the answer vector to obtain the user problem vector corresponding to the answer vector. Among them, the calculation of the user problem vector corresponding to the answer vector is also a vector calculation based on the attention mechanism. First, determine the attention weight matrix according to the answer vector and the user problem vector, and perform combined calculation on the m problem feature vectors according to the m attention weights in the attention weight matrix to obtain the user problem vector corresponding to the answer vector. Based on the vector processing of the double-layer attention mechanism, and when performing the second vector calculation based on the attention mechanism, interact and fuse the answer vector with the m problem feature vectors to obtain the user problem vector, which improves the correlation between the user problem vector and the answer vector and is convenient for matching.

[0052] S50: Determine the correct answer to the user's problem among multiple candidate answers according to the matching value between the corresponding user problem vector and the answer vector.

[0053] After obtaining the user problem vector corresponding to the answer vector, use the dot product operation method to determine the matching value between the corresponding user problem vector and the answer vector, and then determine the correct answer to the user's problem among multiple candidate answers according to the matching value between the corresponding user problem vector and the answer vector.

[0054] Among them, the matching value (matching score) between the user problem vector corresponding to the answer vector and the answer vector is calculated by the following formula:

[0055]

[0056] Among them, q is the user's problem; yq The user question vector for the user question; a i is the i-th candidate answer; is the answer vector of the i-th candidate answer; s(q, a i ) is the matching value between the answer vector of the i-th candidate answer and the user question vector.

[0057] In this embodiment, by obtaining the user question input by the user, and determining multiple candidate answers for the user question and the corresponding answer vectors of the candidate answers; then converting the user question into a sequence of word vectors, and inputting the sequence of word vectors of the user question into the BERT model to obtain multiple hidden state question vectors output by the BERT model; and based on the attention mechanism, converting the multiple hidden state question vectors to obtain m question feature vectors for characterizing the user question, where m is an integer greater than 1; then performing weighted summation transformation on the m question feature vectors according to the answer vectors to obtain the user question vector corresponding to the answer vector; finally, according to the matching value between the corresponding user question vector and the answer vector, determining the correct answer to the user question among multiple candidate answers; in this embodiment, based on the attention mechanism, the user question vector is improved to obtain m question feature vectors, which can represent more global features, and then according to the answer vectors, the m question feature vectors are combined into the corresponding user question vector, which can make the candidate answers and the user question interact with each other, improve the accuracy of the matching value between the user question vector and the answer vector, and further improve the matching effect between the candidate answers and the user question. On this basis, there is no need to perform a large number of interactions between the answers and the questions, reducing the amount of data processing in the matching process, thereby improving the efficiency of question-answer matching.

[0058] In one embodiment, as Figure 3 shown, in step S30, that is, based on the attention mechanism, converting the multiple hidden state question vectors to obtain m question feature vectors for characterizing the user question, specifically includes the following steps:

[0059] a. Determine an initial weight, and determine the first weight matrix according to the multiple hidden state question vectors and the initial weight.

[0060] After obtaining the multiple hidden state question vectors, it is necessary to randomly determine an initial weight to determine the first weight matrix according to the multiple hidden state question vectors and the initial weight. Among them, since there are m hidden state question vectors, the initial weight needs to be cyclically taken m times. The initial weight taken each time is denoted as c i , then finally after m times of initial weight taking, the initial weight matrix composed of m initial weights is (c1..c m), the initial weight matrix includes multiple initial weights, and each initial weight is used to calculate and generate a first weight matrix with multiple hidden state problem vectors. Among them, the initial weight matrix (c1..c m ), is the weight matrix used to measure the hidden state of the last layer of the transformer in the BERT model.

[0061] Among them, each weight in the first weight matrix is the product of the initial weight and each hidden state problem vector, that is, the first weight matrix is: (c i ·h1,..., c i ·h j ), i ∈ [0, m].

[0062] Among them, c i is the i-th initial weight in the initial weight matrix, h j is the j-th hidden state problem vector, and m is the number of initial weights.

[0063] b. Use the softmax function to normalize the first weight matrix to obtain multiple attention weights.

[0064] After determining the first weight matrix, use the softmax function to normalize the first weight matrix to obtain multiple attention weights corresponding to the initial weight.

[0065] Among them, the multiple attention weights are represented by the following formula:

[0066]

[0067] Among them, softmax is the softmax function, c i is the i-th initial weight, h j is the j-th hidden state problem vector, is the j-th attention weight corresponding to the i-th initial weight.

[0068] c. Weighted sum the multiple hidden state problem vectors according to the multiple attention weights to obtain a problem feature vector for representing the user's problem.

[0069] After obtaining the multiple attention weights, weighted sum the multiple hidden state problem vectors according to the multiple attention weights to obtain a problem feature vector for representing the user's problem.

[0070] The calculation formula of the problem feature vector is as follows:

[0071]

[0072] Among them, is the j-th attention weight corresponding to the i-th initial weight, where i ∈ [0, m]; h j is the j-th hidden state problem vector; is the i-th problem feature vector used to represent the user's question.

[0073] d. Repeat steps a - c to obtain m problem feature vectors.

[0074] By repeating steps a - c, m problem feature vectors can be obtained.

[0075] In this embodiment, through: a. determining an initial weight and determining the first weight matrix according to multiple hidden state problem vectors and the initial weight; b. performing normalization processing on the first weight matrix using the softmax function to obtain multiple attention weights; c. performing weighted summation on multiple hidden state problem vectors according to multiple attention weights to obtain a problem feature vector used to represent the user's question; d. repeating steps a - c to obtain m problem feature vectors, the specific process of converting multiple hidden state problem vectors based on the attention mechanism to obtain m problem feature vectors used to represent the user's question is clarified. The attention weights are determined according to the hidden state problem vectors to filter out m problem feature vectors, providing a basis for subsequent calculations.

[0076] In one embodiment, as Figure 4 shown, in step S40, that is, converting m problem feature vectors according to the answer vector to obtain the user question vector corresponding to the answer vector, specifically including the following steps:

[0077] S41: Determine the second weight matrix corresponding to the answer vector according to the answer vector and m problem feature vectors.

[0078] After obtaining the answer vector of the candidate answer and m problem feature vectors, determine the second weight matrix corresponding to the answer vector according to the answer vector of the candidate answer and m problem feature vectors.

[0079] Among them, in the second weight matrix corresponding to the answer vector, each weight is the product of the answer vector of the candidate answer and the problem feature vector. The second weight matrix is expressed as: Among them, is the i-th problem feature vector among m problem feature vectors, is the answer vector corresponding to the i-th candidate answer among multiple candidate answers.

[0080] S42: Perform normalization processing on the second weight matrix using the softmax function to obtain multiple target weights.

[0081] After determining the second weight matrix corresponding to the answer vector, the normalized exponential function softmax is used to normalize the second weight matrix to obtain multiple target weights.

[0082] Among them, the multiple target weights are represented by the following formula:

[0083]

[0084] Among them, softmax is the normalized exponential function, is the i-th question feature vector among m question feature vectors, is the answer vector corresponding to the i-th candidate answer among multiple candidate answers, w i is the i-th target weight, i ∈ [0, m].

[0085] S43: Sum the m question feature vectors according to the multiple target weights to obtain the user question vector corresponding to the answer vector.

[0086] After obtaining the multiple target weights, perform weighted summation on the m question feature vectors according to the multiple target weights to obtain the user question vector corresponding to the answer vector.

[0087] The user question vector corresponding to the answer vector is calculated by the following formula:

[0088]

[0089] Among them, y q is the user question vector corresponding to the answer vector, is the i-th question feature vector among m question feature vectors, w i is the i-th target weight, i ∈ [0, m].

[0090] In this embodiment, according to the answer vector and m question feature vectors, the second weight matrix corresponding to the answer vector is determined; the normalized exponential function is used to normalize the second weight matrix to obtain multiple target weights; the m question feature vectors are summed according to the multiple target weights to obtain the user question vector corresponding to the answer vector, which clarifies the process of transforming the m question feature vectors according to the answer vector to obtain the user question vector corresponding to the answer vector. Based on the attention mechanism, the user question vector corresponding to the answer vector of the candidate answer is determined, which provides a basis for the matching value of the user question vector and the answer vector of the candidate answer. Moreover, the process of introducing the answer vector of the candidate answer into the user question vector increases the interaction between the user question and the candidate answer, improves the accuracy of the user question vector, and further improves the accuracy of the candidate matching value, which is beneficial to enhancing the matching effect.

[0091] In one embodiment, asFigure 5 As shown in Figure 5 , in step S50, that is, according to the matching value between the corresponding user question vector and the answer vector, the correct answer to the user question is determined among multiple candidate answers, which specifically includes the following steps:

[0092] S51: Calculate the matching value between the corresponding user question vector and the answer vector.

[0093] After obtaining the user question vector corresponding to the answer vector, calculate the matching value between the user question vector corresponding to the answer vector and the answer vector through the dot product operation. Among them, the matching value between the user question vector corresponding to the answer vector and the answer vector is calculated by the following formula:

[0094]

[0095] where s(q, a i ) is the matching value between the i-th candidate answer a and the user question q; y q is the user question vector of the user question q; is the answer vector of the i-th candidate answer a.

[0096] S52: Take the matching value as the target matching value between the candidate answer and the user question.

[0097] After calculating the matching value between the corresponding user question vector and the answer vector, take the matching value as the target matching value between the candidate answer and the user question.

[0098] S53: Sort the multiple candidate answers in ascending order according to the size of the target matching value to obtain a candidate answer list.

[0099] After taking the matching value as the target matching value between the candidate answer and the user question, sort the multiple candidate answers in ascending order according to the size of the target matching value to obtain a candidate answer list, that is, the multiple candidate answers in the candidate answer list are sorted according to the size of the corresponding target matching value.

[0100] S54: Take the candidate answer ranked first in the candidate answer list as the correct answer to the user question.

[0101] After sorting the multiple candidate answers in ascending order according to the size of the target matching value to obtain a candidate answer list, the target matching value of the candidate answer ranked first in the candidate answer list is the largest, indicating that the candidate answer ranked first is the most matched with the user question. Then take the candidate answer ranked first in the candidate answer list as the correct answer to the user question, so that the user can obtain the most accurate answer and improve the user experience.

[0102] In this embodiment, the matching value between the corresponding user question vector and the answer vector is calculated; the matching value is used as the target matching value between the candidate answer and the user question; according to the magnitude of the target matching value, multiple candidate answers are sorted in ascending order to obtain a candidate answer list; the candidate answer ranked first in the candidate answer list is used as the correct answer to the user question, clarifying the process of determining the correct answer to the user question from multiple candidate answers according to the matching value between the corresponding user question vector and the answer vector.

[0103] In one embodiment, as Figure 6 shown, in step S10, that is, determining multiple candidate answers to the user question and the answer vectors corresponding to the candidate answers, specifically includes the following steps:

[0104] S11: Obtain multiple standard answers stored in the database.

[0105] After obtaining the user question input by the user, it is necessary to obtain multiple standard answers stored in the database.

[0106] S12: Determine the named entities in the user question, and use the standard answers containing the named entities as the candidate answers to the user question.

[0107] After obtaining multiple standard answers stored in the database, multiple candidate answers are determined from the multiple standard answers according to the user question. Among them, multiple candidate answers are determined from the multiple standard answers by means of keyword (entity) matching.

[0108] First, it is necessary to determine the named entities in the user question, determine whether the standard answer contains the named entities in the user question. If the standard answer contains the named entities in the user question, then use the standard answer containing the named entities in the user question as the candidate answer to the user question. By determining the standard answers with the same named entities as candidate answers from multiple standard answers, the amount of candidate matching calculations can be reduced, thereby improving the Q&A matching efficiency, and then being able to quickly reply to the user and improve the user experience.

[0109] S13: Obtain the answer vectors corresponding to the candidate answers in the database to obtain the answer vectors corresponding to multiple candidate answers.

[0110] In this embodiment, the database stores the standard answers and the answer vectors of the standard answers, and the standard answers and the answer vectors of the standard answers are in one-to-one correspondence. After determining the candidate answers to the user's question, the answer vectors corresponding to the candidate answers are obtained from the database to obtain the answer vectors corresponding to multiple candidate answers. The standard answers are pre-converted into answer vectors and stored in an offline manner, facilitating the subsequent server to quickly determine the answer vectors of the candidate answers according to the actual user questions, without calculating the answer vectors of the candidate answers online, reducing the computational load of the server, and achieving the effects of reducing the server load, improving the question-answer matching efficiency, and the server response speed.

[0111] In this embodiment, by obtaining multiple standard answers stored in the database, the named entities in the user's question are determined, and the standard answers containing the named entities are used as the candidate answers to the user's question. The answer vectors corresponding to the candidate answers are obtained from the database to obtain the answer vectors corresponding to multiple candidate answers, clarifying the specific process of obtaining the answer vectors corresponding to multiple candidate answers, without calculating the answer vectors of the candidate answers online, reducing the computational load of the server, and achieving the effects of reducing the server load, improving the question-answer matching efficiency, and the server response speed.

[0112] In one embodiment, as Figure 7 shown, that is, to obtain the answer vector corresponding to the candidate answer, which specifically includes the following steps:

[0113] S01: Segment the candidate answer into single characters to obtain the character vector sequence of the candidate answer.

[0114] After obtaining the candidate answer, segment the candidate answer into single characters to obtain the character vector sequence of the candidate answer. Using the character vector sequence as the input of the BERT model to obtain the character vectors of the candidate answer is more accurate than the traditional word vector division, and the obtained vectors are more accurate.

[0115] S02: Input the character vector sequence of the candidate answer into the BERT model for encoding to obtain multiple hidden state answer vectors output by the BERT model.

[0116] After obtaining the character vector sequence of the candidate answer, input the character vector sequence of the candidate answer into the BERT model for encoding to obtain the hidden state of the last layer of the transformer in the BERT model as the multiple hidden state answer vectors of the candidate answer output by the BERT model.

[0117] S03: Aggregate the multiple hidden state answer vectors based on a preset aggregation method to obtain the answer vector corresponding to the candidate answer.

[0118] After obtaining multiple hidden state answer vectors output by the BERT model, aggregate the multiple hidden state answer vectors based on a preset aggregation method to obtain the answer vector corresponding to the candidate answer.

[0119] Among them, the word vector sequence of candidate answer a is Then the calculation formula for the answer vector corresponding to the candidate answer is:

[0120]

[0121] Among them, BERT is the function of the BERT model, Concat represents the aggregation function (preset aggregation method), N y is the length of the word vector sequence of candidate answer a, is the Nth x word vector in the user question, is the answer vector corresponding to the i-th candidate answer, i ∈ [0, N y .

[0122] For example, the user question is: What is the capital of China? The candidate answer (correct answer) is: The capital of China is Beijing. Then split "The capital of China is Beijing" into single characters to obtain the word vector sequence of the candidate answer, input the word vector sequence of the candidate answer into the BERT model for encoding to obtain multiple hidden state answer vectors output by the BERT model, and aggregate the multiple hidden state answer vectors based on a preset aggregation method to obtain the answer vector corresponding to the candidate answer. According to the above process, the normalized answer vector for the candidate answer - The capital of China is Beijing is: 0.2341, 0.4353, 0.2352, 0.6436,..., 0.3453; the length of the answer vector is 1 * 128.

[0123] In this embodiment, by splitting the candidate answer into single characters to obtain the word vector sequence of the candidate answer, then inputting the word vector sequence of the candidate answer into the BERT model for encoding to obtain multiple hidden state answer vectors output by the BERT model; finally, aggregating the multiple hidden state answer vectors based on a preset aggregation method to obtain the answer vector corresponding to the candidate answer, the process of obtaining the answer vector is clarified, providing a basis for subsequently determining the correct answer according to the matching value between the answer vector corresponding to the candidate answer and the user question vector.

[0124] In one embodiment, as Figure 8 shown, in step S03, that is, aggregating the multiple hidden state answer vectors based on a preset aggregation method to obtain the answer vector corresponding to the candidate answer, specifically includes the following steps:

[0125] S031: Among multiple hidden state answer vectors, determine the first hidden state answer vector output by the BERT model.

[0126] After obtaining multiple hidden state answer vectors output by the BERT model, among the multiple hidden state answer vectors, determine the first hidden state answer vector output by the BERT model. The first hidden state answer vector is the vector corresponding to [CLS] in the BERT model.

[0127] S032: Use the first hidden state answer vector output by the BERT model as the answer vector corresponding to the candidate answer.

[0128] After determining to use the first hidden state answer vector output by the BERT model, use the first hidden state answer vector output by the BERT model as the answer vector corresponding to the candidate answer. In the BERT model, the vector corresponding to [CLS] can more fairly integrate the semantic information of each word in the text, can better represent the candidate answer than other words, and directly using the first hidden state answer vector output by the BERT model as the answer vector corresponding to the candidate answer can reduce the amount of calculation and improve the acquisition speed of the answer vector corresponding to the candidate answer, thereby improving the question-answer matching efficiency.

[0129] In other embodiments, the preset aggregation method can also be: determine the mean of multiple hidden state answer vectors, and use the mean of multiple hidden state answer vectors as the answer vector corresponding to the candidate answer to improve the accuracy of the answer vector corresponding to the candidate answer; or, determine the first m hidden state answer vectors among the multiple hidden state answer vectors output by the BERT model, and determine the mean of the first m hidden state answer vectors, and use the mean of the first m hidden state answer vectors as the answer vector corresponding to the candidate answer, reducing the processing of meaningless vectors and reducing the amount of calculation.

[0130] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0131] In one embodiment, a question-answer matching device based on the attention mechanism is provided. The question-answer matching device based on the attention mechanism corresponds one-to-one with the question-answer matching method based on the attention mechanism in the above embodiment. As Figure 9 shown, the question-answer matching device based on the attention mechanism includes an acquisition module 901, an encoding module 902, a conversion module 903, a second conversion module 904, and a determination module 905. The detailed description of each functional module is as follows:

[0132] An acquisition module 901, configured to acquire a user question input by a user, and determine a plurality of candidate answers to the user question and answer vectors corresponding to the candidate answers;

[0133] An encoding module 902, configured to convert the user question into a sequence of word vectors, and input the sequence of word vectors of the user question into a BERT model to obtain a plurality of hidden state question vectors output by the BERT model;

[0134] A first conversion module 903, configured to perform conversion on the plurality of hidden state question vectors based on an attention mechanism to obtain m question feature vectors for characterizing the user question, where m is an integer greater than 1;

[0135] A second conversion module 904, configured to perform conversion on the m question feature vectors according to the answer vectors to obtain a user question vector corresponding to the answer vectors;

[0136] A determination module 905, configured to determine the correct answer to the user question from the plurality of candidate answers according to the matching value between the corresponding user question vector and the answer vector.

[0137] Further, the first conversion module 903 is specifically configured to:

[0138] a. Determine an initial weight, and determine a first weight matrix according to the plurality of hidden state question vectors and the initial weight;

[0139] b. Perform normalization processing on the first weight matrix by using a softmax function to obtain a plurality of attention weights;

[0140] c. Perform weighted summation on the plurality of hidden state question vectors according to the plurality of attention weights to obtain a question feature vector for characterizing the user question;

[0141] d. Repeat steps a - c to obtain m question feature vectors.

[0142] Further, the second conversion module 904 is specifically configured to:

[0143] Determine a second weight matrix corresponding to the answer vectors according to the answer vectors and the m question feature vectors;

[0144] Perform normalization processing on the second weight matrix by using a softmax function to obtain a plurality of target weights;

[0145] Perform summation on the m question feature vectors according to the plurality of target weights to obtain a user question vector corresponding to the answer vectors.

[0146] Further, the determination module 905 is specifically configured to:

[0147] Calculate the matching value between the corresponding user question vector and the answer vector;

[0148] Take the matching value as the target matching value between the candidate answer and the user question;

[0149] Sort multiple candidate answers in ascending order according to the size of the target matching value to obtain a candidate answer list;

[0150] Take the candidate answer ranked first in the candidate answer list as the correct answer to the user question.

[0151] Further, the obtaining module 901 is specifically configured to:

[0152] Obtain multiple standard answers stored in the database;

[0153] Determine the named entities in the user question, and take the standard answers containing the named entities as candidate answers to the user question;

[0154] Obtain the answer vectors corresponding to the candidate answers in the database to obtain the answer vectors corresponding to multiple candidate answers.

[0155] Further, the obtaining module 901 is also specifically configured to obtain the answer vectors corresponding to the candidate answers in the following manner:

[0156] Segment the candidate answers into single characters to obtain a character vector sequence of the candidate answers;

[0157] Input the character vector sequence of the candidate answers into the BERT model for encoding to obtain multiple hidden state answer vectors output by the BERT model;

[0158] Aggregate the multiple hidden state answer vectors based on a preset aggregation method to obtain the answer vectors corresponding to the candidate answers.

[0159] Further, the obtaining module 901 is also specifically configured to:

[0160] Among the multiple hidden state answer vectors, determine the first hidden state answer vector output by the BERT model;

[0161] Take the first hidden state answer vector output by the BERT model as the answer vector corresponding to the candidate answer.

[0162] For the specific limitations of the question-answering matching device based on the attention mechanism, reference can be made to the limitations of the question-answering matching method based on the attention mechanism in the above text, which will not be elaborated here. Each module in the above question-answering matching device based on the attention mechanism can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0163] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as candidate answers and answer vectors corresponding to subsequent answers. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a question-answering matching method based on the attention mechanism.

[0164] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0165] Obtain a user question input by a user, and determine multiple candidate answers to the user question and answer vectors corresponding to the candidate answers;

[0166] Convert the user question into a sequence of word vectors, and input the sequence of word vectors of the user question into the BERT model to obtain multiple hidden state question vectors output by the BERT model;

[0167] Based on the attention mechanism, convert the multiple hidden state question vectors to obtain m question feature vectors for characterizing the user question, where m is an integer greater than 1;

[0168] According to the answer vectors, convert the m question feature vectors to obtain user question vectors corresponding to the answer vectors;

[0169] According to the matching values of the corresponding user question vectors and the answer vectors, determine the correct answer to the user question among the multiple candidate answers.

[0170] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0171] Obtain a user question input by a user, and determine multiple candidate answers to the user question and answer vectors corresponding to the candidate answers;

[0172] Convert the user question into a sequence of word vectors, and input the sequence of word vectors of the user question into a BERT model to obtain multiple hidden state question vectors output by the BERT model;

[0173] Based on an attention mechanism, convert the multiple hidden state question vectors to obtain m question feature vectors for characterizing the user question;

[0174] According to the answer vectors, convert the m question feature vectors to obtain user question vectors corresponding to the answer vectors, where m is an integer greater than 1;

[0175] Determine the correct answer to the user question from the multiple candidate answers according to the matching value between the corresponding user question vector and the answer vector.

[0176] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0177] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the division of the above functional units and modules is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0178] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A question-answering matching method based on an attention mechanism, characterized in that Including: Obtain the user question input by the user, and determine multiple candidate answers to the user question and the answer vectors corresponding to the candidate answers; Convert the user question into a sequence of word vectors, and input the sequence of word vectors of the user question into the BERT model to obtain multiple hidden state question vectors output by the BERT model; Based on the attention mechanism, convert the multiple hidden state question vectors to obtain m question feature vectors for characterizing the user question, where m is an integer greater than 1; Convert the m question feature vectors according to the answer vectors to obtain the user question vector corresponding to the answer vectors; Determine the correct answer to the user question from the multiple candidate answers according to the matching value between the corresponding user question vector and the answer vectors; Among them, converting the m question feature vectors according to the answer vectors to obtain the user question vector corresponding to the answer vectors includes: Determine the second weight matrix corresponding to the answer vectors according to the answer vectors and the m question feature vectors; Perform normalization processing on the second weight matrix using the softmax function to obtain multiple target weights; Sum the m question feature vectors according to the multiple target weights to obtain the user question vector corresponding to the answer vectors; The matching value between the corresponding user question vector and the answer vectors is calculated by the following formula: Among them, is the user's question; is the user question vector of the user's question; is the i-th candidate answer; is the answer vector of the i-th candidate answer; is the matching value between the answer vector of the i-th candidate answer and the user question vector; Among them, the answer vectors corresponding to the candidate answers are obtained in the following manner: Perform single-character segmentation on the candidate answers to obtain a sequence of word vectors of the candidate answers; Input the sequence of word vectors of the candidate answers into the BERT model for encoding to obtain multiple hidden state answer vectors output by the BERT model; Aggregate the multiple hidden state answer vectors based on a preset aggregation method to obtain the answer vectors corresponding to the candidate answers.

2. The question-answering matching method based on the attention mechanism according to claim 1, wherein The conversion of the multiple hidden state question vectors based on the attention mechanism to obtain m question feature vectors for characterizing the user question includes: a. Determine an initial weight, and determine the first weight matrix according to the multiple hidden state question vectors and the initial weight; b. Perform normalization processing on the first weight matrix using the softmax function to obtain multiple attention weights; c. Perform weighted summation on the multiple hidden state question vectors according to the multiple attention weights to obtain a question feature vector for characterizing the user question; d. Repeat steps a - c to obtain the m question feature vectors.

3. The question-answering matching method based on the attention mechanism according to claim 1, wherein Determining the correct answer to the user question from the multiple candidate answers according to the matching value between the corresponding user question vector and the answer vectors includes: Calculate the matching value between the corresponding user question vector and the answer vectors; Use the matching value as the target matching value between the candidate answer and the user question; Sort the multiple candidate answers in ascending order according to the magnitude of the target matching value to obtain a list of candidate answers; Take the candidate answer ranked first in the candidate answer list as the correct answer to the user's question.

4. The question-answering matching method based on the attention mechanism according to claim 1, wherein The determining of multiple candidate answers to the user's question and the answer vectors corresponding to the candidate answers includes: Obtain multiple standard answers stored in the database; Determine the named entities in the user's question, and use the standard answers containing the named entities as the candidate answers to the user's question; Obtain the answer vectors corresponding to the candidate answers in the database to obtain the answer vectors corresponding to multiple candidate answers.

5. The question-answering matching method based on the attention mechanism according to claim 1, characterized in that The aggregating of the multiple hidden state answer vectors based on a preset aggregation method to obtain the answer vectors corresponding to the candidate answers includes: Among the multiple hidden state answer vectors, determine the first hidden state answer vector output by the BERT model; Take the first hidden state answer vector output by the BERT model as the answer vector corresponding to the candidate answer.

6. A question-answering matching device based on an attention mechanism, characterized in that, Includes: An acquisition module, configured to acquire a user's question input by a user, and determine multiple candidate answers to the user's question and the answer vectors corresponding to the candidate answers; An encoding module, configured to convert the user's question into a sequence of word vectors, and input the sequence of word vectors of the user's question into the BERT model to obtain multiple hidden state question vectors output by the BERT model; A first conversion module, configured to convert the multiple hidden state question vectors based on an attention mechanism to obtain m question feature vectors for characterizing the user's question; A second conversion module, configured to convert the m question feature vectors according to the answer vectors to obtain a user question vector corresponding to the answer vector, where m is an integer greater than 1; A determination module, configured to determine the correct answer to the user's question among the multiple candidate answers according to the matching value between the corresponding user question vector and the answer vector; Among them, converting the m question feature vectors according to the answer vectors to obtain a user question vector corresponding to the answer vector includes: Determine a second weight matrix corresponding to the answer vector according to the answer vector and the m question feature vectors; Perform normalization processing on the second weight matrix by using a normalized exponential function to obtain multiple target weights; Sum the m question feature vectors according to the multiple target weights to obtain a user question vector corresponding to the answer vector; The matching value between the corresponding user question vector and the answer vector is calculated by the following formula: Among them, is the user's question; is the user question vector of the user's question; is the i-th candidate answer; is the answer vector of the i-th candidate answer; is the matching value between the answer vector of the i-th candidate answer and the user question vector; Among them, the answer vector corresponding to the candidate answer is obtained in the following manner: Perform single-character segmentation on the candidate answer to obtain a sequence of word vectors of the candidate answer; Input the sequence of word vectors of the candidate answer into the BERT model for encoding to obtain multiple hidden state answer vectors output by the BERT model; Aggregate the multiple hidden state answer vectors based on a preset aggregation method to obtain the answer vector corresponding to the candidate answer.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the question-answering matching method based on the attention mechanism according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the question-answering matching method based on the attention mechanism according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Question answering matching degree calculation method, question answering automatic matching method and device

    CN109376222A

  • Question classification method and application thereof

    CN112597304A