A method and device for question-and-answer interaction
By constructing and integrating the matching feature matrix of consulting context and consulting statements, the problem of failure to extract multiple rounds of dialogue interaction characteristics in the prior art is solved, and the accuracy of response statements is improved.
Patent Information
- Application Number
- CN202210709686.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-22
AI Technical Summary
The existing multi-round dialogue response statement selection method fails to extract the interaction characteristics in multiple rounds of dialogue to the greatest extent, which affects the accuracy of response statements in multi-round dialogue Q&A of smart human-computer.
By obtaining the semantic representation of the consulting context and the semantic representation of each consulting statement, the first matching feature matrix and the second matching feature matrix are constructed, and interactively fused, the matching score value of the candidate response statement is calculated, and the response statement with the largest matching score is finally selected.
The accuracy of the response statements in multi-wheel dialogue Q&A of smart human-computer has been improved, and the needs of practical applications have been met.
Smart Images

Figure CN114996426B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for question-and-answer interaction. Background Art
[0002] Intelligent human-machine question-and-answer interaction technology refers to the technology that enables a machine to understand and use natural language to achieve human-machine communication. In recent years, with the rise of artificial intelligence, the research on intelligent human-machine question-and-answer interaction technology has become more and more in-depth, from the question-and-answer type of one question and one answer, to the task completion type, and then to the multi-round dialogue question-and-answer type. For multi-round dialogue question-and-answer, it is necessary to select the response that best matches the question from the given candidate responses. The existing multi-round dialogue response selection methods can be divided into two categories: one is to splice the multi-round dialogue and encode it, and then match it with the candidate response sentences; the other is to model based on the interaction features between each sentence in the multi-round dialogue and the candidate response sentences, and then screen the candidate response sentences.
[0003] In the process of implementing the present invention, the inventors found the following problems in the prior art:
[0004] The currently adopted method for selecting multi-round dialogue response sentences either only considers the matching relationship between the overall multi-round dialogue context and the candidate response sentences, or only considers the interaction features between each sentence in the multi-round dialogue and the candidate response sentences, and cannot extract the interaction features in the multi-round dialogue to the greatest extent, which affects the accuracy of the response sentences in intelligent human-machine multi-round dialogue question-and-answer and cannot well meet the actual application needs. Summary of the Invention
[0005] In view of this, an embodiment of the present invention provides a method and device for question-and-answer interaction. Based on the semantic representation of each consultation sentence in the consultation context and the semantic representation of the consultation context, a first matching feature matrix representing each consultation sentence and each candidate response sentence and a second matching feature matrix representing the consultation context and each candidate response sentence are obtained. By interacting and fusing the first matching feature matrix and the second matching feature matrix, the candidate response sentence with the largest matching score in the candidate response sentences is calculated and selected as the response sentence, realizing the screening of candidate response sentences that consider both the matching features at the consultation sentence level and the matching features at the consultation context level, improving the accuracy of the response sentences in intelligent human-machine multi-round dialogue question-and-answer, and better meeting the actual application.
[0006] To achieve the above object, according to one aspect of the embodiments of the present invention, a method for question-and-answer interaction is provided, including:
[0007] According to the received consultation statement, obtain the consultation context related to the consultation statement, encode the consultation context, and obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context;
[0008] According to the semantic representations of each consultation statement in the consultation context and the semantic representations of each candidate response statement, obtain the first matching feature matrix of each consultation statement and the candidate response statement, and according to the semantic representation of the consultation context and the semantic representations of each candidate response statement, obtain the second matching feature matrix of the consultation context and the candidate response statement. The semantic representation of the candidate response statement is obtained by encoding the candidate response statement;
[0009] Perform interactive fusion on the first matching feature matrix and the second matching feature matrix, and calculate the matching score of each candidate response statement;
[0010] Output the candidate response statement with the highest matching score as the response statement for question-and-answer interaction.
[0011] Optionally, encoding the consultation context to obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context includes: based on the vocabulary of each consultation statement in the consultation context, through a word embedding layer, obtain the word vector representation of each consultation statement in the consultation context; splice the word vector representations of each consultation statement in the consultation context to obtain the word vector representation of the consultation context; according to the word vector representation of each consultation statement and the word vector representation of the consultation context, through a semantic analysis model, respectively obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context.
[0012] Optionally, obtaining the first matching feature matrix of each consultation statement and the candidate response statement according to the semantic representations of each consultation statement in the consultation context and the semantic representations of each candidate response statement includes: according to the semantic representations of each consultation statement in the consultation context and the semantic representations of each candidate response statement, by calculating the lexical matching degree of each consultation statement and the candidate response statement, obtain the interaction matrix of each consultation statement and the candidate response statement; through convolution and pooling operations on the interaction matrix, obtain the first matching feature matrix of each consultation statement and the candidate response statement.
[0013] Optionally, obtaining a second matching feature matrix of the consultation context and the candidate response statement according to the semantic representation of the consultation context and the semantic representation of each candidate response statement includes: according to the semantic representation of the consultation context and the semantic representation of each candidate response statement, using an attention mechanism, by extracting the key information matching degree between the consultation context and each candidate response statement, obtaining an attention weight; according to the attention weight, performing a weighted representation on the consultation context to obtain the second matching feature matrix of the consultation context and the candidate response statement.
[0014] Optionally, performing interactive fusion on the first matching feature matrix and the second matching feature matrix, and calculating the matching score of each candidate response statement includes: according to the first matching feature matrix and the second matching feature matrix, calculating an interactive fusion attention weight matrix; according to the interactive fusion attention weight matrix, performing interactive fusion on the first matching feature matrix and the second matching feature matrix to obtain a fused feature representation corresponding to each candidate response statement; performing logistic regression processing on the fused feature representation to obtain the matching score of each candidate response statement.
[0015] Optionally, before performing interactive fusion on the first matching feature matrix and the second matching feature matrix, it further includes: using an activation function to perform transformations on the first matching feature matrix and the second matching feature matrix respectively to obtain a first transformed matching feature matrix and a second transformed matching feature matrix; the performing interactive fusion on the first matching feature matrix and the second matching feature matrix includes: performing interactive fusion on the first transformed matching feature matrix and the second transformed matching feature matrix.
[0016] Optionally, calculating an interactive fusion attention weight matrix according to the first matching feature matrix and the second matching feature matrix includes: according to the first matching feature matrix and the second matching feature matrix, obtaining a second-order interactive feature representation; performing rank reduction splitting on the interactive weight matrix in the second-order interactive feature representation to represent the interactive weight matrix in the form of the multiplication of two two-dimensional interactive sub-weight matrices; based on the rank reduction split interactive weight matrix, performing second-order interaction on the corresponding elements in the first matching feature matrix and the second matching feature matrix pairwise, and performing pooling processing to obtain the interactive fusion attention weight matrix.
[0017] Optionally, the method is implemented based on a pre-trained question-answering interaction model, and the cross entropy is used as the loss function when the question-answering interaction model is trained.
[0018] According to the second aspect of the embodiments of the present invention, a question-answering interaction device is provided, including:
[0019] A semantic representation acquisition module, configured to obtain a consultation context related to the received consultation statement according to the received consultation statement, encode the consultation context, and obtain semantic representations of each consultation statement in the consultation context and a semantic representation of the consultation context;
[0020] A matching feature vector acquisition module, configured to obtain a first matching feature matrix between each consultation statement and the candidate response statement according to the semantic representations of each consultation statement in the consultation context and the semantic representations of each candidate response statement, and obtain a second matching feature matrix between the consultation context and the candidate response statement according to the semantic representation of the consultation context and the semantic representations of each candidate response statement, where the semantic representation of the candidate response statement is obtained by encoding the candidate response statement;
[0021] An interaction fusion module, configured to interactively fuse the first matching feature matrix and the second matching feature matrix, and calculate a matching score for each candidate response statement;
[0022] A response statement acquisition module, configured to output the candidate response statement with the maximum matching score as the response statement for question-and-answer interaction.
[0023] According to a third aspect of an embodiment of the present invention, there is provided an electronic device for question-and-answer interaction, characterized by including:
[0024] One or more processors;
[0025] A storage device, configured to store one or more programs,
[0026] When the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the first aspect of the embodiment of the present invention.
[0027] According to a fourth aspect of an embodiment of the present invention, there is provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method provided in the first aspect of the embodiment of the present invention is implemented.
[0028] One embodiment of the described invention has the following advantages or beneficial effects: By obtaining the consultation context related to the received consultation statement according to the received consultation statement, encoding the consultation context, obtaining the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context; obtaining the first matching feature matrix of each consultation statement and each candidate response statement according to the semantic representation of each consultation statement in the consultation context and the semantic representation of each candidate response statement, obtaining the second matching feature matrix of the consultation context and each candidate response statement according to the semantic representation of the consultation context and the semantic representation of each candidate response statement, and the semantic representation of the candidate response statement is obtained by encoding the candidate response statement; interacting and fusing the first matching feature matrix and the second matching feature matrix, calculating the matching score of each candidate response statement; taking the candidate response statement with the largest matching score as the response statement to output for question-and-answer interaction. The technical solution realizes obtaining the first matching feature matrix representing each consultation statement and each candidate response statement and the second matching feature matrix representing the consultation context and each candidate response statement based on the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context. By interacting and fusing the first matching feature matrix and the second matching feature matrix, calculating and selecting the candidate response statement with the largest matching score in the candidate response statements as the response statement, realizing the screening of candidate response statements that consider both the matching features at the consultation statement level and the matching features at the consultation context level, improving the accuracy of the response statement in intelligent human-machine multi-round dialogue question and answer, and better meeting the actual application. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0030] Figure 1 is a schematic diagram of the main process of the question-and-answer interaction method according to an embodiment of the present invention;
[0031] Figure 2 is a schematic diagram of the overall framework of the question-and-answer interaction model according to an embodiment of the present invention;
[0032] Figure 3 is a schematic diagram of the main modules of the question-and-answer interaction device according to an embodiment of the present invention;
[0033] Figure 4 is an exemplary system architecture diagram to which an embodiment of the present invention can be applied;
[0034] Figure 5 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0036] The currently adopted method for selecting multi-turn dialogue response statements either only considers the interaction features between each statement in the multi-turn dialogue and the candidate response statements, or only considers the matching relationship between the overall multi-turn dialogue context and the candidate response statements. It cannot extract the interaction features in the multi-turn dialogue to the greatest extent, affecting the accuracy of the response statements in intelligent human-machine multi-turn dialogue Q&A and not being able to well meet the actual application needs.
[0037] To solve the above problems existing in the prior art, the present invention proposes a question-and-answer interaction method. Based on the semantic representations of each consultation statement in the consultation context and the semantic representation of the consultation context, a first matching feature matrix representing each consultation statement and each candidate response statement and a second matching feature matrix representing the consultation context and each candidate response statement are obtained. By interacting and fusing the first matching feature matrix and the second matching feature matrix, the candidate response statement with the largest matching score in the candidate response statements is calculated and selected as the response statement, realizing the screening of candidate response statements that consider both the matching features at the consultation statement level and the matching features at the consultation context level, improving the accuracy of the response statements in intelligent human-machine multi-turn dialogue Q&A and better meeting the actual application.
[0038] In the description of the embodiments of the present invention, the nouns involved and their meanings are as follows:
[0039] Transformer: A classic NLP model that uses the Self-Attention mechanism and does not adopt the sequential structure of RNN, enabling the model to be trained in parallel and having global information;
[0040] max-pooling: The maximum pooling operation in CNN;
[0041] Adam optimizer: A first-order optimization algorithm that can replace the traditional stochastic gradient descent process and iteratively updates the neural network weights based on the training data;
[0042] softmax function: A normalized exponential function used in the multi-classification process that maps the outputs of multiple neurons to the interval (0, 1);
[0043] Softmax layer: One of the most commonly used methods for solving multi-classification problems through neural networks;
[0044] ES: Elasticsearch, a search server based on Lucene. It provides a distributed multi-user full-text search engine and is a popular enterprise-level search engine.
[0045] Figure 1 It is a schematic diagram of the main process of the question-and-answer interaction method according to an embodiment of the present invention. As Figure 1 shown, the question-and-answer interaction method according to the embodiment of the present invention includes the following steps S101 to S104.
[0046] Step S101: According to the received consultation statement, obtain the consultation context related to the consultation statement, and encode the consultation context to obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context.
[0047] Specifically, in intelligent human-machine multi-turn question-and-answer interaction, since the consultation statement may be fragmented and may also change and be modified as the conversation progresses, it is necessary to obtain the consultation context related to the received consultation statement according to the received consultation statement. The system forms the consultation context by combining all the consultation statements received before the consultation statement and the currently received consultation statement in the current multi-turn question-and-answer interaction.
[0048] According to an embodiment of the present invention, encoding the consultation context to obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context includes: based on the vocabulary of each consultation statement in the consultation context, obtaining the word vector representation of each consultation statement in the consultation context through a word embedding layer; concatenating the word vector representations of each consultation statement in the consultation context to obtain the word vector representation of the consultation context; according to the word vector representation of each consultation statement and the word vector representation of the consultation context, respectively obtaining the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context through a semantic analysis model.
[0049] Specifically, assume that the above-mentioned consultation context is represented by and s k ={w1, w2,..., w n}, where s k represents the kth consultation statement in the consultation context, and w n represents in s kThe nth word in the consultation statement. For each consultation statement in the consultation context s, through the word embedding layer, the word vector representation of each consultation statement is obtained, and the word vector representation of one of the consultation statements is set as E S =[e1,e2…e n ; According to the word vector representation of each consultation statement, the word vector representations of each consultation statement in the consultation context are concatenated. Assuming there are a total of L words in the consultation context, the word vector representation of the consultation context E C =[e1,e2…e L ; According to the word vector representation of each consultation statement and the word vector representation of the consultation context, through the Transformer semantic analysis model based on the self-attention mechanism, the semantic representation of each consultation statement and the semantic representation of the consultation context are obtained. Taking the word vector representation E S of the above-mentioned one consultation statement as an example, its semantic representation S is:
[0050] S = Transformer Encoder (E S ),
[0051] Correspondingly, the semantic representation C of the consultation context is:
[0052] C = Transformer Encoder (E C ).
[0053] Step S102: Obtain the first matching feature matrix between each consultation statement and the candidate response statement according to the semantic representation of each consultation statement in the consultation context and the semantic representation of each candidate response statement, and obtain the second matching feature matrix between the consultation context and the candidate response statement according to the semantic representation of the consultation context and the semantic representation of each candidate response statement. The semantic representation of the candidate response statement is obtained by encoding the candidate response statement.
[0054] Specifically, according to the above-mentioned consultation context, multiple candidate response statements can be retrieved through ES. Each candidate response statement is encoded according to the above-mentioned encoding method to obtain the semantic representation of the candidate response statement. Taking one of the candidate response statements r as an example, r = {w1,w2,…,w q}, w p represents the qth word in the candidate response statement r. After obtaining the word vector representation E R = [e1,e2…e q of this candidate response statement, through the Transformer semantic analysis model based on the self-attention mechanism, the semantic representation R of this candidate response statement is obtained as:
[0055] R = Transformer Encoder (E R )。
[0056] According to an embodiment of the present invention, obtaining a first matching feature matrix of each query statement and each candidate response statement according to the semantic representations of each query statement and each candidate response statement in the query context includes: according to the semantic representations of each query statement and each candidate response statement in the query context, calculating the lexical matching degree of each query statement and candidate response statement to obtain an interaction matrix of each query statement and candidate response statement; performing convolution and pooling operations on the interaction matrix to obtain the first matching feature matrix of each query statement and each candidate response statement.
[0057] Specifically, according to the semantic representations of each query statement and each candidate response statement in the query context, taking one candidate response statement in the candidate response statements as a matching object, after calculating the lexical matching degree of each query statement in the query context and the matching object, according to the method of calculating the lexical matching degree of each query statement and the matching object described above, traversing other candidate response statements in the candidate response statements, calculating the lexical matching degree of each query statement and candidate response statement, and further obtaining an interaction matrix of each query statement and candidate response statement; on this basis, performing convolution and pooling operations on the interaction matrix, splicing and projecting the extracted matching features to obtain the first matching feature matrix of each query statement and each candidate response statement.
[0058] Exemplarily, taking one candidate response statement in the candidate response statements as a matching object, according to the semantic representation S of one query statement and the semantic representation R of one candidate response statement above, the lexical matching degree of the corresponding positions of the two sentences is calculated as:
[0059] M ij = S i· T R j· ,
[0060] where S i· represents the semantic representation of the i-th word in this query statement, and R j·Denote the semantic representation of the j-th word of this candidate response sentence, and thus obtain the interaction matrix M between this query sentence and this candidate response sentence. According to this method, the interaction matrix between each query sentence and this candidate response sentence can be obtained, and then the interaction matrix between each query sentence and all candidate response sentences can be obtained; for the above interaction matrix M, by performing convolution and pooling operations on this interaction matrix, extract the key matching information between this query sentence and the word candidate response sentence, splice the matching features of max-pooling in the pooling operation, and project them into a low-dimensional matching feature space to obtain the first matching feature vector between the i-th query sentence and this candidate response sentence It can be represented by a convolutional neural network as:
[0061]
[0062] In a similar manner as above, calculate the first matching feature vectors for the remaining m - 1 query sentences in the query context in turn to obtain the first matching feature matrix between each query sentence and this candidate response sentence For all candidate response sentences, similarly, by performing convolution and pooling operations on the interaction matrix, calculate the first matching feature matrix between each query sentence and the candidate response sentence, and then obtain the set of first matching feature matrices between each query sentence and all candidate response sentences
[0063] According to another embodiment of the present invention, obtaining the second matching feature matrix between the query context and the candidate response sentence according to the semantic representation of the query context and the semantic representation of each candidate response sentence includes: according to the semantic representation of the query context and the semantic representation of each candidate response sentence, adopting an attention mechanism, by extracting the key information matching degree between the query context and each candidate response sentence, obtaining the attention weight; according to the attention weight, performing a weighted representation on the query context to obtain the second matching feature matrix between the query context and the candidate response sentence
[0064] Specifically, since the scale of the query context is generally relatively large, for the extraction of the matching degree between the query context and the candidate response sentence, it is preferably to adopt an attention mechanism to extract the matching degree of key information. According to the semantic representation of the query context and the semantic representation of each candidate response sentence, adopt an attention mechanism to extract the key information matching degree between the query context and each candidate response sentence to obtain the attention weight. For example, for a candidate response sentence response, the key information matching degree d between the query context context and response i,j is:
[0065] d i,j =tanh(C j·T WR i· +b),
[0066] where C j· is the semantic representation of the j-th word in the consultation context context, R i· is the semantic representation of the i-th word in the candidate response sentence response, W is the weight matrix, b is the bias term, and correspondingly, its attention weight α i,j is:
[0067]
[0068] Similarly, following this method, traverse all candidate response sentences, extract the key information matching degree between the consultation context and each candidate response sentence, and obtain the attention weight; according to the above attention weight, perform weighted representation on the consultation context context to obtain the second matching feature matrix of the consultation context context and the candidate response sentence response, and then obtain the set of second matching feature matrices of the consultation context context and all candidate response sentences. For example, for a candidate response sentence response as above, the consultation context context can be represented as the weighted representation ∑ j α i,j C j· to obtain the second matching feature vector of the i-th word of the consultation context context and the candidate response sentence response where denotes element-wise multiplication, and finally obtain the second matching feature matrix of each word in the consultation context context and the candidate response sentence response n represents the total number of words in response. Similarly, following this method, traverse all candidate response sentences to obtain the set of second matching feature matrices of the consultation context context and all candidate response sentences.
[0069] Through the calculation of the above first matching feature matrix and second matching feature matrix, both the matching features at the sentence level between the consultation context and the candidate response sentence and the finer-grained lexical matching features between the consultation context and the candidate response sentence are obtained, realizing the extraction of multi-level matching features of the consultation sentence and the candidate response sentence.
[0070] Step S103, interactively fuse the first matching feature matrix and the second matching feature matrix, and calculate the matching score of each candidate response sentence.
[0071] Specifically, each candidate response statement in the candidate response statements corresponds to a first matching feature matrix and a second matching feature matrix. The first matching feature matrix and the second matching feature matrix of each candidate response statement are interactively fused, and the matching score of each candidate response statement is calculated.
[0072] According to an embodiment of the present invention, interactively fusing the first matching feature matrix and the second matching feature matrix, and calculating the matching score of each candidate response statement includes: calculating an attention weight matrix of interactive fusion according to the first matching feature matrix and the second matching feature matrix; interactively fusing the first matching feature matrix and the second matching feature matrix according to the attention weight matrix of interactive fusion to obtain a fusion feature representation corresponding to each candidate response statement; performing logistic regression processing on the fusion feature representation to obtain the matching score of each candidate response statement.
[0073] According to another embodiment of the present invention, calculating an attention weight matrix of interactive fusion according to the first matching feature matrix and the second matching feature matrix includes: obtaining a second-order interactive feature representation according to the first matching feature matrix and the second matching feature matrix; performing rank reduction splitting on the interactive weight matrix in the second-order interactive feature representation to represent the interactive weight matrix in the form of the multiplication of two two-dimensional interactive sub-weight matrices; based on the interactive weight matrix after rank reduction splitting, performing second-order interaction on the corresponding elements in the first matching feature matrix and the second matching feature matrix pairwise, and performing pooling processing to obtain an attention weight matrix of interactive fusion.
[0074] Exemplarily, for the second matching feature vector of the i-th word in the above second matching feature matrix and the first matching feature vector of the j-th query statement in the first matching feature matrix its second-order interactive feature representation is:
[0075]
[0076] where V i is the i-th matrix of the interactive weight matrix To reduce the number of parameters, perform rank reduction splitting on the interactive weight matrix V, and use the multiplication of two
[0077] low-rank two-dimensional matrices to replace it, where s ≤ min(d w , d c ); based on the interactive weight matrix after rank reduction splitting, the second matching feature vector of the i-th word in the second matching feature matrix and the corresponding elements of the first matching feature vector of the j-th consultation statement in the first matching feature matrix are multiplied pairwise, and the pooling matrix is used to also generate the second-order interaction feature vector with the effect of obtaining:
[0078]
[0079] According to the above representation of the second-order interaction feature vector, each element in the attention weight matrix A for interaction fusion can be obtained as:
[0080]
[0081] where a scalar A ij is obtained. For the convenience of subsequent operation efficiency, based on the element A ij , the calculation method of the attention weight matrix A for interaction fusion can be obtained:
[0082]
[0083] H′ t = u(H t ; V, b v )
[0084]
[0085] where the g function represents element-wise softmax, and the u function is: u(X; θ, b θ ) = σ(Xθ + b θ ), σ represents the activation function, and b θ represents the bias term.
[0086] Based on the attention weight matrix A for interaction fusion obtained above, the first matching feature matrix and the second matching feature matrix are interactively fused to obtain the fused feature representation corresponding to each candidate response statement.
[0087] According to another embodiment of the present invention, before the first matching feature matrix and the second matching feature matrix are interactively fused, it further includes: using the activation function to respectively transform the first matching feature matrix and the second matching feature matrix to obtain the first transformed matching feature matrix and the second transformed matching feature matrix; the interactive fusion of the first matching feature matrix and the second matching feature matrix includes: performing interactive fusion on the first transformed matching feature matrix and the second transformed matching feature matrix.
[0088] Specifically, before the interaction and fusion of the first matching feature matrix and the second matching feature matrix, the activation function u function is used to transform the first matching feature matrix H t to obtain the first transformed matching feature matrix H″ t :
[0089] H″ t = u(H t ; V1′, b′ V );
[0090] where V1′ is the parameter of the activation function u function, and b′ V is the bias term. Similarly, the second matching feature matrix H r is also transformed to obtain the second transformed matching feature matrix H″ r :
[0091] H″ r = u(H r ; U1′, b′ U );
[0092] where U1′ is the parameter of the activation function u function, and b′ U is the bias term.
[0093] According to the interaction and fusion attention weight matrix A calculated above, combined with the first transformed matching feature matrix and the second transformed matching feature matrix, the first transformed matching feature matrix and the second transformed matching feature matrix are interactively fused.
[0094] Exemplarily, for the first transformed matching feature matrix H″ t and the second transformed matching feature matrix H″ r , take the i-th dimension feature, that is, the i-th column vector of H″ t and H″ r : According to the interaction and fusion attention weight matrix A, perform second-order interaction and fusion, that is:
[0095]
[0096] where f′ i is a scalar. To make it a vector, the pooling matrix is still introduced.
[0097] f F = P T f′;
[0098] where, f′ i is an element of f F and is one of the elements in According to the above method, perform second-order interaction on the first matching feature matrix and the second matching feature matrix corresponding to each candidate response statement to obtain the fused feature representation corresponding to each candidate response statement.
[0099] Based on the fused feature representation corresponding to each candidate response statement above, input the fused feature representation into the softmax layer for logistic regression processing to obtain the matching score of each candidate response statement. For example, based on the fused feature representation f F of the candidate response statement response above, its matching score z is:
[0100] z = softmax(W1f F )
[0101] where W1 is an operation parameter.
[0102] Step S104, output the candidate response statement with the highest matching score as the response statement for question-and-answer interaction.
[0103] According to an embodiment of the present invention, the method is implemented based on a pre-trained question-and-answer interaction model, and cross-entropy is used as the loss function when training the question-and-answer interaction model.
[0104] Specifically, the implementation of the method of the present invention is based on a pre-trained question-and-answer interaction model. According to the above question-and-answer interaction method, a question-and-answer interaction model is established, and the cross-entropy function is used as the loss function L:
[0105] L = CE(p, y);
[0106] where CE represents the cross-entropy function, y is the true label, and gradient descent training is performed using an Adam-based optimizer.
[0107] Through the first matching feature matrix at the sentence level between the above consultation context and the candidate response statement, and the second matching feature matrix at a finer-grained lexical level between the consultation context and the candidate response statement, the established question-and-answer interaction model takes into account matching information at multiple levels. Based on the matching features at multiple levels, a second-order interaction fusion method is used to effectively fuse the matching features at different levels, making the predicted matching score more accurate and improving the accuracy of the response statement in multi-turn dialogue question-and-answer.
[0108] Figure 2It is a schematic diagram of the overall framework of the question - answering interaction model according to an embodiment of the present invention. In the figure, taking a candidate response statement "response" as an example, the consultation context is obtained based on the consultation statement. The consultation context consists of consultation statement 1, consultation statement 2, and consultation statement 3. The candidate response statement "response", consultation statement 1, consultation statement 2, consultation statement 3, and the consultation context composed of the concatenation of the 3 consultation statements are encoded to obtain their respective semantic representations, such as the hollow circles in the figure; the matching degrees between consultation statement 1, 2, 3 and the candidate response statement "response" are calculated respectively to obtain the first matching feature matrix As the dark solid circles in the figure, the matching degree between the consultation context and the candidate response statement "response" is calculated to obtain the second matching feature matrix As the light solid circles in the figure; the first matching feature matrix is combined to obtain The second matching feature matrix is combined to obtain Regarding H t and H r The interaction weight matrix in the second - order interaction feature representation of is rank - reduced and split into two low - rank two - dimensional matrices V and U. Through the second - order interaction calculation of H t V and H r U, the attention weight matrix A of interaction fusion is obtained. The activation function is used to transform H t and H r to obtain H t V′ and H r U′. For the transformed H t V′ and H r U′, combined with the attention weight matrix A of interaction fusion, second - order interaction fusion is performed to obtain the fusion feature representation corresponding to the candidate response statement "response". Finally, the fusion feature representation is input into the softmax layer to predict the final matching score.
[0109] Figure 3 is a schematic diagram of the main modules of the question - answering interaction device according to an embodiment of the present invention. As Figure 3 shown, the question - answering interaction device 300 mainly includes a semantic representation acquisition module 301, a matching feature vector acquisition module 302, an interaction fusion module 303, and a response statement acquisition module 304.
[0110] The semantic representation acquisition module 301 is used to obtain the consultation context related to the received consultation statement according to the received consultation statement, encode the consultation context, and obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context;
[0111] A matching feature vector obtaining module 302, configured to obtain a first matching feature matrix between each query statement and each candidate response statement according to the semantic representations of each query statement in the query context and the semantic representations of each candidate response statement, and obtain a second matching feature matrix between the query context and each candidate response statement according to the semantic representation of the query context and the semantic representations of each candidate response statement, where the semantic representation of the candidate response statement is obtained by encoding the candidate response statement;
[0112] An interaction and fusion module 303, configured to interactively fuse the first matching feature matrix and the second matching feature matrix, and calculate the matching scores of each candidate response statement;
[0113] A response statement obtaining module 304, configured to output the candidate response statement with the highest matching score as the response statement for question-and-answer interaction.
[0114] Specifically, the semantic representation obtaining module 301 may further be configured to: based on the vocabulary of each query statement in the query context, obtain the word vector representation of each query statement in the query context through a word embedding layer; splice the word vector representations of each query statement in the query context to obtain the word vector representation of the query context; and respectively obtain the semantic representation of each query statement in the query context and the semantic representation of the query context through a semantic analysis model according to the word vector representation of each query statement and the word vector representation of the query context.
[0115] Specifically, the matching feature vector obtaining module 302 may further be configured to: according to the semantic representations of each query statement in the query context and the semantic representations of each candidate response statement, obtain an interaction matrix between each query statement and each candidate response statement by calculating the lexical matching degree between each query statement and the candidate response statement; and obtain the first matching feature matrix between each query statement and the candidate response statement by performing convolution and pooling operations on the interaction matrix.
[0116] Specifically, the matching feature vector obtaining module 302 may further be configured to: according to the semantic representation of the query context and the semantic representations of each candidate response statement, adopt an attention mechanism to obtain attention weights by extracting the key information matching degree between the query context and each candidate response statement; and obtain the second matching feature matrix between the query context and the candidate response statement by performing a weighted representation on the query context according to the attention weights.
[0117] Specifically, the interaction fusion module 303 can also be used to: calculate an attention weight matrix for interaction fusion based on the first matching feature matrix and the second matching feature matrix; perform interaction fusion on the first matching feature matrix and the second matching feature matrix according to the attention weight matrix for interaction fusion to obtain a fused feature representation corresponding to each candidate response statement; perform logistic regression processing on the fused feature representation to obtain a matching score for each candidate response statement.
[0118] Specifically, the question-and-answer interaction device 300 may further include a transformation module (not shown in the figure) for: before performing interaction fusion on the first matching feature matrix and the second matching feature matrix, respectively using an activation function to transform the first matching feature matrix and the second matching feature matrix to obtain a first transformed matching feature matrix and a second transformed matching feature matrix; the interaction fusion module 303 can also be used to: perform interaction fusion on the first transformed matching feature matrix and the second transformed matching feature matrix.
[0119] Specifically, the interaction fusion module 303 can also be used to: obtain a second-order interaction feature representation based on the first matching feature matrix and the second matching feature matrix; perform rank reduction splitting on the interaction weight matrix in the second-order interaction feature representation to represent the interaction weight matrix in the form of the product of two two-dimensional interaction sub-weight matrices; based on the interaction weight matrix after rank reduction splitting, perform pairwise second-order interaction on the corresponding elements in the first matching feature matrix and the second matching feature matrix, and perform pooling processing to obtain an attention weight matrix for interaction fusion.
[0120] Specifically, the question-and-answer interaction device 300 is implemented based on a pre-trained question-and-answer interaction model, and the cross entropy is used as the loss function when the question-and-answer interaction model is trained.
[0121] Figure 4 It is an exemplary system architecture diagram in which the embodiments of the present invention can be applied.
[0122] As Figure 4 shown, the system architecture 400 may include terminal devices 401, 402, 403, a network 404, and a server 405. The network 404 is used to provide a medium for communication links between the terminal devices 401, 402, 403 and the server 405. The network 404 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0123] Users can use terminal devices 401, 402, and 403 to interact with server 405 via network 404 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 401, 402, and 403, such as question-and-answer interaction applications, human-machine dialogue applications, etc. (for example only).
[0124] Terminal devices 401, 402, and 403 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop portable computers, desktop computers, and so on.
[0125] Server 405 can be a server that provides various services, such as a background management server that supports the question-and-answer interaction performed by users using terminal devices 401, 402, and 403 (for example only). The background management server can, according to the received consultation statement, obtain the consultation context related to the consultation statement, encode the consultation context, and obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context; according to the semantic representation of each consultation statement in the consultation context and the semantic representation of each candidate response statement, obtain the first matching feature matrix of each consultation statement and the candidate response statement, and according to the semantic representation of the consultation context and the semantic representation of each candidate response statement, obtain the second matching feature matrix of the consultation context and the candidate response statement, where the semantic representation of the candidate response statement is obtained by encoding the candidate response statement; perform interactive fusion on the first matching feature matrix and the second matching feature matrix, calculate the matching score of each candidate response statement; use the candidate response statement with the maximum matching score as the response statement for output to perform question-and-answer interaction and other processing, and feedback the processing result (such as the response statement, etc. - for example only) to the terminal device.
[0126] It should be noted that the method for question-and-answer interaction provided in the embodiments of the present invention is generally executed by server 405. Correspondingly, the device for question-and-answer interaction is generally arranged in server 405.
[0127] It should be understood that Figure 4 the numbers of terminal devices, networks, and servers in
[0128] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 5 Shown below with reference to Figure 5 is a schematic structural diagram of a computer system 500 of a terminal device or a server suitable for implementing the embodiments of the present invention.
[0129] As Figure 5As shown, computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or programs loaded from a storage section 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0130] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 510 as needed so that a computer program read therefrom can be installed into the storage section 508 as needed.
[0131] Specifically, according to an embodiment disclosed by the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment disclosed by the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by a central processing unit (CPU) 501, the above functions defined in the system of the present invention are executed.
[0132] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0134] The units involved in the embodiments of the present invention can be implemented in software or in hardware. The described units can also be provided in a processor. For example, a processor can be described as including: a semantic representation acquisition module, a matching feature vector acquisition module, an interaction fusion module, and a snapshot module.
[0135] Among them, the names of these modules do not constitute a limitation to the modules themselves in some cases. For example, the interaction fusion module can also be described as "a module for interactively fusing the first matching feature matrix and the second matching feature matrix and calculating the matching scores of each candidate response statement".
[0136] On the other hand, the present invention also provides a computer-readable medium, which can be included in the device described in the embodiments; or can exist separately without being assembled into the device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: obtaining an inquiry context related to the received inquiry statement, encoding the inquiry context to obtain the semantic representation of each inquiry statement in the inquiry context and the semantic representation of the inquiry context; obtaining a first matching feature matrix of each inquiry statement and each candidate response statement according to the semantic representation of each inquiry statement in the inquiry context and the semantic representation of each candidate response statement, obtaining a second matching feature matrix of the inquiry context and the candidate response statement according to the semantic representation of the inquiry context and the semantic representation of each candidate response statement, and the semantic representation of the candidate response statement is obtained by encoding the candidate response statement; interactively fusing the first matching feature matrix and the second matching feature matrix and calculating the matching scores of each candidate response statement; and outputting the candidate response statement with the highest matching score as the response statement for question-answering interaction.
[0137] According to the technical solution of the embodiment of the present invention, it has the following advantages or beneficial effects: by obtaining the consultation context related to the received consultation statement according to the received consultation statement, encoding the consultation context, obtaining the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context; obtaining the first matching feature matrix of each consultation statement and the candidate response statement according to the semantic representation of each consultation statement in the consultation context and the semantic representation of each candidate response statement, obtaining the second matching feature matrix of the consultation context and the candidate response statement according to the semantic representation of the consultation context and the semantic representation of each candidate response statement, and the semantic representation of the candidate response statement is obtained by encoding the candidate response statement; interacting and fusing the first matching feature matrix and the second matching feature matrix to calculate the matching score of each candidate response statement; taking the candidate response statement with the largest matching score as the response statement to output for question-and-answer interaction. The technical solution realizes obtaining the first matching feature matrix representing each consultation statement and each candidate response statement and the second matching feature matrix representing the consultation context and each candidate response statement based on the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context. By interacting and fusing the first matching feature matrix and the second matching feature matrix, calculating and selecting the candidate response statement with the largest matching score among the candidate response statements as the response statement, it realizes the screening of candidate response statements that consider both the matching features at the consultation statement level and the matching features at the consultation context level, improves the accuracy of the response statement in intelligent human-machine multi-round dialogue question and answer, and better meets the actual application.
[0138] The specific implementation manners do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for question-and-answer interaction, characterized in that, Including: According to the received consultation statement, obtain the consultation context related to the consultation statement, encode the consultation context to obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context; Obtain a first matching feature matrix between each consultation statement and the candidate response statement according to the semantic representation of each consultation statement in the consultation context and the semantic representation of each candidate response statement, and obtain a second matching feature matrix between the consultation context and the candidate response statement according to the semantic representation of the consultation context and the semantic representation of each candidate response statement. The semantic representation of the candidate response statement is obtained by encoding the candidate response statement; Interactively fuse the first matching feature matrix and the second matching feature matrix, and calculate the matching score of each candidate response statement, including: obtaining a second-order interaction feature representation according to the first matching feature matrix and the second matching feature matrix; performing rank reduction splitting on the interaction weight matrix in the second-order interaction feature representation to represent the interaction weight matrix in the form of the product of two two-dimensional interaction sub-weight matrices; based on the interaction weight matrix after rank reduction splitting, perform second-order interaction on the corresponding elements in the first matching feature matrix and the second matching feature matrix pairwise, and perform pooling processing to obtain an interaction-fused attention weight matrix; according to the interaction-fused attention weight matrix, perform interactive fusion on the first matching feature matrix and the second matching feature matrix to obtain a fused feature representation corresponding to each candidate response statement; perform logistic regression processing on the fused feature representation to obtain the matching score of each candidate response statement; Output the candidate response statement with the maximum matching score as the response statement for question-and-answer interaction.
2. The method according to claim 1, characterized in that, Encoding the consultation context to obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context includes: Based on the vocabulary of each consultation statement in the consultation context, obtain the word vector representation of each consultation statement in the consultation context through a word embedding layer; Concatenate the word vector representations of each consultation statement in the consultation context to obtain the word vector representation of the consultation context; According to the word vector representation of each consultation statement and the word vector representation of the consultation context, respectively obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context through a semantic analysis model.
3. The method according to claim 1, characterized in that Obtaining a first matching feature matrix between each consultation statement and the candidate response statement according to the semantic representation of each consultation statement in the consultation context and the semantic representation of each candidate response statement includes: According to the semantic representation of each consultation statement in the consultation context and the semantic representation of each candidate response statement, calculate the lexical matching degree between each consultation statement and the candidate response statement to obtain an interaction matrix between each consultation statement and the candidate response statement; Through convolution and pooling operations on the interaction matrix, obtain a first matching feature matrix between each consultation statement and the candidate response statement.
4. The method according to claim 1, characterized in that, Obtaining a second matching feature matrix between the consultation context and the candidate response statement according to the semantic representation of the consultation context and the semantic representation of each candidate response statement, including: According to the semantic representation of the consultation context and the semantic representation of each candidate response statement, using an attention mechanism, by extracting the key information matching degree between the consultation context and each candidate response statement, obtaining attention weights; According to the attention weights, performing weighted representation on the consultation context to obtain the second matching feature matrix between the consultation context and the candidate response statement.
5. The method according to claim 1, wherein Before performing interactive fusion on the first matching feature matrix and the second matching feature matrix, it further includes: Using activation functions to transform the first matching feature matrix and the second matching feature matrix respectively to obtain a first transformed matching feature matrix and a second transformed matching feature matrix; The performing interactive fusion on the first matching feature matrix and the second matching feature matrix includes: Performing interactive fusion on the first transformed matching feature matrix and the second transformed matching feature matrix.
6. The method according to claim 1, wherein The method is implemented based on a pre-trained question-answering interaction model, and the cross-entropy is used as the loss function when the question-answering interaction model is trained.
7. An apparatus for question-and-answer interaction, characterized in that, Including: A semantic representation acquisition module, configured to obtain a consultation context related to the received consultation statement according to the received consultation statement, encode the consultation context to obtain the semantic representation of each consultation statement in the consultation context and the semantic representation of the consultation context; A matching feature vector acquisition module, configured to obtain a first matching feature matrix between each consultation statement and the candidate response statement according to the semantic representation of each consultation statement in the consultation context and the semantic representation of each candidate response statement, and obtain a second matching feature matrix between the consultation context and the candidate response statement according to the semantic representation of the consultation context and the semantic representation of each candidate response statement, and the semantic representation of the candidate response statement is obtained by encoding the candidate response statement; An interactive fusion module, configured to perform interactive fusion on the first matching feature matrix and the second matching feature matrix and calculate the matching score of each candidate response statement; The interactive fusion module is further configured to obtain a second-order interaction feature representation according to the first matching feature matrix and the second matching feature matrix; perform rank reduction decomposition on the interaction weight matrix in the second-order interaction feature representation to represent the interaction weight matrix in the form of the product of two two-dimensional interaction sub-weight matrices; Based on the rank reduction decomposed interaction weight matrix, perform second-order interaction on the corresponding elements in the first matching feature matrix and the second matching feature matrix pairwise and perform pooling processing to obtain an attention weight matrix for interactive fusion; according to the attention weight matrix for interactive fusion, perform interactive fusion on the first matching feature matrix and the second matching feature matrix to obtain a fusion feature representation corresponding to each candidate response statement; Performing logistic regression processing on the fusion feature representation to obtain the matching score of each candidate response statement; A response statement acquisition module, configured to output the candidate response statement with the largest matching score as the response statement for question-and-answer interaction.
8. A mobile electronic device terminal, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-6 is implemented.