Semantic matching method, device and medium
Through word segmentation and splicing processing, and feature extraction is performed using embedded networks and transform networks, the semantic matching prediction results of shallow transform layer are output in advance, solving the problems of expensive model computing resources and memory tightness in the prior art, and efficient semantic matching prediction is achieved.
Patent Information
- Application Number
- CN202110073897.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-20
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-01-20
AI Technical Summary
In the semantic matching task, due to too many parameters, the model is expensive to calculate resources and memory tight when it is launched for service, making it difficult to achieve efficient prediction.
Generative word sequences are processed through word segmentation and splicing, and converted into word vectors using the embedded network, and layer-by-layer feature extraction and semantic matching prediction are performed through the transformation network and the classification network, and the prediction results of the shallow transformation layer are output in advance to reduce the computational burden.
On the premise of ensuring that the model performance does not decline, accelerated prediction of semantic matching models is realized, reducing the consumption of computing resources and improving the model's inference speed.
Smart Images

Figure CN113407664B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of artificial intelligence, and more specifically, to a semantic matching method, device and medium. Background Art
[0002] The semantic matching problem between user questions and standard question libraries in search scenarios has undergone a technical evolution from unsupervised learning to supervised learning, and from traditional machine learning to deep learning. The earliest algorithms such as TF-IDF (termfrequency–inverse document frequency), LD (Levenshtein Distance), and LCS (Longest Common Subsequence) calculated the lexical overlap between two sentences to obtain semantic matching. However, since these algorithms determine semantic matching based on word overlap and co-occurrence information, they do not mine enough semantic information of the sentence itself and cannot achieve a deep understanding of user questions. At the same time, when user questions are long, it is impossible to locate keywords or key phrases.
[0003] In recent years, deep learning algorithms have made breakthrough progress, and the application of deep learning in semantic matching tasks has attracted more and more attention. Although the semantic matching model based on deep learning can mine deeper semantic information, the model introduces too many parameters (for example, more than 100 million parameters), making it difficult to put the model into service. Summary of the invention
[0004] In view of the above situation, it is expected to provide a new semantic matching method, device and medium, which can achieve the purpose of accelerating prediction without significantly decreasing the model performance, and solve the problems of expensive computing resources and tight memory.
[0005] According to one aspect of the present disclosure, a semantic matching method is provided, comprising: performing word segmentation and concatenation processing on a first input text and a second input text to obtain a first word sequence; providing the first word sequence to an embedding network, and converting the first word sequence into a first word vector through the embedding network; providing the first word vector to a transformation network, wherein the transformation network also includes first to Nth transformation layers connected in series, wherein N is an integer greater than 1, the first transformation layer receives the first word vector as an input vector and other transformation layers receive feature vectors generated by the previous transformation layer connected in series as their input vectors, each transformation layer performs feature extraction on the input vector and generates a feature vector, and each transformation layer has a classification network corresponding thereto; starting from the first transformation layer, performing the following operations layer by layer until a semantic matching result of the first text and the second text is generated: providing the feature vector generated by the transformation layer to the classification network corresponding thereto; using the classification network corresponding to the transformation layer to generate a semantic matching prediction result based on the feature vector received by it; and generating a semantic matching result of the first text and the second text based on the semantic matching prediction result when the semantic matching prediction result meets a predetermined condition.
[0006] In addition, in the method according to the embodiment of the present disclosure, the semantic matching prediction result includes a probability value indicating whether the first text and the second text match, and the predetermined condition includes: the probability value is greater than a predetermined threshold.
[0007] In addition, in the method according to an embodiment of the present disclosure, when the semantic matching prediction result of the classification network corresponding to the i-th transformation layer meets a predetermined condition, the operation of other transformation layers in the transformation network and their corresponding classification networks is stopped, where i is an integer greater than or equal to 1.
[0008] In addition, in the method according to the embodiment of the present disclosure, each of the multiple classification networks includes: a fully connected layer, a classification transformation layer and a normalization layer, and wherein the classification network generates a semantic matching prediction result based on the feature vector it receives, including: the fully connected layer receives the feature vector output by the transformation layer corresponding to the classification network, and the fully connected layer transforms the feature vector into a feature vector of a dimension corresponding to the number of categories of the semantic matching prediction result; the feature vector output by the fully connected layer is provided to the classification transformation layer, and the classification transformation layer outputs the transformed feature vector; the transformed feature vector is provided to the normalization layer, and the normalization layer performs normalization on each element therein, and the normalized feature vector is used as the semantic matching prediction result.
[0009] In addition, in the method according to the embodiment of the present disclosure, each network is trained by the following processing: using a first training data set to train the embedding network, the transformation network and the classification network corresponding to the Nth transformation layer; while keeping the parameters of the embedded network, the transformation network and the classification network corresponding to the Nth transformation layer fixed, using a second training data set to train the classification networks corresponding to the first to (N-1)th transformation layers.
[0010] In addition, in the method according to the embodiment of the present disclosure, the first training data set includes multiple training data, each training data includes a third text and a fourth text and a real semantic matching result of the third text and the fourth text, wherein, using the first training data set, training the embedding network, the transformation network and the classification network corresponding to the Nth transformation layer includes: in at least a part of the training data in the first training data set, performing word segmentation and splicing processing on the third text and the fourth text of each training data to obtain a second word sequence; providing the second word sequence to the embedding network, and converting the second word sequence into a second word vector through the embedding network; providing the second word vector to the transformation network, and providing the feature vector output by the Nth transformation layer in the transformation network to the classification network corresponding thereto; calculating the first loss function between the semantic matching prediction result output by the classification network corresponding to the Nth transformation layer and the real semantic matching result; based on the first loss function, training the embedding network, the transformation network and the classification network corresponding to the Nth transformation layer.
[0011] In addition, in the method according to the embodiment of the present disclosure, the second training data set includes multiple training data, each training data includes a fifth text and a sixth text, wherein the second training data set is used to train the classification network corresponding to the first to (N-1)th transformation layers: in at least a part of the training data in the second training data set, the fifth text and the sixth text of each training data are segmented and concatenated to obtain a third word sequence; the third word sequence is provided to the embedding network, and the third word sequence is converted into a third word vector through the embedding network; the third word vector is provided to the transformation network, and the feature vectors output by the first to (N-1)th transformation layers in the transformation network are respectively provided to the corresponding classification networks; a second loss function is calculated between multiple semantic matching prediction results output by multiple classification networks corresponding to the first to (N-1)th transformation layers and the semantic matching prediction results output by the classification network corresponding to the Nth transformation layer; based on the second loss function, the classification network corresponding to the first to (N-1)th transformation layers is trained.
[0012] In addition, in the method according to the embodiment of the present disclosure, a second loss function is calculated between multiple semantic matching prediction results output by multiple classification networks corresponding to the first to (N-1)th transformation layers and the semantic matching prediction results output by the classification network corresponding to the Nth transformation layer, including: calculating the KL divergence between the semantic matching prediction results output by each of the multiple classification networks corresponding to the first to (N-1)th transformation layers and the semantic matching prediction results output by the classification network corresponding to the Nth transformation layer; and taking the sum of all calculated KL divergences as the second loss function.
[0013] According to another aspect of the present disclosure, a semantic matching method is provided, comprising: receiving an input user question; determining, according to the method described above, for each of at least a portion of standard questions in a standard question and answer library, a semantic matching result with the user question; and displaying a standard question that best matches the semantics of the user question.
[0014] According to another aspect of the present disclosure, a semantic matching device is provided, comprising: a memory for storing a computer program thereon; a processor for performing the following processing when executing the computer program: performing word segmentation and concatenation processing on an input first text and a second text to obtain a first word sequence; providing the first word sequence to an embedding network, and converting the first word sequence into a first word vector through the embedding network; providing the first word vector to a transformation network, wherein the transformation network also includes first to Nth transformation layers connected in series, wherein the first transformation layer receives the first word vector as an input vector and other transformation layers receive feature vectors generated by the previous transformation layer connected in series with it as their input vectors, each transformation layer performs feature extraction on the input vector and generates a feature vector, and each transformation layer has a classification network corresponding thereto; and performing the following operations layer by layer starting from the first transformation layer until a semantic matching result of the first text and the second text is generated: providing the feature vector generated by the transformation layer to a classification network corresponding thereto; using the classification network corresponding to the transformation layer to generate a semantic matching prediction result based on the feature vector received by it; and generating a semantic matching result of the first text and the second text based on the semantic matching prediction result when the semantic matching prediction result meets a predetermined condition.
[0015] In addition, in the device according to the embodiment of the present disclosure, the semantic matching prediction result includes a probability value indicating whether the first text and the second text match, and the predetermined condition includes: the probability value is greater than a predetermined threshold.
[0016] In addition, in the device according to the embodiment of the present disclosure, each of the multiple classification networks includes: a fully connected layer, a classification transformation layer and a normalization layer, and the classification network corresponding to the transformation layer is used to generate a semantic matching prediction result based on the feature vector received by it, including: the fully connected layer receives the feature vector output by the transformation layer corresponding to the classification network, and the fully connected layer transforms the feature vector into a feature vector of a dimension corresponding to the number of categories of the semantic matching prediction result; the feature vector output by the fully connected layer is provided to the classification transformation layer, and the classification transformation layer outputs the transformed feature vector; and the transformed feature vector is provided to the normalization layer, and the normalization layer performs normalization on each element therein, and the normalized feature vector is used as the semantic matching prediction result.
[0017] In addition, in the device according to the embodiment of the present disclosure, each network is trained by the following processing: using a first training data set to train the embedding network, the transformation network and the classification network corresponding to the Nth transformation layer; and using a second training data set to train the classification networks corresponding to the first to (N-1)th transformation layers while keeping the parameters of the embedded network, the transformation network and the classification network corresponding to the Nth transformation layer fixed.
[0018] According to another aspect of the present disclosure, a semantic matching device is provided, comprising: a word segmentation and splicing unit, configured to perform word segmentation and splicing processing on a first text and a second text input to obtain a first word sequence; an embedding unit, comprising an embedding network, and converting the first word sequence into a first word vector through the embedding network; a transformation unit, comprising a transformation network, wherein the transformation network comprises first to Nth transformation layers connected in series, wherein N is an integer greater than 1, the first transformation layer receives the first word vector as an input vector and other transformation layers receive a feature vector generated by the previous transformation layer connected in series with it as its input vector, and each transformation layer extracts features from the input vector and generates a feature vector; and a classification unit, comprising N classification networks. , each transformation layer has a one-to-one corresponding classification network, each classification network is used to receive feature vectors from the corresponding transformation layer, and generate semantic matching prediction results of the first text and the second text corresponding to the transformation layer; wherein the transformation network and N classification networks are configured to: start from the first transformation layer, perform the following operations layer by layer until the semantic matching results of the first text and the second text are generated: provide the feature vector generated by the transformation layer to the corresponding classification network; use the classification network corresponding to the transformation layer to generate semantic matching prediction results based on the feature vector received; and generate semantic matching results of the first text and the second text based on the semantic matching prediction results when the semantic matching prediction results meet predetermined conditions.
[0019] In addition, according to another aspect of the present disclosure, a semantic matching device is provided, including: a receiving device for receiving an input user question; a semantic matching device according to the above text, for determining, for each of at least a part of the standard questions in the standard question and answer library, a semantic matching result with the user question; and a display device for displaying a standard question that best matches the semantics of the user question.
[0020] According to another aspect of the present disclosure, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described above is performed.
[0021] By using the semantic matching method, device and medium according to the embodiments of the present disclosure, once the prediction result output by the classification network corresponding to the shallow transformation layer meets the predetermined conditions, such as a high enough confidence level, the subsequent deep transformation layer will no longer be processed. Through such a setting, when the user question is relatively simple, the satisfactory semantic matching prediction result of the classification network corresponding to the shallow transformation layer can be output in advance without using the semantic matching prediction result of the classification network corresponding to the last transformation layer. Thereby, the computational burden of the semantic matching model is reduced, and the reasoning speed of the model is improved, and the performance of the semantic matching calculation will not be significantly reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a flowchart illustrating a process of a semantic matching method according to an embodiment of the present disclosure;
[0023] Figure 2 is a schematic structural diagram illustrating a semantic matching model according to an embodiment of the present disclosure;
[0024] Figure 3 is a flowchart illustrating a first stage training process according to an embodiment of the present disclosure;
[0025] Figure 4 is a flowchart illustrating a training process of the second stage according to an embodiment of the present disclosure;
[0026] Figure 5 is a schematic diagram illustrating a first example of an application scenario to which the semantic matching method according to an embodiment of the present disclosure is applied;
[0027] Figure 6 is a schematic diagram illustrating a second example of an application scenario to which the semantic matching method according to an embodiment of the present disclosure is applied;
[0028] Figure 7 is a functional block diagram illustrating a configuration of a semantic matching apparatus according to an embodiment of the present disclosure; and
[0029] Figure 8 is a schematic diagram of an exemplary architecture of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] The various preferred embodiments of the present invention will be described below with reference to the accompanying drawings. The following description with reference to the accompanying drawings is provided to assist in the understanding of the exemplary embodiments of the present invention as defined by the claims and their equivalents. It includes various specific details to assist in understanding, but they are to be viewed as exemplary only. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Moreover, in order to make the specification more clear and concise, detailed descriptions of functions and configurations well known in the art will be omitted.
[0031] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0032] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. For example, natural language processing technology can include text processing, semantic understanding, semantic matching and other technologies.
[0033] First, refer to Figure 1 A semantic matching method according to an embodiment of the present disclosure is described. As shown in FIG1 , the method includes the following steps.
[0034] First, in step S101, word segmentation and concatenation are performed on the input first text and second text to obtain a first word sequence.
[0035] For example, when the semantic matching method according to an embodiment of the present disclosure is applied to the semantic matching problem of a user question and a standard question library in a search scenario, the first text may be a user question (query) and the second text may be a standard question (question). Assume that the first text is represented as And the second text is represented as P and M represent the length of the word sequence of query and question respectively. Respectively represent the words obtained after performing word segmentation on the first text, and They respectively represent words obtained after performing word segmentation processing on the second text.
[0036] Then, in order to construct the standard input provided to the subsequent network, the two word sequences of the first text and the second text are concatenated into the first word sequence, which can be represented as The [CLS] mark is placed at the beginning of the first sentence, and the [SEP] mark is used to separate two input sentences. Since the first text and the second text are input, the [SEP] mark is added between the first text and the second text.
[0037] Next, the process proceeds to step S102. In step S102, the first word sequence is provided to the embedding network, and the first word sequence is converted into a first word vector through the embedding network. Through the processing of step S102, the character data of the input text can be converted into numerical data. For example, assuming that the first word sequence is converted into a first word vector e through the embedding network, then there is the following formula:
[0038]
[0039] in, is represented as the embedding vector of the i-th word and is a vector of dimension 1×d, and because one [CLS] flag and two [SEP] flags are added, the embedding vector of the last word after concatenation is
[0040] Then, in step S103, the first word vector is provided to a transformation network. Figure 2 FIG. 4 shows a semantic matching model according to an embodiment of the present disclosure. Figure 2 As shown, the semantic matching model includes the embedding network 201 described above, the transformation network 202 and multiple classification networks 203 to be described here. 1 , 203 2 ,……203 N The transformation network 202 includes first to Nth transformation layers 202 connected in series. 1 , 202 2 ,……202 N Wherein, N is an integer greater than 1, and the first transformation layer 202 1 The first word vector is received as an input vector, and the other transformation layer 202 2,……202 N The feature vector generated by the previous transformation layer connected in series is received as its input vector, each transformation layer extracts features from the input vector and generates a feature vector, and each transformation layer has a classification network corresponding to it. Figure 2 It can be seen that the output of each transformation layer is connected to the corresponding classification network. And multiple classification networks 203 1 , 203 2 ,……203 N The network structures are the same, but the parameters of each network are different.
[0041] The transformation network 202 is used to extract semantic information layer by layer from the input first word vector. For example, as a possible implementation, each transformation layer in the transformation network 202 can obtain feature information in the input word vector through a self-attention mechanism.
[0042] The first step in calculating self-attention is to create three vectors based on the input vector of the transformation layer: Q vector, K vector, and V vector. These vectors are obtained by multiplying the input vector with three transformation matrices. Take the (i+1)th transformation layer as an example. Assume that the feature vector output from the i-th transformation layer is H i , and the feature vector is input to the next transformation layer, namely the (i+1)th transformation layer.
[0043] In the (i+1)th transformation layer, based on the input vector H, the following formula is used i (For example, the dimension is (P+M+3)×d k ) to create Q i Vector, K i Vector and V i vector:
[0044] Q i =H i W i Q , K i =H i W i K , V i =H i W i V (2)
[0045] Among them, W i Q , W i K , W i V Respectively represent and Q iVector, K i Vector and V i The parameter matrix corresponding to the vector. For example, W i Q , W i K , W i V The dimension can be d k ×d k , and the dimensions of Qi vector, Ki vector and Vi vector can be (P+M+3)×d k .
[0046] The second step in calculating self-attention is to calculate the attention score. Assuming that we want to calculate the self-attention score of the first word's feature, we need to score the features of each word in the input vector. This score determines the degree of attention to other words when transforming the word at a certain position. For example, this score is calculated by Q i Vector and K i The dot product of the vectors.
[0047] The third step in calculating self-attention is to divide the score by K i The dimension of the vector (d k ). This makes the gradient update process more stable.
[0048] The fourth step in calculating the self-attention is to then operate the result through the Softmax function. The Softmax function normalizes the scores so that they are all positive and add up to 1.
[0049] The fifth step in calculating self-attention is to convert V i The purpose of this is to keep the feature values of the words you want to focus on unchanged as much as possible, while masking out those irrelevant words (for example, multiplying them by a very small number).
[0050] Assume H i represents the feature vector output from the (i+1)th transformation layer, then the processing from the second to the fifth steps described above can be implemented by the following formula:
[0051]
[0052] Return to reference Figure 1After step S103, the process proceeds to step S104. In step S104, the following operations are performed layer by layer starting from the first transformation layer until the semantic matching result of the first text and the second text is generated: the feature vector generated by the transformation layer is provided to the classification network corresponding to it; the classification network corresponding to the transformation layer is used to generate a semantic matching prediction result based on the feature vector received by it; if the semantic matching prediction result meets a predetermined condition, the semantic matching result of the first text and the second text is generated based on the semantic matching prediction result.
[0053] As described above, the network structures of the multiple classification networks corresponding to the first to Nth transformation layers are the same, but the specific network parameters are different. Specifically, each of the multiple classification networks may include: a fully connected layer, a classification transformation layer, and a normalization layer.
[0054] Wherein, the classification network generates a semantic matching prediction result based on the feature vector it receives, including the following steps.
[0055] First, the feature vector output by the transformation layer corresponding to the classification network is provided to the fully connected layer, and the fully connected layer transforms the feature vector output by the transformation layer into a feature vector with a dimension corresponding to the number of categories of the semantic matching prediction result. Assume that the feature vector output from the i-th transformation layer is H i , and the vector output from the fully connected layer is Y, then there is the following formula:
[0056] Y=W Y H i +b Y (4)
[0057] Where W Y represents the corresponding parameter matrix, and b Y Represents a constant.
[0058] For example, if the semantic matching prediction result output by the classification network is set as a 2-class prediction result, the feature vector Y output from the fully connected layer is a 1×2-dimensional vector.
[0059] Then, the feature vector output by the fully connected layer is provided to the classification transformation layer, and the classification transformation layer outputs the transformed feature vector. For example, the classification transformation layer here can have a similar network structure to the first to Nth transformation layers described above, but with different network parameters. Let Y' represent the feature vector output from the classification transformation layer, then the following formula exists:
[0060] Y′=Transformer(Y) (5)
[0061] For example, when the prediction result is a 2-class result, where y 0 ′ and y 1 ' is not necessarily a number in the range of 0 to 1, and the sum of the two is not necessarily equal to 1.
[0062] Finally, the transformed feature vector is provided to the normalization layer, which performs normalization on each element therein and uses the normalized feature vector as the semantic matching prediction result. The normalization is performed by a Softmax function. For example, Representing the feature vector output from the normalized layer, the following formula exists:
[0063]
[0064] Specifically,
[0065] Among them, with y 0 ′ and y 1 'different, and is a number between 0 and 1 whose sum is equal to 1.
[0066] In addition, the semantic matching prediction result may include a probability value indicating whether the first text and the second text match, and the predetermined condition includes: the probability value is greater than a predetermined threshold. The closer the probability value output by the classification network is to 1, the higher its confidence is considered. For example, the predetermined threshold can be set to 0.8, 0.7, etc.
[0067] exist Figure 2 In FIG. 1 , an example of outputting 2 classification results as semantic prediction results is shown. The classification network output includes P 0 , P 1 The prediction results of these two elements, where P 0 represents the probability that the first text and the second text do not match each other semantically, and P 1 represents the probability of semantic matching between the first text and the second text, and P 0 With P 1 The sum is 1. In this case, as long as P 0 and P 1 If one of the two is greater than a predetermined threshold, it is considered that the semantic matching prediction result meets the predetermined condition and can be output as the final semantic matching prediction result.
[0068] However, in the present disclosure, the semantic matching prediction result output by the classification network is not limited to a 2-classification result. For example, the semantic matching prediction result output by the classification network may also be a 3-classification result. In this case, the classification network may output a P 0 , P 1 , P 2 The prediction results of these three elements, among which, P 0 represents the probability that the first text and the second text do not match each other semantically, P 1 represents the probability of semantic matching between the first text and the second text, and P 2 represents the probability that it is not possible to determine whether the first text and the second text are semantically matched, and P 0 , P 1 , P 2 The sum is 1. Of course, depending on different application scenarios, other classification results may appear, such as 4-classification results, 5-classification results, etc.
[0069] In addition, it should be pointed out here that when the semantic matching prediction result of the classification network corresponding to the i-th transformation layer meets the predetermined conditions, the operation of other transformation layers and their corresponding classification networks in the transformation network is stopped, where i is an integer greater than or equal to 1. In other words, once the prediction result output by the classification network corresponding to the shallow transformation layer meets the predetermined conditions, such as the confidence level is high enough, the subsequent deep transformation layer will no longer be processed. Through such a setting, when the user question is relatively simple, the satisfactory semantic matching prediction result of the classification network corresponding to the shallow transformation layer can be output in advance without using the semantic matching prediction result of the classification network corresponding to the last transformation layer. Thereby, the computational burden of the semantic matching model is reduced, and the reasoning speed of the model is improved, and the performance of the semantic matching calculation will not be significantly reduced. This enables the semantic matching model according to the present disclosure to be truly put into service.
[0070] In the above, refer to Figure 1 Flowchart of Figure 2 The model structure diagram of FIG. 1 describes in detail the specific process of the semantic matching method according to the embodiment of the present disclosure. The method described above is performed when the embedding network, the transformation network and the classification network have all been trained. Next, the training process of each network included in the semantic matching model will be described in detail.
[0071] According to an embodiment of the present disclosure, the training process of each network included in the semantic matching model may include two stages.
[0072] In the first stage, the embedding network, the transformation network and the classification network corresponding to the Nth transformation layer are trained using a first training data set.
[0073] In the second stage, while keeping the parameters of the embedded network, the transformation network and the classification network corresponding to the Nth transformation layer fixed, the classification networks corresponding to the first to (N-1)th transformation layers are trained using the second training data set.
[0074] Return to reference Figure 2 , it can be seen that multiple classification networks 203 1 , 203 2 ,……203 N Divide into two parts. N As the first part of the classification network, for example, this part of the classification network can be called a teacher classification network, and the parameters are adjusted in the first stage of training. 1 , 203 2 ,……203 N-1 As the classification network of the second part, for example, this part of the classification network can be called a student classification network, and parameters are adjusted in the second stage of training.
[0075] The training process of these two stages can be considered as a training process based on knowledge distillation. Knowledge distillation refers to the use of knowledge learned from a large model to train a small model, so that the small model has the generalization ability of the large model. Generalization ability refers to the adaptability of a machine learning algorithm to fresh samples. The purpose of learning is to learn the rules hidden behind the data, and the trained network can also give appropriate outputs for data outside the learning set with the same rules. This ability is called generalization ability. In the present disclosure, since the teacher classification network is added on the basis of the last layer of transformation layers, it must correspond to a large network. In contrast, since the student classification network is added on the basis of each layer of transformation layers, it corresponds to a small network. The student classification network is trained by using the teacher classification network trained with the semantic matching task-related data set, that is, the student classification network is used to distill the probability distribution of the teacher classification network.
[0076] Below, we will first refer to Figure 3 Describe the training process of the first stage. The first training data set may include multiple training data, each training data includes the third text and the fourth text and the true semantic matching result of the third text and the fourth text. Figure 3 As shown, using the first training data set, training the embedding network, the transformation network and the classification network corresponding to the Nth transformation layer includes the following steps.
[0077] First, in step S301, in at least a portion of the training data in the first training data set, the third text and the fourth text of each training data are segmented and concatenated to obtain a second word sequence. This step is similar to the step described above with reference to Figure 1 The step S101 is similar, except that the input data is different. Figure 1 The actual user questions and standard questions are entered in Figure 3 The inputs are user questions and standard questions for training. Compared with the actual user questions and standard questions, the standard questions for training user questions have true semantic matching results as correct answers.
[0078] Then, in step S302, the second word sequence is provided to the embedding network, and the second word sequence is converted into a second word vector by the embedding network. This step is similar to step S102 described above with reference to FIG. 1 in terms of processing method.
[0079] Next, in step S303, the second word vector is provided to the transformation network, and the feature vector output by the Nth transformation layer in the transformation network is provided to the corresponding classification network.
[0080] Then, in step S304, a first loss function between the semantic matching prediction result output by the classification network corresponding to the Nth transformation layer and the actual semantic matching result is calculated.
[0081] Next, in step S305, based on the first loss function, the embedding network, the transformation network and the classification network corresponding to the Nth transformation layer are trained.
[0082] Specifically, in at least a portion of the training data in the first training data set, the above steps S301 to S305 are repeatedly performed for each training data, so as to continuously adjust the network parameters in the embedding network, the transformation network, and the classification network corresponding to the Nth transformation layer based on the first loss function. When the first loss function converges, the first stage of the training process ends, and the network parameters of the embedding network, the transformation network, and the classification network corresponding to the Nth transformation layer are fixed at this time for use in the subsequent second stage of the training process.
[0083] Next, we will refer to Figure 4 Describe the training process of the second stage. The second training data set includes a plurality of training data, each of which includes a fifth text and a sixth text. Figure 4 As shown, using the second training data set, training the classification network corresponding to the first to (N-1)th transformation layers includes the following steps.
[0084] First, in step S401, in at least a portion of the training data in the second training data set, the fifth text and the sixth text of each training data are segmented and concatenated to obtain a third word sequence. This step is similar to the step described above with reference to Figure 1 The step S101 is similar, except that the input data is different. Figure 1 The actual user questions and standard questions are entered in Figure 4 The inputs are user questions and standard questions for training. Compared with the actual user questions and standard questions, the standard questions for training user questions have semantic matching results as correct answers. However, Figure 3 The difference from the first stage of training in is that the semantic matching result as the correct answer here is not the real semantic matching result, but the semantic matching prediction result output by the trained teacher classification network.
[0085] Then, in step S402, the third word sequence is provided to the embedding network, and the third word sequence is converted into a third word vector by the embedding network. This step is similar to step S102 described above with reference to FIG. 1 in terms of processing method.
[0086] Next, in step S403, the third word vector is provided to the transformation network, and the feature vectors output by the first to (N-1)th transformation layers in the transformation network are respectively provided to the corresponding classification networks.
[0087] Then, in step S404, a second loss function is calculated between multiple semantic matching prediction results output by multiple classification networks corresponding to the first to (N-1)th transformation layers and the semantic matching prediction results output by the classification network corresponding to the Nth transformation layer.
[0088] For example, as a possible implementation, calculating the second loss function between multiple semantic matching prediction results output by multiple classification networks corresponding to the first to (N-1)th transformation layers and the semantic matching prediction results output by the classification network corresponding to the Nth transformation layer may include the following steps.
[0089] First, the KL divergence (Kullback-Leibler Divergence), also known as relative entropy, between the semantic matching prediction results output by each of the multiple classification networks corresponding to the first to (N-1)th transformation layers and the semantic matching prediction results output by the classification network corresponding to the Nth transformation layer is calculated. is the prediction result of the i-th student classification network, and the case of 2 classification is used as an example to illustrate.
[0090] in this case, As shown above, P 0 and P 1 are the probabilities of predicting whether query and question are semantically matched. In addition, assuming that p t is the prediction result of the teacher classification network, and similarly, It can be measured using the KL divergence using the following formula With p t The difference between:
[0091] p t = Teacher_Classifier(H N ) (8)
[0092]
[0093]
[0094] Among them, Teacher_Classifier represents the teacher classification network, Student_Classifier_i represents the i-th student classification network, H N represents the feature vector output by the Nth transformation layer, and H i Represents the feature vector output by the i-th transformation layer.
[0095] Then, the sum of all calculated KL divergences is used as the second loss function. In the case where there are (N-1) student classification networks in the semantic matching model described above, the second loss function is the sum of the KL divergences of all student classification networks. Let Represents the second loss function, then there is the following formula:
[0096]
[0097] Next, in step S405, based on the second loss function, the classification network corresponding to the first to (N-1)th transformation layers is trained.
[0098] Specifically, in at least a portion of the training data in the second training data set, the above steps S401 to S405 are repeatedly performed for each training data, so as to continuously adjust the network parameters in the classification network corresponding to the first to (N-1)th transformation layers based on the second loss function. When the second loss function converges, the training process of the second stage ends.
[0099] In the above, refer to Figure 3 and Figure 4The training process of each network in the semantic matching model according to the embodiment of the present disclosure is described in detail. Figure 5 and Figure 6 Describe the application scenario of the semantic matching method according to the embodiment of the present disclosure.
[0100] Figure 5 A schematic diagram of a search scenario of a browser to which the semantic matching method according to an embodiment of the present disclosure is applied is shown. Specifically, in a search scenario, a question-and-answer library including multiple standard questions and multiple associated answers may be pre-stored. When a user enters a user question (e.g., "How to play Li Bai in King of Glory") in the search box 501, the user question (as the first text) and each of the multiple standard questions in the question-and-answer library (as the second text) may be matched according to the reference above. Figures 1 to 4 The semantic matching method is processed to generate a semantic matching prediction result of the two. Of course, alternatively, a part of standard questions can be screened out from the question and answer library in advance through other processing for semantic matching with the user question. Then, a standard question with the highest semantic matching degree is selected as the standard question corresponding to the user question and displayed in box 502. Here, in the case where the semantic matching prediction result includes a probability value, it can be considered that the larger the probability value indicating the semantic matching of the two texts, the higher the matching degree of the two. The user can directly browse the answer corresponding to the question by clicking box 502.
[0101] In addition, Figure 5 It can be seen that after the user inputs the question and before clicking the search button 503, the standard question that best matches the user question is displayed. In addition, after the user inputs the question, other questions associated with the user question are further displayed in box 504, such as "How is the King of Glory skin?", "What is the relationship between King of Glory Li Bai and Angela?", etc. When the user wishes to switch the input question to an associated question, the switch can be performed by clicking the corresponding question in box 504. After the user question is switched, the standard question and its answer that best matches the switched user question can also be further displayed accordingly.
[0102] Figure 6 A schematic diagram of a browser search scenario in which the semantic matching method according to an embodiment of the present disclosure is applied is shown. Figure 6 In the search box 601, the user enters the user question "Why is the sky blue?", and then based on the above reference Figure 5In a similar process, a standard question "Why is the sky blue, and what is the principle" that best matches the user's question is determined in the question-answer database and is displayed in box 602. The user can directly browse the answer corresponding to the question by clicking on box 602.
[0103] It can be seen that Figure 5 The interface shown is different in that Figure 6 The interface shown is after the user clicks the search button. The web page search results associated with the user's question are displayed in box 603. At this time, the standard question and its answer that best matches the user's question are still displayed in box 602.
[0104] Therefore, in a search scenario, as long as the user completes inputting the user question, the most matching standard question and its answer will be displayed, regardless of whether the user clicks the search button.
[0105] In the case where the semantic matching method according to the embodiment of the present disclosure is applied to the browser scenario, being able to match high-quality question-answer pairs for user questions is a key solution for searching and directly recalling question-answer pairs. In the business standard evaluation set, compared with the prior art, by applying the semantic matching method according to the present disclosure, various indicators used to evaluate classification have been significantly improved, such as AUC (Area Under Curve) increased by 8%, ACC (Accuracy) increased by 7.7%, and TOP1 recall rate (which reflects the proportion of correctly judged positive examples to total positive examples) increased by 33.4%. After the service was deployed online, the question and answer exposure PV (Page View) increased by 23%, and the click PV increased by 15%. Therefore, based on these indicators, it can be seen that in the search scenario where the semantic matching method according to the present disclosure is applied, the question and answer pairs can be recalled with high quality and user needs can be accurately met. Moreover, in the business standard evaluation set, the recall rate can be significantly improved while the accuracy rate remains unchanged.
[0106] Despite Figure 5 and Figure 6 The figure shows the case where the semantic matching method according to the embodiment of the present disclosure is applied to the search scenario of the browser, but those skilled in the art can understand that the semantic matching method according to the embodiment of the present disclosure can also be similarly applied to any other scenario that requires semantic analysis and matching of text.
[0107] In the above, reference has been made to Figures 1 to 6The semantic matching method according to the embodiment of the present disclosure is described in detail. Through the semantic matching method according to the embodiment of the present disclosure, once the prediction result output by the classification network corresponding to the shallow transformation layer meets the predetermined conditions, such as the confidence is high enough, the subsequent deep transformation layer will no longer be processed. Through such a setting, when the user question is relatively simple, the satisfactory semantic matching prediction result of the classification network corresponding to the shallow transformation layer can be output in advance, without using the semantic matching prediction result of the classification network corresponding to the last transformation layer. Thereby, the computational burden of the semantic matching model is reduced, and the reasoning speed of the model is improved, and the performance of the semantic matching calculation will not be significantly reduced.
[0108] Next, we will refer to Figure 7 The semantic matching device according to the embodiment of the present disclosure is described. Figure 7 As shown, the semantic matching device 700 includes: a word segmentation and splicing unit 701, an embedding unit 702, a transformation unit 703 and a classification unit 704.
[0109] The word segmentation and concatenation unit 701 is used to perform word segmentation and concatenation processing on the input first text and the second text to obtain a first word sequence. For example, when the semantic matching method according to the embodiment of the present disclosure is applied to the semantic matching problem of a user question and a standard question library in a search scenario, the first text can be a user question (query) and the second text can be a standard question (question). Assume that the first text is represented as And the second text is represented as P and M represent the length of the word sequence of query and question respectively. represent the words obtained after the word segmentation processing is performed on the first text by the word segmentation and splicing unit 901, and Then, in order to construct the standard input provided to the subsequent network, the word segmentation concatenation unit 701 concatenates the two word sequences of the first text and the second text into a first word sequence, which can be represented as The [CLS] mark is placed at the beginning of the first sentence, and the [SEP] mark is used to separate two input sentences. Since the first text and the second text are input, the [SEP] mark is added between the first text and the second text.
[0110] The embedding unit 702 includes an embedding network, and converts the first word sequence into a first word vector through the embedding network. Through the processing performed by the embedding unit 702, character data of the input text can be converted into numerical data.
[0111] The transformation unit 703 includes a transformation network, wherein the transformation network also includes first to Nth transformation layers connected in series, wherein the first transformation layer receives the first word vector as an input vector and other transformation layers receive feature vectors generated by the previous transformation layer connected in series therewith as their input vectors, each transformation layer extracts features from the input vector and generates a feature vector, and each transformation layer has a classification network corresponding thereto.
[0112] The transformation network is used to extract semantic information from the input first word vector layer by layer. For example, as a possible implementation, each transformation layer in the transformation network can obtain feature information in the input word vector through a self-attention mechanism. The specific processing process has been described above.
[0113] The classification unit 704 includes N classification networks, each transformation layer has a one-to-one corresponding classification network, each classification network is used to receive a feature vector from its corresponding transformation layer and generate a semantic matching prediction result of the first text and the second text corresponding to the transformation layer.
[0114] Among them, the transformation network and N classification networks are configured as follows: starting from the first transformation layer, the following operations are performed layer by layer until a semantic matching result of the first text and the second text is generated: the feature vector generated by the transformation layer is provided to the classification network corresponding to it; the classification network corresponding to the transformation layer is used to generate a semantic matching prediction result based on the feature vector received by it; when the semantic matching prediction result meets a predetermined condition, a semantic matching result of the first text and the second text is generated based on the semantic matching prediction result.
[0115] As described above, the network structures of the multiple classification networks corresponding to the first to Nth transformation layers are the same, but the specific network parameters are different. Specifically, each of the multiple classification networks may include: a fully connected layer, a classification transformation layer, and a normalization layer.
[0116] The classification unit 704 can be further configured to: provide the feature vector output by the transformation layer corresponding to the classification network to the fully connected layer, and the fully connected layer transforms the feature vector output by the transformation layer into a feature vector of a dimension corresponding to the number of categories of the semantic matching prediction result; provide the feature vector output by the fully connected layer to the classification transformation layer, and the classification transformation layer outputs the transformed feature vector; provide the transformed feature vector to the normalization layer, and the normalization layer normalizes each element therein, and uses the normalized feature vector as the semantic matching prediction result.
[0117] In addition, the semantic matching prediction result includes a probability value indicating whether the first text and the second text match, and the predetermined condition includes: the probability value is greater than a predetermined threshold. The closer the probability value output by the classification network is to 1, the higher its confidence is considered. For example, the predetermined threshold can be set to 0.8, 0.7, etc.
[0118] In addition, it should be pointed out here that the transformation network and the classification network can be further configured to: when the semantic matching prediction result of the classification network corresponding to the i-th transformation layer meets the predetermined conditions, stop the operation of other transformation layers and their corresponding classification networks in the transformation network, where i is an integer greater than or equal to 1. In other words, once the prediction result output by the classification network corresponding to the shallow transformation layer meets the predetermined conditions, such as the confidence is high enough, the subsequent deep transformation layer will no longer be processed. Through such a setting, when the user question is relatively simple, the satisfactory semantic matching prediction result of the classification network corresponding to the shallow transformation layer can be output in advance without using the semantic matching prediction result of the classification network corresponding to the last transformation layer. Thereby, the computational burden of the semantic matching model is reduced, the reasoning speed of the model is improved, and the performance of the semantic matching calculation will not be significantly reduced. This enables the semantic matching model according to the present disclosure to be truly put into service.
[0119] In addition, the specific training process of the embedding network, the transformation network and the classification network used by the embedding unit 702, the transformation unit 703 and the classification unit 704 when performing the processing has been described in detail above, so for the sake of simplicity, it will not be repeated here.
[0120] In addition, the method or device according to the embodiment of the present disclosure can also be used by Figure 8 The architecture of the computing device 800 shown in FIG. Figure 8 As shown, the computing device 800 may include a bus 810, one or more CPUs 820, a read-only memory (ROM) 830, a random access memory (RAM) 840, a communication port 850 connected to a network, an input / output component 860, a hard disk 870, etc. The storage device in the computing device 800, such as the ROM 830 or the hard disk 870, may store various data or files used for processing and / or communication of the information processing method provided by the present disclosure, as well as program instructions executed by the CPU. Of course, Figure 8 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 8 One or more components of a computing device are shown.
[0121] In addition, according to another aspect of the present disclosure, a semantic matching device is provided, which may include: a receiving device for receiving an input user question; a semantic matching device according to the above text, for determining, for each of at least a part of the standard questions in the standard question and answer library, a semantic matching result with the user question; and a display device, for displaying a standard question that best matches the semantics of the user question.
[0122] The embodiments of the present disclosure may also be implemented as a computer-readable storage medium. Computer-readable instructions are stored on a computer-readable storage medium according to an embodiment of the present disclosure. When the computer-readable instructions are executed by a processor, the semantic matching method according to the embodiment of the present disclosure described with reference to the above figures may be executed. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may, for example, include a random access memory (RAM) and / or a cache memory (cache), etc. The non-volatile memory may, for example, include a read-only memory (ROM), a hard disk, a flash memory, etc.
[0123] In addition, the embodiments of the present disclosure may also be implemented as a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device performs the above-mentioned semantic matching method.
[0124] So far, reference has been made to Figures 1 to 8 The semantic matching method and device according to the embodiment of the present disclosure are described in detail. Through the semantic matching method and device according to the embodiment of the present disclosure, once the prediction result output by the classification network corresponding to the shallow transformation layer meets the predetermined conditions, such as the confidence is high enough, the subsequent deep transformation layer will no longer be processed. Through such a setting, when the user question is relatively simple, the satisfactory semantic matching prediction result of the classification network corresponding to the shallow transformation layer can be output in advance, without using the semantic matching prediction result of the classification network corresponding to the last transformation layer. Thereby, the computational burden of the semantic matching model is reduced, and the reasoning speed of the model is improved, and the performance of the semantic matching calculation will not be significantly reduced.
[0125] It should be noted that, in this specification, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "includes..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0126] Finally, it should be noted that the above series of processing includes not only processing executed in time series in the order described here, but also processing executed in parallel or separately rather than in time series.
[0127] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary hardware platform, and of course can also be implemented entirely by software. Based on such an understanding, all or part of the contribution of the technical solution of the present invention to the background technology can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention or certain parts of the embodiments.
[0128] The present invention has been introduced in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A semantic matching method, include: Perform word segmentation and concatenation processing on the input first text and the second text to obtain a first word sequence; Providing the first word sequence to an embedding network, and converting the first word sequence into a first word vector through the embedding network; Providing the first word vector to a transformation network, wherein the transformation network further includes first to Nth transformation layers connected in series, wherein N is an integer greater than 1, the first transformation layer receives the first word vector as an input vector and other transformation layers receive a feature vector generated by a previous transformation layer connected in series therewith as their input vectors, each transformation layer performs feature extraction on the input vector and generates a feature vector, and each transformation layer has a classification network corresponding thereto; as well as The following operations are performed layer by layer starting from the first transformation layer until a semantic matching result between the first text and the second text is generated: Providing the feature vector generated by the transformation layer to the corresponding classification network; Utilize the classification network corresponding to the transformation layer to generate a semantic matching prediction result based on the feature vector received by the classification network; In a case where the semantic matching prediction result meets a predetermined condition, a semantic matching result of the first text and the second text is generated based on the semantic matching prediction result.
2. The method according to claim 1, in, The semantic matching prediction result includes a probability value indicating whether the first text and the second text match, and the predetermined condition includes: the probability value is greater than a predetermined threshold.
3. The method according to claim 1, in, When the semantic matching prediction result of the classification network corresponding to the i-th transformation layer meets a predetermined condition, the operations of other transformation layers in the transformation network and their corresponding classification networks are stopped, where i is an integer greater than or equal to 1.
4. The method according to claim 1, wherein each of the plurality of classification networks include: Fully connected layers, classification transformation layers, and normalization layers, and The classification network generates a semantic matching prediction result based on the received feature vector, including: The fully connected layer receives the feature vector output by the transformation layer corresponding to the classification network, and the fully connected layer transforms the feature vector into a feature vector with a dimension corresponding to the number of categories of the semantic matching prediction result; Providing the feature vector output by the fully connected layer to the classification transformation layer, and the classification transformation layer outputs a transformed feature vector; and The transformed feature vector is provided to the normalization layer, and the normalization layer performs normalization on each element therein, and uses the normalized feature vector as the semantic matching prediction result.
5. The method according to claim 1, wherein each network is trained by: Using the first training data set, training the embedding network, the transformation network, and the classification network corresponding to the Nth transformation layer; and While keeping the parameters of the embedded network, the transformation network and the classification network corresponding to the Nth transformation layer fixed, the classification networks corresponding to the first to (N-1)th transformation layers are trained using a second training data set.
6. The method according to claim 5, wherein the first training data set comprises a plurality of training data, each training data comprises a third text, a fourth text, and a true semantic matching result between the third text and the fourth text, in, Using the first training data set, training the embedding network, the transformation network, and the classification network corresponding to the Nth transformation layer includes: In at least a portion of the training data in the first training data set, performing word segmentation and concatenation processing on the third text and the fourth text of each training data to obtain a second word sequence; Providing the second word sequence to the embedding network, and converting the second word sequence into a second word vector through the embedding network; Providing the second word vector to the transformation network, and providing the feature vector output by the Nth transformation layer in the transformation network to the corresponding classification network; Calculating a first loss function between a semantic matching prediction result output by the classification network corresponding to the Nth transformation layer and a true semantic matching result; and Based on the first loss function, the embedding network, the transformation network and the classification network corresponding to the Nth transformation layer are trained.
7. The method according to claim 5, wherein the second training data set comprises a plurality of training data, each training data comprises a fifth text and a sixth text, in, Using the second training data set, train the classification network corresponding to the first to (N-1)th transformation layers: In at least a portion of the training data in the second training data set, performing word segmentation and concatenation processing on the fifth text and the sixth text of each training data to obtain a third word sequence; Providing the third word sequence to the embedding network, and converting the third word sequence into a third word vector through the embedding network; Providing the third word vector to the transformation network, and providing the feature vectors output by the first to (N-1)th transformation layers in the transformation network to the corresponding classification networks respectively; Calculating a second loss function between a plurality of semantic matching prediction results output by a plurality of classification networks corresponding to the first to (N-1)th transformation layers and a semantic matching prediction result output by the classification network corresponding to the Nth transformation layer; and Based on the second loss function, the classification network corresponding to the first to (N-1)th transformation layers is trained.
8. The method according to claim 7, wherein a second loss function is calculated between a plurality of semantic matching prediction results output by a plurality of classification networks corresponding to the first to (N-1)th transformation layers and a semantic matching prediction result output by the classification network corresponding to the Nth transformation layer, include: Calculating the KL divergence between the semantic matching prediction result output by each of the plurality of classification networks corresponding to the first to (N-1)th transformation layers and the semantic matching prediction result output by the classification network corresponding to the Nth transformation layer; as well as The sum of all calculated KL divergences is used as the second loss function.
9. A semantic matching method, include: Receive input from user questions; The method according to any one of claims 1 to 8, for each of at least a portion of the standard questions in the standard question and answer library, determining a semantic matching result between the standard question and the user question; A standard question that best matches the semantics of the user question is displayed.
10. A semantic matching device, include: a memory for storing a computer program thereon; A processor, configured to perform the following processing when executing the computer program: Perform word segmentation and concatenation processing on the input first text and the second text to obtain a first word sequence; Providing the first word sequence to an embedding network, and converting the first word sequence into a first word vector through the embedding network; Providing the first word vector to a transformation network, wherein the transformation network further includes first to Nth transformation layers connected in series, wherein the first transformation layer receives the first word vector as an input vector and other transformation layers receive feature vectors generated by the previous transformation layer connected in series with it as their input vectors, each transformation layer performs feature extraction on the input vector and generates a feature vector, and each transformation layer has a classification network corresponding thereto; as well as The following operations are performed layer by layer starting from the first transformation layer until a semantic matching result between the first text and the second text is generated: Providing the feature vector generated by the transformation layer to the corresponding classification network; Utilize the classification network corresponding to the transformation layer to generate a semantic matching prediction result based on the feature vector received by the classification network; In a case where the semantic matching prediction result meets a predetermined condition, a semantic matching result of the first text and the second text is generated based on the semantic matching prediction result.
11. The device according to claim 10, in, The semantic matching prediction result includes a probability value indicating whether the first text and the second text match, and the predetermined condition includes: the probability value is greater than a predetermined threshold.
12. The apparatus of claim 10, wherein each of the plurality of classification networks include: Fully connected layers, classification transformation layers, and normalization layers, and The classification network corresponding to the transformation layer is used to generate a semantic matching prediction result based on the received feature vector, including: The fully connected layer receives the feature vector output by the transformation layer corresponding to the classification network, and the fully connected layer transforms the feature vector into a feature vector with a dimension corresponding to the number of categories of the semantic matching prediction result; Providing the feature vector output by the fully connected layer to the classification transformation layer, and the classification transformation layer outputs a transformed feature vector; and The transformed feature vector is provided to the normalization layer, and the normalization layer performs normalization on each element therein, and uses the normalized feature vector as the semantic matching prediction result.
13. The apparatus according to claim 10, wherein each network is trained by: Using the first training data set, training the embedding network, the transformation network, and the classification network corresponding to the Nth transformation layer; and While keeping the parameters of the embedded network, the transformation network and the classification network corresponding to the Nth transformation layer fixed, the classification networks corresponding to the first to (N-1)th transformation layers are trained using a second training data set.
14. A semantic matching device, include: A word segmentation and splicing unit, used for performing word segmentation and splicing processing on the input first text and the second text to obtain a first word sequence; An embedding unit, comprising an embedding network, and converting the first word sequence into a first word vector through the embedding network; A transformation unit, comprising a transformation network, wherein the transformation network comprises first to Nth transformation layers connected in series, wherein N is an integer greater than 1, the first transformation layer receives the first word vector as an input vector and other transformation layers receive a feature vector generated by a previous transformation layer connected in series therewith as their input vectors, and each transformation layer extracts features from the input vector and generates a feature vector; and A classification unit, comprising N classification networks, each transformation layer having a one-to-one corresponding classification network, each classification network being used to receive a feature vector from its corresponding transformation layer and generate a semantic matching prediction result of the first text and the second text corresponding to the transformation layer; The transformation network and the N classification networks are configured to perform the following operations layer by layer starting from the first transformation layer until a semantic matching result of the first text and the second text is generated: Providing the feature vector generated by the transformation layer to the corresponding classification network; Utilize the classification network corresponding to the transformation layer to generate a semantic matching prediction result based on the feature vector received by the classification network; In a case where the semantic matching prediction result meets a predetermined condition, a semantic matching result of the first text and the second text is generated based on the semantic matching prediction result.
15. A semantic matching device, include: A receiving device for receiving input user questions; The semantic matching device according to any one of claims 10 to 14, used to determine, for each of at least a portion of the standard questions in the standard question-answer library, a semantic matching result with the user question; as well as The display device is used to display a standard question that best matches the semantics of the user question.
16. A computer readable medium having a computer program stored thereon, which, when executed by a processor, performs the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method for semantic matching, learning method for semantic matching model, and server
CN108932342A
Text sentiment classification method and device, electronic equipment and storage medium
CN111930940A