Intention recognition method, electronic device, storage medium, and program product

The self-attention algorithm is used to extract the word meaning and word boundary information of Chinese text, generate probability distribution vectors and screen candidate vector sequences, which solves the problem of inaccurate intent recognition caused by fuzzy Chinese word boundaries and improves the accuracy of intent recognition.

CN119740584BActive Publication Date: 2025-10-10KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411943847.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-10-10
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

The fuzzy word boundaries in Chinese input make it difficult for machines to accurately identify intent, affecting the accuracy of entity recognition tasks.

Method used

The self-attention algorithm is used to extract word meaning and word boundary information from Chinese text, generate multiple probability distribution vectors, screen out candidate vector sequences that conform to grammatical rules, and generate intent summary text.

Benefits of technology

The accuracy of Chinese intent recognition is improved and the impact of word boundary ambiguity on intent recognition is alleviated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119740584B_ABST
    Figure CN119740584B_ABST
Patent Text Reader

Abstract

The present disclosure provides an intention recognition method, an electronic device, a readable storage medium and a computer program product. The intention recognition method of the present disclosure comprises: extracting word sense information and word boundary information of characters in the recognized text; converting the word sense information and the word boundary information into a plurality of probability distribution vectors based on a self-attention algorithm, the probability distribution vectors representing the probability of the predicted characters appearing in the predicted sequence; screening at least one candidate vector sequence from the plurality of probability distribution vectors, the candidate vector sequence being a probability distribution vector whose relevance to any prototype vector sequence in a prototype vector sequence library is higher than a relevance threshold; and generating an intention summary text based on the candidate vector sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and more particularly to an intention recognition method, electronic device, storage medium, and program product. Background Art

[0002] In scenarios like question-answering and tool invocation in AI conversational bots, it's often necessary to identify the intent and entity information in user input to determine which tool to invoke and what parameters to pass in. However, compared to English input, Chinese input has more ambiguous word boundaries, which affects the accuracy of entity recognition tasks. Summary of the Invention

[0003] The present disclosure provides an intention recognition method, an electronic device, a storage medium, and a program product.

[0004] According to one aspect of the present disclosure, there is provided a method for identifying intent, comprising:

[0005] Extracting word meaning information and word boundary information of characters in the recognized text;

[0006] Based on a self-attention algorithm, the word meaning information and the word boundary information are converted into a plurality of probability distribution vectors, wherein the probability distribution vectors represent the probability of the predicted character appearing in the predicted sequence;

[0007] Screening out at least one candidate vector sequence from the plurality of probability distribution vectors, wherein the candidate vector sequence is the probability distribution vector whose correlation with any prototype vector sequence in the prototype vector sequence library is higher than a correlation threshold;

[0008] Generate intent summary text based on the candidate vector sequence.

[0009] According to at least one embodiment of the present disclosure, the intent recognition method extracts word meaning information and word boundary information from the recognized text, including:

[0010] Performing character decomposition processing and word segmentation processing on the recognized text to obtain a plurality of recognized characters and a plurality of recognized words respectively;

[0011] Performing character embedding and position embedding processing on each of the recognized characters to obtain a query vector corresponding to each of the recognized characters, and using the multiple query vectors corresponding to the recognized characters as the word meaning information;

[0012] Each of the recognized words is subjected to word embedding processing and position embedding processing to obtain a key vector and a value vector corresponding to each of the recognized words, and multiple key vectors and multiple value vectors corresponding to each of the recognized words are used as the word boundary information.

[0013] According to at least one embodiment of the present disclosure, the intent recognition method converts the word meaning information and the word boundary information into multiple probability distribution vectors based on a self-attention algorithm, including:

[0014] Based on the self-attention algorithm, calculating the similarity between each query vector in the word meaning information and each key vector in the word boundary information as an attention score for each recognized word;

[0015] Performing a weighted summation on the attention score of the recognized word and each of the value vectors in the word boundary information to obtain a context vector for each of the recognized words, and generating a context vector sequence based on the context vector of each of the recognized words;

[0016] A probability distribution vector is generated based on the attention weight of each of the attention heads in the multiple attention heads and the context vector sequence to obtain multiple probability distribution vectors.

[0017] According to at least one embodiment of the present disclosure, the intent recognition method generates a probability distribution vector based on the attention weight of each of the multiple attention heads and the context vector sequence, thereby obtaining multiple probability distribution vectors, including:

[0018] Calculating the occurrence probability of the kth predicted character based on the 1st to k-1th context vectors in the context vector sequence and the attention weights of the attention heads to obtain a kth probability vector, where k is an integer greater than 1;

[0019] Generating a k-th predicted character based on the k-th probability vector;

[0020] Determining the kth context vector in the context vector sequence corresponding to the kth predicted character;

[0021] Calculating the occurrence probability of the k+1th predicted character based on the 1st to kth context vectors in the context vector sequence and the attention weights of the attention heads to obtain a k+1th probability vector;

[0022] The probability distribution vector is generated based on each of the probability vectors corresponding to the attention head.

[0023] According to at least one embodiment of the present disclosure, the intention recognition method of screening out at least one candidate vector sequence from the plurality of probability distribution vectors includes:

[0024] Calculating a correlation score between each probability distribution vector and each prototype vector sequence respectively, and generating a correlation matrix based on the calculated correlation scores, wherein the correlation matrix represents a corresponding relationship between the probability distribution vector and the correlation score;

[0025] screening the correlation scores in the correlation matrix based on the correlation threshold, and screening the correlation scores higher than the correlation threshold as target scores;

[0026] The probability distribution vector corresponding to the target score is searched in the correlation matrix, and the searched probability distribution vector is determined as the candidate vector sequence.

[0027] According to at least one embodiment of the present disclosure, the intent recognition method generates an intent summary text based on the candidate vector sequence, including:

[0028] Calculating an adaptive weight of the candidate vector sequence based on the correlation score of the candidate vector sequence;

[0029] Calculating a weighted average sum of the products of each candidate vector sequence and the corresponding adaptive weight to obtain a fused vector sequence;

[0030] Converting the fused vector sequence into a fused probability sequence, where the fused probability sequence represents a probability distribution of multiple intent labels;

[0031] The intent label with the highest probability in the fused probability sequence is output as the intent summary text.

[0032] According to at least one embodiment of the present disclosure, the intention recognition method further includes:

[0033] Obtain entity tags corresponding to the fused vector sequence, and output entity words in the recognized text based on the entity tags.

[0034] According to another aspect of the present disclosure, an electronic device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, so that the processor executes the intention recognition method of any embodiment of the present disclosure.

[0035] According to another aspect of the present disclosure, a readable storage medium is provided, in which execution instructions are stored. When the execution instructions are executed by a processor, they are used to implement the intention recognition method of any embodiment of the present disclosure.

[0036] According to yet another aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the intent recognition method of any one of the embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0037] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and together with the description serve to explain the principles of the present disclosure.

[0038] Figure 1 is a flowchart of an intent recognition method of one embodiment of the present disclosure.

[0039] Figure 2 is a process diagram of an intent recognition method of one embodiment of the present disclosure.

[0040] Figure 3 is a flowchart of an intent recognition method of one embodiment of the present disclosure.

[0041] Figure 4 is a flowchart of an intent recognition method of one embodiment of the present disclosure.

[0042] Figure 5 is a flowchart of an intent recognition method of one embodiment of the present disclosure.

[0043] Figure 6 is a flowchart of an intent recognition method of one embodiment of the present disclosure.

[0044] Figure 7 is a flowchart of an intent recognition method of one embodiment of the present disclosure.

[0045] Figure 8 is a flowchart of an intent recognition method of one embodiment of the present disclosure.

[0046] Figure 9 is a structural schematic block diagram of an intent recognition apparatus of one embodiment of the present disclosure.

[0047] Figure 10 is a structural schematic block diagram of an electronic device of one embodiment of the present disclosure. DETAILED DESCRIPTION

[0048] The present disclosure will be described in further detail by the following examples in conjunction with the accompanying drawings. It is to be understood that the specific examples described herein are intended to be illustrative only and are not to be limiting of the present disclosure. In addition, it is to be understood that the present disclosure is not limited to the particular examples described herein but includes any alternatives, modifications, and equivalents falling within the scope of the present disclosure.

[0049] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure can be combined with each other. The technical solution of the present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0050] Currently, commonly used intent recognition models are generally developed for the English language environment. In English text sequences, word boundaries are often defined by individual characters, and individual letters rarely represent meaning. Therefore, word boundaries are relatively clear. Compared to English text sequences, individual Chinese characters in Chinese text sequences can express a wide range of meanings, and phrases composed of these characters can have diverse meanings. However, when machines (computers that can run software or programs, hereinafter referred to as computers) perform intent recognition on Chinese text sequences, they struggle to distinguish word boundaries, significantly impacting the accuracy of intent recognition.

[0051] An embodiment of the present disclosure provides an intent recognition method, which can be executed by an electronic device such as a server, and is used to recognize the intent of a question text input by a user terminal to obtain a machine-understandable intent vector.

[0052] Figure 1 FIG. 1 shows a schematic diagram of the overall process of the intention recognition method M10 according to an embodiment of the present disclosure. Figure 1 The method shown includes steps S11 to S14.

[0053] Specifically, Figure 1 The methods shown include:

[0054] S11. Extract word meaning information and word boundary information of characters in the recognized text.

[0055] The recognized text is the text that is instructed to perform intent recognition. For example, the intent recognition method M10 of the embodiment is generally used to perform intent recognition on text input by the client to obtain the intent represented by the text input by the client. In this scenario, the recognized text refers to the text input by the client.

[0056] Character semantics refers to the meaning represented by a single character or a combination of characters. The semantic information of a character or word is typically obtained by using word embedding techniques to find the corresponding address of the character or word in the vocabulary and map it to the real number space of a vector, representing it as a vector. The resulting vector captures the relationship between the word's meaning and grammar, serving as the semantic information for the character.

[0057] Character word boundary information refers to the boundaries of words composed of characters and is used to determine whether multiple characters form an independent word. In the English language environment, words are separated by spaces, and word boundaries are relatively clear. By finding the location of spaces in the English text, word boundaries between words in the English text can be determined. In the Chinese language environment, different words are not separated by spaces, and word boundaries are relatively vague. Based on this, Chinese word segmentation technology can be used to separate an entire Chinese text into multiple Chinese words. The separated Chinese words can be used as character word boundary information. Chinese word segmentation technology can be used with Chinese word segmentation tools such as IKAnalyzer, Jieba, HanLP, and THULAC.

[0058] S12. Based on the self-attention algorithm, word meaning information and word boundary information are converted into multiple probability distribution vectors.

[0059] The self-attention algorithm refers to an algorithm related to the self-attention mechanism. The self-attention mechanism is a computational mechanism used in machine deep learning and natural language processing. It can be used to process text sequences, enabling each element in the sequence to pay attention to other elements in the sequence. As described above, word meaning information and word boundary information describe the Chinese characters in the recognized text. Each Chinese character in the recognized text is an element in the recognized text sequence. Based on this, the self-attention mechanism can help machines understand the dependencies between characters in the recognized text, thereby deriving the intent expressed by the recognized text.

[0060] The self-attention algorithm may specifically include an algorithm for generating a query vector, a key vector, and a value vector based on the elements in the sequence, an algorithm for calculating an attention score based on the query vector and the key vector, and an algorithm for calculating a probability distribution based on the attention score and the value vector. The elements in the recognized text sequence calculated using the self-attention algorithm can be obtained based on the word meaning information and word boundary information extracted in step S11. The self-attention algorithm is used to calculate the character elements represented by the word meaning information and word boundary information in order to obtain a probability distribution vector.

[0061] A probability distribution vector represents the probability that a machine-predicted character will appear at a specific position in the predicted sequence. For example, when a machine uses an intent recognition method to identify the intent of the text being recognized, the machine identifies the intent of the recognized text and generates a predicted character sequence to represent the intent understood by the machine. Suppose the predicted character sequence includes five characters, and suppose the machine predicts that the first character is the Chinese character "我" based on the self-attention mechanism. At this time, a vector will be generated to represent the probability that the first character in the predicted character sequence is the Chinese character "我". Similarly, for each character in the predicted character sequence, a corresponding vector can be obtained to represent the probability that the character at that position is a certain Chinese character. The sequence composed of these probability vectors is a probability distribution vector, which represents the probability of a predicted character appearing at each character position in a predicted character sequence.

[0062] Furthermore, to more comprehensively capture the associations between characters in the recognized text, a multi-head self-attention mechanism can be used to identify the intent of the recognized text. The multi-head self-attention mechanism is an attention mechanism that performs multiple sets of self-attention algorithm-based calculations in parallel. Each set of self-attention algorithm-based calculations is called an attention head. It performs a series of independent calculations, and different attention heads focus on different dependencies. By configuring different parameters for the self-attention algorithm in different attention heads, different attention heads can focus on different dependencies. In other words, when using the multi-head self-attention mechanism to convert word meaning information and word boundary information into multiple probability distribution vectors, each attention head can output a probability distribution vector. Thus, the self-attention algorithm based on the multi-head self-attention mechanism can convert word meaning information and word boundary information into multiple probability distribution vectors. The number of probability distribution vectors is the same as the number of attention heads in the multi-head self-attention mechanism.

[0063] S13. Filter out at least one candidate vector sequence from the multiple probability distribution vectors.

[0064] Among them, the candidate vector sequence is a probability distribution vector whose correlation with any prototype vector sequence in the prototype vector sequence library is higher than the correlation threshold. The prototype vector sequence is vector data pre-stored before executing the intent recognition method M10. A prototype vector sequence represents a character sequence that conforms to the grammatical rules. The correlation between the candidate vector sequence and the prototype vector sequence can be obtained by similarity calculation or distance calculation. The correlation threshold is used to measure the level of correlation to determine whether the character sequence represented by the probability distribution vector conforms to the grammatical rules. If the correlation between the probability distribution vector and any prototype vector sequence in the prototype vector sequence library is higher than the correlation threshold, it means that the character sequence represented by this probability distribution vector conforms to the grammatical rules and can be determined as a candidate vector as the intent expression of the recognized text, thereby improving the accuracy of intent recognition. Conversely, if the correlation between the probability distribution vector and all prototype vector sequences in the prototype vector sequence library is lower than the correlation threshold, it means that the character sequence represented by the probability distribution vector does not conform to the grammatical rules and should be filtered out to avoid using this probability distribution vector as the intent expression of the recognized text.

[0065] For example, multiple probability distribution vectors each represent the probability distribution of each character in a character sequence consisting of 5 characters. Assume that the predicted character sequence represented by each probability distribution vector is obtained based on the maximum probability distribution, where one probability distribution vector is recorded as vector A, and the predicted character sequence represented by vector A is "I / like / eat / noodles", and the other probability distribution vector is recorded as vector B, and the predicted character sequence represented by vector B is "I / like / eat / noodles". There is only one prototype vector sequence in the prototype vector sequence library, and the character arrangement represented by the prototype vector sequence is "like to eat noodles". Assuming that the correlation between vector A and the prototype vector sequence is 0.90, the correlation between vector B and the prototype vector sequence is 0.50, and the correlation threshold is 0.80, it means that the predicted character sequence represented by vector B does not conform to the grammatical rules, and vector A can be determined as a candidate vector sequence as the intended expression of the recognized text.

[0066] S14. Generate intent summary text based on the candidate vector sequence.

[0067] Based on the above description, the candidate vector sequence represents the intent of the recognized text in vector form. Based on this, converting the candidate vector sequence into a natural language text sequence can produce a natural language intent summary text. The intent summary text represents the summary result of summarizing the intent of the recognized text.

[0068] In summary, the disclosed embodiments can extract the semantic information and word boundary information of the characters in the recognized text for use in the calculation of the self-attention mechanism to obtain a probability distribution vector, and use the candidate vector as the compliance screening gate of the probability distribution vector to screen out a candidate vector sequence. Finally, based on the candidate vector sequence, an intent summary text is generated as the intent expression of the recognized text. In this way, the impact of word boundary ambiguity in the Chinese language environment on intent recognition can be alleviated, thereby improving the accuracy of intent recognition.

[0069] See also Figure 2 , Figure 2 It is a schematic diagram of application scenarios of the intent recognition method of some embodiments of the present disclosure.

[0070] exist Figure 2 In the illustrated application scenario, the intention recognition method M10 of the present disclosure can be implemented through an intention recognition neural network (neural network model). The intention recognition neural network includes an embedding module, a compilation module, a memory module and an output module. Among them, the embedding module is used to implement the method in step S11, that is, to extract the word meaning information and word boundary information of the characters in the recognized text. The compilation module is used to implement the method in step S12, that is, based on the self-attention algorithm, converting the word meaning information and word boundary information into multiple probability distribution vectors. The memory module is used to store the candidate vector sequence library, and to implement the method in step S13, that is, to screen out at least one candidate vector sequence from multiple probability distribution vectors. The output module is used to generate an intention summary text based on the candidate vector sequence.

[0071] Regarding step S11, in some embodiments of the present disclosure, it may include the following Figure 3 Steps S111 to S113 are shown.

[0072] S111 , performing character decomposition processing and word segmentation processing on the recognized text to obtain a plurality of recognized characters and a plurality of recognized words respectively.

[0073] The text to be recognized is decomposed into characters to separate sentences consisting of consecutive Chinese characters into multiple Chinese characters. The separated individual Chinese characters are the recognized characters. The text to be recognized is segmented into words to separate sentences consisting of consecutive Chinese characters into multiple Chinese words. The separated individual Chinese words are the recognized words.

[0074] In one example, during word segmentation processing, in addition to obtaining the recognized words, the part-of-speech annotation of the recognized words can also be obtained based on the vocabulary, indicating the entity type of the recognized characters to help the machine understand the meaning of the recognized characters as part of the word meaning information.

[0075] S112 , performing character embedding processing and position embedding processing on each recognized character to obtain a query vector corresponding to each recognized character, and using multiple query vectors corresponding to each recognized character as word meaning information.

[0076] Character embedding refers to word embedding at the character level. Word embedding maps individual word segments to a continuous vector space, placing semantically similar segments closer together in this vector space to aid machine understanding of the segmented words. Word embedding can be implemented using word embedding tools. Commonly used word embedding tools include Word2Vec, GloVe, and BERT. Character-level word embedding uses a single character as the smallest embedding unit for word embedding.

[0077] Position embedding is used to represent the relative or absolute position of each word in a sentence, that is, the order of the words in the sentence, so that each word can retain the structural information of the sentence. In one example, position embedding can be implemented by encoding the word embedding words by position based on sine and cosine functions, for example, as shown in the following formulas 1 and 2.

[0078] Formula 1: ;

[0079] Formula 2: .

[0080] in, represents the position code, It represents the position of the word in the sentence, i is the dimension index of the word embedding vector of the word segmentation, and d is the total dimension of the word embedding vector of the word segmentation. i and d can be obtained by word embedding processing of the word segmentation.

[0081] Finally, the character embedding and position embedding results for the same character are superimposed to produce a vector containing position information. This vector, derived from the recognized character, is called the character embedding vector. The dot product of the character embedding vector and the query vector weight matrix is ​​then calculated to obtain the query vector corresponding to the recognized character, which is used in the self-attention mechanism calculations. The query vector corresponding to each recognized character in the recognized text represents the semantic information corresponding to the recognized text. The query vector weight matrix can be obtained by training the intent recognition neural network.

[0082] S113. Perform word embedding and position embedding processing on each recognized word to obtain a key vector and a value vector corresponding to each recognized word, and use the multiple key vectors and multiple value vectors corresponding to each recognized word as word boundary information.

[0083] Among them, the word embedding processing performed on the recognized words is a word embedding method that uses words as the smallest embedding unit. Similar to the method of obtaining the word meaning information corresponding to the recognized text, the word embedding processing results and the position embedding processing results of the same word are superimposed to obtain a vector containing position information. This vector obtained based on the recognized word is called a word embedding vector. Then, the dot product of the word embedding vector and the key vector weight matrix is ​​calculated to obtain the key vector corresponding to the recognized word, and the dot product of the word embedding vector and the value vector weight matrix is ​​calculated to obtain the value vector corresponding to the recognized word. The key vector and value vector are used for the calculation of the self-attention mechanism. The key vector and value vector corresponding to each recognized word in the recognized text are the word boundary information corresponding to the recognized text. Among them, the key vector weight matrix and the value vector weight matrix can be obtained by training the intent recognition neural network.

[0084] Regarding step S12, in some embodiments of the present disclosure, it may include the following Figure 4 Steps S121 to S123 are shown.

[0085] S121. Based on the self-attention algorithm, the similarity between each query vector in the word meaning information and each key vector in the word boundary information is calculated as the attention score of each recognized word.

[0086] Among them, the query vector represents the focus in the self-attention mechanism. Based on the correlation between the query vector and the key vector, it can be determined which value vector is most relevant to the current focus. In the embodiment of the present disclosure, the query vector is a vector obtained based on the recognized characters, which is mainly used to represent the meaning of the characters; the key vector is a vector obtained based on the recognized words, which is mainly used to represent the word boundaries. Based on this, when converting the word meaning information and word boundary information into a probability distribution vector based on the self-attention algorithm, the word meaning information and word boundary information can be fused to enhance the characteristics of Chinese words in the Chinese language environment and improve the accuracy of intent recognition.

[0087] In one example, the similarity between a query vector and a key vector is the normalized representation of the dot product between the query vector and the key vector. That is, after normalizing the dot product between the query vector and the key vector, the similarity between the two can be expressed as a score in the range [0, 1]. In one example, the similarity between the query vector and the key vector can be normalized using the softmax function.

[0088] S122. Perform weighted summation on the attention score of the recognized word and each value vector in the word boundary information to obtain a context vector for each recognized word, and generate a context vector sequence based on the context vector of each recognized word.

[0089] The value vector contains the specific content information of each word in the recognized text. The similarity calculated in step S121 can represent the correlation between the key vector and the focus point, and the key vector contains the context information of the recognized text. Based on this, the attention score of the recognized word is used as the weight value of the value vector, and the attention score of the recognized word and each value vector in the word boundary information are weighted and summed. The obtained vector is the context vector that represents the focus of attention, which can help the machine understand the context information and the focus of attention of the context. Each value vector corresponds to a recognized word, so the corresponding context vector can be obtained for each recognized word. According to the order of the recognized words in the recognized text, the context vectors corresponding to each recognized word are combined into a vector sequence, and the obtained vector sequence is the context vector sequence.

[0090] Please combine Figure 2 In one example, the compilation module includes an encoder and a decoder. The encoder receives word meaning information and word boundary information provided by the embedding module. The encoder is configured to execute steps 121 and 122 above and output a context vector sequence. The decoder is configured to execute step S123 and generate a probability distribution vector based on the context vector sequence output by the encoder.

[0091] S123. Generate a probability distribution vector based on the attention weight of each attention head in the multiple attention heads and the context vector sequence to obtain multiple probability distribution vectors.

[0092] In the multi-head self-attention mechanism, each calculation based on the multi-head self-attention algorithm is called an attention head. Different attention heads can focus on calculating different types of dependencies. Different attention heads are configured with different attention weights. These weights can be assigned to different attention heads to focus on different types of dependencies. The attention weights of the attention heads can be obtained by training the intent recognition neural network.

[0093] In the multi-head self-attention mechanism, the representation of each word is updated based on the dependencies between it and other words. These dependencies are quantitatively represented by attention scores. For example, for adjacent or close words, the attention scores between them are generally higher, indicating a stronger dependency. In contrast, for words that are farther apart, the attention scores between them are generally lower, indicating a weaker dependency.

[0094] Please combine Figure 2In one example, the probability distribution vector is obtained by decoding the context vector sequence by the decoder based on the attention weight. The architecture of the encoder and decoder can adopt the architecture based on the transformer model to realize the encoding and decoding of the context vector sequence. Figure 2 As shown in the figure, the encoder and decoder architecture based on the Transformer model consists of a self-attention layer and a feedforward neural network layer, each linked by residual connections and layer normalization units (add&norm). Compared to the encoder, the decoder adds an encoder-decoder attention layer between the encoder's self-attention layer and the feedforward neural network layer. The encoder-decoder attention layer inherits parameters from the encoder's self-attention layer, such as attention weights. This allows the decoder to not only rely on its own generated content but also refer to the intent expression extracted by the encoder to understand the intent of the recognized text. The decoder outputs a probability distribution vector, representing the probability of each character in the predicted character sequence.

[0095] For example, the decoder output is "[D1, D2, D3, D4, D5, D6, D7, D8]", where each Di represents a vector indicating the probability of occurrence of the character predicted by the current step size.

[0096] Furthermore, during decoding, the decoder uses a masked multi-head self-attention mechanism to shield future context information and decode the predicted next character one by one.

[0097] Specifically, regarding step S123, in some embodiments of the present disclosure, it may include the following: Figure 5 Steps S1231 to S1235 are shown.

[0098] S1231. Calculate the occurrence probability of the kth predicted character based on the 1st to k-1th context vectors in the context vector sequence and the attention weight of the attention head to obtain the kth probability vector.

[0099] Where k is an integer greater than 1. The kth predicted character is the character predicted by the decoder at the kth prediction step position. The decoder first predicts the probability of the character corresponding to a step position and then outputs the character with the highest probability as the predicted character for that step position from the character table. During this process, the probability of the kth predicted character is calculated based on the 1st to k-1th context vectors in the context vector sequence and the attention weights of the attention head. The 1st to k-1th context vectors are the currently obtained predicted characters (the context vectors corresponding to the 1st to k-1th predicted characters).

[0100] That is to say, when calculating the probability of occurrence of the k-th predicted character, the context information corresponding to the predicted characters at the step position that did not appear will be blocked, that is, the future context information will be blocked, and the historical context information (the 1st to the k-1th context vectors) will be fused through the attention weight of the attention head to predict the next character, thereby improving the robustness of the prediction.

[0101] S1232. Generate the kth predicted character based on the kth probability vector.

[0102] A probability vector represents the probability of one or more characters appearing at a step position. Given the probability vector for the k-th step position, the softmax function is used to convert the k-th probability vector into probability values ​​corresponding to one or more characters. The character with the maximum probability value is then used as the k-th predicted character.

[0103] S1233. Determine the kth context vector corresponding to the kth predicted character in the context vector sequence.

[0104] Once the k-th predicted character has been obtained, the context information up to the current step size can be determined. The k-th context vector corresponding to the k-th predicted character in the context vector sequence is the cutoff point of the context information for the current step size. Generally, the predicted character sequence composed of predicted characters generated by the decoder has the same dimension as the context vector sequence provided by the encoder. Based on this, the context vector with the same dimension as the k-th predicted character is the k-th context vector corresponding to the k-th predicted character.

[0105] S1234. Calculate the occurrence probability of the k+1th predicted character based on the 1st to kth context vectors in the context vector sequence and the attention weight of the attention head to obtain the k+1th probability vector.

[0106] Step S1234 is basically the same as step S1232, except that the kth probability vector obtained in step S1232 is the prediction result of the kth prediction step, and the k+1th probability vector obtained in step S1234 is the prediction result of the k+1th prediction step, which represents the probability distribution of the next adjacent character after the kth predicted character.

[0107] S1235. Generate a probability distribution vector based on each probability vector corresponding to the attention head.

[0108] Calculate the probability vector for each prediction step one by one according to the method of steps S1231 to S1234. Assuming that the dimension of the context vector sequence is n, the calculation of the probability vector for an attention head is considered complete when k+1=n. At this time, the sequence of probability vectors corresponding to this attention head is used as the probability distribution vector corresponding to this attention head. Where n is the dimension of the context vector sequence.

[0109] Regarding step S13, in some embodiments of the present disclosure, it may include the following Figure 6 Steps S131 to S133 are shown.

[0110] S131. Calculate the correlation score between each probability distribution vector and each prototype vector sequence respectively, and generate a correlation matrix based on the calculated correlation scores.

[0111] The correlation matrix represents the correspondence between the probability distribution vector and the correlation score. In one example, the correlation score between the probability distribution vector and the prototype vector sequence is the normalized similarity between the probability distribution vector and the prototype vector sequence.

[0112] S132: Filter the correlation scores in the correlation matrix based on the correlation threshold, and filter the correlation scores that are higher than the correlation threshold as target scores.

[0113] In one example, the memory module includes a memory matrix and an adaptive gating layer. The memory matrix includes a library of prototype vector sequences, and the output of the memory matrix is ​​m vectors and the correlation matrix corresponding to the m vectors. The adaptive gating layer is a fully connected layer after the correlation matrix, which is used to filter the vectors output by the memory matrix. Here, m is the number of heads in the multi-head attention mechanism. The adaptive gating layer filters the vectors output by the memory matrix based on a correlation threshold. Specifically, the adaptive gating layer traverses each correlation score in the correlation matrix and selects the correlation score above the correlation threshold as the target score.

[0114] S133: Query the probability distribution vector corresponding to the target score in the correlation matrix, and determine the queried probability distribution vector as a candidate vector sequence.

[0115] When the adaptive gating layer filters out the target scores, the probability distribution vector corresponding to each target score can be queried in the correlation matrix. The queried probability distribution vector is the sequence that satisfies the grammatical rules and is output as the candidate vector sequence. This can avoid outputting non-existent character sequence combinations.

[0116] Regarding step S14, in some embodiments of the present disclosure, it may include the following Figure 7 Steps S141 to S144 are shown.

[0117] S141. Calculate the adaptive weight of the candidate vector sequence based on the correlation score of the candidate vector sequence.

[0118] In one example, the correlation scores of multiple filtered candidate vector sequences can be normalized, and an adaptive weight can be calculated for each normalized correlation score, ensuring that the sum of these calculated adaptive weights is 1. The adaptive weight obtained after each normalized correlation score is the adaptive weight of the candidate vector sequence corresponding to that correlation score. This allows for dynamic adjustments based on the correlation score of each sequence, ensuring that the final output sequence both conforms to the rule prototype and exhibits contextual relevance.

[0119] S142: Calculate the weighted average sum of the products of each candidate vector sequence and the corresponding adaptive weight to obtain a fused vector sequence.

[0120] By weighted fusion of multiple candidate sequences, the uncertainty brought by a single sequence can be reduced, and the robustness and accuracy of intent recognition can be improved.

[0121] S143. Convert the fused vector sequence into a fused probability sequence, where the fused probability sequence represents the probability distribution of multiple intent labels.

[0122] In one example, a softmax function can be used to convert a fused vector sequence into a fused probability sequence. The fused probability sequence includes multiple intent probability values, each corresponding to an intent label in the intent label library. Based on this, the fused probability sequence can represent the probability distribution of multiple intent labels. Intent labels are words or phrases used to profile intent. For example, intent labels can be the word "query," the word "calculation," and so on.

[0123] S144: Output the intent label with the highest probability in the fused probability sequence as the intent summary text.

[0124] For example, in the probability distribution of the fused probability sequence representation, the probability of the intent label "query" is 95%, and the probability of the intent label "calculation" is 83%. Then the intent label "query" is output as the intent summary text.

[0125] In some embodiments of the present disclosure, the intention recognition method M10 further includes: Figure 8 Step S15 is shown.

[0126] S15. Obtain entity tags corresponding to the fused vector sequence, and output entity words in the identified text based on the entity tags.

[0127] Please combine Figure 2The intent recognition method M10 can be performed by the trained intent recognition neural network. When training the intent recognition neural network, the original text in which the user asks a question in the question and answer task is taken as an input sequence of a training set, and the labeled entity information is taken as a target sequence of the training set. The labeled entity information includes a labeled intent label and an entity label. In this way, the target sequence contains the labels of the intent recognition and entity recognition tasks, and the intent recognition task and the entity recognition task can share the weights of the intent recognition neural network, so that the trained intent recognition neural network can not only recognize the intent of the recognized text, but also extract the entity words in the recognized text, thereby improving the robustness and accuracy of the intent recognition neural network.

[0128] For example, one original text of the training set is "give me the row data of CA Zhang San in Jiangbei District, Chongqing City", and the corresponding target sequence is "calculate / intent give / O me / O Chongqing City / B- city name Jiangbei District / B- district name CA / B- position Zhang San / B- personal name's / O row / B- indicator name data / O". Among them, "calculate / intent" is the intent label of the target sequence, and " / intent" indicates that the character or word labeled is the intent summary of the target sequence. "Chongqing City / B- city name" is an entity label of the target sequence, indicating that "Chongqing City" is an entity word in the target sequence, and the classification of the entity word is "city name". The word "data" in the target sequence does not have the classification of the entity word, so the word "data" does not have a corresponding entity label and is labeled as "number / O" and "data / O" two separate characters in the target sequence. Based on this, the trained intent recognition neural network can generate a corresponding fusion vector sequence based on the recognized text when processing the recognized text, and output the fusion vector sequence as a text sequence with intent label annotation and entity label annotation in the same format as the target sequence, called a predicted text sequence. In this way, the words or phrases with intent label annotations in the predicted text sequence are output as intent summary text, and the words or characters with each entity label annotation in the predicted text sequence are output as entity words of the recognized text.

[0129] Based on any one of the above embodiments, the present disclosure also provides an intent recognition device. Figure 9 is a structural schematic block diagram of an intent recognition device according to an embodiment of the present disclosure.

[0130] As shown in Figure 9 , the intent recognition device includes:

[0131] The embedding module 110 is configured to extract the word sense information and the word boundary information of the characters in the recognized text.

[0132] The compilation module 120 is used to convert word meaning information and word boundary information into multiple probability distribution vectors based on the self-attention algorithm, where the probability distribution vectors represent the probability of the predicted character appearing in the predicted sequence.

[0133] The memory module 130 is configured to select at least one candidate vector sequence from a plurality of probability distribution vectors, where the candidate vector sequence is a probability distribution vector having a correlation with any prototype vector sequence in the prototype vector sequence library that is higher than a correlation threshold.

[0134] The output module 140 is configured to generate an intent summary text based on the candidate vector sequence.

[0135] The above-mentioned intention recognition device can be computer software, and its various modules can be computer software modules. The implementation process of the functions and effects of each module in the above-mentioned intention recognition device can be specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.

[0136] The execution subject of the intention recognition method in the specific implementation of the present disclosure may be an electronic device such as a server.

[0137] Based on any of the above embodiments, the present disclosure further provides an electronic device, which can execute the intention recognition method of any of the embodiments described above in the present disclosure.

[0138] Figure 10 1 is a schematic block diagram of the structure of an electronic device 1000 according to an embodiment of the present disclosure.

[0139] The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnecting buses and bridges, depending on the specific application and overall design constraints of the hardware. The bus 1100 connects various circuits including one or more processors 1200, memory 1300, and / or hardware modules. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0140] Bus 1100 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component Architecture (EISA) bus. Buses can be classified as address buses, data buses, control buses, and the like. For ease of illustration, this figure shows only one connecting line, but this does not imply that there is only one bus or only one type of bus.

[0141] The present disclosure also provides a readable storage medium having a computer program stored therein, which is used to implement the above-mentioned method when the computer program is executed by a processor. "Readable storage medium" can be any device that can contain, store, communicate, propagate or transmit a program for use in an instruction execution system, device or equipment or in combination with these instruction execution systems, devices or equipment. More specific examples of readable storage media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.

[0142] The present disclosure also provides a computer program product. The methods of the present disclosure can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer programs or instructions are loaded and executed, the processes or functions of the present disclosure are performed in whole or in part.

[0143] A computer program or instruction can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, a computer program or instruction can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. A readable storage medium can be any accessible medium or a data storage device such as a server or data center that integrates one or more accessible media. The accessible medium can be a magnetic medium such as a floppy disk, hard disk, or magnetic tape; an optical medium such as a digital video disk; or a semiconductor medium such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.

[0144] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0145] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable intent recognition device to produce a machine, so that the instructions executed by the processor of the computer or other programmable intent recognition device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0146] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable intent recognition device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0147] These computer program instructions can also be loaded onto a computer or other programmable intent recognition device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0148] In the description of this specification, the description with reference to the terms "one embodiment / method", "some embodiments / methods", "example", "specific example", or "some examples" means that the specific features, structures, or characteristics described in conjunction with the embodiment / method or example are included in at least one embodiment / method or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / method or example. Moreover, the specific features, structures, or characteristics described may be combined in a suitable manner in any one or more embodiments / methods or examples. In addition, those skilled in the art may combine and combine different embodiments / methods or examples described in this specification and the features of different embodiments / methods or examples, unless they are contradictory.

[0149] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0150] Those skilled in the art will appreciate that the above embodiments are merely intended to clearly illustrate the present disclosure and are not intended to limit the scope of the present disclosure. Other changes or modifications may be made based on the above disclosure, and such changes or modifications are still within the scope of the present disclosure.

Claims

1. A method for identifying intention, characterized in that: include: Extracting word meaning information and word boundary information of characters in the recognized text; Based on a self-attention algorithm, the word meaning information and the word boundary information are converted into a plurality of probability distribution vectors, wherein the probability distribution vectors represent the probability of the predicted character appearing in the predicted sequence; Screening out at least one candidate vector sequence from the plurality of probability distribution vectors, wherein the candidate vector sequence is the probability distribution vector whose correlation with any prototype vector sequence in the prototype vector sequence library is higher than a correlation threshold; as well as Generate an intent summary text based on the candidate vector sequence; The extracting word meaning information and word boundary information from the recognized text includes: Performing character decomposition processing and word segmentation processing on the recognized text to obtain a plurality of recognized characters and a plurality of recognized words respectively; Performing character embedding and position embedding processing on each of the recognized characters to obtain a query vector corresponding to each of the recognized characters, and using the multiple query vectors corresponding to the recognized characters as the word meaning information; and Each of the recognized words is subjected to word embedding processing and position embedding processing to obtain a key vector and a value vector corresponding to each of the recognized words, and multiple key vectors and multiple value vectors corresponding to each of the recognized words are used as the word boundary information.

2. The intention recognition method according to claim 1, characterized in that Based on the self-attention algorithm, the word meaning information and the word boundary information are converted into multiple probability distribution vectors, including: Based on the self-attention algorithm, calculating the similarity between each query vector in the word meaning information and each key vector in the word boundary information as an attention score for each recognized word; Performing a weighted summation on the attention score of the recognized word and each of the value vectors in the word boundary information to obtain a context vector for each of the recognized words, and generating a context vector sequence based on the context vector for each of the recognized words; and A probability distribution vector is generated based on the attention weight of each of the multiple attention heads and the context vector sequence to obtain multiple probability distribution vectors.

3. The intention recognition method according to claim 2, characterized in that Generating a probability distribution vector based on the attention weight of each of the plurality of attention heads and the context vector sequence to obtain the plurality of probability distribution vectors includes: Calculating the occurrence probability of the kth predicted character based on the 1st to k-1th context vectors in the context vector sequence and the attention weights of the attention heads to obtain a kth probability vector, where k is an integer greater than 1; Generating a k-th predicted character based on the k-th probability vector; Determining the kth context vector in the context vector sequence corresponding to the kth predicted character; Calculating the occurrence probability of the k+1th predicted character based on the 1st to kth context vectors in the context vector sequence and the attention weights of the attention heads to obtain a k+1th probability vector; and The probability distribution vector is generated based on each of the probability vectors corresponding to the attention head.

4. The intention recognition method according to claim 1, characterized in that Screening out at least one candidate vector sequence from the plurality of probability distribution vectors includes: Calculating a correlation score between each probability distribution vector and each prototype vector sequence respectively, and generating a correlation matrix based on the calculated correlation scores, wherein the correlation matrix represents a corresponding relationship between the probability distribution vector and the correlation score; Filtering the correlation scores in the correlation matrix based on the correlation threshold, and filtering the correlation scores higher than the correlation threshold as target scores; and The probability distribution vector corresponding to the target score is searched in the correlation matrix, and the searched probability distribution vector is determined as the candidate vector sequence.

5. The intention recognition method according to claim 4, characterized in that: Generating an intent summary text based on the candidate vector sequence, including: Calculating an adaptive weight of the candidate vector sequence based on the correlation score of the candidate vector sequence; Calculating a weighted average sum of the products of each candidate vector sequence and the corresponding adaptive weight to obtain a fused vector sequence; Converting the fused vector sequence into a fused probability sequence, where the fused probability sequence represents a probability distribution of multiple intent labels; and The intent label with the highest probability in the fused probability sequence is output as the intent summary text.

6. The intention recognition method according to claim 5, characterized in that: Also includes: Obtain entity tags corresponding to the fused vector sequence, and output entity words in the recognized text based on the entity tags.

7. An electronic device, characterized in that: include: a memory storing execution instructions; as well as A processor, wherein the processor executes the execution instructions stored in the memory, so that the processor executes the intention recognition method according to any one of claims 1 to 6.

8. A readable storage medium, characterized in that: The readable storage medium stores execution instructions, which, when executed by a processor, are used to implement the intention recognition method according to any one of claims 1 to 6.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the intention recognition method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Push information generation method and device

    CN110427617A

  • Data processing method and device, equipment and storage medium

    CN113821592A