An intention recognition method, device and electronic device
By obtaining and processing session data in the robot customer service system, calculating session superposition features and using recurrent neural network models for intent recognition, the problem of low user intent recognition accuracy in the prior art is solved, and a higher intent recognition accuracy is achieved.
Patent Information
- Application Number
- CN202210152930.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-02-18
AI Technical Summary
The existing robot customer service system has low user intention recognition accuracy and cannot meet user needs.
By obtaining session data, the feature information and implicit association expression information of session subdata are determined, the session superposition features are calculated, and the intention recognition is performed using the recurrent neural network model.
In patterned chat scenarios or open chat scenarios, user intentions can be accurately identified and the accuracy of intention recognition can be improved.
Smart Images

Figure CN114547265B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and more particularly, to a method, apparatus, and electronic device for intent recognition. Background Art
[0002] With the continuous increase in labor costs, in order to save on customer service centers, many enterprises have introduced robot customer service systems to enable self-service communication between the robot customer service systems and users.
[0003] During the communication process between the robot customer service system and the user, the robot customer service system often adopts a patterned process, that is, only when the user replies with a pre-set word can the robot customer service system recognize the user's intent. For example, the robot customer service system asks the user, "Are you consulting about the package fee?", and only when the user replies "yes" or "no" can the robot customer service system recognize whether the user wants to consult about the package fee.
[0004] With this processing method of the robot customer service system, the accuracy of user intent recognition is low, and thus the user's needs cannot be met. Summary of the Invention
[0005] In view of this, the present invention provides a method, apparatus, and electronic device for intent recognition to solve the problem that the accuracy of user intent recognition is low and thus the user's needs cannot be met.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] An intent recognition method, comprising:
[0008] Obtaining session data to be subjected to intent recognition; the session data includes at least one session sub-data;
[0009] Determining the feature information of the session sub-data, and determining the implicit association expression information corresponding to the feature information of the session sub-data;
[0010] Determining the session stacking feature of the last session sub-data in the session data; the session stacking feature of the session sub-data is determined by the implicit association expression information of the session sub-data, the session stacking feature of the previous session sub-data of the session sub-data, and the hidden state information;
[0011] Performing intent recognition on the session stacking feature to obtain the intent recognition result of the session data.
[0012] Optionally, determining the feature information of the session sub-data includes:
[0013] Performing dictionary encoding on each word in the session sub-data to obtain the dictionary encoding result corresponding to each word;
[0014] Combine the dictionary encoding results of each word in the session sub-data to obtain the dictionary encoding result of the session sub-data;
[0015] Based on the position of the session sub-data in the session data, perform position encoding on the session sub-data to obtain the position encoding result of the session sub-data;
[0016] Combine the dictionary encoding result of the session sub-data and the position encoding result to obtain the feature information of the session sub-data.
[0017] Optionally, determining the implicit association expression information corresponding to the feature information of the session sub-data includes:
[0018] Perform implicit association expression analysis on the feature information of the session sub-data to obtain the implicit association expression information of the session sub-data.
[0019] Optionally, determining the session superposition feature of the last session sub-data in the session data includes:
[0020] Obtain the implicit association expression information of the last session sub-data;
[0021] Obtain the session superposition feature of the previous session sub-data of the session sub-data, and obtain the hidden state information corresponding to the session superposition feature obtained based on the recurrent neural network sub-model;
[0022] Use the recurrent neural network sub-model to process the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information corresponding to the session superposition feature to obtain the session superposition feature of the last session sub-data in the session data.
[0023] Optionally, performing intent recognition on the session superposition feature to obtain the intent recognition result of the session data includes:
[0024] Perform dimensionality reduction processing on the session superposition feature to obtain a dimensionality reduction result;
[0025] Determine the probability values of the dimensionality reduction result under different intent category information;
[0026] Use the intent category information with the largest probability value as the intent recognition result of the session data.
[0027] An intent recognition device, comprising:
[0028] A data acquisition module, configured to acquire session data to be subjected to intent recognition; the session data includes at least one session sub-data;
[0029] An information determination module, configured to determine the characteristic information of the session sub-data and determine the implicit association expression information corresponding to the characteristic information of the session sub-data;
[0030] A feature determination module, configured to determine the session superposition feature of the last session sub-data in the session data; the session superposition feature of the session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information;
[0031] An intention recognition module, configured to perform intention recognition on the session superposition feature to obtain the intention recognition result of the session data.
[0032] Optionally, the information determination module includes:
[0033] A dictionary encoding sub-module, configured to perform dictionary encoding on each word in the session sub-data to obtain the dictionary encoding result corresponding to each word;
[0034] A first combination sub-module, configured to combine the dictionary encoding results of each word in the session sub-data to obtain the dictionary encoding result of the session sub-data;
[0035] A position encoding sub-module, configured to perform position encoding on the session sub-data based on the position of the session sub-data in the session data to obtain the position encoding result of the session sub-data;
[0036] A second combination sub-module, configured to combine the dictionary encoding result of the session sub-data and the position encoding result to obtain the characteristic information of the session sub-data.
[0037] Optionally, when the information determination module is configured to determine the implicit association expression information corresponding to the characteristic information of the session sub-data, it is specifically configured to:
[0038] Perform implicit association expression analysis on the characteristic information of the session sub-data to obtain the implicit association expression information of the session sub-data.
[0039] Optionally, the feature determination module includes:
[0040] An information acquisition sub-module, configured to acquire the implicit association expression information of the last session sub-data;
[0041] An information processing sub-module, configured to acquire the session superposition feature of the previous session sub-data of the session sub-data and acquire the hidden state information corresponding to the session superposition feature obtained based on the recurrent neural network sub-model;
[0042] A feature determination sub-module, configured to use the recurrent neural network sub-model to process the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information corresponding to the session superposition feature, so as to obtain the session superposition feature of the last session sub-data in the session data.
[0043] An electronic device, comprising: a memory and a processor;
[0044] Wherein, the memory is used to store programs;
[0045] The processor calls the program and is used to execute the above-mentioned intention recognition method.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] The present invention provides an intention recognition method, device and electronic device, which obtain session data to be subjected to intention recognition, the session data includes at least one session sub-data, determine the feature information of the session sub-data, and determine the implicit association expression information corresponding to the feature information of the session sub-data, determine the session superposition feature of the last session sub-data in the session data, the session superposition feature of the session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information, and perform intention recognition on the session superposition feature to obtain the intention recognition result of the session data. In the present invention, by determining the session superposition feature of the last session sub-data and performing intention recognition on the session superposition feature, the intention recognition result of the session data can be obtained, and the user intention can be accurately recognized in both patterned chat scenarios and open chat scenarios. Further, in the present invention, the session superposition feature of the last session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information, so that when performing intention recognition, the session context information is considered, thereby improving the accuracy of user intention recognition. Description of the Drawings
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings according to the provided drawings without creative efforts.
[0049] Figure 1 It is a flowchart of a method for an intention recognition method provided by an embodiment of the present invention;
[0050] Figure 2 A method flow chart of another intent recognition method provided by an embodiment of the present invention;
[0051] Figure 3 A schematic diagram of a partial model of an intent recognition model provided by an embodiment of the present invention;
[0052] Figure 4 A method flow chart of another intent recognition method provided by an embodiment of the present invention;
[0053] Figure 5 A schematic diagram of another part of a model of an intent recognition model provided by an embodiment of the present invention;
[0054] Figure 6 A schematic diagram of the structure of an intention recognition device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0055] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0056] In the process of communication between the robot customer service system and the user, the robot customer service system often adopts a patterned process, that is, the robot customer service system can only recognize the user's intention when the user replies with pre-set words.
[0057] Specifically, we use the method of designing dialogue logic to realize multi-round dialogue. It can solve many industrial problems. In particular, some relatively fixed processes, such as: telemarketing, the robot asks the user if he is interested; here the most important thing for the robot is not to facilitate the order, but to screen the interested users. For example, if the user says that he is interested, or even chats with the robot for a few more sentences, he will be marked as interested, and the tree structure will be used to judge the current dialogue intention layer by layer and select the corresponding dialogue response.
[0058] For example, the robot customer service system may ask the user, "Do you want to inquire about the package fees?" The user replies "yes" or "no," and the robot customer service system can then recognize whether the user intends to inquire about the package fees.
[0059] However, this method is only applicable to patterned processes. For open chat or customer service robots, the processing method of this type of robot customer service system has low accuracy in user intent recognition, cannot meet its intelligent requirements, and thus cannot meet user needs.
[0060] To solve this technical problem, the inventors found that for patterned chat scenarios or open-ended chat scenarios, the speech intent of a user can be recognized through a deep neural network model, which requires constructing a spatial representation form of user text. Currently, the common representation form is to use the current text of the user's speech as a separate representation to recognize the intent expressed in the current sentence. However, when a sentence of the user does not carry key information, the user's intent cannot be recognized.
[0061] Furthermore, the inventors found that when performing intent recognition, the context information of the user can be considered, so that more key information in the context can be extracted to recognize the user's intent and improve the accuracy of user intent recognition.
[0062] Specifically, in the model training stage, first, multiple conversation contents representing the same intent are used as a training sample, and two splicing methods, namely intra-sentence (dictionary encoding) and inter-sentence (position encoding), are constructed. After splicing and fusion, a multi-layer deep neural network is used to obtain the feature expressions within and between multiple conversations. After fusing the two feature expressions, an intent recognition model is trained with the intent category as the target. In the model prediction stage, taking advantage of the characteristics of the redis database for fast caching and backup of data, the first N sentence dialogue texts of the current text in the same conversation are obtained in real time, and through the trained intent recognition model, an intent type of the current content is judged, and the redis database storing historical conversation data is updated. This method ensures data consistency in both the training and prediction stages, not only utilizes the current conversation, but also does not lose the influence of historical conversation content on the current intent judgment, and also ensures the real-time nature of the data.
[0063] Specifically, the present invention provides an intention recognition method, apparatus, and electronic device, which obtain session data to be recognized for intention, where the session data includes at least one session sub-data, determine the feature information of the session sub-data, and determine the implicit association expression information corresponding to the feature information of the session sub-data, determine the session superposition feature of the last session sub-data in the session data, where the session superposition feature of the session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information, perform intention recognition on the session superposition feature, and obtain the intention recognition result of the session data. In the present invention, by determining the session superposition feature of the last session sub-data and performing intention recognition on the session superposition feature, the intention recognition result of the session data can be obtained, and the user intention can be accurately recognized in both patterned chat scenarios and open chat scenarios. Further, in the present invention, the session superposition feature of the last session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information, so that when performing intention recognition, the session context information is considered, thereby improving the accuracy of user intention recognition.
[0064] On the basis of the above content, another embodiment of the present invention provides an intention recognition method, which can be applied to intention recognition devices such as processors and servers. Referring to Figure 1 , the intention recognition method may include:
[0065] S11. Obtain session data to be recognized for intention.
[0066] In this embodiment, the session data to be recognized for intention may be the same session during the user's chat with the customer service robot. Generally speaking, when the user and the customer service robot adopt a question-and-answer method, at this time, the user's session can be extracted. Each sentence of the session is used as a session sub-data.
[0067] For example, the customer session content is as follows:
[0068] The first sentence. Hello,
[0069] The second sentence. I would like to ask what the insurance coverage of this product is?
[0070] The third sentence. It is the XX product.
[0071] It should be noted that when the user finishes answering a sentence, the method of the present invention can be used for intention recognition. If it can be recognized, then when performing intention recognition next time, it directly starts from the second sentence. If it cannot be recognized, then when the user finishes the second sentence, the first and second sentences are used for intention recognition. If it still cannot be recognized, then the third, fourth, and up to the Nth sentence are combined, where N can generally be 10.
[0072] In this embodiment, the conversation for which intention recognition needs to be performed is referred to as conversation data, and the conversation data includes at least one conversation sub-data. Specifically, each sentence in the conversation data is used as a conversation sub-data, and the conversation data includes at most the above N conversation sub-datas. The characteristics of fast caching and backup of data in the redis database can be utilized in advance to cache the conversation sub-datas and update the cached conversation sub-datas in real time. This ensures data consistency in both the training and prediction phases, not only utilizes the current conversation, but also does not lose the influence of historical conversation content on the current intention judgment, and also ensures the real-time nature of the data.
[0073] It should be noted that if the number of conversation sub-datas in the conversation data for which intention recognition is performed this time is less than N, it can be supplemented by adding blank conversation sub-datas until the number of conversation sub-datas is N.
[0074] S12. Determine the feature information of the conversation sub-data, and determine the implicit association expression information corresponding to the feature information of the conversation sub-data.
[0075] In practical applications, the feature information includes: dictionary encoding result and position encoding result.
[0076] Specifically, referring to Figure 2 , step S12 may include:
[0077] S21. Perform dictionary encoding on each word in the conversation sub-data to obtain the dictionary encoding result corresponding to each word.
[0078] Specifically, referring to Figure 3 , the conversation sub-data ( Figure 3 the first sentence, the second sentence, the third sentence... in it) can be used as the input, and by performing dictionary encoding on each word in each sentence of the conversation sub-data, the dictionary encoding result corresponding to each word is obtained.
[0079] S22. Combine the dictionary encoding results of each word in the conversation sub-data to obtain the dictionary encoding result of the conversation sub-data.
[0080] Specifically, the dictionary encoding results of each character in the session sub-data are combined according to the character identifier (token_id, specifically referring to the position in the session sub-data) of each character.
[0081] During the combination process, the number of dictionary encoding results of characters in the dictionary encoding result of the session sub-data is limited. For example, it can be 128, and 128 means that each sentence can encode at most 128 characters. When it is less than 128, it is padded with 0s. When it exceeds 128, the excess part is truncated, and the dictionary encoding result token_embedding of the session sub-data can be obtained.
[0082] S23. Based on the position of the session sub-data in the session data, perform position encoding on the session sub-data to obtain the position encoding result of the session sub-data.
[0083] Specifically, referring to Figure 3 , perform position encoding on the sentence order (segment_id) of the session sub-data in sequence. For example, when the session sub-data is the first sentence in the session data, the position encoding result of the session sub-data is 000000……, where the number of 0s is equal to the number of dictionary encoding results of characters in the dictionary encoding result of the session sub-data, such as 128. Then the position encoding result of the second sentence is 128 1s, and the position encoding result of the third sentence is 128 0s, and so on until the last sentence. The position encoding result in this embodiment can be called segment_id_embedding.
[0084] S24. Combine the dictionary encoding result and the position encoding result of the session sub-data to obtain the feature information of the session sub-data.
[0085] Specifically, for each session sub-data, add its dictionary encoding result token_embedding and the position encoding result segment_id_embedding, and the feature information of the session sub-data can be obtained (that is, Figure 3 the combine_embedding in
[0086] token_embedding + segment_id_embedding = combine_embedding.
[0087] After determining the feature information of the session sub-data, it is necessary to determine the implicit association expression information corresponding to the feature information of the session sub-data.
[0088] Specifically, perform implicit association expression analysis on the feature information of the session sub-data to obtain the implicit association expression information of the session sub-data.
[0089] In detail, for the session sub-data, input the feature information of the session sub-data into Figure 3 the multi-head-self-attention network structure in, the function of this structure is to allow the model to learn relevant information in different representation sub-spaces, that is, to learn the implicit association expression between words within a single-sentence session. Each sentence (session sub-data) in a session data will pass through this network layer to obtain a matrix of text representations containing the implicit association expression within the sentence (i.e., implicit association expression information): H m , in a session data, we can obtain where n is the number of session sub-data. When n is less than N, pad with 0 to obtain the implicit association expression matrix of the session data.
[0090] S13. Determine the session superposition feature of the last session sub-data in the session data; the session superposition feature of the session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information.
[0091] S31. Obtain the implicit association expression information of the last session sub-data.
[0092] The implicit association expression information can be determined through step S12, and the specific implementation process refers to the corresponding description above.
[0093] S32. Obtain the session superposition feature of the previous session sub-data of the session sub-data, and obtain the hidden state information corresponding to the session superposition feature based on the recurrent neural network sub-model.
[0094] S33. Use the recurrent neural network sub-model to process the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information corresponding to the session superposition feature to obtain the session superposition feature of the last session sub-data in the session data.
[0095] Specifically, refer to Figure 3 , input the implicit association expression information corresponding to each session sub-data above into the neural network layer of the GRU (Gate Recurrent Unit, referred to as the recurrent neural network sub-model in this embodiment) structure. This neural network structure is a type of recurrent neural network. The advantage of the recurrent neural network is that it has a current input, that is, the implicit association expression information and the hidden state information passed down from the previous session sub-data This hidden state information contains relevant information of the previous session sub-data.
[0096] Then, combined with and the GRU will obtain the current hidden node, that is, the output of the current session sub-data, namely the session stacked feature and the hidden state information passed to the next session sub-data Thus, when calculating each loop point, the hidden state in the previous loop point can be taken into account. For example, the transmission of a sample with three sentences in this part is as follows:
[0097] The representation vector of the first sentence (Initial hidden state, which can be configured according to the actual situation, such as a zero matrix), through the GRU neural network layer, the output of the current node is obtained and the hidden state passed from the current node to the next node
[0098] The representation vector of the second sentence ( are the session stacked feature and the hidden state of the previous node respectively), through the GRU neural network layer, the output of the current node is obtained and the hidden state passed from the current node to the next node
[0099] The representation vector of the third sentence (Hidden state of the previous node), through the GRU neural network layer, the output of the current node is obtained and the hidden state passed from the current node to the next node
[0100] Two points need to be explained here:
[0101] (1) During each transmission, using the residual structure, the output of the previous node is added to the input of the current node and combined as the input of the current node. Here, is used for the operation. The advantage of the residual network is that it can solve part of the degradation problem of the deep neural network.
[0102] (2) ⊙ represents the Hadamard Product of the input and the hidden state inside the GRU. There are multiple calculations here.
[0103] Through the calculations of these three parts, including initial encoding, multi-head attention mechanism, and GRU neural network, the session stacking features of the sample are finally output. The session stacking features include the character information of the text itself (token_embdinng), the sequential relationship of session sentences (segment_id_embedding), the implicit association expression information within a single-sentence session, and the hidden state information between sentences. The session stacking feature final_embedding before the final intention category classification is obtained.
[0104] S14. Perform intention recognition on the session stacking features to obtain the intention recognition result of the session data.
[0105] Specifically, perform dimensionality reduction processing on the session stacking features to obtain a dimensionality reduction result, and then determine the probability values of the dimensionality reduction result under different intention category information. Finally, take the intention category information with the largest probability value as the intention recognition result of the session data.
[0106] Refer to Figure 5 , Figure 5 The classification sub-model mainly passes the feature vector of the session stacking feature final_embedding encoded and generated through a (GlobalAveragePooling) global average layer to reduce the dimension to obtain a low-dimensional expression, followed by a fully connected layer with the activation function sofmax to obtain a K-dimensional probability distribution with values between 0 and 1. K represents the number of intention category information to be classified. Finally, select the dimension with the highest probability as the final output category, that is, the intention recognition result of the session data.
[0107] It should be noted that Figure 3 the model of Figure 5 and the classification sub-model of
[0108] can constitute the intention recognition model in the present invention. The intention recognition model is obtained through training. When determining the training samples, with N (N is at most 10) as the window size, select N consecutive session sub-data within a session as a sample, and use the intention category information corresponding to this sample, that is, the intention type, as the sample label to construct multiple text pairs of sample-label as the training data.
[0109] After determining the training data, through Figure 3For the token_embedding (dictionary encoding) in , an N×128-dimensional matrix of a sample can be obtained, where N can be 10 as described above. For those exceeding this number, they will be truncated, and for those insufficient, they will be padded with the default value 0.
[0110] After that, processing is performed according to the Figure 3 network architecture to obtain the session superimposed feature final_embedding of the last session sub-data, which is input into the Figure 5 classification sub-model, and then the final intention recognition result can be obtained.
[0111] To enable those skilled in the art to understand the present invention more clearly, examples are given below:
[0112] For example, the customer's conversation content is as follows:
[0113] 1. Hello,
[0114] 2. I would like to ask what the insurance coverage of this product is?
[0115] 3. It is the xx product.
[0116] From these three sentences, we can obtain the following points:
[0117] A. The customer's intention is the insurance coverage of xx;
[0118] B. This intention can only be obtained after the customer finishes saying these three sentences;
[0119] C. The clue of this intention is obtained from the second and third sentences.
[0120] Because with the above model structure design, when the customer finishes saying the third sentence, the first two sentences can be combined as a session data, and clues A, B, and C can be obtained in the above model. Through the GRU, it can be obtained that the second sentence has the greatest influence on the current sentence (the third sentence), and thus the correct intention result can be obtained.
[0121] In this embodiment, session data to be subject to intent recognition is obtained. The session data includes at least one session sub-data. Feature information of the session sub-data is determined, and implicit association expression information corresponding to the feature information of the session sub-data is determined. A session superposition feature of the last session sub-data in the session data is determined. The session superposition feature of the session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and hidden state information. Intent recognition is performed on the session superposition feature to obtain an intent recognition result of the session data. In the present invention, by determining the session superposition feature of the last session sub-data and performing intent recognition on the session superposition feature, the intent recognition result of the session data can be obtained, and the user intent can be accurately recognized in both patterned chat scenarios and open chat scenarios. Further, in the present invention, the session superposition feature of the last session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and hidden state information, so that session context information is considered during intent recognition, thereby improving the accuracy of user intent recognition.
[0122] In addition, the present invention uses complete multi-turn session data and fully retains the complete content expression within a session. Through text character encoding, session sentence order encoding, implicit association expression analysis within a single-sentence session, and hidden state information between sentences, the intent judgment of one sentence or a paragraph of dialogue is more complete. In addition, by utilizing the characteristics of GRU, the sentence features of multiple sentences are associated, so that the final text representation before classification takes into account the implicit expression between the original orders of the texts.
[0123] Optionally, based on the above embodiment of the intent recognition method, another embodiment of the present invention provides an intent recognition device. Referring to Figure 6 , it may include:
[0124] A data acquisition module 11, configured to acquire session data to be subject to intent recognition; the session data includes at least one session sub-data;
[0125] An information determination module 12, configured to determine feature information of the session sub-data and determine implicit association expression information corresponding to the feature information of the session sub-data;
[0126] A feature determination module 13, configured to determine a session superposition feature of the last session sub-data in the session data; the session superposition feature of the session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and hidden state information;
[0127] An intent recognition module 14, configured to perform intent recognition on the session superimposed features to obtain an intent recognition result of the session data.
[0128] Further, the information determination module includes:
[0129] A dictionary encoding sub-module, configured to perform dictionary encoding on each word in the session sub-data to obtain a dictionary encoding result corresponding to each word;
[0130] A first combination sub-module, configured to combine the dictionary encoding results of each word in the session sub-data to obtain a dictionary encoding result of the session sub-data;
[0131] A position encoding sub-module, configured to perform position encoding on the session sub-data based on the position of the session sub-data in the session data to obtain a position encoding result of the session sub-data;
[0132] A second combination sub-module, configured to combine the dictionary encoding result of the session sub-data and the position encoding result to obtain feature information of the session sub-data.
[0133] Further, when the information determination module is configured to determine implicit association expression information corresponding to the feature information of the session sub-data, it is specifically configured to:
[0134] Perform implicit association expression analysis on the feature information of the session sub-data to obtain implicit association expression information of the session sub-data.
[0135] Further, the feature determination module includes:
[0136] An information acquisition sub-module, configured to acquire implicit association expression information of the last session sub-data;
[0137] An information processing sub-module, configured to acquire the session superimposed features of the previous session sub-data of the session sub-data, and acquire hidden state information corresponding to the session superimposed features obtained based on the recurrent neural network sub-model;
[0138] A feature determination sub-module, configured to use the recurrent neural network sub-model to process the implicit association expression information of the session sub-data, the session superimposed features of the previous session sub-data of the session sub-data, and the hidden state information corresponding to the session superimposed features to obtain the session superimposed features of the last session sub-data in the session data.
[0139] Further, the intent recognition module includes:
[0140] A dimensionality reduction sub-module, configured to perform dimensionality reduction processing on the session superimposed features to obtain a dimensionality reduction result;
[0141] A probability value determination sub-module, configured to determine the probability values of the dimensionality reduction results under different intent category information;
[0142] An intent determination sub-module, configured to use the intent category information with the largest probability value as the intent recognition result of the session data.
[0143] In this embodiment, session data to be subjected to intent recognition is obtained. The session data includes at least one session sub-data. The feature information of the session sub-data is determined, and the implicit association expression information corresponding to the feature information of the session sub-data is determined. The session superposition feature of the last session sub-data in the session data is determined. The session superposition feature of the session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information. Intent recognition is performed on the session superposition feature to obtain the intent recognition result of the session data. In the present invention, by determining the session superposition feature of the last session sub-data and performing intent recognition on the session superposition feature, the intent recognition result of the session data can be obtained, and the user intent can be accurately recognized in both patterned chat scenarios and open chat scenarios. Further, in the present invention, the session superposition feature of the last session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information, so that session context information is considered during intent recognition, thereby improving the accuracy of user intent recognition.
[0144] In addition, the present invention uses complete multi-round session data and fully retains the complete content expression within a session. Through the character encoding of the text itself, the session sentence order encoding, the analysis of the implicit association expression within a single-sentence session, and the hidden state information between sentences, the intent judgment of one sentence or a paragraph of dialogue is more complete. In addition, by utilizing the characteristics of GRU, the sentence features of multiple sentences are associated, so that the final text representation before classification considers the implicit expression between the original orders of the texts.
[0145] It should be noted that for the working processes of the various modules and sub-modules in this embodiment, please refer to the corresponding descriptions in the above embodiments, and will not be elaborated here.
[0146] Optionally, based on the above embodiments of the intent recognition method and apparatus, another embodiment of the present invention provides an electronic device, including: a memory and a processor;
[0147] Wherein, the memory is used to store a program;
[0148] The processor calls the program and is used to execute the above-mentioned intent recognition method.
[0149] In this embodiment, session data to be subjected to intent recognition is obtained. The session data includes at least one session sub-data. Feature information of the session sub-data is determined, and implicit association expression information corresponding to the feature information of the session sub-data is determined. A session superposition feature of the last session sub-data in the session data is determined. The session superposition feature of the session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and hidden state information. Intent recognition is performed on the session superposition feature to obtain an intent recognition result of the session data. In the present invention, by determining the session superposition feature of the last session sub-data and performing intent recognition on the session superposition feature, an intent recognition result of the session data can be obtained, and the user intent can be accurately recognized in both patterned chat scenarios and open chat scenarios. Further, in the present invention, the session superposition feature of the last session sub-data is determined by the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and hidden state information, so that session context information is considered during intent recognition, thereby improving the accuracy of user intent recognition.
[0150] In addition, the present invention adopts complete multi-round session data and fully retains the complete content expression within a session. Through character encoding of the text itself, session sentence order encoding, analysis of implicit association expressions within a single-sentence session, and hidden state information between sentences, the intent judgment of one sentence or a paragraph of dialogue is more complete. In addition, by utilizing the characteristics of GRU, the sentence features of multiple sentences are associated, so that the final text representation before classification takes into account the implicit expressions between the original orders of the texts.
[0151] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intention recognition method, characterized in that, Including: Obtain session data to be subjected to intent recognition; the session data includes at least one session sub-data; Determine the feature information of the session sub-data, and determine the implicit association expression information corresponding to the feature information of the session sub-data; Obtain the implicit association expression information of the last session sub-data; Obtain the session superposition feature of the previous session sub-data of the session sub-data, and obtain the hidden state information corresponding to the session superposition feature obtained based on the recurrent neural network sub-model; Use the recurrent neural network sub-model to process the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information corresponding to the session superposition feature, to obtain the session superposition feature of the last session sub-data in the session data, where the session superposition feature includes text itself character information, session sentence order relationship, implicit association expression information within a single-sentence session, and hidden state information between sentences; Perform intent recognition on the session superposition feature to obtain the intent recognition result of the session data.
2. The intention recognition method according to claim 1, characterized in that, Determining the feature information of the session sub-data includes: Perform dictionary encoding on each character in the session sub-data to obtain the dictionary encoding result corresponding to each character; Combine the dictionary encoding results of each character in the session sub-data to obtain the dictionary encoding result of the session sub-data; Perform position encoding on the session sub-data based on the position of the session sub-data in the session data to obtain the position encoding result of the session sub-data; Combine the dictionary encoding result and the position encoding result of the session sub-data to obtain the feature information of the session sub-data.
3. The intention recognition method according to claim 1, characterized in that, Determining the implicit association expression information corresponding to the feature information of the session sub-data includes: Perform implicit association expression analysis on the feature information of the session sub-data to obtain the implicit association expression information of the session sub-data.
4. The intention recognition method according to claim 1, characterized in that, Performing intent recognition on the session superposition feature to obtain the intent recognition result of the session data includes: Perform dimensionality reduction processing on the session superposition feature to obtain a dimensionality reduction result; Determine the probability value of the dimensionality reduction result under different intent category information; Use the intent category information with the largest probability value as the intent recognition result of the session data.
5. An intention recognition device, characterized in that, Including: A data acquisition module, configured to obtain session data to be subjected to intent recognition; the session data includes at least one session sub-data; An information determination module, configured to determine the feature information of the session sub-data and determine the implicit association expression information corresponding to the feature information of the session sub-data; An information acquisition sub-module, configured to obtain the implicit association expression information of the last session sub-data; An information processing sub-module, configured to obtain the session superposition feature of the previous session sub-data of the session sub-data, and obtain the hidden state information corresponding to the session superposition feature obtained based on the recurrent neural network sub-model; A feature determination sub-module, configured to use the recurrent neural network sub-model to process the implicit association expression information of the session sub-data, the session superposition feature of the previous session sub-data of the session sub-data, and the hidden state information corresponding to the session superposition feature, so as to obtain the session superposition feature of the last session sub-data in the session data, where the session superposition feature includes text itself character information, session sentence order relationship, implicit association expression information within a single-sentence session, and hidden state information between sentences; An intention recognition module, configured to perform intention recognition on the session superposition feature to obtain an intention recognition result of the session data.
6. The intention recognition device according to claim 5, characterized in that, The information determination module includes: A dictionary encoding sub-module, configured to perform dictionary encoding on each character in the session sub-data to obtain a dictionary encoding result corresponding to each character; A first combination sub-module, configured to combine the dictionary encoding results of each character in the session sub-data to obtain a dictionary encoding result of the session sub-data; A position encoding sub-module, configured to perform position encoding on the session sub-data based on the position of the session sub-data in the session data to obtain a position encoding result of the session sub-data; A second combination sub-module, configured to combine the dictionary encoding result of the session sub-data and the position encoding result to obtain feature information of the session sub-data.
7. The intention recognition device according to claim 5, characterized in that, When the information determination module is used to determine the implicit association expression information corresponding to the feature information of the session sub-data, it is specifically used for: Performing implicit association expression analysis on the feature information of the session sub-data to obtain the implicit association expression information of the session sub-data.
8. An electronic device, characterized in that, Including: A memory and a processor; Wherein, the memory is used to store a program; The processor calls the program and is used to execute the intention recognition method according to any one of claims 1-4.
Citation Information
Patent Citations
Record question and answer classification method based on ERNIE and DPCNN
CN111813938A
Emotion recognition method and device based on recurrent neural network and storage medium
CN111950275A
Generative chatting robot based on deep learning method
CN112364148A