A method and device for identifying intention and related equipment
By acquiring phoneme sequences and using the cross attention mechanism to fuse the features of text sequences and phoneme sequences, and combining with a fully connected neural network for intention prediction, the problem of low accuracy of oral intention recognition in the prior art is solved, and a higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111506575.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The prior art has low accuracy in oral intention recognition tasks and is difficult to effectively improve.
By obtaining phoneme sequences and feature extraction and fusion of text sequences and phoneme sequences based on the cross attention mechanism, sentence feature vectors are obtained, and intention prediction is made by combining fully connected neural networks.
It improves the accuracy of oral intention recognition and can reconstruct the sentence feature vector representations corresponding to the real text when facing pronunciation and tone errors.
Smart Images

Figure CN114187894B_ABST
Abstract
Claims
1. A method for identifying intent, It is characterized in that The method includes: Acquire a phoneme sequence corresponding to the audio to be detected, and determine a text sequence corresponding to the audio to be detected based on the phoneme sequence; Based on the cross attention mechanism, feature extraction and fusion are performed on the text sequence and the phoneme sequence to obtain a sentence feature vector corresponding to the audio to be detected; Based on the sentence feature vector, a prediction result for characterizing the intention of the audio to be detected is obtained, wherein the step of extracting and fusing features of the text sequence and the phoneme sequence based on a cross attention mechanism to obtain a sentence feature vector corresponding to the audio to be detected includes: Encoding the text sequence and the phoneme sequence respectively to obtain a text sequence feature vector corresponding to the text sequence and a phoneme sequence feature vector corresponding to the phoneme sequence; The text sequence feature vector and the phoneme sequence feature vector are concatenated according to the text sequence weight parameter and the phoneme sequence weight parameter to obtain the sentence feature vector, wherein the text sequence weight parameter and the phoneme sequence weight parameter are obtained by training a plurality of pre-stored audio samples.
2. The method for identifying intention according to claim 1, It is characterized in that The step of obtaining a prediction result for characterizing the intention of the audio to be detected based on the sentence feature vector specifically includes: The sentence feature vector is input into a fully connected neural network for classification prediction to obtain the prediction result.
3. The method for identifying intention according to claim 1, It is characterized in that The step of encoding the text sequence and the phoneme sequence to obtain a text sequence feature vector corresponding to the text sequence and a phoneme sequence feature vector corresponding to the phoneme sequence includes: A bidirectional neural network is used to encode the text sequence and the phoneme sequence to obtain a text sequence feature vector corresponding to the text sequence and a phoneme sequence feature vector corresponding to the phoneme sequence.
4. The method for identifying intention according to claim 1, It is characterized in that The text sequence weight parameter and the phoneme sequence weight parameter are obtained by training a plurality of pre-stored audio samples, specifically including: Encoding a sample text sequence and a sample phoneme sequence to obtain a sample text sequence feature vector corresponding to the sample text sequence and a sample phoneme sequence feature vector corresponding to the sample phoneme sequence; wherein the sample text sequence and the sample phoneme sequence are obtained by performing speech recognition on the audio sample; Calculating the cosine similarity between the embedding vector of each text in the sample text sequence feature vector and the sample phoneme sequence feature vector to obtain the weight of each text, and normalizing the weights corresponding to all texts through a normalized exponential function to obtain the text sequence weight parameter; The phoneme sequence weight parameter is obtained according to the text sequence weight parameter.
5. The method for identifying intention according to claim 4, It is characterized in that The text sequence weight parameter and the phoneme sequence weight parameter are adjusted in the following manner, specifically including: determining a true text transcribed from said audio sample listened to by a human; Based on a cross entropy loss function, calculating the difference between a prediction result representing the intent of the audio sample and the real text; Based on the difference, the text sequence weight parameter and the phoneme sequence weight parameter are adjusted.
6. An intention recognition device, It is characterized in that The device is used to implement the intention recognition method described in any one of claims 1 to 5, and the device includes: A speech recognition module, used to obtain a phoneme sequence corresponding to the audio to be detected, and determine a text sequence corresponding to the audio to be detected based on the phoneme sequence; A cross attention module, used for extracting and fusing features of the text sequence and the phoneme sequence based on a cross attention mechanism, and obtaining a sentence feature vector corresponding to the audio to be detected; A fully connected neural network classification module is used to obtain a prediction result for characterizing the intention of the audio to be detected based on the sentence feature vector.
7. The intention recognition device according to claim 6, It is characterized in that The fully connected neural network classification module includes: The sentence feature vector is input into a fully connected neural network for classification prediction to obtain the prediction result.
8. The intention recognition device according to claim 7, It is characterized in that The cross attention module specifically includes: An encoding unit encodes the text sequence and the phoneme sequence respectively to obtain a text sequence feature vector corresponding to the text sequence and a phoneme sequence feature vector corresponding to the phoneme sequence; A concatenation unit concatenates the text sequence feature vector and the phoneme sequence feature vector according to the text sequence weight parameter and the phoneme sequence weight parameter to obtain the sentence feature vector.
9. The intention recognition device according to claim 8, It is characterized in that The encoding unit specifically includes: A bidirectional neural network is used to encode the text sequence and the phoneme sequence to obtain a text sequence feature vector corresponding to the text sequence and a phoneme sequence feature vector corresponding to the phoneme sequence.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the steps of the intention recognition method as described in any one of claims 1 to 5 are implemented.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the intention recognition method as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Polyphone recognition method and device, electronic equipment and storage medium
CN112348073A
Chinese text intention recognition method based on Bert and full-connection neural network fusion
CN113239690A
Data processing method and device
CN113515586A