A method and device for identifying intention and related equipment

By acquiring phoneme sequences and using the cross attention mechanism to fuse the features of text sequences and phoneme sequences, and combining with a fully connected neural network for intention prediction, the problem of low accuracy of oral intention recognition in the prior art is solved, and a higher recognition accuracy is achieved.

CN114187894BActive Publication Date: 2025-05-23JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111506575.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-05-23
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

The prior art has low accuracy in oral intention recognition tasks and is difficult to effectively improve.

Method used

By obtaining phoneme sequences and feature extraction and fusion of text sequences and phoneme sequences based on the cross attention mechanism, sentence feature vectors are obtained, and intention prediction is made by combining fully connected neural networks.

Benefits of technology

It improves the accuracy of oral intention recognition and can reconstruct the sentence feature vector representations corresponding to the real text when facing pronunciation and tone errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187894B_ABST
    Figure CN114187894B_ABST
Patent Text Reader

Abstract

The present disclosure provides an intention recognition method, apparatus and related equipment, the method comprising: obtaining a phoneme sequence corresponding to the audio to be detected, and determining a text sequence corresponding to the audio to be detected based on the phoneme sequence; extracting and fusing the text sequence and the phoneme sequence based on a cross-attention mechanism to obtain a sentence feature vector corresponding to the audio to be detected; and obtaining a prediction result for characterizing the intention of the audio to be detected based on the sentence feature vector. The method can improve the accuracy of spoken intention recognition.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A method for identifying intent, It is characterized in that The method includes: Acquire a phoneme sequence corresponding to the audio to be detected, and determine a text sequence corresponding to the audio to be detected based on the phoneme sequence; Based on the cross attention mechanism, feature extraction and fusion are performed on the text sequence and the phoneme sequence to obtain a sentence feature vector corresponding to the audio to be detected; Based on the sentence feature vector, a prediction result for characterizing the intention of the audio to be detected is obtained, wherein the step of extracting and fusing features of the text sequence and the phoneme sequence based on a cross attention mechanism to obtain a sentence feature vector corresponding to the audio to be detected includes: Encoding the text sequence and the phoneme sequence respectively to obtain a text sequence feature vector corresponding to the text sequence and a phoneme sequence feature vector corresponding to the phoneme sequence; The text sequence feature vector and the phoneme sequence feature vector are concatenated according to the text sequence weight parameter and the phoneme sequence weight parameter to obtain the sentence feature vector, wherein the text sequence weight parameter and the phoneme sequence weight parameter are obtained by training a plurality of pre-stored audio samples.

2. The method for identifying intention according to claim 1, It is characterized in that The step of obtaining a prediction result for characterizing the intention of the audio to be detected based on the sentence feature vector specifically includes: The sentence feature vector is input into a fully connected neural network for classification prediction to obtain the prediction result.

3. The method for identifying intention according to claim 1, It is characterized in that The step of encoding the text sequence and the phoneme sequence to obtain a text sequence feature vector corresponding to the text sequence and a phoneme sequence feature vector corresponding to the phoneme sequence includes: A bidirectional neural network is used to encode the text sequence and the phoneme sequence to obtain a text sequence feature vector corresponding to the text sequence and a phoneme sequence feature vector corresponding to the phoneme sequence.

4. The method for identifying intention according to claim 1, It is characterized in that The text sequence weight parameter and the phoneme sequence weight parameter are obtained by training a plurality of pre-stored audio samples, specifically including: Encoding a sample text sequence and a sample phoneme sequence to obtain a sample text sequence feature vector corresponding to the sample text sequence and a sample phoneme sequence feature vector corresponding to the sample phoneme sequence; wherein the sample text sequence and the sample phoneme sequence are obtained by performing speech recognition on the audio sample; Calculating the cosine similarity between the embedding vector of each text in the sample text sequence feature vector and the sample phoneme sequence feature vector to obtain the weight of each text, and normalizing the weights corresponding to all texts through a normalized exponential function to obtain the text sequence weight parameter; The phoneme sequence weight parameter is obtained according to the text sequence weight parameter.

5. The method for identifying intention according to claim 4, It is characterized in that The text sequence weight parameter and the phoneme sequence weight parameter are adjusted in the following manner, specifically including: determining a true text transcribed from said audio sample listened to by a human; Based on a cross entropy loss function, calculating the difference between a prediction result representing the intent of the audio sample and the real text; Based on the difference, the text sequence weight parameter and the phoneme sequence weight parameter are adjusted.

6. An intention recognition device, It is characterized in that The device is used to implement the intention recognition method described in any one of claims 1 to 5, and the device includes: A speech recognition module, used to obtain a phoneme sequence corresponding to the audio to be detected, and determine a text sequence corresponding to the audio to be detected based on the phoneme sequence; A cross attention module, used for extracting and fusing features of the text sequence and the phoneme sequence based on a cross attention mechanism, and obtaining a sentence feature vector corresponding to the audio to be detected; A fully connected neural network classification module is used to obtain a prediction result for characterizing the intention of the audio to be detected based on the sentence feature vector.

7. The intention recognition device according to claim 6, It is characterized in that The fully connected neural network classification module includes: The sentence feature vector is input into a fully connected neural network for classification prediction to obtain the prediction result.

8. The intention recognition device according to claim 7, It is characterized in that The cross attention module specifically includes: An encoding unit encodes the text sequence and the phoneme sequence respectively to obtain a text sequence feature vector corresponding to the text sequence and a phoneme sequence feature vector corresponding to the phoneme sequence; A concatenation unit concatenates the text sequence feature vector and the phoneme sequence feature vector according to the text sequence weight parameter and the phoneme sequence weight parameter to obtain the sentence feature vector.

9. The intention recognition device according to claim 8, It is characterized in that The encoding unit specifically includes: A bidirectional neural network is used to encode the text sequence and the phoneme sequence to obtain a text sequence feature vector corresponding to the text sequence and a phoneme sequence feature vector corresponding to the phoneme sequence.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the steps of the intention recognition method as described in any one of claims 1 to 5 are implemented.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the intention recognition method as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Polyphone recognition method and device, electronic equipment and storage medium

    CN112348073A

  • Chinese text intention recognition method based on Bert and full-connection neural network fusion

    CN113239690A

  • Data processing method and device

    CN113515586A