Role identification method, device, computer equipment and storage medium

By correcting and extracting features from the target audio text and combining it with emotion recognition results, the problem of low accuracy in role recognition in voice interaction is solved and the accuracy of role recognition is improved.

CN115376558BActive Publication Date: 2025-09-26CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211004872.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-22
Publication Date
2025-09-26
Estimated Expiration
2042-08-22

AI Technical Summary

Technical Problem

The accuracy of role recognition in voice interaction in the existing technology is low, especially when multiple speakers speak alternately, it is difficult to accurately identify the identity of the speaker.

Method used

By performing text detection on the target audio text, correcting the erroneous parts, extracting audio and text feature vectors, and combining the emotion recognition results for auxiliary recognition of role categories.

Benefits of technology

The accuracy of role recognition is improved, and by correcting erroneous audio and text data, an accurate data basis is provided for subsequent steps. The judgment ability of role recognition is strengthened by combining text and audio feature vectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115376558B_ABST
    Figure CN115376558B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for character identification, which includes obtaining a target audio text, performing text detection on the target audio text, and obtaining a text detection result; performing correction processing on the target audio text corresponding to the detection failure result to obtain a corrected audio text; obtaining corrected audio data corresponding to the corrected audio text, performing voiceprint feature extraction on the corrected audio data, and obtaining an audio voiceprint feature; determining a text feature vector corresponding to the corrected audio text, and determining an audio feature vector corresponding to the audio voiceprint feature; determining an emotion recognition result corresponding to the corrected audio text based on the audio feature vector and the text feature vector, and determining a character category corresponding to the corrected audio text based on the emotion recognition result, the audio feature vector, and the text feature vector. In this way, the present invention assists in identifying the character category corresponding to the corrected audio text through the emotion recognition result, thereby improving the accuracy of character identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice interaction technology, and in particular to a role recognition method, device, computer equipment and storage medium. Background Art

[0002] In the application of intelligent voice, the scenario of speaker identification in voice interaction is very typical and common, such as speaker identification in intelligent conferences and customer service / customer identification in intelligent customer service.

[0003] Existing technologies often use voiceprint recognition models built on voice data to identify speakers. In scenarios like smart conferences or customer service calls, where multiple speakers are speaking alternately, the rapid switching between voices can lead to lower speaker identification accuracy. Summary of the Invention

[0004] Embodiments of the present invention provide a role identification method, apparatus, computer equipment, and storage medium to solve the problem of low accuracy in role identification of voice data in the prior art.

[0005] A role identification method, comprising:

[0006] Acquire a target audio text, perform text detection on the target audio text, and obtain a text detection result; the text detection result includes a detection failure result; the detection failure result indicates that there is an error in the target audio text;

[0007] Correcting the target audio text corresponding to the detection failure result to obtain a corrected audio text;

[0008] Acquire corrected audio data corresponding to the corrected audio text, perform voiceprint feature extraction on the corrected audio data, and obtain audio voiceprint features;

[0009] Determining a text feature vector corresponding to the corrected audio text, and determining an audio feature vector corresponding to the audio voiceprint feature;

[0010] Based on the audio feature vector and the text feature vector, an emotion recognition result corresponding to the corrected audio text is determined, and based on the emotion recognition result, the audio feature vector and the text feature vector, a character category corresponding to the corrected audio text is determined.

[0011] A role recognition device, comprising:

[0012] A text detection module is used to obtain a target audio text, perform text detection on the target audio text, and obtain a text detection result; the text detection result includes a detection failure result; the detection failure result indicates that there is an error in the target audio text;

[0013] A text correction module is used to correct the target audio text corresponding to the detection failure result to obtain a corrected audio text;

[0014] A feature extraction module is used to obtain the corrected audio data corresponding to the corrected audio text, perform voiceprint feature extraction on the corrected audio data, and obtain audio voiceprint features;

[0015] A feature vector module, configured to determine a text feature vector corresponding to the corrected audio text, and an audio feature vector corresponding to the audio voiceprint feature;

[0016] The role recognition module is used to determine the emotion recognition result corresponding to the corrected audio text based on the audio feature vector and the text feature vector, and to determine the role category corresponding to the corrected audio text based on the emotion recognition result, the audio feature vector and the text feature vector.

[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned role identification method when executing the computer program.

[0018] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for identifying a role is implemented.

[0019] The present invention provides a method, apparatus, computer device, and storage medium for character recognition. The method performs text detection on a target audio text, and when it is determined that the target audio text contains errors, the target audio text containing errors is corrected. This provides an accurate data basis for character recognition in subsequent steps, thereby improving the accuracy of character recognition. Furthermore, the present invention combines text feature vectors and audio feature vectors to determine emotion recognition results, and uses the emotion recognition results to assist in identifying the character category corresponding to the corrected audio text. In this way, the ability to identify and judge characters can be enhanced by combining the emotional states of different character categories, further improving the accuracy of character recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0021] Figure 1 Schematic diagram of an application environment of a role identification method according to an embodiment of the present invention;

[0022] Figure 2 is a flow chart of a role identification method according to an embodiment of the present invention;

[0023] Figure 3 is a flow chart of step S50 in the role identification method in one embodiment of the present invention;

[0024] Figure 4 is another flow chart of step S50 in the role identification method in one embodiment of the present invention;

[0025] Figure 5 is a functional block diagram of a role recognition device according to an embodiment of the present invention;

[0026] Figure 6 is a schematic diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] The role identification method provided by the embodiment of the present invention can be applied as follows: Figure 1 Specifically, the role identification method is applied in a role identification device, which includes Figure 1The client and server shown communicate with each other over a network to solve the problem of low accuracy in character recognition of voice data in the prior art. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The client, also known as the user end, refers to a program that corresponds to the server and provides local services to customers. The client can be installed on, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices.

[0029] In one embodiment, if Figure 2 As shown, a role identification method is provided, which is applied in Figure 1 The server in the example is used as an example, and the steps are as follows:

[0030] S10: Acquire a target audio text, perform text detection on the target audio text, and obtain a text detection result; the text detection result includes a detection failure result; the detection failure result indicates that there is an error in the target audio text.

[0031] Understandably, the target audio text can be the audio data that the user recognizes through the voice recognition software of the mobile terminal (Lightning text-to-speech software or Fengyun text-to-speech software, etc.) and then sends to the server, or the audio data that the user recognizes through the automatic speech recognition technology (ASR, Automatic Speech Recognition) on the client and then sends to the server. Text detection is performed on the target audio text, that is, whether there are words or sentences with recognition errors in the target audio text is detected, so as to obtain the text detection result. The text detection result includes a successful detection result and a failed detection result. The successful detection result is used to indicate that the content of the target audio text is completely correct, and the failed detection result is used to indicate that there are errors in the content of the target audio text.

[0032] S20: Correcting the target audio text corresponding to the detection failure result to obtain a corrected audio text.

[0033] It can be understood that the corrected audio text is the text after the wrong words or wrong sentences in the target audio text corresponding to the detection failure result are corrected.

[0034] Specifically, after obtaining the text detection results, the detection failure results are screened out from the text detection results, and the target audio text corresponding to the detection failure results is determined. The target audio text corresponding to the detection failure results is corrected, that is, the wrong words or wrong sentences in the target audio text corresponding to the detection failure results are replaced. The target audio text corresponding to the detection failure results is predicted based on the context information to obtain at least one predicted replacement word or predicted replacement sentence. All predicted replacement words or all predicted replacement sentences are replaced to the positions corresponding to the wrong words or wrong sentences in the target audio text corresponding to the detection failure results, that is, the wrong words or wrong sentences are replaced with predicted replacement words or predicted replacement sentences, thereby obtaining a corrected audio text.

[0035] S30: Obtain corrected audio data corresponding to the corrected audio text, perform voiceprint feature extraction on the corrected audio data, and obtain audio voiceprint features.

[0036] Understandably, the corrected audio data can be recorded by the user through the recording software of the mobile terminal (such as Recording Wizard or Jinzhou computer recording software) and then sent to the server, or it can be recorded by the user through other recording devices (such as a voice recorder) and then sent to the server through the client. For example, when an insurance company salesperson explains a certain insurance, he can use the mobile phone recording software to record the conversation with the customer and send the recording to the server. Alternatively, the conversation with the customer can be recorded by a voice recorder and sent to the server through the client, so that the corrected audio data can be obtained. Voiceprint features of the corrected audio data can be extracted by MFCC feature extraction technology (Mel-scale Frequency Cepstral Coefficients), or by using a pre-trained voiceprint feature extraction model or a convolutional neural network model to extract the voiceprint features of the corrected audio data, thereby obtaining the audio voiceprint features corresponding to the corrected audio text.

[0037] S40: Determine a text feature vector corresponding to the corrected audio text, and determine an audio feature vector corresponding to the audio voiceprint feature.

[0038] It can be understood that the text feature vector is obtained by converting the corrected audio text into a vector, and is used to characterize the text features of the corrected audio text. The audio feature vector is obtained by converting the audio voiceprint feature into a vector, and is used to characterize the audio voiceprint feature of the corrected audio text.

[0039] Specifically, after obtaining the audio voiceprint features, the corrected audio text is segmented to obtain segmentation results. Part-of-speech tagging is performed on all segmentation results to obtain segmentation results for a specified part of speech (e.g., adjectives, nouns, and verbs). The segmentation results for the specified part of speech are then weighted to obtain weights for all segmentation results for the specified part of speech. The weights are sorted from highest to lowest, and the segmentation results corresponding to a preset number (e.g., two, three, or four) of the weighted values ​​are selected as keywords. Alternatively, a term frequency-inverse text frequency method can be used to select the segmentation results corresponding to a preset number (e.g., two, three, or four) of the TF-IDF values ​​as keywords. All keywords are then encoded to obtain a text feature vector corresponding to the corrected audio text. Furthermore, the audio voiceprint features are segmented to obtain audio signal segments. Each audio signal segment is vectorized, i.e., each segment is encoded and represented numerically, and the numerical values ​​are then represented as vectors to obtain an audio vector corresponding to each segment. All audio vectors are fused according to the preset weights to obtain a fused audio vector. The fused audio vector is then determined as an audio feature vector, thereby obtaining an audio feature vector corresponding to the audio voiceprint feature.

[0040] S50: Determine an emotion recognition result corresponding to the corrected audio text based on the audio feature vector and the text feature vector, and determine a character category corresponding to the corrected audio text based on the emotion recognition result, the audio feature vector and the text feature vector.

[0041] Understandably, the emotion recognition result is the corrected audio text and the emotions contained in the corrected audio data. The role category is the role in the corrected audio text. For example, in a sales scenario, the role can include sales and customer, in a customer service scenario, the role can include customer service and client, etc.

[0042] Specifically, after obtaining the text feature vector and the audio feature vector, the audio feature vector and the text feature vector are fused through an attention mechanism to obtain a fused feature vector, and emotion recognition is performed on the corrected audio text based on the fused feature vector to obtain an emotion recognition result corresponding to the corrected audio text. The character recognition of the corrected audio text is performed using the emotion recognition result, the audio feature vector, and the text feature vector. That is, the character in the corrected audio text is identified using the audio feature vector and the text feature vector, and auxiliary recognition is performed based on the emotion recognition result, thereby obtaining the character category corresponding to the corrected audio text.

[0043] In an embodiment of the present invention, a method for character recognition is provided. The method performs text detection on a target audio text, and when it is determined that an error exists in the target audio text, the target audio text with an error is corrected. This provides an accurate data basis for character recognition in subsequent steps, thereby improving the accuracy of character recognition. Furthermore, the present invention combines text feature vectors and audio feature vectors to determine the emotion recognition result, and uses the emotion recognition result to assist in identifying the character category corresponding to the corrected audio text. In this way, the emotional state of different character categories can be combined to enhance the ability of character recognition judgment, further improving the accuracy of character recognition.

[0044] In one embodiment, in step S10, text detection is performed on the target audio text to obtain a text detection result, including:

[0045] S101: Perform text detection on the target audio text to obtain a text detection value corresponding to the target audio text.

[0046] It can be understood that the text detection value is obtained by detecting incorrect words or incorrect sentences in the target audio text.

[0047] Specifically, after obtaining the target audio text, the scene information corresponding to the target audio text is obtained, and the content of the target audio text is judged based on the scene information to determine whether the content in the target audio text belongs to the scene information. If it is determined that the content in the target audio text does not belong to the scene information, the content is deleted. If it is determined that the content in the target audio text belongs to the scene information, the target audio text is input into a preset detection model, and text detection is performed on the target audio text using the preset detection model, that is, the target audio text is first segmented to obtain the segmentation results corresponding to the target audio text. All segmentation results are then predicted using the preset detection model based on context information to obtain prediction results. A test is performed to determine whether the prediction results and segmentation results are the same. If the prediction results and segmentation results are the same, the segmentation result is determined to be correct, and the detection value is increased by 1. If the prediction results and segmentation results are different, the segmentation result is determined to be incorrect, the detection value is decreased by 1, and the word corresponding to the incorrect segmentation result is determined as the word to be corrected. The word to be corrected is then marked, that is, the word to be corrected and its position in the target audio text are recorded. The detection values ​​corresponding to all word segmentation results are calculated to determine the text detection value corresponding to the target audio text.

[0048] S102: Obtain a preset threshold, and determine the text detection result according to the preset threshold and the text detection value.

[0049] It can be understood that the preset threshold is set in advance to determine whether there are errors in the target audio text.

[0050] Specifically, after obtaining the text detection value, a preset threshold is retrieved from the server or a preset threshold is obtained from a third-party platform, and the preset threshold and the text detection value are compared. When the text detection value is greater than or equal to the preset threshold, the text detection result corresponding to the text detection value greater than or equal to the preset threshold is determined as a successful detection result, and the successful detection result is used to characterize that there are no errors in the target audio text, that is, the target audio text is all correct. When the text detection value is less than the preset threshold, the text detection result corresponding to the text detection value less than the preset threshold is determined as a failed detection result, and the failed detection result is used to characterize that there are errors in the target audio text, that is, the target audio text needs to be corrected.

[0051] The embodiment of the present invention performs text detection on the target audio text to determine a text detection value corresponding to the target audio text, and compares the text detection value with a preset threshold to determine whether there is an error in the target audio text.

[0052] In one embodiment, in step S20, correcting the target audio text corresponding to the detection failure result to obtain a corrected audio text includes:

[0053] S201: Determine the target audio text corresponding to the detection failure result as an erroneous audio text, and determine the words to be corrected contained in the erroneous audio text.

[0054] S202: Masking the words to be corrected in the erroneous audio text to obtain masked text to be corrected.

[0055] Understandably, the erroneous audio text is the target audio text corresponding to the detection failure result. The word to be corrected is the word that was incorrectly recognized in the target audio text. Masking refers to using special characters to cover up the word to be corrected. The masked text to be corrected is the text that needs to be corrected and predicted for the masked word to be corrected.

[0056] Specifically, after obtaining the detection failure result, the target audio text corresponding to the detection failure result is determined as the erroneous audio text, and the erroneous audio text is re-detected through the preset detection model, and the erroneous words in the erroneous audio text are marked, that is, the erroneous words and the corresponding positions of the erroneous words in the erroneous audio text are recorded, so that all the words to be corrected contained in the erroneous audio text can be determined. All the words to be corrected in the erroneous audio text are masked, that is, all the words to be corrected in the erroneous audio text are covered with special characters. After masking all the words to be corrected in turn, the erroneous audio text after masking all the words to be corrected is determined as the masked text to be corrected.

[0057] S203: Inputting the masked text to be corrected into a preset language model, performing correction prediction on the masked text to be corrected by the preset language model, and obtaining a predicted replacement word corresponding to the word to be corrected.

[0058] It is understandable that the preset language model is a model set in advance for correcting the masked text to be corrected. The predicted replacement word is a word predicted by the preset language model for the masked word to be corrected.

[0059] Specifically, after obtaining the masked text to be corrected, the masked text to be corrected is input into the preset language model, and the masked text to be corrected is converted into code by the encoding module in the preset language model to obtain the corresponding corrected coding vector. The corrected coding vector is linearly mapped, and the corrected coding vector is matched with each coding vector in the preset character dictionary. For example, the corrected coding vector can be matched by determining the Euclidean distance between the corrected coding vector and each coding vector in the preset character dictionary, and then the matching score of each word to be replaced in the preset character dictionary can be determined based on the Euclidean distance. Each matching score is normalized to obtain a matching probability corresponding to each matching score, and then the word to be replaced with the largest matching probability among the matching probabilities and greater than the preset matching threshold is recorded as the predicted replacement word. In this way, the predicted replacement words corresponding to all words to be corrected are obtained in the above manner.

[0060] Furthermore, if the maximum matching probability is less than a preset matching threshold, the word corresponding to the maximum matching probability is sent to the third-party platform. The staff then retrieves the word corresponding to the maximum matching probability from the third-party platform. The staff then uploads the correct word to the third-party platform, which then sends the feedback to the server and replaces the correct word in the corresponding position.

[0061] S204: Replace the word to be corrected with the predicted replacement word, and record the masked text to be corrected after the replacement as the corrected audio text.

[0062] Specifically, after obtaining the predicted replacement word, the masked word to be corrected is replaced with the corresponding predicted replacement word. That is, the masked word to be corrected is replaced with the predicted replacement word predicted by the preset language model to correct the masked word to be corrected. In this way, all masked words to be corrected are replaced with the corresponding predicted replacement words, and the masked text to be corrected after the replacement is recorded as the corrected audio text.

[0063] The embodiment of the present invention facilitates the subsequent correction of the erroneous audio text by determining the words to be corrected contained in the erroneous audio text, and speeds up the subsequent masking of the words to be corrected. A preset language model is used to predict the correction of the masked text to be corrected, and all predicted replacement words are replaced with the corresponding words to be corrected. This provides an accurate data basis for character recognition in the subsequent steps, thereby improving the accuracy of character recognition.

[0064] In one embodiment, in step S30, that is, obtaining the corrected audio data corresponding to the corrected audio text, performing voiceprint feature extraction on the corrected audio data to obtain the audio voiceprint feature, includes:

[0065] S301, preprocessing the corrected audio data corresponding to the corrected audio text to obtain target voice data.

[0066] It can be understood that the target speech data is obtained by performing pre-emphasis, framing and windowing processing on the corrected audio data.

[0067] Specifically, after obtaining the corrected audio data, the corrected audio data is first pre-emphasized, that is, to enhance the high-frequency part of the corrected audio data so that the corrected audio data can use the same signal-to-noise ratio to calculate the spectrum in the entire frequency band from low frequency to high frequency. The corrected audio data is then framed, that is, the corrected audio data is divided into fixed time periods (such as 25 milliseconds) to obtain multiple frame units. In order to avoid excessive changes in adjacent frame units, there is an overlapping area between two adjacent frame units. Each frame unit is multiplied by a window function so that the discontinuous audio signal after division becomes continuous and exhibits the characteristics of a periodic function, and the left and right ends of each frame unit are continuous to obtain a continuous time window signal, thereby obtaining the target speech data.

[0068] S302: Perform voiceprint feature extraction on the target speech data to obtain the audio voiceprint feature corresponding to the corrected audio data.

[0069] It can be understood that the audio voiceprint feature is the feature of the target voice data. The voiceprint is the sound wave spectrum of the carrier's speech information displayed by an electroacoustic instrument.

[0070] Specifically, after obtaining the target speech data, a fast Fourier transform (FFT) is performed on the target speech data. Specifically, a FFT is performed on the continuous time window signal, transforming the signal distribution in the time domain into an energy distribution in the frequency domain. This yields the spectrum corresponding to the target speech data. The spectrum is then input into a Mel filter, where it is transformed to obtain a Mel spectrum. This converts the linear natural spectrum into a Mel spectrum that reflects human auditory characteristics. The logarithm of the Mel spectrum is obtained to obtain its logarithmic energy. This logarithmic energy is then inversely transformed using a discrete cosine transform (DCT). The second to thirteenth coefficients after the DCT are taken as Mel-frequency cepstral coefficients, which are then determined as audio voiceprint features.

[0071] The embodiment of the present invention preprocesses the target audio text and extracts the voiceprint features of the target speech data, thereby determining the audio feature vector, improving the accuracy of emotion recognition, and improving the accuracy of role recognition.

[0072] In one embodiment, step S40, i.e., determining the text feature vector corresponding to the corrected audio text and determining the audio feature vector corresponding to the audio voiceprint feature, includes:

[0073] S401: Perform word segmentation processing on the corrected audio text to obtain audio words corresponding to the corrected audio text.

[0074] It can be understood that the audio words are the result of word segmentation of the corrected audio text.

[0075] Specifically, after obtaining the audio voiceprint features, the corrected audio text is segmented using the Chinese word segmentation algorithm, and the corrected audio text is segmented using the full segmentation path selection method based on the connection between the contextual features to obtain the audio words corresponding to the corrected audio text. The full segmentation path selection segmentation process is to list all possible segmentation results, select the best segmentation path from them, and form a directed acyclic graph of all segmentation results. The segmentation results can be used as nodes, and the edges between words are weighted. The path with the smallest weight is the final result. For example, the word frequency can be used as the weight, and a path with the largest total word frequency can be considered the best path. In this way, the segmentation process can be completed through the best path to obtain the audio words corresponding to the corrected audio text. Among them, the segmentation results are the audio words obtained after segmentation, and the directed acyclic graph is a graph with no loops and a direction.

[0076] S402, performing vector conversion on the audio words to obtain word vectors corresponding to the audio words, and determining the text feature vector corresponding to the corrected audio text based on all the word vectors.

[0077] Specifically, after obtaining the audio words, the preset words in the preset word bag model are matched with the audio words for similarity. When the preset words and the audio words are matched successfully, the vector corresponding to the successfully matched preset words is determined as the word vector corresponding to the audio words. In this way, the word vectors corresponding to all audio words are determined in sequence. When all preset words and audio words fail to match, the audio words are sent to a third-party platform, and the staff performs vector conversion on the audio words to obtain the word vector corresponding to the audio words, and feeds it back to the word bag model. Furthermore, the TF-IDF values ​​of all audio words are calculated by the word frequency-inverse document frequency method, and all TF-IDF values ​​are sorted from large to small, and a preset number (such as 2, 3 or 4) of audio words with the largest TF-IDF values ​​are determined as keywords. The word vectors corresponding to this preset number (such as 2, 3 or 4) of audio words are spliced ​​to obtain the text feature vector corresponding to the corrected audio text. Among them, term frequency (TF) = the number of times a word appears in an article / the total number of times in the article, and inverse document frequency (IDF) = log (total number of documents / number of documents containing the word + 1).

[0078] S403: Perform segmentation processing on the audio voiceprint feature to obtain a segmented audio signal, and perform vector conversion on the segmented audio signal to obtain the audio feature vector corresponding to the audio voiceprint feature.

[0079] Specifically, after obtaining the text feature vector, the audio voiceprint feature is signal segmented, that is, the audio voiceprint feature is divided into segments of segmented audio signals with a fixed duration (e.g., 30 milliseconds). Each segment of the segmented audio signal is vectorized, that is, each segment of the segmented audio signal is encoded and represented by a number, and then the number is represented by a vector to obtain the audio vector corresponding to each segment of the segmented audio signal. All audio vectors are fused according to preset weights to obtain a fused audio vector. The fused audio vector is then determined as an audio feature vector, thus obtaining the audio feature vector corresponding to the audio voiceprint feature.

[0080] This embodiment of the present invention determines the text feature vector by calculating the TF-IDF values ​​of all audio words and concatenating the word vectors with the top three TF-IDF values. All audio vectors are then fused according to their weights to determine the audio feature vector, improving the accuracy of emotion recognition for corrected audio text.

[0081] In one embodiment, if Figure 3 As shown, in step S50, that is, based on the audio feature vector and the text feature vector, determining the emotion recognition result corresponding to the corrected audio text includes:

[0082] S501 , obtaining a preset emotion set, and clustering all the target emotions in the preset emotion set to obtain at least one emotion cluster group; one emotion cluster group includes a plurality of target emotions.

[0083] Understandably, a preset emotion set is constructed by acquiring all target emotions. An emotion cluster group is a collection of emotions with the same meaning, such as the emotion cluster group of "anger" includes anger, annoyance, and boredom, and the emotion cluster group of "joy" includes ecstasy, happiness, and joy.

[0084] Specifically, after obtaining the audio feature vector and the text feature vector, a preset emotion set is obtained, which includes multiple target emotions. All target emotions are clustered using the k-means clustering algorithm, and k cluster centers are randomly selected from the preset emotion set. The distance between each target emotion and each cluster center is calculated using Euclidean distance or cosine similarity, and each target emotion is assigned to the cluster center closest to it. The cluster center and the assigned target emotion represent an emotion cluster group. Each time a target emotion is assigned, the cluster center is recalculated based on the existing target emotions in the emotion cluster group. This process will be repeated until a certain termination condition is met. The termination condition can be that no (or a minimum number) of target emotions are reassigned to different clusters, that no (or a minimum number) of cluster centers change again, or that the sum of squared errors is locally minimized. In this way, an emotion cluster group can be obtained.

[0085] S502: Perform vector conversion on all the emotion cluster groups to obtain emotion cluster vectors.

[0086] S503: Fusing the audio feature vector and the text feature vector to obtain a fused feature vector.

[0087] It can be understood that the emotion cluster vector is obtained by vectorizing the emotion cluster group, such as vectorizing the emotion cluster group of "joy" to obtain the corresponding emotion cluster vector. The fusion feature vector is a vector containing text features and audio voiceprint features.

[0088] Specifically, after obtaining the emotion cluster groups, vector conversion is performed on all emotion cluster groups, and all emotion cluster groups are encoded through the encoding layer in the preset conversion model to obtain digital codes corresponding to all emotion cluster groups. All digital codes are then vectorized through the vector layer in the preset conversion model to obtain digital vectors corresponding to each emotion cluster group. Normalization is performed through the normalization layer in the preset conversion model to obtain emotion cluster vectors. Among them, all target emotion vectors are determined through the preset bag-of-words model. The specific process is the same as the above step S402 and will not be repeated here.

[0089] Furthermore, the audio feature vector and the text feature vector are fused through the attention mechanism, that is, the audio feature vector and the text feature vector can be spliced ​​through the attention mechanism to obtain a fused feature vector, and the audio feature vector and the text feature vector can also be fused through the preset weights in the attention mechanism to obtain a fused feature vector.

[0090] S504 , performing similarity matching on the fused feature vector and the emotion clustering vector to obtain a vector matching result, and determining an emotion recognition result corresponding to the corrected audio text based on the vector matching result.

[0091] It can be understood that the vector matching result is used to represent the similarity between the fused feature vector and the emotion clustering vector. The emotion recognition result is the emotion contained in the corrected audio text.

[0092] Specifically, after obtaining the fused feature vector, similarity matching is performed between the fused feature vector and all emotion cluster vectors. This involves calculating the similarity between the fused feature vector and all emotion cluster vectors using Euclidean distance or cosine similarity, resulting in the Euclidean distance or cosine similarity between the fused feature vector and all emotion cluster vectors. All calculated results are compared, and the emotion cluster vector with the highest similarity—that is, the one with the smallest Euclidean distance or the highest cosine similarity between the fused feature vector and all emotion cluster vectors—is recorded as the vector matching result. This determines the emotion cluster group corresponding to the fused feature vector.

[0093] Furthermore, similarity matching is performed between the fused feature vector and the target emotion vectors corresponding to all target emotions in the emotion cluster group, that is, the similarity between the fused feature vector and all target emotion vectors is calculated by Euclidean distance or cosine similarity, and the Euclidean distance or cosine similarity between the fused feature vector and all target emotion vectors is obtained. All calculation results are compared with a preset similarity threshold, and all target emotions corresponding to the target emotion vectors corresponding to the Euclidean distance or cosine similarity greater than or equal to the preset similarity threshold are determined as emotion recognition results, so that all emotion recognition results corresponding to the corrected audio text can be obtained. Among them, the preset similarity threshold is set in advance for judging the similarity between the fused feature vector and the target emotion vector.

[0094] By clustering target emotions, the present invention identifies emotion clusters representing the same meaning, and thus determines emotion cluster vectors. The audio feature vectors and text feature vectors are combined to determine the emotion recognition result, further improving the accuracy of character recognition.

[0095] In one embodiment, if Figure 4As shown, in step S50, that is, based on the emotion recognition result, the audio feature vector and the text feature vector, determining the character category corresponding to the target audio text includes:

[0096] S505: Determine scene information corresponding to the corrected audio data, and obtain a role recognition model corresponding to the scene information.

[0097] It can be understood that scenario information includes one or more of conversation scene characteristics and conversation character characteristics. Common scenarios for character identification include customer service calls, conference calls, interrogations, and fraudulent harassment calls. The vocabulary, sentences, and templates used in conversations vary in different scenarios. Scenario information can be information about the templates used in the conversation. Each scenario information corresponds to a character identification model.

[0098] Specifically, after obtaining the emotion recognition result, the corrected audio data is matched against all stored preset audio data for similarity. If a match is successful, the preset scene information corresponding to the preset audio data with the highest similarity is determined as the scene information corresponding to the corrected audio data. The scene information corresponding to the corrected audio data is matched against all stored preset scene information for similarity. If a match is successful, the preset recognition model corresponding to the preset scene information with the highest similarity is determined as the character recognition model that matches the scene information corresponding to the corrected audio data.

[0099] S506 , performing role recognition on the text feature vector and the audio feature vector using the role recognition model to obtain a role recognition result.

[0100] It can be understood that the role recognition model is obtained by training a preset model with a large amount of scene information, and is used to recognize roles in the corresponding scene information.

[0101] Specifically, the text feature vector and the audio feature vector are input into the role recognition model, and the role recognition model is used to perform role recognition on the text feature vector and the audio feature vector. That is, the convolution layer in the role recognition model performs convolution processing on the text feature vector and the audio feature vector to obtain a text convolution vector and an audio convolution vector. The text convolution vector and the audio convolution vector are pooled by the pooling layer, that is, the text convolution vector and the audio convolution vector are dimensionally compressed to obtain a text pooling vector and an audio pooling vector. The role probability of the text pooling vector and the audio pooling vector is calculated by a fully connected layer and a normalized layer to determine the role recognition result.

[0102] S507: Determine the role category corresponding to the corrected audio text according to the emotion recognition result and the role recognition result.

[0103] Understandably, the role category is to correct the roles in the audio text, such as sales and customer in a sales scenario, customer service and client in a customer service scenario, etc.

[0104] Specifically, after obtaining the role recognition result, the role recognition result is assisted by the emotion recognition result corresponding to the corrected audio text, that is, the role in the corrected audio text is accurately identified by correcting the emotion in the audio text, thereby determining the role category in the corrected audio text. For example, in a sales scenario, when explaining a product to a customer, the salesperson will maintain a leisurely mood, while the customer will maintain a questioning mood. When the customer decides to buy the product, the customer will maintain an expectant mood, while the salesperson will maintain a happy mood. In this way, the text feature vector and the audio feature vector can be assisted by the emotion recognition result to determine the role category.

[0105] The present invention uses a character recognition model to identify characters using text and audio feature vectors, thereby confirming the character recognition results. Emotion recognition is used to supplement the character recognition results, thereby enhancing the character recognition capabilities by combining the emotional states of different character categories and further improving the accuracy of character recognition.

[0106] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0107] In one embodiment, a role identification device is provided, which corresponds one-to-one with the role identification method in the above embodiment. Figure 5 As shown, the role recognition device includes a text detection module 11, a text correction module 12, a feature extraction module 13, a feature vector module 14 and a role recognition module 15. The functional modules are described in detail as follows:

[0108] The text detection module 11 is used to obtain a target audio text, perform text detection on the target audio text, and obtain a text detection result; the text detection result includes a detection failure result; the detection failure result indicates that there is an error in the target audio text;

[0109] A text correction module 12 is used to correct the target audio text corresponding to the detection failure result to obtain a corrected audio text;

[0110] A feature extraction module 13 is configured to obtain corrected audio data corresponding to the corrected audio text, perform voiceprint feature extraction on the corrected audio data, and obtain audio voiceprint features;

[0111] A feature vector module 14 is configured to determine a text feature vector corresponding to the corrected audio text and an audio feature vector corresponding to the audio voiceprint feature;

[0112] The role recognition module 15 is used to determine the emotion recognition result corresponding to the corrected audio text based on the audio feature vector and the text feature vector, and determine the role category corresponding to the corrected audio text based on the emotion recognition result, the audio feature vector and the text feature vector.

[0113] In one embodiment, the text detection module 11 includes:

[0114] A text detection unit, configured to perform text detection on the target audio text to obtain a text detection value corresponding to the target audio text;

[0115] The result determination unit is configured to obtain a preset threshold value and determine the text detection result according to the preset threshold value and the text detection value.

[0116] In one embodiment, the text correction module 12 includes:

[0117] A determination unit, configured to determine the target audio text corresponding to the detection failure result as an erroneous audio text, and determine the words to be corrected contained in the erroneous audio text

[0118] a masking unit, configured to perform masking processing on the words to be corrected in the erroneous audio text to obtain masked text to be corrected;

[0119] a correction prediction unit, configured to input the masked text to be corrected into a preset language model, perform correction prediction on the masked text to be corrected using the preset language model, and obtain a predicted replacement word corresponding to the word to be corrected;

[0120] The recording unit is used to replace the word to be corrected with the predicted replacement word, and record the masked text to be corrected after the replacement as the corrected audio text.

[0121] In one embodiment, the extraction module 13 includes:

[0122] A preprocessing unit, configured to preprocess the corrected audio data corresponding to the corrected audio text to obtain target voice data;

[0123] The feature extraction unit is used to extract the voiceprint feature of the target speech data to obtain the audio voiceprint feature corresponding to the corrected audio data.

[0124] In one embodiment, the feature vector module 14 includes:

[0125] A text word segmentation unit, configured to perform word segmentation processing on the corrected audio text to obtain audio words corresponding to the corrected audio text;

[0126] a text feature vector unit, configured to perform vector conversion on the audio words to obtain word vectors corresponding to the audio words, and determine the text feature vector corresponding to the corrected audio text based on all the word vectors;

[0127] The audio feature vector unit is used to perform segmentation processing on the audio voiceprint feature to obtain a segmented audio signal, and perform vector conversion on the segmented audio signal to obtain the audio feature vector corresponding to the audio voiceprint feature.

[0128] In one embodiment, the role identification module 15 includes:

[0129] An emotion acquisition unit, configured to acquire a preset emotion set and cluster all the target emotions in the preset emotion set to obtain at least one emotion cluster group; one emotion cluster group includes a plurality of target emotions;

[0130] A vector conversion unit, configured to perform vector conversion on all the emotion cluster groups to obtain emotion cluster vectors;

[0131] a vector fusion unit, configured to fuse the audio feature vector and the text feature vector to obtain a fused feature vector;

[0132] The vector matching unit is used to perform similarity matching on the fusion feature vector and the emotion clustering vector to obtain a vector matching result, and determine the emotion recognition result corresponding to the corrected audio text based on the vector matching result.

[0133] In one embodiment, the role identification module 15 further includes:

[0134] a model acquisition unit, configured to determine scene information corresponding to the corrected audio data, and acquire a character recognition model corresponding to the scene information;

[0135] a role recognition result unit, configured to perform role recognition on the text feature vector and the audio feature vector using the role recognition model to obtain a role recognition result;

[0136] The role category unit is used to determine the role category corresponding to the corrected audio text according to the emotion recognition result and the role recognition result.

[0137] The specific definition of the role identification device can be found in the definition of the role identification method above and will not be repeated here. The various modules in the above-mentioned role identification device can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor of the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0138] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data used in the role identification method in the above-mentioned embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a role identification method is implemented.

[0139] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the role recognition method in the above embodiment is implemented.

[0140] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the role recognition method in the above embodiment is implemented.

[0141] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0142] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0143] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A role identification method, characterized in that: include: Acquire a target audio text, perform text detection on the target audio text, and obtain a text detection result; The text detection result includes a detection failure result; The detection failure result indicates that there is an error in the target audio text; Correcting the target audio text corresponding to the detection failure result to obtain a corrected audio text; Acquire corrected audio data corresponding to the corrected audio text, perform voiceprint feature extraction on the corrected audio data, and obtain audio voiceprint features; Determining a text feature vector corresponding to the corrected audio text, and determining an audio feature vector corresponding to the audio voiceprint feature; Determining an emotion recognition result corresponding to the corrected audio text based on the audio feature vector and the text feature vector, and determining a character category corresponding to the corrected audio text based on the emotion recognition result, the audio feature vector, and the text feature vector; The determining, based on the emotion recognition result, the audio feature vector, and the text feature vector, of a character category corresponding to the corrected audio text includes: Determining scene information corresponding to the corrected audio data, and obtaining a role recognition model corresponding to the scene information; wherein one scene information corresponds to one role recognition model; Performing role recognition on the text feature vector and the audio feature vector using the role recognition model to obtain a role recognition result; Determining the role category corresponding to the corrected audio text according to the emotion recognition result and the role recognition result; The determining of the text feature vector corresponding to the corrected audio text and the determining of the audio feature vector corresponding to the audio voiceprint feature include: Performing word segmentation processing on the corrected audio text to obtain audio words corresponding to the corrected audio text; Performing vector conversion on the audio words to obtain word vectors corresponding to the audio words, and determining the text feature vector corresponding to the corrected audio text based on all the word vectors; The audio voiceprint feature is segmented to obtain a segmented audio signal, and the segmented audio signal is vector-converted to obtain the audio feature vector corresponding to the audio voiceprint feature.

2. The role recognition method according to claim 1, wherein: The performing text detection on the target audio text to obtain a text detection result includes: Performing text detection on the target audio text to obtain a text detection value corresponding to the target audio text; A preset threshold is obtained, and the text detection result is determined according to the preset threshold and the text detection value.

3. The role recognition method according to claim 1, wherein: The correcting the target audio text corresponding to the detection failure result to obtain a corrected audio text includes: Determining the target audio text corresponding to the detection failure result as an erroneous audio text, and determining the words to be corrected contained in the erroneous audio text; Performing masking processing on the words to be corrected in the erroneous audio text to obtain masked text to be corrected; Inputting the masked text to be corrected into a preset language model, performing correction prediction on the masked text to be corrected by the preset language model, and obtaining a predicted replacement word corresponding to the word to be corrected; The predicted replacement word replaces the word to be corrected, and the erroneous audio text after the replacement is recorded as the corrected audio text.

4. The role recognition method according to claim 1, wherein: The obtaining of the corrected audio data corresponding to the corrected audio text, and performing voiceprint feature extraction on the corrected audio data to obtain the audio voiceprint feature, includes: Preprocessing the corrected audio data corresponding to the corrected audio text to obtain target voice data; Voiceprint features are extracted from the target speech data to obtain the audio voiceprint features corresponding to the corrected audio data.

5. The role recognition method according to claim 1, wherein: The determining, based on the audio feature vector and the text feature vector, an emotion recognition result corresponding to the corrected audio text includes: Obtaining a preset emotion set, and clustering all target emotions in the preset emotion set to obtain at least one emotion cluster group; one emotion cluster group includes a plurality of target emotions; Performing vector conversion on all the emotion cluster groups to obtain emotion cluster vectors; Fusing the audio feature vector and the text feature vector to obtain a fused feature vector; The fusion feature vector and the emotion clustering vector are similarly matched to obtain a vector matching result, and the emotion recognition result corresponding to the corrected audio text is determined based on the vector matching result.

6. A role recognition device, characterized in that: include: A text detection module is used to obtain a target audio text, perform text detection on the target audio text, and obtain a text detection result; The text detection result includes a detection failure result; the detection failure result indicates that there is an error in the target audio text; A text correction module is used to correct the target audio text corresponding to the detection failure result to obtain a corrected audio text; A feature extraction module is used to obtain the corrected audio data corresponding to the corrected audio text, perform voiceprint feature extraction on the corrected audio data, and obtain audio voiceprint features; A feature vector module, configured to determine a text feature vector corresponding to the corrected audio text, and an audio feature vector corresponding to the audio voiceprint feature; a role identification module, configured to determine an emotion recognition result corresponding to the corrected audio text based on the audio feature vector and the text feature vector, and determine a role category corresponding to the corrected audio text based on the emotion recognition result, the audio feature vector, and the text feature vector; The role identification module also includes: A model acquisition unit, configured to determine scene information corresponding to the corrected audio data, and acquire a role recognition model corresponding to the scene information; wherein one scene information corresponds to one role recognition model; a role recognition result unit, configured to perform role recognition on the text feature vector and the audio feature vector using the role recognition model to obtain a role recognition result; A role classification unit, configured to determine the role category corresponding to the corrected audio text according to the emotion recognition result and the role recognition result; The feature vector module includes: A text word segmentation unit, configured to perform word segmentation processing on the corrected audio text to obtain audio words corresponding to the corrected audio text; a text feature vector unit, configured to perform vector conversion on the audio words to obtain word vectors corresponding to the audio words, and determine the text feature vector corresponding to the corrected audio text based on all the word vectors; The audio feature vector unit is used to perform segmentation processing on the audio voiceprint feature to obtain a segmented audio signal, and perform vector conversion on the segmented audio signal to obtain the audio feature vector corresponding to the audio voiceprint feature.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the role identification method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the role identification method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Speaker role recognition method and device thereof, electronic equipment and storage medium

    CN112233680A

  • User emotion recognition method and system

    CN113506586A