Emotion recognition method, device, equipment, medium and program product
By fusing speech and text features, using deep learning networks to extract rich features and process them in segments, the problem of low emotion recognition accuracy in existing technologies is solved, and more efficient emotion recognition effects are achieved.
Patent Information
- Application Number
- CN202210964637.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-11
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-08-11
AI Technical Summary
Existing technologies have low accuracy in emotion recognition, especially the classification accuracy through machine learning methods. Deep learning methods are limited by the speech-to-text format and cannot be effectively applied and implemented.
By determining the text features and speech features of the target speech data, feature fusion is performed, and the long short-term memory network and convolutional neural network in deep learning are used to extract rich fusion features. Combined with speech segmentation processing, emotion categories can be identified.
It improves the accuracy and efficiency of emotion recognition, can adapt to various emotional expression habits, and is suitable for a variety of application scenarios.
Smart Images

Figure CN115331700B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the field of deep learning and speech processing technology, and more specifically to an emotion recognition method, device, equipment, medium and program product. Background Art
[0002] With the development of computer technology and Internet technology, artificial intelligence can assist in all aspects of production and life. Among them, emotion recognition is an application scenario of artificial intelligence. How to accurately perform emotion recognition has become a technical problem that needs to be solved urgently. Summary of the Invention
[0003] In view of the above problems, the present disclosure provides an emotion recognition method, apparatus, device, medium and program product.
[0004] According to one aspect of the present disclosure, an emotion recognition method is provided, comprising: determining text features and speech features of target speech data; determining fusion features of the text features and the speech features based on the text features and the speech features; and determining a target emotion category of the target speech data based on the fusion features.
[0005] According to an embodiment of the present disclosure, text features include initial text features and target text features, speech features include initial speech features and target speech features, and fusion features include initial fusion features, intermediate fusion features, and target fusion features, wherein the intermediate fusion features are obtained based on the initial fusion features. Determining fusion features of text features and speech features based on the text features and speech features includes: determining initial fusion features based on the initial text features and the initial speech features; and determining target fusion features based on the intermediate fusion features, the target text features, and the target speech features.
[0006] According to an embodiment of the present disclosure, determining the text features and speech features of target speech data includes: determining the initial text features and initial speech features of the target speech data according to a first feature extraction network; and inputting the initial text features, initial speech features, and initial fusion features into a second feature extraction network, respectively, to obtain target text features, target speech features, and intermediate fusion features, respectively.
[0007] According to an embodiment of the present disclosure, the second feature extraction network includes m long short-term memory network hidden layers, where m is a positive integer greater than or equal to 1.
[0008] According to an embodiment of the present disclosure, the target emotion category includes a target emotion category sequence. Determining the target emotion category of the target speech data based on the fusion features includes: segmenting the target speech data to obtain at least one target speech segment data; determining the segment emotion category of the target speech segment data based on the fusion features of each target speech segment data; and determining the target emotion category sequence of the target speech data based on the entire target speech segment data.
[0009] According to an embodiment of the present disclosure, the emotion category of a segment is represented by at least one of a positive label, a negative label, and a neutral label. The emotion recognition method further includes: determining the emotion categories of the segments corresponding to the target ratio at the end of the target emotion category sequence as a target segment emotion category sequence; and determining feedback data based on the target speech data when multiple emotion categories in the target segment emotion category sequence do not include a negative label and include at least one positive label.
[0010] According to the embodiment of the present disclosure, the emotion recognition method further includes: separating target object voice data from the initial voice data to obtain target voice data.
[0011] Another aspect of the present disclosure provides an emotion recognition device, comprising: a first determination module, a second determination module, and a target emotion category determination module. The first determination module is configured to determine text features and speech features of target speech data; the second determination module is configured to determine fusion features of the text features and speech features based on the text features and speech features; and the target emotion category determination module is configured to determine a target emotion category for the target speech data based on the fusion features.
[0012] Another aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned emotion recognition method.
[0013] On the other hand, the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the above-mentioned emotion recognition method.
[0014] On the other hand, the present disclosure further provides a computer program product, including a computer program, wherein the computer program is stored on at least one of a readable storage medium and an electronic device, and when the computer program is executed by a processor, the above-mentioned emotion recognition method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0016] Figure 1 Schematically illustrates an architecture diagram of an emotion recognition method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;
[0017] Figure 2 The flowchart of the emotion recognition method according to the embodiment of the present disclosure is schematically shown;
[0018] Figure 3 Schematically shows a flow chart of determining fusion features of text features and speech features in an emotion recognition method according to another embodiment of the present disclosure;
[0019] Figure 4A Schematically shows a flow chart of determining text features and speech features of target speech data in a method for emotion recognition according to yet another embodiment of the present disclosure;
[0020] Figure 4B A schematic diagram schematically shows an emotion recognition network model according to an embodiment of the present disclosure executing the emotion recognition method of the present disclosure;
[0021] Figure 5 Schematically shows a flow chart of determining a target emotion category of target speech data in an emotion recognition method according to yet another embodiment of the present disclosure;
[0022] Figure 6 Schematically shows a flow chart of an emotion recognition method according to yet another embodiment of the present disclosure;
[0023] Figure 7 Schematically shows a structural block diagram of an emotion recognition device according to an embodiment of the present disclosure; and
[0024] Figure 8 The block diagram schematically shows an electronic device suitable for implementing the emotion recognition method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0026] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0028] When expressions such as "at least one of A, B and C, etc." are used, they should generally be interpreted in accordance with the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0029] With the development of computer technology and Internet technology, artificial intelligence can assist in all aspects of production and life. Among them, emotion recognition is an application scenario of artificial intelligence. How to accurately perform emotion recognition has become a technical problem that needs to be solved urgently.
[0030] In some embodiments, emotion recognition of speech data is performed through machine learning methods. In machine learning methods, sound features are a type of manual features that contain rich prior knowledge, such as spectrograms, Mel spectra, and Mel-frequency cepstral coefficients. Prosodic features are used as the basis for distinguishing emotional states, and the differences in sound characteristics are captured to achieve classification tasks. For example, decision trees, SVMs, and random forests are used for classification. However, this emotion recognition method has the defects of low classification accuracy and a single classification category.
[0031] In some implementations, emotion recognition from speech data is performed using deep learning methods. These methods can overcome feature limitations and, with more flexible networking methods and layers, better extract high-level features from speech, thereby achieving improved classification metrics. However, most deep learning algorithms are based on converting speech into text or spectral signals, and then using the textual data or spectral signals as input. This limits model accuracy and hinders effective application and implementation.
[0032] It should be noted that the emotion recognition method and device determined in the embodiments of the present disclosure can be used in the financial field, and can also be used in any field other than the financial field. The embodiments of the present disclosure do not limit the application field of the emotion recognition method and device.
[0033] Taking the application of the emotion recognition method of the embodiment of the present disclosure in application scenarios such as banks as an example, by performing emotion recognition on the user's voice data, for example, information recommendations can be made to the user.
[0034] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0035] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0036] Figure 1 The following schematically shows an architecture diagram of an emotion recognition method according to an embodiment of the present disclosure.
[0037] like Figure 1 As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0038] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0039] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0040] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0041] It should be noted that the emotion recognition method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the emotion recognition device provided in the embodiment of the present disclosure can generally be set in the server 105. The emotion recognition method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the emotion recognition device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0042] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0043] The following will be based on Figure 1 The scene described by Figures 2 to 6 The emotion recognition method of the disclosed embodiment is described in detail.
[0044] Figure 2 The flowchart of the emotion recognition method 200 according to the embodiment of the present disclosure is schematically shown.
[0045] like Figure 2 As shown, the emotion recognition method 200 of this embodiment includes operations S210 to S230.
[0046] In operation S210 , text features and voice features of target voice data are determined.
[0047] The target speech data is a type of speech data, which is an audio form of a language. Therefore, for example, the corresponding target text can be determined based on the target speech data, and the text features of the target speech data can be determined based on the target text.
[0048] The text features of the target speech data can be understood as text-related features obtained after feature parameter extraction of the target speech data. The text features can be in the form of word vectors, for example.
[0049] The speech features of the target speech data can be understood as speech-related features obtained after feature parameter extraction of the target speech data. The speech features can be in the form of feature vectors, for example.
[0050] In operation S220 , a fusion feature of the text feature and the speech feature is determined based on the text feature and the speech feature.
[0051] For example, text features and speech features may be concatenated or superimposed to obtain fused features.
[0052] In operation S230 , a target emotion category of the target speech data is determined based on the fused features.
[0053] In reality, emotional expression habits vary. For example, the tone of voice reflected in speech data is more correlated with emotional categories, or the language reflected in speech data is more correlated with emotional categories. Therefore, determining the target emotional category of target speech data based solely on text features or speech features is less accurate and cannot cover application scenarios with a wide range of emotional expression habits.
[0054] According to the emotion recognition method of the disclosed embodiment, the fusion features determined by the text features and speech features of the target speech data can comprehensively characterize the target speech data from both text and speech dimensions, with the fusion features having superior representational properties. For emotion recognition tasks, the target emotion category of the target speech data determined based on the fusion features has higher accuracy in various application scenarios based on emotional expression habits, resulting in higher emotion recognition efficiency.
[0055] Figure 3 The flowchart of determining the fusion features of text features and speech features in the emotion recognition method according to another embodiment of the present disclosure is schematically shown.
[0056] The text features include initial text features and target text features, the speech features include initial speech features and target speech features, and the fusion features include initial fusion features, intermediate fusion features, and target fusion features. The intermediate fusion features are obtained based on the initial fusion features.
[0057] like Figure 3 As shown, for example, the following embodiment can be used to implement a specific example of determining the fusion feature of the text feature and the speech feature according to the text feature and the speech feature in operation S320.
[0058] In operation S321 , an initial fusion feature is determined based on the initial text feature and the initial speech feature.
[0059] Text features include initial text features and target text features. Initial text features and target text features can be used as shallow text features and deep text features, respectively, to enrich the representation of text features. Speech features include initial speech features and target speech features. Initial speech features and target speech features can also be used as shallow speech features and deep speech features, respectively, to enrich the representation of speech features.
[0060] In operation S322 , a target fusion feature is determined according to the intermediate fusion feature, the target text feature, and the target speech feature.
[0061] According to the emotion recognition method of the disclosed embodiment, the initial fused features obtained by the first feature fusion using initial text features and initial speech features have a higher feature breadth. The target fused features obtained by the second feature fusion using intermediate fused features, target text features, and target speech features have both feature breadth and feature depth, and are more representative.
[0062] Exemplarily, operations S321 to S322 may be performed between operations S210 to S230 , and operation S320 is similar to operation S220 .
[0063] Figure 4A The flowchart of determining text features and speech features of target speech data in an emotion recognition method according to yet another embodiment of the present disclosure is schematically shown.
[0064] like Figure 4A As shown, for example, a specific example of determining the text features and voice features of the target voice data in operation S410 can be implemented using the following embodiments.
[0065] In operation S411, initial text features and initial speech features of target speech data are determined according to a first feature extraction network.
[0066] For example, the target speech data may be preprocessed, and the preprocessed target speech data may be input into a first feature extraction network, through which the initial text features and initial speech features of the target speech data may be determined respectively.
[0067] Exemplarily, the target speech data may be subjected to pre-processing such as pre-emphasis and spectrogram extraction to obtain a spectrogram of the target speech data, which may be used as input data of a first feature extraction network to determine the initial speech features of the target speech data.
[0068] Specifically, for example, the target speech data can be passed through a first-order high-pass filter to achieve pre-emphasis by increasing the weight and resolution of the high-frequency signal part. Pre-emphasis can increase the energy of the high-frequency part of the signal, eliminate the influence of the rapid attenuation of the speech signal caused by some influencing factors, and make the signal spectrum relatively stable when transmitting information.
[0069] Exemplarily, spectrogram extraction may include, for example, sound framing, sound windowing, and Fourier transform.
[0070] For example, the target speech data can be divided into several speech frames by overlapping segmentation, for example, the frame length is 25ms, the frame shift is 10ms, and the frame shift refers to the overlapping part of the previous frame and the next frame of the speech signal.
[0071] For example, sound windowing can use a Hamming window as the window function to reduce the impact of the Gibbs effect. The Gibbs effect occurs when a periodic function with discontinuities is expanded using a Fourier series and then synthesized using a finite number of terms. As more terms are selected, the peaks in the synthesized waveform become closer to the discontinuities in the original signal. When a large number of terms are selected, the peaks approach a constant, approximately equal to 9% of the total jump value.
[0072] It can be understood that Fourier transform can convert the data in each window from a time domain signal to a frequency domain signal.
[0073] For example, each speech frame may be connected in time series to obtain a complete spectrogram of the target speech data.
[0074] For example, image normalization may be performed to normalize the spectrograms to the same size, for example, the spectrograms may be 256×256 in size.
[0075] Exemplarily, the target speech data may be subjected to preprocessing such as speech-to-text conversion, text cleaning, and text feature vectorization.
[0076] For example, text cleaning can remove meaningless symbols, remove non-Chinese characters, etc.
[0077] For example, a word vector generation model such as Word to Vector (word-to-vector conversion) may be used to vectorize text features.
[0078] In operation S412, the initial text features, the initial speech features, and the initial fusion features are respectively input into a second feature extraction network to obtain target text features, target speech features, and intermediate fusion features, respectively.
[0079] According to the emotion recognition method of the embodiment of the present disclosure, feature extraction based on deep learning can be achieved through the first feature extraction network and the second feature extraction network, with higher feature extraction efficiency.
[0080] In addition, the emotion recognition method of the embodiment of the present disclosure can extract speech features and text features through the first feature extraction network and the second feature extraction network respectively. The depths of the speech features and text features extracted by the first feature extraction network and the second feature extraction network are different, and the extracted features can be fused. This enables the emotion recognition method of the embodiment of the present disclosure to extract rich features from the target speech data, and the features have better representativeness, which facilitates the subsequent accurate determination of the target emotion category of the target speech data.
[0081] For example, the emotion recognition method of the disclosed embodiments can be implemented using an emotion recognition network model, which includes a first feature extraction network and a second feature extraction network. The emotion recognition network model can simultaneously extract both textual and speech features, and the extracted features are richer and more representative, resulting in superior performance.
[0082] Exemplarily, operations S411 to S412 may be performed before the above-mentioned operation S220 , and operation S410 is similar to the above-mentioned operation S210 .
[0083] Exemplarily, the second feature extraction network may include m long short-term memory network hidden layers, where m is a positive integer greater than or equal to 1.
[0084] Long Short-Term Memory (LSTM) network: Each hidden layer of the LSTM network includes an input gate, an output gate, and a forget gate. The forget gate is used to control which data of the input data of the current hidden layer (i.e., the output data of the previous hidden layer) is forgotten.
[0085] According to the emotion recognition method of the embodiment of the present disclosure, the long short-term memory network hidden layer of the second feature extraction network can selectively "memorize" historical features (historical features can be understood as features determined by the hidden layer before the current hidden layer), so that the second feature extraction network has temporal characteristics, which matches the temporal characteristics of emotions in the emotion recognition application scenario, and the accuracy of the target emotion category determined thereby is higher.
[0086] Exemplarily, the first feature extraction network may include n convolutional neural network hidden layers, convolutional neural network (CNN), where n is a positive integer greater than or equal to 1.
[0087] Figure 4B The diagram schematically shows a diagram of an emotion recognition method according to an embodiment of the present disclosure executed by an emotion recognition network model according to an embodiment of the present disclosure.
[0088] like Figure 4B As shown, the emotion recognition network model for executing the emotion recognition method according to the embodiment of the present disclosure may include a first feature extraction network N1 and a second feature extraction network N2.
[0089] like Figure 4BAs shown, for example, target speech data 401 can be input into a first feature extraction network N1. The first feature extraction network N1 includes two convolutional neural network hidden layers, CNN-L1 and CNN-L2. Initial text features 402 and initial speech features 403 can be obtained through the first feature extraction network N1. Initial fused features 404 can be obtained based on the initial text features 402 and initial speech features 403. Initial text features 402 are input into a second feature extraction network N2 to obtain target text features 405. Initial fused features 404 are input into the second feature extraction network N2 to obtain intermediate fused features 406. Initial speech features 403 are input into the second feature extraction network N2 to obtain target speech features 407. The second feature extraction network N2 includes two long short-term memory network hidden layers, LSTM-L1 and LSTM-L2. The second extraction network N2 can also include a convolutional neural network hidden layer, CNN-L3. Target fused features 408 can be determined based on the target text features 405, intermediate fused features 406, and target speech features 407. And according to the target fusion feature 408 , the target emotion category 409 can be determined.
[0090] Figure 5 The flowchart of determining the target emotion category of target speech data in the emotion recognition method according to another embodiment of the present disclosure is schematically shown. The target emotion category includes a target emotion category sequence.
[0091] like Figure 5 As shown, for example, the following embodiments may be used to implement a specific example of determining the target emotion category of the target speech data according to the fusion features in operation S530 .
[0092] In operation S531 , the target speech data is segmented to obtain at least one target speech segment data.
[0093] In operation S532 , the segment emotion category of each target speech segment data is determined according to the fusion feature of each target speech segment data.
[0094] In operation S533 , a target emotion category sequence of the target speech data is determined based on the full amount of target speech segment data.
[0095] For example, based on the target speech data, A target speech segments can be determined. For each target speech segment, a corresponding segment emotion category can be determined, resulting in A segment emotion categories. In this example, operation S533 of "determining a target emotion category sequence for the target speech data based on all target speech segment data" can be understood as determining a target emotion category sequence for the target speech data based on the A segment emotion categories of the A target speech segments. The A segment emotion categories are sorted in time series.
[0096] When the target speech data has a long duration, determining the target emotion category based on the target speech data will be more complicated and may also result in reduced accuracy.
[0097] The emotions reflected in the target speech data may change. Therefore, using only a single emotion category as the result of emotion recognition may be one-sided and inaccurate.
[0098] According to the emotion recognition method of the embodiment of the present disclosure, by segmenting the target speech data, at least one target speech segment data obtained has a finer granularity. By fusion features based on each target speech segment data, the segment emotion category of the target speech segment determined also has a correspondingly finer granularity. Based on the full amount of target speech segment data, the target emotion category sequence of the target speech data determined can characterize the emotional changes reflected by the target speech data with higher accuracy.
[0099] For example, the target speech segment data may be segmented based on each sentence to obtain at least one target speech segment data, where each target speech segment is a sentence.
[0100] Generally speaking, the meaning of a sentence is relatively complete. When the target speech segment data is segmented based on each sentence, the target speech segment data obtained can be both fine-grained and complete, so the emotion category of the segment determined thereby is more accurate.
[0101] Exemplarily, for example, the following embodiment can be used to implement a specific example of determining the segment emotion category of the target speech segment data based on the fusion features of each target speech segment data: using a multi-classification classifier, the fusion features of the target speech segment data are classified to obtain the segment emotion category of the target speech segment data.
[0102] The target speech segment data is classified according to the fusion features of the multi-classification classifier, and the target speech segment data obtained can have multiple segment emotion categories, the emotion category classification is more detailed, and the emotion recognition is more accurate.
[0103] Exemplarily, operations S531 to S533 may be performed after the above-mentioned operation S220 , and operation S530 is similar to the above-mentioned operation S230 .
[0104] Figure 6 The flowchart of the emotion recognition method according to another embodiment of the present disclosure is schematically shown.
[0105] like Figure 6 As shown, the emotion recognition method 600 according to another embodiment of the present disclosure may include operations S640 to S650.
[0106] The emotion category of the segment is represented by at least one of a positive label, a negative label, and a neutral label.
[0107] In operation S640 , the segment emotion categories corresponding to the target ratio at the tail of the target emotion category sequence are determined as the target segment emotion category sequence.
[0108] For example, the target emotion category sequence including 5 segment emotion categories is: negative-negative-negative-neutral-positive. When the target ratio is 40%, the target segment emotion category sequence corresponding to the target ratio at the tail of the target emotion category sequence is the last 2 (2 / 5=40%) segment emotion categories of the target emotion category sequence, namely neutral-positive.
[0109] In operation S650 , feedback data is determined based on the target speech data when the plurality of segment emotion categories of the target segment emotion category sequence do not include a negative label and include at least one positive label.
[0110] The emotions reflected by the target speech data may change. Relatively speaking, the emotions reflected by the tail part of the target speech data (i.e., the target segment emotion category sequence) are more meaningful for reference in application scenarios.
[0111] According to the emotion recognition method of an embodiment of the present disclosure, when multiple segment emotion categories in the target segment emotion category sequence do not include negative labels and include at least one positive label, the emotion reflected by the target speech data is characterized as positive. Feedback data determined based on the target speech data can, for example, be fed back to relevant personnel. In application scenarios such as banks, the feedback data can, for example, be fed back to bank staff, who can then recommend information to the user corresponding to the target speech data. Because the user's emotion is positive, the information recommendation is less likely to be rejected, resulting in higher information recommendation efficiency.
[0112] Exemplarily, operations S640 to S650 may be performed after the above-mentioned operation S230.
[0113] Exemplarily, when multiple segment emotion categories in the target segment emotion category sequence do not include a positive label and include at least one negative label, prompt data may also be determined based on the target speech data.
[0114] The prompt data can be used, for example, to prompt relevant personnel of the negative or other abnormal emotions of the user corresponding to the target voice data, and then the relevant personnel can communicate with the user with negative or other abnormal emotions and other processing procedures.
[0115] For example, negative labels may include: angry, sad, and afraid, and positive labels may include: happy. Segment emotion categories may also include: surprised. This allows for detailed emotion recognition.
[0116] Exemplarily, the emotion recognition method according to another embodiment of the present disclosure may further include:
[0117] Target object voice data is separated from the initial voice data to obtain target voice data.
[0118] In some application scenarios, a robot typically initiates a voice call with a user to generate initial voice data. This initial voice data includes the target user's voice data and the robot's voice data. After emotion recognition is performed based on the initial voice data, if the user's emotion corresponding to the initial voice data is positive, a relevant person can communicate with the user to improve accuracy in application scenarios such as information recommendations.
[0119] According to the emotion recognition method of the embodiment of the present disclosure, by separating the target object voice data from the initial voice data, irrelevant voice data such as robots can be eliminated, and the target object voice data can be retained. The target object voice data is used as the target voice data, and the accuracy of emotion recognition based on the target voice data is higher.
[0120] Based on the above emotion recognition method, the present disclosure also provides an emotion recognition device. Figure 7 The device is described in detail.
[0121] Figure 7 The figure schematically shows a structural block diagram of an emotion recognition device according to an embodiment of the present disclosure.
[0122] like Figure 7 As shown, the emotion recognition device 700 of this embodiment includes a first determination module 710 , a second determination module 720 and a target emotion category determination module 730 .
[0123] The first determination module 710 is configured to determine text features and speech features of target speech data.
[0124] The second determining module 720 is configured to determine a fusion feature of the text feature and the speech feature based on the text feature and the speech feature.
[0125] The target emotion category determination module 730 is configured to determine the target emotion category of the target speech data based on the fusion features.
[0126] According to an embodiment of the present disclosure, text features include initial text features and target text features, speech features include initial speech features and target speech features, and fusion features include initial fusion features, intermediate fusion features, and target fusion features, where the intermediate fusion features are derived based on the initial fusion features. The second determination module may include an initial fusion feature determination submodule and a target fusion feature determination submodule.
[0127] The initial fusion feature determination submodule is used to determine the initial fusion feature based on the initial text feature and the initial speech feature.
[0128] The target fusion feature determination submodule is used to determine the target fusion feature based on the intermediate fusion feature, the target text feature and the target speech feature.
[0129] According to an embodiment of the present disclosure, the first determining module may include: a first determining submodule and a second determining submodule.
[0130] The first determination submodule is used to determine the initial text features and initial speech features of the target speech data according to the first feature extraction network.
[0131] The second determination submodule is used to input the initial text features, the initial speech features and the initial fusion features into the second feature extraction network respectively to obtain the target text features, the target speech features and the intermediate fusion features respectively.
[0132] According to an embodiment of the present disclosure, the second feature extraction network includes m long short-term memory network hidden layers, where m is a positive integer greater than or equal to 1.
[0133] According to an embodiment of the present disclosure, the target emotion category includes a target emotion category sequence. The target emotion category determination module may include: a segmentation submodule, a segment emotion category determination submodule, and a target emotion category sequence determination submodule.
[0134] The segmentation submodule is used to segment the target speech data to obtain at least one target speech segment data.
[0135] The segment emotion category determination submodule is used to determine the segment emotion category of the target speech segment data according to the fusion features of each target speech segment data.
[0136] The target emotion category sequence determination submodule is used to determine the target emotion category sequence of the target speech data based on the full amount of target speech segment data.
[0137] According to an embodiment of the present disclosure, the emotion category of a segment is represented by at least one of a positive label, a negative label, and a neutral label. The emotion recognition device may further include: a target segment emotion category sequence determination module and a feedback data determination module.
[0138] The target segment emotion category sequence determination module is configured to determine the segment emotion category corresponding to the target ratio at the tail of the target emotion category sequence as the target segment emotion category sequence.
[0139] The feedback data determination module is configured to determine feedback data according to the target speech data when the plurality of segment emotion categories in the target segment emotion category sequence do not include negative labels and include at least one positive label.
[0140] The emotion recognition module according to the embodiment of the present disclosure may further include: a target voice data determination module.
[0141] The target voice data determination module is used to separate the target object voice data from the initial voice data to obtain the target voice data.
[0142] According to an embodiment of the present disclosure, any multiple modules among the first determination module 710, the second determination module 720, and the target emotion category determination module 730 can be combined into a single module, or any one of them can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present disclosure, at least one of the first determination module 710, the second determination module 720, and the target emotion category determination module 730 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware by any other reasonable means of integrating or packaging circuits, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, at least one of the first determination module 710, the second determination module 720, and the target emotion category determination module 730 can be at least partially implemented as a computer program module, which can perform the corresponding function when the computer program module is executed.
[0143] It should be understood that the embodiments of the device part of the present disclosure are the same or similar to the embodiments of the method part of the present disclosure, and the technical problems solved and the technical effects achieved are also the same or similar, and the present disclosure will not elaborate on them here.
[0144] Figure 8 The block diagram schematically shows an electronic device suitable for implementing the emotion recognition method according to an embodiment of the present disclosure.
[0145] like Figure 8As shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage part 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include an onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0146] Various programs and data required for the operation of the electronic device 800 are stored in the RAM 803. The processor 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The processor 801 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and RAM 803. The processor 801 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0147] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the I / O interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage portion 808 including a hard disk; and a communication portion 809 including a network interface card such as a LAN card or a modem. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in the drive 810 as needed, so that a computer program read therefrom can be installed into the storage portion 808 as needed.
[0148] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0149] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above and / or one or more memories other than ROM 802 and RAM 803.
[0150] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiments of the present disclosure.
[0151] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the computer program is executed by the processor 801. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0152] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0153] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from a removable medium 811. When the computer program is executed by the processor 801, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0154] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0156] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or couplings are intended to fall within the scope of this disclosure.
[0157] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. An emotion recognition method, comprising: Determining text features and speech features of target speech data; Determining, based on the text features and the voice features, a fusion feature of the text features and the voice features; as well as determining a target emotion category of the target speech data based on the fusion features; The text features include initial text features and target text features, the speech features include initial speech features and target speech features, the fusion features include initial fusion features, intermediate fusion features and target fusion features, and the intermediate fusion features are obtained based on the initial fusion features. Determining the text features and voice features of the target voice data includes: Determining the initial text features and the initial speech features of the target speech data according to a first feature extraction network, wherein the first feature extraction network includes n convolutional neural network hidden layers, where n≥1; Inputting the initial text features, the initial speech features, and the initial fusion features into a second feature extraction network respectively to obtain the target text features, the target speech features, and the intermediate fusion features respectively, wherein the second feature extraction network includes m long short-term memory network hidden layers, where m ≥ 1; Determining the fusion feature of the text feature and the voice feature based on the text feature and the voice feature includes: Determining the initial fusion feature according to the initial text feature and the initial speech feature; and The target fusion feature is determined according to the intermediate fusion feature, the target text feature, and the target speech feature.
2. The method according to claim 1, wherein The second feature extraction network also includes a convolutional neural network hidden layer.
3. The method according to any one of claims 1 to 2, wherein The target emotion category includes a target emotion category sequence; and determining the target emotion category of the target speech data according to the fusion feature includes: Segmenting the target speech data to obtain at least one target speech segment data; Determining a segment emotion category of the target speech segment data according to the fusion feature of each target speech segment data; and The target emotion category sequence of the target speech data is determined based on the entire target speech segment data.
4. The method according to claim 3, wherein: The emotion category of the segment is represented by at least one of a positive label, a negative label, and a neutral label; and the emotion recognition method further includes: Determining the segment emotion categories corresponding to the target ratio at the tail of the target emotion category sequence as the target segment emotion category sequence; and In a case where a plurality of the segment emotion categories in the target segment emotion category sequence do not include the negative label and include at least one positive label, feedback data is determined according to the target speech data.
5. The method according to any one of claims 1 to 2, further comprising: Target object voice data is separated from the initial voice data to obtain the target voice data.
6. An emotion recognition device, comprising: A first determination module, configured to determine text features and speech features of target speech data; A second determining module is configured to determine a fusion feature of the text feature and the speech feature based on the text feature and the speech feature; as well as a target emotion category determination module, configured to determine a target emotion category of the target speech data based on the fusion features; The text features include initial text features and target text features, the speech features include initial speech features and target speech features, the fusion features include initial fusion features, intermediate fusion features and target fusion features, and the intermediate fusion features are obtained based on the initial fusion features. The first determination module includes: a first determination submodule and a second determination submodule; The first determination submodule is configured to determine the initial text features and the initial speech features of the target speech data according to a first feature extraction network, wherein the first feature extraction network includes n convolutional neural network hidden layers, where n≥1; The second determination submodule is used to input the initial text features, the initial speech features, and the initial fusion features into a second feature extraction network, respectively, to obtain the target text features, the target speech features, and the intermediate fusion features, respectively, wherein the second feature extraction network includes m long short-term memory network hidden layers, m ≥ 1; The second determination module includes: an initial fusion feature determination submodule and a target fusion feature determination submodule; The initial fusion feature determination submodule is used to determine the initial fusion feature according to the initial text feature and the initial speech feature; The target fusion feature determination submodule is used to determine the target fusion feature according to the intermediate fusion feature, the target text feature and the target speech feature.
7. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 5.
9. A computer program product, comprising a computer program, wherein the computer program is stored in at least one of a readable storage medium and an electronic device, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Emotion recognition method and device based on bimodal combination multi-learning model recognizer
CN114595744A