Voice emotion recognition method and device, storage medium and computer device
By combining logarithmic Mel spectrum and differential feature extraction with Transformer and graph convolutional neural networks, the problem of low accuracy in speech emotion recognition is solved, and more efficient emotion information extraction and classification are achieved.
Patent Information
- Application Number
- CN202310114018.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-13
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-02-13
AI Technical Summary
Existing speech emotion recognition technologies are affected by background noise and speaker voice characteristics, resulting in low recognition accuracy and insufficient ability to extract semantic features in space.
By extracting the log-Mel spectrum and its first and second differences from the speech data, three-dimensional speech features are obtained. The Transformer model encoder is used for feature extraction, combined with graph convolutional neural networks and pooling layers, and finally, a classification network is used for sentiment classification.
It improves the accuracy of speech emotion recognition, retains more effective emotional information, reduces the influence of emotion-irrelevant factors, enhances feature extraction capabilities, and captures the dependencies between frames.
Smart Images

Figure CN116312639B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of voice emotion recognition, specifically to a voice emotion recognition method, apparatus, storage medium, and computer device. Background Technology
[0002] Speech emotion recognition plays an important role in many applications, but factors such as background noise and speaker voice characteristics increase the difficulty of speech emotion recognition. This makes it difficult for existing speech emotion recognition technologies to capture prominent emotional information, and existing related technologies also have the drawback of low ability to extract semantic features in space, resulting in low accuracy of speech emotion recognition results. Summary of the Invention
[0003] The purpose of this application is to overcome the shortcomings and deficiencies in the prior art and provide a voice emotion recognition method, device, storage medium and computer equipment that can improve the accuracy of voice emotion recognition.
[0004] The first aspect of this application provides a voice emotion recognition method, including:
[0005] The log-Mel spectrum of the speech data, as well as the first and second differences of the log-Mel spectrum, are extracted to obtain three-dimensional speech features;
[0006] Feature extraction is performed on the three-dimensional speech features to obtain frame-level global features containing speech context information;
[0007] The frame-level global features are input into a graph convolutional neural network for global information reorganization to obtain graph node features containing global information.
[0008] The graph node features are input into a pooling layer for pooling to obtain the corresponding graph-level features.
[0009] The graph-level features are input into a classification network for emotion classification to obtain the emotion category of the speech data; wherein, the classification network includes a fully connected layer and a softmax layer.
[0010] A second aspect of this application provides a voice emotion recognition device, comprising:
[0011] The three-dimensional speech feature acquisition module is used to extract the log-Mel spectrum of the speech data, as well as the first-order difference and second-order difference of the log-Mel spectrum, to obtain three-dimensional speech features;
[0012] The global feature acquisition module is used to extract features from the three-dimensional speech features to obtain frame-level global features containing speech context information.
[0013] The graph node feature acquisition module is used to input the frame-level global features into the graph convolutional neural network for global information reorganization to obtain graph node features containing global information.
[0014] The graph-level feature acquisition module is used to input the graph node features into the pooling layer for pooling to obtain the corresponding graph-level features;
[0015] The emotion category acquisition module is used to input the graph-level features into a classification network for emotion classification to obtain the emotion category of the speech data; wherein, the classification network includes a fully connected layer and a softmax layer.
[0016] A third aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the speech emotion recognition method described above.
[0017] A fourth aspect of this application provides a computer device including a storage device, a processor, and a computer program stored in the storage device and executable by the processor, wherein the processor executes the computer program to implement the steps of the voice emotion recognition method as described above.
[0018] Compared to related technologies, this application first obtains three-dimensional speech features based on the log-Mel spectrum of the speech data, as well as the first and second differences of the log-Mel spectrum. Then, feature extraction is performed on the three-dimensional speech features to obtain frame-level global features containing speech context information. Next, global information is reorganized from the frame-level global features to obtain graph node features containing global information. Then, pooling is used to obtain the corresponding graph-level features. The graph-level features are input into a classification network for sentiment classification to obtain the sentiment category of the speech data. Because the log-Mel spectrum, as well as the first and second differences of the log-Mel spectrum, are used as three-dimensional speech features, more effective sentiment information can be retained and the impression of factors unrelated to sentiment can be reduced. By extracting features from the three-dimensional speech features, the model's ability to extract global context features can be improved. Furthermore, by using a graph convolutional neural network, the dependencies between frames in the sequence can be better captured, enhancing the concentration of features and further improving the feature extraction capability, thereby improving the accuracy of speech sentiment recognition.
[0019] To provide a clearer understanding of this application, the specific embodiments of this application will be described below in conjunction with the accompanying drawings. Attached Figure Description
[0020] Figure 1 This is a flowchart of a speech emotion recognition method according to an embodiment of this application.
[0021] Figure 2This is a flowchart illustrating the frame-level global feature acquisition process of a speech emotion recognition method according to an embodiment of this application.
[0022] Figure 3 This is an undirected cyclic graph structure of the adjacency matrix of a speech emotion recognition method according to an embodiment of this application.
[0023] Figure 4 This is a schematic diagram of the module connections of a voice emotion recognition device according to an embodiment of this application.
[0024] 100. Voice emotion recognition device; 101. Three-dimensional voice feature acquisition module; 102. Global feature acquisition module; 103. Graph node feature acquisition module; 104. Graph-level feature acquisition module; 105. Emotion category acquisition module. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0026] It should be understood that the described embodiments are merely some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of the embodiments of this application.
[0027] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances. The singular forms "a," "the," and "the" used in this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. The word "if" as used herein can be interpreted as "when," "when," or "in response to determination."
[0028] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0029] Please see Figure 1This is a flowchart of a speech emotion recognition method according to an embodiment of this application. The first embodiment of this application provides a speech emotion recognition method, including:
[0030] S1: Extract the log-Mel spectrum of the speech data, as well as the first and second differences of the log-Mel spectrum, to obtain three-dimensional speech features.
[0031] The log-Mel spectrum of the speech data is obtained by dividing the pre-emphasized speech data into short-time frames, multiplying each frame by a window function, performing a discrete Fourier transform on each frame to obtain the corresponding short-time spectrum, and then squaring the modulus of the short-time spectrum to obtain the corresponding discrete power spectrum. Then, a Mel filter bank is used to convert the linear frequency discrete power spectrum into a nonlinear Mel frequency spectrum, and a logarithmic operation is performed on the Mel frequency spectrum to obtain the log-Mel spectrum, which is used to extract low-level speech features from each frame of the speech data. The Mel filter bank includes multiple triangular filters. For example, if the speech data is divided into 300 short-time frames and there are 40 triangular filters, the resulting log-Mel spectrum matrix can be represented as a [300, 40] matrix, where 300 is the number of frames in the log-Mel spectrum and 40 is the dimension of each frame.
[0032] Since the number of parameters in the first and second difference matrices of the log-Mel spectrum is exactly the same as the number of parameters in the log-Mel spectrum matrix, based on the above example, the first and second differences of the log-Mel spectrum can both be represented as [300, 40]. Based on the log-Mel spectrum and its first and second differences, a three-dimensional Mel spectrum is constructed to obtain a three-dimensional matrix of [300, 40, 3]. This allows for the acquisition of more low-level speech features of different dimensions. The three-dimensional Mel spectrum is then used as the three-dimensional speech feature to more fully and comprehensively acquire low-level speech features from the speech data, thereby acquiring effective emotional information more fully and comprehensively.
[0033] After obtaining the three-dimensional speech features, the three-dimensional speech features are divided into equal-length 3-second segments. Segments with a duration of less than 3 seconds are padded to 3 seconds using the zero-padding method. Then, the content of step S2 is executed on each 3-second segment of the three-dimensional speech features.
[0034] S2: Extract features from the three-dimensional speech features to obtain frame-level global features containing speech context information.
[0035] In this process, feature extraction for 3D speech features is achieved through a Transformer model encoder. The Transformer model encoder can further learn from these low-level speech features to obtain high-level speech features containing global information, i.e., frame-level global features. Using the Transformer model encoder as the main model to replace the traditional RNN network structure for high-dimensional feature extraction provides the ability to focus on different spatiotemporal locations, and it is more capable of sequence modeling the relative dependencies between features at different locations, thus improving the model's ability to extract global contextual features.
[0036] Frame-level global features include the results of emotion feature extraction for each frame of the three-dimensional speech features. In frame-level global features, the emotion feature extraction results of two adjacent frames are used to reflect the corresponding speech context information. Compared with general global features, the frame-level global features of this instance combine the influence of the order of the upper and lower frames on the emotion features. Therefore, the emotion feature information in the frame-level global features includes speech context information.
[0037] S3: Input the frame-level global features into a graph convolutional neural network to reorganize the global information and obtain graph node features containing global information.
[0038] The working principle of graph convolutional neural networks is to propagate information between nodes based on the correlation matrix. Graph convolution consists of nodes and edges, and the weight of the edges is generally calculated from the adjacency matrix.
[0039] S4: Input the graph node features into the pooling layer for pooling to obtain the corresponding graph-level features.
[0040] The main function of the pooling layer is to sample features and reduce parameters. In this embodiment, the pooling layer uses average pooling (or mean pooling).
[0041] S5: Input the graph-level features into a classification network for emotion classification to obtain the emotion category of the speech data; wherein, the classification network includes a fully connected layer and a softmax layer.
[0042] Fully connected layers are often located at the end of the model. Each neuron is connected to all neurons in the upper layer. They can combine local information with category recognition in convolutional or pooling layers. The feature scores output by the fully connected layer can be obtained by weighted summation of the inputs, and the feature scores are used to indicate the emotion category corresponding to the speech data.
[0043] The softmax layer maps the feature scores to the probability interval (0,1), and then takes the sentiment category corresponding to the dimension with the highest probability as the final output, thus obtaining the sentiment category corresponding to the speech data.
[0044] Compared to related technologies, this application first obtains three-dimensional speech features based on the log-Mel spectrum of the speech data, as well as the first and second differences of the log-Mel spectrum. Then, feature extraction is performed on the three-dimensional speech features to obtain frame-level global features containing speech context information. Next, global information is reorganized from the frame-level global features to obtain graph node features containing global information. Then, pooling is used to obtain the corresponding graph-level features. The graph-level features are input into a classification network for sentiment classification to obtain the sentiment category of the speech data. Because the log-Mel spectrum, as well as the first and second differences of the log-Mel spectrum, are used as three-dimensional speech features, more effective sentiment information can be retained and the impression of factors unrelated to sentiment can be reduced. By extracting features from the three-dimensional speech features, the model's ability to extract global context features can be improved. Furthermore, by using a graph convolutional neural network, the dependencies between frames in the sequence can be better captured, enhancing the concentration of features and further improving the feature extraction capability, thereby improving the accuracy of speech sentiment recognition.
[0045] Please see Figure 2 In one feasible embodiment, step S2: extracting features from the three-dimensional speech features to obtain frame-level global features containing speech context information includes:
[0046] S21: Add position vectors to the three-dimensional speech features to obtain a speech sequence encoding containing position vectors.
[0047] Position vector addition refers to adding a position vector to each frame of the input 3D speech features through a positional encoding layer. This position vector represents the order of frames (and corresponding sentiment features) within the 3D speech features, allowing subsequent feature extraction processes to be executed according to the order of the encoded position vectors in the speech sequence. This is because each word in a speech has a specific positional relationship, and therefore each frame in the speech sequence also has a specific positional relationship. After positional encoding of each frame, frame-level features at the corresponding position are extracted. Taking a 3-second 3D speech feature as an example, it corresponds to 300 frames, meaning 300 position vectors need to be added. Position vector addition is performed before the 3D speech features are input into the multi-layer Transformer model encoder.
[0048] In this embodiment, the position vector is added using the following formula:
[0049]
[0050]
[0051] Where PE is the result of adding the position vector, pos is the position of the frame, i is the dimension of the frame, and d is the position vector of the frame. modelThis is the preset output dimension.
[0052] S22: Input the 3D speech features containing position vectors into the multilayer Transformer model encoder.
[0053] S23: Each layer of the Transformer model encoder extracts features from the input and uses the feature extraction results as the input to the next layer of the Transformer model encoder; wherein, the input of the first layer of the Transformer model encoder is the three-dimensional speech features, and the feature extraction result of the last layer of the Transformer model encoder is the frame-level global features.
[0054] The multi-layer Transformer model encoder in this embodiment consists of multiple identical encoder layers, with a total of six encoder layers. Therefore, the three-dimensional speech features need to be extracted sequentially through these six encoder layers to obtain frame-level global features. In other embodiments, those skilled in the art can modify the specific number of encoder layers according to their needs.
[0055] In this embodiment, the features of each frame of the three-dimensional speech features can be extracted by a multi-layer Transformer model encoder, thereby obtaining the frame-level global features of the three-dimensional speech features.
[0056] In one feasible embodiment, each layer of the Transformer model encoder includes a multi-head self-attention mechanism layer and a feedforward neural network.
[0057] S23: The step of each Transformer model encoder extracting features from the input and using the feature extraction results as the input to the next Transformer model encoder includes:
[0058] S231: The speech sequence is encoded and input into the multi-head self-attention mechanism layer to perform attention operations on the speech, resulting in multiple attention matrices.
[0059] Specifically, the attention matrix can be obtained using the following formula:
[0060]
[0061] Where Q, K, and V are three vector matrices generated by encoding the speech sequence, T is the transpose symbol, and d k This is a scaling factor.
[0062] Q, K, and V are vector matrices Query(Q), Key(K), and Value(V), respectively. These three vector matrices are generated from the speech sequence encoded into the encoder during the multi-head self-attention mechanism layer's attention operation on the speech. These three vector matrices are obtained by combining the speech sequence encoding with three weight matrices W. Q W K W V The result obtained by multiplying.
[0063] The multi-head self-attention mechanism layer can generate multiple attention weight matrices, and each attention head has three independent weight matrices. Therefore, the vector matrices Q, K, and V generated by each attention head are not exactly the same, and each attention matrix is also different. In this embodiment, the multi-head self-attention mechanism layer uses 8 attention heads, thus obtaining 8 different attention matrices.
[0064] By using a multi-head self-attention mechanism layer, the ability to focus on information in different spatiotemporal emotional subspaces at different spatiotemporal locations can be extended, making the model more capable of sequence modeling the relative dependencies between features at different locations.
[0065] S232: Concatenate and splice the multiple attention matrices to obtain the target attention matrix.
[0066] Since the feedforward neural network receives a single vector matrix as input, while step S31 yields multiple different attention matrices, it is necessary to concatenate these multiple attention matrices obtained in step S231 and then multiply them by an additional weight matrix to obtain a single attention matrix for input into the feedforward neural network. Specifically, the target attention matrix can be obtained using the following formula:
[0067] MultiHead(Q,K,V)=Concat(head1,...,head h W O ;
[0068] head i =Attention(QW i Q ,KW i K VW i V );
[0069] Among them, head i Let W be the i-th attention matrix; h is the total number of attention matrices; O For the additional weight matrix; Q, K, and V are the three vector matrices generated by encoding the speech sequence; Wi Q W is the weight matrix of the vector matrix Q; i K W is the weight matrix of the vector matrix K; i V Let V be the weight matrix of the vector matrix V.
[0070] S233: Input the target attention matrix into the feedforward neural network to extract features from the target attention matrix through two linear transformation layers of the feedforward neural network, and obtain the feature extraction result output by the feedforward neural network.
[0071] The feedforward neural network includes two linear transformation layers. The first linear transformation layer uses the ReLU activation function, while the second linear transformation layer does not use an activation function. Because of the use of the ReLU activation function, non-linear activation can be achieved, which improves the non-linear fitting ability of the feedforward neural network and thus increases the performance of the model.
[0072] The feature extraction result of the feedforward neural network output can be obtained using the following formula:
[0073] FFN(x)=max(0,xE1+b1)E2+b2;
[0074] Wherein, FFN(x) is the feature extraction result, x is the target attention matrix, E1 is the transformation matrix of the first linear transformation layer, b1 is the bias of the first linear transformation layer, E2 is the transformation matrix of the second linear transformation layer, and b2 is the bias of the second linear transformation layer.
[0075] In this embodiment, when each layer of the Transformer model encoder extracts features from 3D speech features, the combination of the multi-head self-attention mechanism layer and the feedforward neural network makes the model more capable of performing sequence modeling of the relative dependencies between features at different locations, and also increases the model's performance, thereby improving the accuracy of the feature extraction results output by the model.
[0076] In one feasible embodiment, the graph convolutional neural network includes at least two graph convolutional layers;
[0077] Step S3: The step of inputting the frame-level global features into a graph convolutional neural network for global information reorganization to obtain graph node features containing global information includes:
[0078] S31: Convert the frame-level global features into graph convolution.
[0079] Graph convolution is generated by a graph convolutional neural network that propagates information between nodes based on frame-level global features. A graph convolution consists of nodes and edges, and the edge weights are typically calculated using an adjacency matrix. Taking 300 frames of frame-level global features as an example, the number of nodes in the graph convolution would be 300.
[0080] S32: Input the graph convolution into the at least two graph convolutional layers to obtain the corresponding graph node-level embedding vector features.
[0081] When there are two graph convolutional layers, the embedding vector features at the graph node level are obtained using the following formula:
[0082]
[0083] Among them, H (l+1) H is the embedding vector feature at the graph node level. (0) Let X be the feature matrix containing the feature vectors of all nodes in the graph convolution, D be a diagonal matrix, l+1 and l be the layer numbers of the corresponding graph convolutional layers, and W be the feature matrix. (l) Here, σ is the trainable weight matrix of the l-th layer, σ(·) is the activation function, and A is the adjacency matrix. In this embodiment, the adjacency matrix used is an undirected cyclic graph structure (e.g., Figure 3 As shown, X is the feature matrix containing the feature vectors of all nodes in the graph convolution, M represents the number of nodes, and V represents the set of M nodes. Figure 3 In the X1, X2...X M (This is the eigenvector of the node), and the adjacency matrix is specifically represented as:
[0084]
[0085] S33: The embedded vector features are activated by two activation functions to obtain the corresponding graph node features.
[0086] The graph node features are obtained using the following formula:
[0087]
[0088] in, The normalized adjacency matrix can be represented as: X is the feature matrix containing the feature vectors of all nodes in the graph convolution.
[0089] In this embodiment, by converting frame-level global features into graph convolutional inputs and feeding them into a graph convolutional neural network to update node information, the frame-level global features can be processed by the graph convolutional neural network to obtain graph node features with enhanced sentiment information.
[0090] Please see Figure 4The second embodiment of this application provides a voice emotion recognition device 100, including:
[0091] The three-dimensional speech feature acquisition module 101 is used to extract the log-Mel spectrum of the speech data, as well as the first-order difference and second-order difference of the log-Mel spectrum, to obtain three-dimensional speech features.
[0092] The global feature acquisition module 102 is used to extract features from the three-dimensional speech features to obtain frame-level global features containing speech context information;
[0093] The graph node feature acquisition module 103 is used to input the frame-level global features into the graph convolutional neural network for global information reorganization to obtain graph node features containing global information.
[0094] The graph-level feature acquisition module 104 is used to input the graph node features into the pooling layer for pooling to obtain the corresponding graph-level features.
[0095] The emotion category acquisition module 105 is used to input the graph-level features into a classification network for emotion classification to obtain the emotion category of the speech data; wherein, the classification network includes a fully connected layer and a softmax layer.
[0096] It should be noted that the speech emotion recognition device provided in the second embodiment of this application is only illustrated by the above-described division of functional modules when executing the speech emotion recognition method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the speech emotion recognition device provided in the second embodiment of this application and the speech emotion recognition method in the first embodiment of this application belong to the same concept, and its implementation process is detailed in the method embodiment, which will not be repeated here.
[0097] A third aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the speech emotion recognition method described above.
[0098] A fourth aspect of this application provides a computer device including a storage device, a processor, and a computer program stored in the storage device and executable by the processor, wherein the processor executes the computer program to implement the steps of the voice emotion recognition method as described above.
[0099] The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without any inventive effort.
[0100] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0101] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function selected in one or more boxes.
[0102] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function selected in one or more boxes.
[0103] In a typical configuration, a computing device includes one or more processors (CPUs), input-to-output interfaces, network interfaces, and memory.
[0104] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0105] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0106] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0107] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for speech emotion recognition, characterized in that, The method comprises the following steps: extracting a log-mel spectrum of voice data and first-order and second-order differences of the log-mel spectrum to obtain three-dimensional voice features; performing feature extraction on the three-dimensional voice features to obtain frame-level global features containing voice context information; inputting the frame-level global features into a graph convolutional neural network to reorganize global information and obtain graph node features containing global information; inputting the graph node features into a pooling layer for pooling to obtain corresponding graph-level features; inputting the graph-level features into a classification network for emotion classification to obtain an emotion category of the voice data; wherein the classification network comprises a fully connected layer and a softmax layer.
2. The voice emotion recognition method of claim 1, wherein, The step of performing feature extraction on the three-dimensional voice features to obtain frame-level global features containing voice context information comprises: adding a position vector to the three-dimensional voice features to obtain voice sequence encoding containing a position vector; inputting the three-dimensional voice features containing the position vector into a multi-layer Transformer model encoder; each layer of the Transformer model encoder performs feature extraction on the input and takes the feature extraction result as the input of the next layer of the Transformer model encoder; wherein the input of the first layer of the Transformer model encoder is the three-dimensional voice features, and the feature extraction result of the last layer of the Transformer model encoder is the frame-level global features. 3.The voice emotion recognition method of claim 2, wherein, Each layer of the Transformer model encoder comprises a multi-head self-attention mechanism layer and a feedforward neural network. The step of each layer of the Transformer model encoder performing feature extraction on the input and taking the feature extraction result as the input of the next layer of the Transformer model encoder comprises: inputting the voice sequence encoding into the multi-head self-attention mechanism layer to perform attention operation on the voice to obtain a plurality of attention matrices; concatenating the plurality of attention matrices to obtain a target attention matrix; inputting the target attention matrix into the feedforward neural network to perform feature extraction on the target attention matrix through two linear transformation layers of the feedforward neural network to obtain a feature extraction result output by the feedforward neural network.
4. The voice emotion recognition method of claim 3, wherein, The step of inputting the voice sequence encoding into the multi-head self-attention mechanism layer to perform attention operation on the voice to obtain a plurality of attention matrices comprises: obtaining the attention matrix by the following formula: wherein Q, K, V are three vector matrices generated by the speech sequence encoding, T is a transpose symbol, d k is a proportional factor.
5. The voice emotion recognition method of claim 3, wherein, The step of concatenating the plurality of attention matrices to obtain a target attention matrix comprises: obtaining the target attention matrix by the following formula: MultiHead(Q, K, V) = Concat(head1,...,head h )W O ; head i = Attention(QW i Q ,KW i K ,VW i V ); wherein head i is the i-th attention matrix; h is the total number of attention matrices; W O is an additional weight matrix; Q, K, V are three vector matrices generated by the speech sequence encoding; W i Q is a weight matrix of the vector matrix Q; W i K is a weight matrix of the vector matrix K; W i V is a weight matrix of the vector matrix V.
6. The speech emotion recognition method of claim 3, wherein, The step of inputting the target attention matrix into the feedforward neural network to perform feature extraction on the target attention matrix through two linear transformation layers of the feedforward neural network to obtain a feature extraction result output by the feedforward neural network comprises: obtaining the feature extraction result output by the feedforward neural network by the following formula: FFN(x)=max(0,xE1+b1)E2+b2; Wherein, FFN(x) is the feature extraction result, x is the target attention matrix, E1 is the change matrix of the first linear transformation layer, b1 is the bias of the first linear transformation layer, E2 is the change matrix of the second linear transformation layer, and b2 is the bias of the second linear transformation layer.
7. The method of speech emotion recognition according to claim 1, wherein: The graph convolutional neural network comprises at least two graph convolutional layers. The step of inputting the frame-level global feature into the graph convolutional neural network for global information reorganization to obtain a graph node feature comprising global information comprises: Converting the frame-level global feature into a graph convolution; Inputting the graph convolution into the at least two graph convolutional layers to obtain a corresponding graph node-level embedding vector feature; Activating the embedding vector feature through two activation functions to obtain a corresponding graph node feature.
8. A voice emotion recognition apparatus, characterized by comprising: Comprise: A three-dimensional speech feature acquisition module configured to extract a log mel spectrum of speech data, a first-order difference of the log mel spectrum, and a second-order difference of the log mel spectrum to obtain three-dimensional speech features; A global feature acquisition module configured to perform feature extraction on the three-dimensional speech features to obtain frame-level global features comprising speech context information; A graph node feature acquisition module configured to input the frame-level global features into a graph convolutional neural network for global information reorganization to obtain graph node features comprising global information; A graph-level feature acquisition module configured to input the graph node features into a pooling layer for pooling to obtain corresponding graph-level features; An emotion category acquisition module configured to input the graph-level features into a classification network for emotion classification to obtain emotion categories of the speech data; wherein the classification network comprises a fully connected layer and a softmax layer.
9. A computer readable storage medium storing a computer program, characterized in that: The computer program, when executed by a processor, implements the steps of the speech emotion recognition method according to any one of claims 1 to 7.
10. A computer device, comprising: A computer program product comprising a storage, a processor, and a computer program stored in the storage and executable by the processor, wherein the processor, when executing the computer program, implements the steps of the speech emotion recognition method according to any one of claims 1 to 7. A computer program product comprising a storage, a processor, and a computer program stored in the storage and executable by the processor, wherein the processor, when executing the computer program, implements the steps of the speech emotion recognition method according to any one of claims 1 to 7.