Voice emotion recognition method, device and equipment and readable storage medium

By sampling, feature extraction and environmental signal-to-noise ratio weighted fusion of speech signals, combined with timing modeling network, the problem of reduced accuracy of speech emotion recognition and insufficient fine-grained feature perception in multi-noise environments is solved, and higher robustness and adaptability of emotion recognition are achieved.

CN120356488APending Publication Date: 2025-07-22ULTIMATE IOT (HENAN) TECHNOLOGY LTD +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510693571.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing speech emotion recognition methods have decreased accuracy in multi-noise environments, making it difficult to perceive fine-grained features such as speech speed and rhythm, and the feature fusion method lacks dynamic adaptability to environmental changes.

Method used

By sampling the target voice signal, multiple acoustic features are extracted and initial feature representations are constructed, channel-weighted fusion is combined with environmental signal-to-noise ratio information, input it to the timing modeling network for context modeling, and finally generate emotion recognition results.

Benefits of technology

It improves the recognition accuracy in a multi-noise environment, enhances the perception of fine-grained emotional characteristics such as speech speed and rhythm, and has the ability to adapt to environmental changes during feature fusion, achieving higher robustness and environmental generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356488A_ABST
    Figure CN120356488A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice signal processing, and discloses a voice emotion recognition method, device and equipment and a readable storage medium, and the voice emotion recognition method comprises the steps: obtaining a target voice signal, and executing sampling processing to obtain voice sampling data; extracting a plurality of acoustic features based on the voice sampling data, and constructing initial feature representation; in combination with environment signal-to-noise ratio information, performing channel weighted fusion processing to generate fusion feature representation; inputting the fusion feature representation into a time sequence modeling network, and extracting context information to obtain time sequence abstract features; and executing emotion recognition processing based on the time sequence abstract features to generate a corresponding emotion recognition result. According to the method, the emotion recognition accuracy in a multi-noise environment is improved, the expression ability of fine-grained emotion features such as speech speed changes and rhythm fluctuations is enhanced, the dynamic adaptability to environment changes in the feature fusion process is achieved, and higher recognition stability and environment adaptability are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech signal processing, and particularly to a method, device, equipment and readable storage medium for speech emotion recognition. Background Art

[0002] In the context of the continuous development of speech interaction technology, emotion recognition, as a key component in natural human-computer interaction, is gradually being widely applied in multiple scenarios such as smart home, education assistance, and health monitoring. For example, in a whole-house intelligent system, by detecting the emotional state in the user's speech, the emotional response of the voice assistant and the adaptive scene linkage function can be realized; in children's emotional education, emotion recognition can be used to evaluate children's willingness to express language, psychological state, and emotional characteristics reflected in their behaviors, and assist in teaching intervention and family communication.

[0003] To achieve the above functions, existing speech emotion recognition methods usually use predefined acoustic features (such as MFCC, Prosody, etc.) and combine classification models to judge the emotion labels of speech signals. However, in actual environments with multiple noises, unstructured, and cross-scenarios, the existing methods generally face the following problems:

[0004] (1) In the whole-house intelligent scenario, limited by the computing power of edge devices and the real-time requirements of the system, traditional emotion recognition methods are difficult to balance noise resistance and computational efficiency, and the recognition accuracy significantly decreases under low signal-to-noise ratio conditions;

[0005] (2) In children's speech analysis, existing methods lack the ability to perceive fine-grained features such as speech rate changes, rhythm dynamics, and gender vocal characteristics, and it is difficult to effectively recognize children's unique pronunciation patterns and emotion expression methods;

[0006] (3) Although some methods introduce multiple acoustic features, they do not dynamically adjust the fusion strategy based on environmental changes during the feature fusion stage, resulting in the model being highly sensitive to changes in the input distribution and insufficient generalization ability.

[0007] Therefore, how to improve the multi-dimensional perception ability of emotion changes in a complex environment has become an important technical problem to be solved urgently. Summary of the Invention

[0008] In view of this, the embodiments of the present application provide a method, device, equipment and readable storage medium for speech emotion recognition, which can effectively solve the problems in the prior art such as the decrease in emotion recognition accuracy in a multi-noise environment, the lack of perception ability for fine-grained emotion features such as speech rate and rhythm, and the lack of dynamic adaptability of the feature fusion method to environmental changes.

[0009] In a first aspect, the embodiments of the present application provide a method for speech emotion recognition, including:

[0010] Obtain a target voice signal, and perform sampling processing on the target voice signal to obtain corresponding voice sampling data;

[0011] Based on the voice sampling data, extract multiple acoustic features, and construct the extracted acoustic features into an initial feature representation;

[0012] Based on the initial feature representation and the environmental signal-to-noise ratio information of the target voice signal, perform channel weighted fusion processing to generate a fused feature representation;

[0013] Input the fused feature representation into a temporal modeling network for processing to obtain a temporal abstract feature;

[0014] Based on the temporal abstract feature, perform emotion recognition processing to generate an emotion recognition result corresponding to the target voice signal.

[0015] In some embodiments, the obtaining a target voice signal, and performing sampling processing on the target voice signal to obtain corresponding voice sampling data includes:

[0016] Perform analog-to-digital conversion processing on the target voice signal at a preset sampling rate to generate an audio sample sequence;

[0017] Perform frame division processing and windowing processing on the audio sample sequence in sequence to obtain the voice sampling data.

[0018] In some embodiments, the based on the voice sampling data, extracting multiple acoustic features, and constructing the extracted acoustic features into an initial feature representation includes:

[0019] Perform first spectral feature extraction, second spectral feature extraction, and modulation spectral feature extraction on the voice sampling data respectively to obtain multiple acoustic features covering different frequency band ranges;

[0020] Perform time-scale alignment processing on the multiple acoustic features, and splice the multiple acoustic features after alignment to obtain an initial feature representation.

[0021] In some embodiments, the based on the initial feature representation and the environmental signal-to-noise ratio information of the target voice signal, performing channel weighted fusion processing to generate a fused feature representation includes:

[0022] Extract the environmental signal-to-noise ratio information corresponding to the target voice signal;

[0023] Based on the environmental signal-to-noise ratio information, calculate the weighted parameters of each acoustic feature in the initial feature representation;

[0024] Perform weighted fusion processing on each of the acoustic features according to the weighted parameter to generate the fused feature representation.

[0025] In some embodiments, the performing weighted fusion processing on each of the acoustic features to generate the fused feature representation includes:

[0026] Perform multi-head attention weighted processing on each of the acoustic features respectively to obtain intermediate fusion results generated by multiple attention heads;

[0027] Perform combination processing on the intermediate fusion results in the channel dimension to generate the fused feature representation.

[0028] In some embodiments, the inputting the fused feature representation into a temporal modeling network for processing to obtain temporal abstract features includes:

[0029] Input the fused feature representation into the temporal modeling network including multiple convolutional units;

[0030] In the temporal modeling network, perform depthwise separable convolution processing and dilated convolution processing on the fused feature representation in sequence to extract context information across the time dimension;

[0031] Generate the temporal abstract features based on the context information.

[0032] In some embodiments, the performing emotion recognition processing based on the temporal abstract features to generate an emotion recognition result corresponding to the target speech signal includes:

[0033] Input the temporal abstract features into an emotion classification network, perform fully connected mapping processing and probability normalization processing to obtain prediction probabilities for multiple emotion categories;

[0034] Determine the target emotion category based on the prediction probabilities as the emotion recognition result corresponding to the target speech signal.

[0035] In a second aspect, an embodiment of the present application provides a speech emotion recognition device, including:

[0036] A sampling device, configured to acquire a target speech signal and perform sampling processing on the target speech signal to obtain corresponding speech sampling data;

[0037] A feature extraction device, configured to extract multiple acoustic features based on the speech sampling data and construct the extracted acoustic features into an initial feature representation;

[0038] A fusion device, configured to perform channel weighted fusion processing based on the initial feature representation and the environmental signal-to-noise ratio information of the target speech signal to generate a fused feature representation;

[0039] A feature processing device, configured to input the fused feature representation into a temporal modeling network, perform context modeling processing, and obtain a temporal abstract feature;

[0040] An emotion recognition device, configured to perform emotion recognition processing based on the temporal abstract feature, and generate an emotion recognition result corresponding to the target voice signal.

[0041] In a third aspect, an embodiment of the present application provides a terminal device, which includes a processor and a memory. The memory stores a computer program, and the processor is configured to execute the computer program to implement the voice emotion recognition method in the first aspect above.

[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. When the computer program is executed on a processor, the voice emotion recognition method in the first aspect above is implemented.

[0043] The embodiments of the present application have the following beneficial effects:

[0044] By performing sampling processing on the target voice signal, the present application generates voice sampling data with a clear frame-level structure and time alignment, which helps to improve the stability of downstream feature extraction; by extracting acoustic features covering multiple frequency bands and modulation intervals and constructing an initial feature representation, the perception ability of diverse emotion signals such as pronunciation rhythm and speech rate changes is enhanced; by combining the environmental signal-to-noise ratio information of the target voice signal, calculating the weighted parameters of each channel and performing channel weighted fusion processing, the center of gravity of feature expression can be dynamically adjusted when the noise level changes, enhancing the robustness of the model in different noise scenarios; by inputting the fused feature into a temporal modeling network including depthwise separable convolution and dilated convolution structures, extracting context dependencies, and improving the consistency and semantic integrity of emotion expression in the time dimension; finally, performing emotion recognition processing based on the modeled features, an emotion classification result corresponding to the input target voice signal can be obtained, realizing a stable and highly adaptable emotion recognition ability.

[0045] The voice emotion recognition method of the present application can significantly improve the recognition accuracy in a multi-noise environment, enhance the modeling ability of fine-grained emotion features such as speech rate and rhythm, and have the ability to adapt to environmental changes during feature fusion, thereby realizing an emotion recognition effect with higher robustness and environmental generalization ability. Description of the Drawings

[0046] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0047] Figure 1 It shows a flowchart of a voice emotion recognition method in an embodiment of the present application;

[0048] Figure 2 It shows another flowchart of a voice emotion recognition method in an embodiment of the present application;

[0049] Figure 3 It shows a schematic structural diagram of a voice emotion recognition method in an embodiment of the present application. Detailed implementation manners

[0050] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0051] Generally, the components of the embodiments of the present application described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0052] In the following text, the terms "including", "having" and their cognates that can be used in various embodiments of the present application are only intended to represent specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be understood as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or increasing the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0053] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which various embodiments of the present application pertain. The terms (such as those defined in a general use dictionary) will be interpreted to have the same meaning as their contextual meaning in the relevant technical field and will not be interpreted to have an idealized meaning or an overly formal meaning, unless clearly defined in various embodiments of the present application.

[0054] The following will, with reference to the accompanying drawings, elaborate on some embodiments of the present application. Without conflict, the following embodiments and the features in the embodiments may be combined with each other.

[0055] Considering that in the prior art, the accuracy of emotion recognition decreases in a multi-noise environment, there is insufficient perception of fine-grained features such as speech rate and rhythm, and the feature fusion method lacks dynamic adaptability to environmental changes, a speech emotion recognition method is proposed. By performing sampling processing on the target speech signal, frame-level speech sampling data is obtained; multiple acoustic features are extracted in parallel based on the speech sampling data and concatenated in the channel dimension to form an initial feature representation; the weights of each channel are calculated using the environmental signal-to-noise ratio information, and channel weighted fusion is performed on the initial feature representation to generate a fused feature representation; the fused feature representation is fed into a lightweight temporal modeling network, and context modeling processing is performed in sequence to extract cross-time context dependencies and obtain temporal abstract features; full connection mapping and probability normalization processing are performed based on the temporal abstract features to determine the final emotion category, and the emotion recognition of the input speech is completed.

[0056] The following will illustrate the speech emotion recognition method with some specific embodiments.

[0057] Figure 1 A flowchart of the speech emotion recognition method according to an embodiment of the present application is shown. Exemplarily, the speech emotion recognition method includes the following steps:

[0058] Step S100, obtain a target speech signal and perform sampling processing on the target speech signal to obtain corresponding speech sampling data.

[0059] Among them, the speech sampling data refers to converting the analog target speech signal into a discrete frame-level digital sequence through sampling, segmentation, window function processing, etc. This sequence retains the short-time variation characteristics of the speech and reflects information such as speech rate, intonation, and energy distribution.

[0060] Exemplarily, it is possible to generate speech sampling data with a continuous frame structure by receiving the user's real-time voice input and performing processing such as format conversion and structure division according to a set audio sampling process. For example, in the scenario of full-house intelligent voice control, natural voice commands of the user can be collected and frame-level voice data can be constructed in real time; in a voice emotion analysis platform, voice segments can be extracted from call records or voice materials and standardized to support subsequent emotion recognition operations.

[0061] In an alternative embodiment, step S100 includes the following sub-steps:

[0062] S101, perform analog-to-digital conversion processing on the target voice signal at a preset sampling rate to generate an audio sample sequence.

[0063] Among them, analog-to-digital conversion means uniformly sampling the target voice signal at a set time interval to form a sequence of equally spaced digital sample points. The sampling rate determines the time resolution and frequency upper limit of the samples.

[0064] Specifically, when the sampling rate is set to 16 kHz, the energy information in the 0 to 8 kHz frequency band of the analog signal can be retained in digital form, and a continuously arranged audio sample data sequence can be obtained after the analog-to-digital conversion is completed. For example, in intelligent voice terminals or in-vehicle voice interaction devices, a 16 kHz sampling standard is often adopted to cover the main frequency band of human voices and balance computational efficiency and expression integrity.

[0065] S102, perform frame division processing and windowing processing on the audio sample sequence in sequence to obtain speech sampling data.

[0066] Among them, frame division processing means segmenting the audio sample sequence according to a set frame length and frame shift so that each segment of audio has a finite time length; windowing processing means applying a window function operation to each frame to suppress the energy mutation at the frame edge.

[0067] Specifically, the audio sample sequence is segmented according to a frame length of 400 points and a frame shift of 160 points. Each frame of sample data forms frame-level data with stable structure and continuous amplitude after windowing processing, and finally forms speech sampling data. For example, in speech expressions with frequent changes in speech rate or significant changes in intonation, frame division and windowing operations can retain sufficient time locality information, which helps to identify emotion-related dynamics in the subsequent feature extraction stage.

[0068] Step S200, based on the speech sampling data, extract multiple acoustic features and construct the extracted acoustic features into an initial feature representation.

[0069] Among them, the acoustic feature refers to a set of numerical vectors extracted from speech sampling data through a specific time-frequency domain transformation method, which is used to characterize the variation characteristics of speech in different dimensions such as different frequency bands and modulation rhythms; the initial feature representation refers to a structured feature set formed by combining multiple acoustic features in the channel dimension.

[0070] Exemplarily, the speech sampling data is organized in frames, and each frame of data is used as the input for feature extraction. Through feature extraction processing in multiple frequency band ranges, feature expression forms in different dimensions are obtained respectively, and combined feature representations are constructed in a specified order to form the initial feature representation. For example, in a whole-house intelligent system, the voice control commands issued by users are usually relatively short, but there are significant individual differences in speech rate and intonation changes. By extracting multiple acoustic features, their pronunciation habits and emotional states can be captured more comprehensively; in the scenario of human voice emotion analysis, emotional fluctuations are often reflected in the energy changes or rhythm fluctuations in specific frequency bands. Constructing a multi-dimensional initial feature representation helps to enhance the model's perception ability of fine-grained changes.

[0071] In an alternative embodiment, as Figure 2 shown, step S200 includes the following sub-steps:

[0072] S201, respectively perform first spectral feature extraction, second spectral feature extraction, and modulation spectral feature extraction on the speech sampling data to obtain multiple acoustic features covering different frequency band ranges.

[0073] Among them, the first spectral feature, the second spectral feature, and the modulation spectral feature respectively represent different-dimensional frequency domain description information extracted from the speech sampling data. There are differences in the frequency band range covered, frequency resolution, and processing method among the three types of features.

[0074] Specifically, the first spectral feature is the Mel spectrum, covering a frequency range of 0–8 kHz; the second spectral feature is GFCC, covering a frequency range of 0–4 kHz; the modulation spectral feature is used to extract speech rhythm change information, covering a range of 0–16 Hz. The above feature extraction takes the speech sampling data as the input, and the numerical expressions of the three acoustic features are obtained through frequency domain transformation and cepstrum calculation.

[0075] S202, perform time-scale alignment processing on the multiple acoustic features, and splice the multiple acoustic features after alignment to obtain the initial feature representation.

[0076] Among them, the time-scale alignment processing refers to adjusting the number of time frames of different acoustic features so that the frame length and frame shift in the time dimension are consistent. The splicing processing refers to connecting the aligned multiple features in the channel dimension to construct a unified data expression structure.

[0077] Specifically, upsampling or downsampling operations are respectively performed on three acoustic features, namely Mel spectrogram, GFCC, and modulation spectrogram, to make their number of frames consistent in the time dimension. After alignment, the Mel features (with 64 channels), GFCC features (with 20 channels), and modulation spectrogram features (with 21 channels) are concatenated in sequence in the channel dimension to form an initial feature representation with 105 channels.

[0078] For example, in a voice interaction task in a multi-noise environment, since some feature dimensions may be greatly affected by local noise interference, maintaining channel consistency and frame-level synchronization can ensure feature alignment and effectiveness during subsequent weighted fusion, thereby enhancing the overall system's processing robustness to unstable inputs.

[0079] Step S300: Based on the initial feature representation and the environmental signal-to-noise ratio information of the target voice signal, perform channel weighted fusion processing to generate a fused feature representation.

[0080] Among them, channel weighted fusion processing refers to dynamically adjusting the relative weights of different acoustic feature dimensions in the initial feature representation according to the signal-to-noise ratio level in the environment where the target voice signal is located, so as to obtain a fused feature representation that adapts to the noise environment.

[0081] Specifically, extract the features of each channel from the initial feature representation, perform multiplication operations according to the corresponding weighting parameters, and then perform summation or combination operations in the channel dimension to finally form a fused feature representation. For example, in a whole-house intelligent control system, the voice input environment often has interference such as background TV sound and overlapping human voices. At this time, adjusting the channel weights through a signal-to-noise ratio perception mechanism can enhance the prominent expression of key emotion channels; in the scenario of voice emotion analysis, identifying the true emotion state of users in a complex conversation environment also relies on this channel adaptation mechanism.

[0082] In an optional embodiment, step S300 includes the following sub-steps:

[0083] S301: Extract the environmental signal-to-noise ratio information corresponding to the target voice signal.

[0084] Among them, the environmental signal-to-noise ratio information is a numerical index that measures the intensity ratio of the voice component to the background noise component in the target voice signal, reflecting the noise level of the current voice environment. The signal-to-noise ratio is usually expressed in dB (decibels), and the larger the value, the clearer the voice.

[0085] Specifically, it can be estimated based on the difference between the background energy and the foreground speech energy recorded during the speech acquisition process, or calculated according to the noise evaluation data provided by external sensors. The extraction result is used as the input for the subsequent calculation of the channel weighting parameters. For example, in the scenario of a home multi-microphone array, the signal-to-noise ratio can be estimated in real time by combining the sound source localization and the far-field speech enhancement module, which is used to drive the generation of the weighting strategy for the acoustic feature channels.

[0086] S302. Calculate the weighting parameters for each acoustic feature in the initial feature representation based on the environmental signal-to-noise ratio information.

[0087] Among them, the weighting parameter refers to the floating-point value calculated for each channel in the initial feature representation, which is used to represent the importance of the channel under the current signal-to-noise conditions. The parameter calculation takes the signal-to-noise ratio as the input variable, and calculates independent weight coefficients for each channel through an expression containing trainable parameters.

[0088] Specifically, the channel weighting parameter is calculated using the following formula:

[0089]

[0090] where: w i : the weighting coefficient of the i-th channel; SNR: environmental signal-to-noise ratio; α i , β i : the weight and bias of the i-th channel; τ: temperature parameter, whose value is adaptively adjusted according to the SNR, and the calculation method is:

[0091] τ = max(0.5, 1 - 0.05·SNR)

[0092] The above formula realizes the dynamic adjustment of the feature importance: making the weight distribution more uniform when the environmental noise is large, and enhancing the response of the important channels when the signal-to-noise ratio is high, which is suitable for robust modeling in variable scenarios.

[0093] S303. Perform weighted fusion processing on each acoustic feature according to the weighting parameters to generate a fused feature representation.

[0094] Among them, the weighted fusion processing refers to applying the channel weighting parameters to each channel of the initial feature representation according to the corresponding relationship, scaling the values of each channel, so as to complete the numerical reconstruction in the channel dimension. The fused feature representation refers to the feature set formed after the weighted processing, which has the same structure as the initial feature representation, but the values have been redistributed according to the current environmental conditions.

[0095] Specifically, weighted parameters are applied to the initial feature representation channel by channel to complete the per-channel weighting operation. All frame data of each channel are scaled according to the corresponding weight coefficients, and the weighted frame data are combined to form a new three-dimensional feature tensor, which is the fused feature representation. For example, in scenarios where the call quality is unstable or in far-field voice pickup, this channel weighting operation can reduce the interference of low-quality channels on the overall model performance and improve the robustness of the fused representation.

[0096] In an alternative embodiment:

[0097] Multi-head attention weighting processing is performed on each acoustic feature respectively to obtain intermediate fusion results generated by multiple attention heads. A combination process is performed on the intermediate fusion results in the channel dimension to generate the fused feature representation.

[0098] Among them, multi-head attention weighting processing refers to introducing multiple attention heads on the basis of calculating weighted parameters. Each attention head focuses on different acoustic features with different weight biases, enhancing the flexibility of fusion and the ability to capture information.

[0099] Specifically, 4 attention heads are used for channel weighted fusion, and the partitioning method is as follows: the first attention head focuses on the first 32 channels of the Mel feature; the second attention head focuses on the last 32 channels of the Mel feature; the third attention head focuses on all channels (20 channels) of the GFCC feature; the fourth attention head focuses on all channels (21 channels) of the modulation spectrum feature. Each attention head calculates local weighted parameters based on its corresponding channel segment and performs channel weighting processing on the feature segment of this segment to generate an intermediate fusion result. The intermediate results generated by multiple attention heads perform a concatenation operation in the channel dimension to form the final fused feature representation. For example, in the face of a speech emotion analysis task, this structure can make specific emotion-related features (such as low-frequency modulation spectrum) receive higher attention, thereby improving the model's ability to distinguish subtle emotion states such as "sadness, tension".

[0100] Step S400, input the fused feature representation into the temporal modeling network for processing to obtain the temporal abstract feature.

[0101] Among them, the temporal modeling network refers to a set of structures that perform context-dependent modeling on the input features in the time dimension and can capture the association patterns between the current frame and adjacent frames. The temporal abstract feature refers to the high-level expression with temporal context semantics formed after modeling, which is used for further emotion discrimination. For example, in a speech emotion analysis task, emotion states such as "anxiety" and "depression" are often expressed as multi-frame joint features, and their changes are not limited to a single frame; through context modeling, the recognition ability for such slow-changing and trailing emotions can be enhanced.

[0102] Specifically, the fused feature representation is input into a modeling unit composed of multiple convolutional structures. On the premise of ensuring the consistent time order, the time dependence of the target speech signal is modeled through time-domain convolution operations. In the interaction scenario of an intelligent speech terminal, this processing can help identify long-range behavioral change features such as "continuous instructions" and "semantic transitions", improving the response accuracy.

[0103] In an alternative embodiment, step S400 includes the following sub-steps:

[0104] S401, input the fused feature representation into a temporal modeling network including multiple convolutional units.

[0105] Among them, a convolutional unit refers to the convolution kernel operation applied at a specific time step, which can extract the local change information of features on the time axis through a sliding window method. Multiple convolutional units are stacked in sequence to form a temporal modeling network for extracting hierarchical time features.

[0106] Specifically, the input fused feature representation is a three-dimensional tensor, including time frames, channel dimensions, and frequency features. In each convolutional unit, a one-dimensional convolution operation is performed with the time frame as the convolution dimension to extract feature dynamics within a local time period. The multi-layer convolutional structure can gradually expand the time receptive field and enhance the context capture ability. For example, in a speech segment where the user has a fast speaking speed or continuous emotional fluctuations, the multi-layer convolution can integrate change clues at different time scales layer by layer, making the final recognition process more stable.

[0107] S402, in the temporal modeling network, perform depthwise separable convolution processing and dilated convolution processing on the fused feature representation in sequence to extract cross-time-dimensional context information.

[0108] Among them, the depthwise separable convolution is divided into two-stage operations:

[0109] The depthwise convolution is used to independently extract features for the time evolution of each channel, and the pointwise convolution is used to fuse information between channels; the dilated convolution expands the receptive field by inserting intervals between convolutional kernel elements to capture longer-range dependencies.

[0110] For example, assume the number of input feature channels is C in , and the number of output channels is C out , and the width of the convolutional kernel is K. Then the number of parameters of the standard one-dimensional convolution is: C in ×C out ×K; while the number of parameters of the depthwise separable convolution is only: C in ×K + C in ×C out; It can be seen that the number of parameters of the depthwise separable convolution is about one eighth of the number of parameters of the standard one-dimensional convolution, which significantly reduces the model size.

[0111] Subsequently, the dilated convolution introduces gaps in the first, second, and fourth dimensions with dilation coefficients of [1,2,4][1,2,4][1,2,4], effectively expanding the network's ability to capture time dependencies, such as being able to recognize the dynamic trend of gradually increasing speech frequency during sustained anger, while maintaining computational efficiency comparable to that of depthwise separable convolution.

[0112] It can be understood that the above two convolution processes act on the feature sequence after channel weighted fusion in turn, generating an intermediate result with both local time domain features and long-range dependency information, which provides an efficient and robust basis for the subsequent generation of temporal abstract features. In complex scenarios such as multi-speaker interaction and long-distance speech recognition, this structure can improve the model's sensitivity to context changes and speaking rhythm.

[0113] S403: Generate temporal abstract features based on the context information.

[0114] Among them, temporal abstract features refer to high-level feature expressions obtained through modeling operations that integrate local and long-range dependencies in the temporal dimension. This feature is no longer limited to the current frame information, but reflects the comprehensive dynamic performance within the context.

[0115] Specifically, after the convolution process is completed, the output three-dimensional tensor has encoded the information of the previous and next frames in each frame. This feature tensor is a temporal abstract feature, whose structural dimension is consistent with the number of time frames, and the numerical expression has integrated the semantic information between multiple time points. For example, in the joint modeling scenario of voice command recognition and emotion classification, the temporal abstract feature can be used as a shared base layer for multi-task learning, while supporting the independent decoding of emotion labels and semantic commands.

[0116] Step S500, performing emotion recognition processing based on the temporal abstract features to generate an emotion recognition result corresponding to the target speech signal.

[0117] Among them, emotion recognition processing refers to performing classification reasoning on the input feature representation through the emotion classification network to generate the emotion category corresponding to the current target voice signal. The emotion recognition result refers to the target emotion label determined by the model, which reflects the emotional state of the current voice. The types include happy, angry, sad, neutral and other categories. For example, in the whole-house intelligent voice system, identifying the user's current emotional state can be used to drive the response strategy adjustment of the voice assistant; and in emotion monitoring applications, emotion recognition results can be used as a key input for psychological state assessment.

[0118] Specifically, the emotion recognition process takes the temporal abstract features as input, first performs feature dimension compression and classification mapping processing, and then outputs the predicted probability values of each emotion category through probability normalization operations, and determines the final classification result accordingly.

[0119] In an alternative embodiment, step S500 includes the following sub-steps:

[0120] S501, input the temporal abstract features into the emotion classification network, perform fully connected mapping processing and probability normalization processing, and obtain the predicted probabilities of multiple emotion categories.

[0121] Among them, the emotion classification network refers to a neural structure with discriminant functions that can map input features to a fixed category set. The fully connected mapping processing refers to inputting the temporal feature vector into the fully connected layer to perform mapping from the feature dimension to the category dimension; the probability normalization processing refers to performing a normalization operation on the mapping result and outputting the probability value corresponding to each category.

[0122] Specifically, the last frame of the temporal abstract features is used as input and enters the fully connected layer for dimension mapping, and the output dimension is equal to the number of emotion categories. Subsequently, the output vector is normalized into a probability distribution through the softmax function to obtain the predicted probability of each category, reflecting the relative possibility that the speech signal belongs to various emotion states.

[0123] S502, based on the predicted probabilities, determine the target emotion category as the emotion recognition result corresponding to the target speech signal.

[0124] Among them, the predicted probability refers to the multi-category probability distribution output by the emotion classification network; the target emotion category refers to the emotion type corresponding to the maximum probability value in this distribution, indicating the final recognition and judgment of the input speech.

[0125] Specifically, perform a maximum value search operation on the softmax output result to obtain the corresponding category index, and determine the discrete emotion label corresponding to this index, such as "happy", "angry", "sad", "joyful", etc., as the final emotion recognition result. For example, in the analysis of children's speech expressions, this result can assist the education system in judging the changes in children's expression intentions, thereby optimizing the curriculum arrangement or family intervention strategies.

[0126] Figure 3 Fig. shows a schematic structural diagram of a speech emotion recognition device according to an embodiment of the present application. Exemplarily, the speech emotion recognition device 100 includes:

[0127] A sampling device 101, configured to acquire a target speech signal and perform sampling processing on the target speech signal to obtain corresponding speech sampling data;

[0128] A feature extraction device 102, configured to extract multiple acoustic features based on the speech sampling data, and construct the extracted acoustic features into an initial feature representation;

[0129] A fusion device 103, configured to perform channel weighted fusion processing based on the initial feature representation and the environmental signal-to-noise ratio information of the target speech signal, and generate a fusion feature representation;

[0130] A feature processing device 104, configured to input the fusion feature representation into a temporal modeling network, perform context modeling processing, and obtain a temporal abstract feature;

[0131] An emotion recognition device 105, configured to perform emotion recognition processing based on the temporal abstract feature, and generate an emotion recognition result corresponding to the target speech signal.

[0132] It can be understood that the device in this embodiment corresponds to the method in the above embodiment, and the optional items in the above embodiment are also applicable to this embodiment, so they will not be repeated here.

[0133] This application also provides a terminal device. Exemplarily, the terminal device includes a processor and a memory. Among them, the memory stores a computer program, and the processor runs the computer program to enable the terminal device to execute the functions of each module in the above method or the above device.

[0134] Among them, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.

[0135] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electric Erasable Programmable Read-Only Memory (EEPROM), etc. Among them, the memory is used to store a computer program, and after receiving an execution instruction, the processor can execute the computer program accordingly.

[0136] This application also provides a computer-readable storage medium for storing the computer program used in the above terminal device. For example, the computer-readable storage medium can include, but is not limited to: various media such as USB flash drives, mobile hard disks, Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disks, or optical discs that can store program codes.

[0137] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the structure diagram and / or flowchart, as well as the combination of blocks in the structure diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0138] In addition, in each embodiment of this application, the various functional modules or units can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0139] When the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application.

[0140] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application.

Claims

1. A method for speech emotion recognition, characterized in that, The method includes: Obtaining a target voice signal, and performing sampling processing on the target voice signal to obtain corresponding voice sampling data; Based on the voice sampling data, extracting multiple acoustic features, and constructing the extracted acoustic features into an initial feature representation; Based on the initial feature representation and the environmental signal-to-noise ratio information of the target voice signal, performing channel weighted fusion processing to generate a fusion feature representation; Inputting the fusion feature representation into a temporal modeling network for processing to obtain temporal abstract features; Based on the temporal abstract features, performing emotion recognition processing to generate an emotion recognition result corresponding to the target voice signal.

2. The speech emotion recognition method according to claim 1, wherein The obtaining of the target voice signal, and performing sampling processing on the target voice signal to obtain corresponding voice sampling data includes: Performing analog-to-digital conversion processing on the target voice signal at a preset sampling rate to generate an audio sample sequence; Successively performing frame division processing and windowing processing on the audio sample sequence to obtain the voice sampling data.

3. The voice emotion recognition method according to claim 1, wherein The based on the voice sampling data, extracting multiple acoustic features, and constructing the extracted acoustic features into an initial feature representation includes: Respectively performing first spectral feature extraction, second spectral feature extraction, and modulation spectral feature extraction on the voice sampling data to obtain multiple acoustic features covering different frequency band ranges; Performing time scale alignment processing on the multiple acoustic features, and splicing the multiple acoustic features after alignment to obtain an initial feature representation.

4. The voice emotion recognition method according to claim 1, characterized in that The based on the initial feature representation and the environmental signal-to-noise ratio information of the target voice signal, performing channel weighted fusion processing to generate a fusion feature representation includes: Extracting the environmental signal-to-noise ratio information corresponding to the target voice signal; Based on the environmental signal-to-noise ratio information, calculating the weighting parameters of each acoustic feature in the initial feature representation; According to the weighting parameters, performing weighted fusion processing on each acoustic feature to generate the fusion feature representation.

5. The voice emotion recognition method according to claim 4, wherein The performing weighted fusion processing on each acoustic feature to generate the fusion feature representation includes: Performing multi-head attention weighting processing on each acoustic feature respectively to obtain intermediate fusion results generated by multiple attention heads; Performing combination processing on the intermediate fusion results in the channel dimension to generate the fusion feature representation.

6. The voice emotion recognition method according to claim 1, wherein The inputting the fusion feature representation into a temporal modeling network for processing to obtain temporal abstract features includes: Inputting the fusion feature representation into the temporal modeling network including multiple convolutional units; In the temporal modeling network, successively performing depthwise separable convolution processing and dilated convolution processing on the fusion feature representation to extract context information across the time dimension; Based on the context information, generating the temporal abstract features.

7. The voice emotion recognition method according to claim 1, characterized in that The based on the temporal abstract features, performing emotion recognition processing to generate an emotion recognition result corresponding to the target voice signal includes: Inputting the temporal abstract features into an emotion classification network, performing fully connected mapping processing and probability normalization processing to obtain prediction probabilities of multiple emotion categories; Based on the predicted probability, determine the target emotion category as the emotion recognition result corresponding to the target speech signal.

8. A voice emotion recognition device, characterized in that, Including: A sampling device for acquiring a target speech signal and performing sampling processing on the target speech signal to obtain corresponding speech sampling data; A feature extraction device for extracting multiple acoustic features based on the speech sampling data and constructing the extracted acoustic features into an initial feature representation; A fusion device for performing channel weighted fusion processing based on the initial feature representation and the environmental signal-to-noise ratio information of the target speech signal to generate a fusion feature representation; A feature processing device for inputting the fusion feature representation into a temporal modeling network to perform context modeling processing to obtain temporal abstract features; An emotion recognition device for performing emotion recognition processing based on the temporal abstract features to generate an emotion recognition result corresponding to the target speech signal.

9. A terminal device, characterized in that, The terminal device includes a processor and a memory, the memory stores a computer program, and the processor is configured to execute the computer program to implement the speech emotion recognition method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed on a processor, it implements the speech emotion recognition method according to any one of claims 1-7.

Citation Information

Cited By

  • Teaching effect evaluation system and teaching speech emotion recognition method based on artificial intelligence

    CN120766726A

  • An artificial intelligence-based teaching effect evaluation system and a teaching speech emotion recognition method

    CN120766726B

  • Data processing method and electronic equipment

    CN120951251A

  • Speech recognition method and device and vehicle

    CN122050376A