Mobile terminal emotion recognition acceleration method and system based on signal compression

By performing voice signal feature extraction and Mel spectrogram compression on mobile devices, and multimodal deep learning on server side, the computing and storage problems of deep neural network deployment on mobile devices are solved, achieving efficient emotion recognition and accuracy improvement.

CN120388585APending Publication Date: 2025-07-29SHANGHAI JIAOTONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510562237.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing deep neural network models are too complex and storage requirements when deployed on mobile devices, making it difficult to meet the computing and storage requirements of real-time sentiment recognition, and the recognition accuracy is limited by the knowledge extracted within the dataset.

Method used

By performing voice signal feature extraction and Mel spectrogram compression on mobile devices, the storage needs are reduced using Fbank encoder and singular value decomposition (SVD), and multimodal deep learning is performed on the server side, combining the collaborative attention module to fuse speech and text information for emotion recognition.

Benefits of technology

It reduces the computing and storage burden of mobile devices, improves the accuracy and efficiency of emotional recognition, adapts to the environmental needs of different bandwidths and computing resources, and improves the recognition accuracy and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388585A_ABST
    Figure CN120388585A_ABST
Patent Text Reader

Abstract

The invention relates to a mobile terminal emotion recognition acceleration method and system based on signal compression, and the method comprises the following steps: collecting a voice signal through a mobile terminal device, carrying out the feature extraction, generating a Mel spectrogram through an Fbank encoder, and carrying out the compression to obtain a compression feature; the mobile terminal device sends the compression feature to a server; the server restores the compression features to obtain a restored Mel spectrogram as voice modal input, and performs text recognition on the restored Mel spectrogram to obtain text information as text modal input; the server inputs the voice modal input and the text modal input into a multi-modal deep learning network at the same time for emotion recognition, and the multi-modal deep learning network achieves fusion of the voice modal input and the text modal input through a collaborative attention module. Compared with the prior art, the method has the advantages of high calculation and storage efficiency, low delay, accurate recognition result and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal emotion recognition, and in particular, to a method and system for accelerating mobile emotion recognition based on signal compression. Background Art

[0002] As the most commonly used interaction form in daily life, speech is widely used in smartphones, speakers, and other Internet of Things devices. Different from text information, speech signals contain richer user information, including emotions, gender, etc. In particular, emotion recognition can help intelligent speech systems provide more user-friendly services and feedback. For example, intelligent customer service can select appropriate words according to the customer's emotion, or conduct business evaluation according to the customer's emotion.

[0003] Although the current rapid development of deep learning has promoted the performance improvement of various speech-related applications. For example, CN118072718A discloses a method, device, equipment, and storage medium for constructing a speech emotion classification model. The method includes: processing the original audio into a spectrogram image, and obtaining the Mel frequency cepstral coefficient features of the original audio; inputting the spectrogram image into an image classification model to obtain high-level image features, inputting the Mel frequency cepstral coefficient features into a Mel frequency cepstral coefficient extraction model to obtain high-level Mel frequency cepstral coefficients; obtaining the key features of the image after interaction through a first attention module; obtaining the key features of the Mel frequency cepstral coefficient after interaction through a second attention module; splicing the key features of the image after interaction and the key features of the Mel frequency cepstral coefficient after interaction into a comprehensive feature, obtaining an emotion classification result through the comprehensive feature, and calculating the loss to update the parameters, so as to obtain a speech emotion classification model composed of an image classification model and a Mel frequency cepstral coefficient extraction model. However, the complex structure and huge parameters of deep learning models pose requirements for computing and storage, and this method does not consider this point, only focusing on recognition accuracy, without taking into account computing efficiency and storage requirements. Therefore, the deployment of these existing technologies on mobile devices (such as smartphones and speakers) is limited in terms of computing power, energy consumption, and even heat dissipation. In addition, interactive application scenarios are sensitive to time delay, and the limited processing power of mobile devices is difficult to meet the delay requirements of applications. Therefore, the deployment of a real-time emotion recognition system based on deep neural networks on mobile devices is an urgent challenge to be solved.

[0004] Researchers have explored methods for deploying deep neural networks on mobile devices to reduce the computational complexity of the model through techniques such as branch pruning, weight sharing, tensor quantization, and knowledge distillation. However, these methods inevitably lead to a reduction in model accuracy. At the same time, researchers use technologies such as GPUs, FPGAs, and ASICs to improve the computing power of devices. However, due to size and power consumption limitations, the deployment of these technologies on existing mobile devices is difficult.

[0005] Implementing real-time voice application deployment on mobile devices faces a series of challenges. First, the computing power of mobile devices is limited, and existing deep neural networks usually require complex operations and are difficult to directly deploy. Second, the storage capacity of Internet of Things devices such as smart speakers is limited, making it difficult to store long-term audio signals and a large number of model parameters. Finally, existing emotion recognition models can only extract knowledge within the dataset, limiting the further improvement of emotion recognition accuracy. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and system for accelerating mobile terminal emotion recognition based on signal compression. By task offloading, the deployment of the deep neural network is migrated from the mobile terminal device to the server, reducing the computing and storage burden of the mobile device, and compressing the original voice signal to further reduce the storage space requirements. Through multi-modal fusion, the accuracy of emotion recognition is improved while ensuring the recognition efficiency.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] A method for accelerating mobile terminal emotion recognition based on signal compression includes the following steps:

[0009] Collect voice signals using a mobile device and perform feature extraction, generate a Mel spectrogram using an Fbank encoder, and compress the Mel spectrogram to obtain compressed features;

[0010] The mobile device sends the compressed features to the server;

[0011] The server restores the compressed features to obtain a restored Mel spectrogram as the voice modality input, and performs text recognition on the restored Mel spectrogram to obtain text information as the text modality input;

[0012] The server inputs the voice modality input and the text modality input into a multi-modal deep learning network for emotion recognition at the same time. Among them, the multi-modal deep learning network uses a collaborative attention module to realize the fusion of the voice modality input and the text modality input.

[0013] The feature extraction includes the following steps:

[0014] Perform pre-emphasis and frame segmentation operations on the collected voice signals;

[0015] Divide the variable-length voice signal into fixed-length segments, with a preset overlap between frames, and use a Hamming window to smooth the voice signal;

[0016] Use the fast Fourier transform to convert the voice signal to the frequency domain and calculate the energy spectrum of the signal;

[0017] Based on the energy spectrum, use an Fbank encoder to generate a Mel spectrogram.

[0018] The compression of the Mel spectrogram is performed using the singular value decomposition method.

[0019] The singular value decomposition method includes the following steps:

[0020] For the Mel spectrogram f, which is an 80×T matrix, where 80 is the feature matrix of the Mel spectrogram and T is a quantity related to the time length of the original speech signal, perform the Mel spectrogram decomposition as follows:

[0021] f = Udiag(S)V H

[0022] where U is an 80×k matrix, S is a vector of length k, V is a T×k matrix, and V H is its transpose matrix, and k = min(80, T);

[0023] Take the first γk bits of the vector S to compress S to get S new , while keeping the size of f unchanged, compress U to U new = an 80×γk matrix, compress V to V new = a T×γk matrix to obtain the compressed features, where γ∈[0, 1] is the compression ratio.

[0024] The server processes the compressed features through an SVD inverse transformation to obtain the restored Mel spectrogram f new :

[0025] The multi-modal deep learning network includes:

[0026] A text feature embedding module for extracting features from the text modality input to obtain text features;

[0027] A speech feature extraction module for further extracting emotional features from the speech modality input using a CNN block and concatenating the context using an LSTM block to obtain speech features;

[0028] A collaborative attention module for realizing deep interaction between speech and text features and outputting fused features by utilizing the complementarity between text features and speech features;

[0029] An output module for processing the fused features output by the collaborative attention module using a self-attention mechanism and outputting the predicted value of emotion classification through a linear layer.

[0030] The text feature embedding module uses RoBERTa to convert the text modality input into a tensor embedding.

[0031] The described collaborative attention module adopts an encoder-decoder architecture, stacking multiple attention modules to achieve deep feature interaction.

[0032] The described collaborative attention module performs the following steps:

[0033] Perform self-attention calculation on the input of a single modality to extract the deep features of the corresponding modality itself;

[0034] Apply guided attention, using the features of one modality as a guide to align the features of the two modalities in terms of time sequence and semantics;

[0035] Fuse the features extracted by self-attention and guided attention through concatenation to obtain fused features.

[0036] A mobile emotion recognition acceleration system based on signal compression includes a mobile device and a server. Among them, the mobile device collects voice signals and performs feature extraction, generates a Mel spectrogram using an Fbank encoder, compresses the Mel spectrogram to obtain compressed features, and sends the compressed features to the server; the server restores the compressed features to obtain a restored Mel spectrogram as the voice modality input, performs text recognition on the restored Mel spectrogram to obtain text information as the text modality input, and inputs the voice modality input and the text modality input into a multi-modal deep learning network for emotion recognition at the same time. Among them, the multi-modal deep learning network uses a collaborative attention module to achieve the fusion of the voice modality input and the text modality input.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] (1) The present invention introduces Fbank feature extraction and SVD (Singular Value Decomposition) compression, which can perform preliminary voice signal processing and feature extraction on mobile devices. By efficiently compressing the Mel spectrogram, it reduces the storage requirements and data transmission burden, and improves the overall computing and storage efficiency. Using the SVD algorithm to decompose and compress the acoustic spectrogram can effectively remove redundant information, and while maintaining the accuracy of emotion recognition, significantly reduce the computing and storage costs.

[0039] (2) By transmitting the compressed acoustic spectrogram features instead of the original audio data, the present invention reduces the amount of data transmission, reduces the server bandwidth consumption, improves the transmission speed at the same time, and shortens the overall delay of data processing.

[0040] (3) The present invention first performs speech-to-text conversion and uses the RoBERTa model for text embedding to fully combine text and speech information and enhance the speech emotion recognition ability. The co-attention mechanism is introduced to establish a deep connection between the text modality input and the speech modality input, and the linear layer is combined to adjust the feature dimension, optimize the emotion classification performance, and achieve more accurate capture and recognition of emotion information.

[0041] (4) The present invention uses a CNN block to further extract Mel spectrogram features to obtain deeper emotion features, and at the same time uses LSTM to model the context to improve the ability to capture temporal information.

[0042] (5) The present invention adopts an efficient compression strategy and multi-modal deep fusion technology, enabling the system to maintain a high recognition accuracy at different compression rates and adapt to the environmental requirements of different bandwidths and computing resources. Description of the Drawings

[0043] Figure 1 is the flowchart of the method of the present invention;

[0044] Figure 2 is the schematic diagram of the multi-modal deep learning network structure of the present invention;

[0045] Figure 3 is the schematic diagram of the co-attention module structure of the present invention;

[0046] Figure 4 is the schematic diagram of the system structure of the present invention. Detailed Embodiment

[0047] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0048] Embodiment 1

[0049] This embodiment provides a mobile emotion recognition acceleration method based on signal compression, as Figure 1 shown, including the following steps:

[0050] S1. Use a mobile device to collect a speech signal and perform feature extraction, use an Fbank encoder to generate a Mel spectrogram, and compress the Mel spectrogram to obtain compressed features.

[0051] In this embodiment, the feature extraction includes the following steps:

[0052] S101. Perform pre-emphasis and frame segmentation operations on the collected speech signal;

[0053] S102. Divide the variable-length speech signal into fixed-length segments. To avoid signal omission caused by window boundaries, a certain overlap is required between frames. Then, add a window to each frame and use a Hamming window to smooth the speech signal.

[0054] S103. Use the Fast Fourier Transform (FFT) to convert the speech signal to the frequency domain and calculate the energy spectrum of the signal.

[0055] S104. Based on the energy spectrum, use the Fbank encoder to generate a Mel spectrogram.

[0056] The purpose of feature extraction is to obtain the discriminative components in the speech signal, where pre-emphasis is used to enhance high-frequency signals.

[0057] In this embodiment, singular value decomposition (SVD) is used to compress and denoise the Mel spectrogram, and text and emotional information in the features are retained as much as possible. SVD compression helps reduce the storage requirements of data, lower the overhead during transmission, and maintain low-frequency features at a certain compression rate.

[0058] Specifically, for the Mel spectrogram f, which is an 80×T matrix, where 80 is the feature matrix of the Mel spectrogram and T is a quantity related to the time length of the original speech signal, the Mel spectrogram decomposition is as follows:

[0059] f = Udiag(S)V H

[0060] where U is an 80×k matrix, S is a vector of length k, V is a T×k matrix, and V H is its transpose matrix, and k = min(80, T);

[0061] Take the first γk bits of the vector S to compress S to get S new , while keeping the size of f unchanged, compress U to U new = an 80×γk matrix, compress V to V new = a T×γk matrix, and save the matrices U new , V new and the vector S new to obtain the compressed features, where γ ∈ [0, 1] is the compression rate. The compressed features are compressed to γ times of the original Mel spectrogram features.

[0062] S2. The mobile device sends the compressed features to the server.

[0063] S3. The server restores the compressed features to obtain the restored Mel spectrogram as the speech modality input, and performs text recognition on the restored Mel spectrogram to obtain the text information as the text modality input.

[0064] The server processes the compressed features through SVD inverse transformation to obtain the restored Mel spectrogram f new :

[0065] In this embodiment, the restored Mel spectrogram is transcribed into text information through an ASR system.

[0066] As Figure 2 shown, the multi-modal deep learning network includes:

[0067] A text feature embedding module, which is used to extract features from the text modality input to obtain text features;

[0068] A speech feature extraction module, which is used to further extract emotional features from the speech modality input using a CNN block, and concatenate the context using an LSTM block to obtain speech features;

[0069] A co-attention module, which is used to utilize the complementarity between text features and speech features to achieve deep interaction between speech and text features and output fused features;

[0070] An output module, which is used to process the fused features output by the co-attention module using a self-attention mechanism, and then output the predicted value of emotion classification through a linear layer.

[0071] In this embodiment, the text feature embedding module uses RoBERTa to convert the text modality input into a tensor embedding, and then connects a linear layer after it to adjust the dimension of the RoBERTa embedding features for subsequent fusion with audio features.

[0072] The co-attention module is a key component of multi-modal feature fusion, mainly used to achieve deep interaction between speech and text features, thereby improving the accuracy of speech emotion recognition. The main purpose of this module is to utilize the complementarity between speech features and text features, so that the model can more comprehensively understand the user's emotional information. In this embodiment, the co-attention module adopts an encoder-decoder architecture, and stacks multiple attention modules to achieve deep feature interaction. As Figure 3 shown, it performs the following steps:

[0073] A1. Perform self-attention calculation on the input of a single modality to extract the deep features of the corresponding modality itself; this step ensures that each modality can self-understand the internal temporal and structural information and lays a foundation for subsequent interaction.

[0074] A2. Apply Guided-Attention, using the features of one modality as guidance to align the features of the two modalities both temporally and semantically. For example, the text modality can be used as guidance to help the speech features capture key information related to emotions. Establish cross-modal interaction between the text and speech features, and use the text features to guide the speech features to ensure information complementarity and improve the fusion effect.

[0075] A3. Fuse the features extracted by self-attention and guided-attention through concatenation to maximize information retention and enhance the generalization ability of the model, obtaining fused features.

[0076] S4. The server inputs the speech modality input and the text modality input into the multi-modal deep learning network simultaneously for emotion recognition.

[0077] Embodiment 2

[0078] The above is an introduction to the method embodiments. The following further illustrates the solution of the present invention through system embodiments.

[0079] As Figure 4 shown, a mobile emotion recognition acceleration system based on signal compression includes a mobile device and a server. Among them, the mobile device collects speech signals and performs feature extraction, generates a Mel spectrogram using an Fbank encoder, compresses the Mel spectrogram to obtain compressed features, and sends the compressed features to the server. The server restores the compressed features to obtain a restored Mel spectrogram as the speech modality input, performs text recognition on the restored Mel spectrogram to obtain text information as the text modality input, and inputs the speech modality input and the text modality input into the multi-modal deep learning network simultaneously for emotion recognition. Among them, the multi-modal deep learning network uses a collaborative attention module to achieve the fusion of the speech modality input and the text modality input.

[0080] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0081] Based on the above methods and systems, through actual verification, the following conclusions can be drawn:

[0082] (1) The present invention can improve data processing efficiency and reduce energy consumption.

[0083] After feature extraction and compression are completed on the mobile device side, the required storage space is only γ times that of the original data, effectively reducing the storage pressure and power consumption of the device, enabling it to still run smoothly in resource-constrained environments.

[0084] (2) The present invention can reduce bandwidth consumption and communication latency.

[0085] Through signal compression, the load on the server side is reduced. Meanwhile, during the data transmission process, the transmission rate is increased by an average of 2.24 times, and the energy consumption is reduced by 55.35%, thus achieving more efficient real-time processing capabilities.

[0086] (3) The present invention can improve the accuracy of emotion recognition

[0087] Through multimodal fusion, making full use of the complementarity of speech and text features, a recognition accuracy of 69.83% is achieved on the IEMOCAP dataset, and the F1 score reaches 0.698, which is significantly better than unimodal methods.

[0088] (4) Modularity and scalability of the system

[0089] Due to the modular design architecture, the system can flexibly adapt to the computing capabilities of different devices and is easy to integrate into other voice interaction applications, enhancing the scalability of the system.

[0090] (5) Adapt to the mobile environment and improve the user experience

[0091] Through the distributed architecture design, the present invention reduces the computing pressure on mobile devices, enabling voice emotion recognition to operate efficiently on mobile platforms such as smartphones and Internet of Things devices, and enhancing the user experience.

[0092] In summary, the improvements in this method and system have achieved significant enhancements in data processing, transmission efficiency, multimodal fusion, emotion recognition accuracy, and mobile applicability, and have high application value.

[0093] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. A method for accelerating mobile emotion recognition based on signal compression, characterized in that It includes the following steps: Collect voice signals using a mobile device and perform feature extraction, generate a Mel spectrogram using an Fbank encoder, compress the Mel spectrogram to obtain compressed features; The mobile device sends the compressed features to the server; The server restores the compressed features to obtain a restored Mel spectrogram as the voice modality input, and performs text recognition on the restored Mel spectrogram to obtain text information as the text modality input; The server simultaneously inputs the voice modality input and the text modality input into a multi-modal deep learning network for emotion recognition, where the multi-modal deep learning network uses a collaborative attention module to achieve the fusion of the voice modality input and the text modality input.

2. The mobile emotion recognition acceleration method based on signal compression according to claim 1, wherein The feature extraction includes the following steps: Perform pre-emphasis and frame segmentation operations on the collected voice signals; Divide the variable-length voice signals into fixed-length segments, where there is a preset overlap between frames, and use a Hamming window to smooth the voice signals; Use the fast Fourier transform to convert the voice signals to the frequency domain and calculate the energy spectrum of the signals; Based on the energy spectrum, use an Fbank encoder to generate a Mel spectrogram.

3. A method for accelerating mobile emotion recognition based on signal compression according to claim 1, characterized in that, The compression of the Mel spectrogram uses the singular value decomposition method.

4. A method for accelerating mobile emotion recognition based on signal compression according to claim 3, characterized in that, The singular value decomposition method includes the following steps: For the Mel spectrogram f, which is an 80×T matrix, 80 is the feature matrix of the Mel spectrogram, and T is a quantity related to the time length of the original voice signal, perform the Mel spectrogram decomposition as follows: f = Udiag(S)V H Among them, U is an 80×k matrix, S is a vector of length k, V is a T×k matrix, and V H is its transpose matrix, and k = min(80, T); Take the first γk bits of vector S and compress S to obtain S new , while keeping the magnitude of f unchanged, compress U to U new = a matrix of 80×γk, compress V to V new = a matrix of T×γk, to obtain the compressed feature, where γ ∈ [0,1] is the compression ratio.

5. The method for accelerating mobile emotion recognition based on signal compression according to claim 4, wherein, The described server processes the compressed features through SVD inverse transformation to obtain the restored Mel spectrogram f new :

6. A method for accelerating mobile emotion recognition based on signal compression according to claim 1, characterized in that The multi-modal deep learning network includes: A text feature embedding module for extracting features from the text modality input to obtain text features; A voice feature extraction module for further extracting emotional features from the voice modality input using a CNN block and obtaining voice features by concatenating contexts using an LSTM block; A collaborative attention module for using the complementarity between text features and voice features to achieve deep interaction between voice and text features and output fused features; An output module for processing the fused features output by the collaborative attention module using a self-attention mechanism and outputting the predicted value of emotion classification through a linear layer.

7. A method for accelerating mobile emotion recognition based on signal compression according to claim 6, characterized in that, The text feature embedding module uses RoBERTa to convert the text modality input into a tensor embedding.

8. An acceleration method for mobile emotion recognition based on signal compression according to claim 6, characterized in that, The collaborative attention module adopts an encoder-decoder architecture and stacks multiple attention modules to achieve deep feature interaction.

9. A method for accelerating mobile emotion recognition based on signal compression according to claim 8, characterized in that The collaborative attention module performs the following steps: Perform self-attention calculation on the input of a single modality to extract the deep features of the corresponding modality itself; Apply guided attention, using the features of one modality as a guide to align the features of the two modalities in time series and semantics; Fuse the features extracted by self-attention and guided attention through concatenation to obtain fused features.

10. A mobile emotion recognition acceleration system based on signal compression, characterized in that, It includes a mobile device and a server. Among them, the mobile device collects voice signals and performs feature extraction, generates a Mel spectrogram using an Fbank encoder, compresses the Mel spectrogram to obtain compressed features, and sends the compressed features to the server; the server restores the compressed features to obtain a restored Mel spectrogram as the voice modality input, performs text recognition on the restored Mel spectrogram to obtain text information as the text modality input, and simultaneously inputs the voice modality input and the text modality input into a multi-modal deep learning network for emotion recognition. Among them, the multi-modal deep learning network uses a collaborative attention module to achieve the fusion of the voice modality input and the text modality input.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method based on attention enhancing mechanism

    CN112489635A

  • Multi-modal emotion recognition method based on voiceprint and text

    CN118486333A

  • Learning method of apparatus for emotion estimating using multi-modal model

    KR102786748B1

  • Crowd-information-fused speech emotion recognition method and system

    WO2022199215A1