Speech recognition method, system and electronic device

CN122821933APending Publication Date: 2026-09-25ANHUI KAIYANG TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611123066.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本申请实施例提供一种语音识别方法、系统和电子设备,以至少解决相关技术进行语音识别的准确性较差的技术问题

Benefits of technology

[0023]根据本申请实施例的另一方面,还提供了一种计算机程序,计算机程序被处理器执行时实现本申请各个实施例中的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821933A_ABST
    Figure CN122821933A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a speech recognition method, system and electronic device, the method comprising: in response to receiving a speech recognition instruction, obtaining a speech signal corresponding to the speech recognition instruction; using an acoustic encoder to perform acoustic feature extraction on the speech signal to obtain an acoustic hidden state sequence of the speech signal; using a projection mapping model to project the acoustic hidden state sequence from an acoustic feature space to a hidden vector space corresponding to a speech recognition model to obtain an acoustic embedding sequence; inputting the acoustic embedding sequence and the speech recognition instruction to the speech recognition model, and using the speech recognition model to perform classification recognition on the speech signal to obtain a classification recognition result. The present application solves the technical problem of poor accuracy of speech recognition in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition, and more specifically, to a speech recognition method, system, and electronic device. Background Technology

[0002] With the rapid development of technologies such as artificial intelligence, human-computer interaction, and speech signal processing, speech emotion recognition research has gradually become a fundamental capability for smart terminals to achieve natural interaction, emotion perception, and intelligent emotional feedback.

[0003] However, the commonly used voice emotion recognition method currently uses classifiers to classify and recognize acoustic features. However, classifiers are not adaptable to task expression, candidate categories, and output formats, and it is difficult to output appropriate recognition results for different occasions. Therefore, the related technologies have the drawback of low recognition accuracy.

[0004] There is currently no good solution to the above problems. Summary of the Invention

[0005] This application provides a speech recognition method, system, and electronic device to at least solve the technical problem of poor accuracy in speech recognition using related technologies.

[0006] According to one aspect of the embodiments of this application, a speech emotion recognition method is provided, comprising: in response to receiving a speech recognition instruction, acquiring a speech signal corresponding to the speech recognition instruction; extracting acoustic features from the speech signal using an acoustic encoder to obtain an acoustic hidden state sequence of the speech signal; projecting the acoustic hidden state sequence from the acoustic feature space to the hidden vector space corresponding to the speech emotion recognition model using a projection mapping model to obtain an acoustic embedding sequence, wherein the acoustic embedding sequence is aligned with the text embedding sequence corresponding to the speech recognition instruction; inputting the acoustic embedding sequence and the speech recognition instruction into the speech emotion recognition model, and classifying and recognizing the speech signal using the speech emotion recognition model to obtain a classification and recognition result.

[0007] Optionally, before extracting acoustic features from the speech signal using an acoustic encoder to obtain the acoustic hidden state sequence of the speech signal, the above method further includes: extracting features from the speech signal using a preprocessing unit to obtain an acoustic feature sequence; the acoustic encoder includes a single-modal feedforward temporal memory coding unit, a memory-enhanced self-attention temporal modeling unit, and a projection and causal convolution enhancement unit; extracting acoustic features from the speech signal using the acoustic encoder to obtain the acoustic hidden state sequence of the speech signal includes: using the single-modal feedforward temporal memory coding unit to perform temporal enhancement on the acoustic feature sequence corresponding to the speech signal to obtain a context feature sequence; using the memory-enhanced self-attention temporal modeling unit to perform attention processing on the context feature sequence to obtain an attention feature sequence; and using the projection and causal convolution enhancement unit to project and enhance the attention feature sequence to obtain the acoustic hidden state sequence.

[0008] Optionally, a single-modal feedforward temporal memory coding unit is used to temporally enhance the acoustic feature sequence to obtain a context feature sequence. This includes: for the acoustic features at any time step in the acoustic feature sequence, performing a linear transformation and activation function processing on the acoustic features at any time step to obtain an intermediate feature at any time step; using the historical intermediate features of historical time steps, enhancing the intermediate features at any time step to obtain an enhanced feature at any time step, where the historical time step is the time step preceding any time step; and fusing the intermediate feature at any time step with the enhanced feature at any time step to obtain the context feature at any time step.

[0009] Optionally, by utilizing the historical intermediate features of historical time steps, the intermediate features of any time step can be enhanced to obtain the enhanced features of any time step. This includes: performing a weighted summation of the intermediate features of any time step and the historical intermediate features to obtain the enhanced features of any time step.

[0010] Optionally, the memory-enhanced self-attention temporal modeling unit includes a global enhancement unit, a temporal memory unit, and a fusion unit. The memory-enhanced self-attention temporal modeling unit performs attention processing on the context feature sequence to obtain an attention feature sequence, including: performing global attention processing on the context feature sequence using the global enhancement unit to obtain a global feature sequence of the context feature sequence; performing temporal enhancement processing on the value matrix of the context feature sequence using the temporal memory unit to obtain a local feature sequence of the context feature sequence, wherein the memory parameters of the temporal memory unit are different from those of the unimodal feedforward temporal memory encoding unit; and fusing the global feature sequence of the context feature sequence and the local feature sequence of the context feature sequence using the fusion unit to obtain the attention feature sequence.

[0011] Optionally, the projection and causal convolution enhancement unit includes a linear projection layer and a causal one-dimensional convolution layer; using the projection and causal convolution enhancement unit, the attention feature sequence is projected and enhanced to obtain the acoustic hidden state sequence, including: using the linear projection layer to perform a linear transformation on the attention feature sequence to obtain the hidden dimension sequence of the speech signal; using the causal one-dimensional convolution layer to perform convolution processing on the hidden dimension sequence to obtain the acoustic hidden state sequence.

[0012] Optionally, the acoustic hidden state at any time step can be determined by a causal one-dimensional convolutional layer using the hidden dimension of any time step and the hidden dimension of the time step preceding any time step.

[0013] Optionally, the projection mapping model is a multilayer perceptron; the acoustic hidden state sequence is projected from the acoustic feature space to the hidden vector space corresponding to the speech perception recognition model using the projection mapping model to obtain the acoustic embedding sequence, including: using the multilayer perceptron, the acoustic hidden state sequence is processed by linear processing, activation processing, regularization processing and linear processing respectively to obtain the acoustic embedding sequence, wherein the multilayer perceptron processes each time step of the acoustic hidden state sequence with the same network parameters, and the time steps of the acoustic hidden state sequence and the acoustic embedding sequence are the same.

[0014] Optionally, the preprocessing unit includes a time-frequency transformation unit and a filter; the preprocessing unit performs initial feature extraction on the speech signal to obtain an acoustic feature sequence, including: using the time-frequency transformation unit to perform Fourier transform on the speech signal to obtain the spectral features of the speech signal; and using the filter to compress the spectral features according to the time sequence to obtain the acoustic feature sequence.

[0015] Optionally, the acoustic embedding sequence and the speech recognition command are input into the speech emotion recognition model, and the speech signal is classified and recognized by the speech emotion recognition model to obtain the classification and recognition result. This includes: performing word embedding processing on the speech recognition command to obtain the text embedding sequence corresponding to the speech recognition command; concatenating the acoustic embedding sequence and the text embedding sequence to obtain the recognition sequence; and inputting the recognition sequence into the speech emotion recognition model to obtain the classification and recognition result.

[0016] Optionally, the above method further includes: acquiring training data, wherein the training data includes a speech training signal, a classification label corresponding to the speech training signal, and a text training instruction; obtaining a training embedding sequence based on the speech training signal and the text training instruction; and training an initial model based on the training embedding sequence and the classification label to obtain a speech emotion recognition model.

[0017] Optionally, the initial model is trained based on the training embedding sequence and classification labels to obtain a speech emotion recognition model, including: inputting the training embedding sequence into the initial model to obtain an initial classification result; constructing a loss function based on the initial classification result and classification labels; and adjusting the parameters of the initial model based on the loss function to obtain the speech emotion recognition model.

[0018] According to another aspect of the embodiments of this application, a speech emotion recognition system is also provided, comprising: a speech sampling module, configured to acquire a speech signal corresponding to a speech recognition command in response to receiving a speech recognition command; an acoustic coding module, configured to extract acoustic features from the speech signal using an acoustic encoder to obtain an acoustic hidden state sequence of the speech signal; a projection mapping module, configured to project the acoustic hidden state sequence from the acoustic feature space to the hidden vector space corresponding to the speech emotion recognition model using a projection mapping model to obtain an acoustic embedding sequence, wherein the acoustic embedding sequence is aligned with the text embedding sequence corresponding to the speech recognition command; and an emotion recognition output module, configured to input the acoustic embedding sequence and the speech recognition command into the speech emotion recognition model, and classify and recognize the speech signal using the speech emotion recognition model to obtain a classification and recognition result.

[0019] According to another aspect of the embodiments of this application, a vehicle is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.

[0020] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0021] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0022] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.

[0023] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.

[0024] In this embodiment, in response to receiving a speech recognition command, the speech signal corresponding to the speech recognition command is acquired; acoustic features are extracted from the speech signal using an acoustic encoder to obtain an acoustic hidden state sequence of the speech signal; the acoustic hidden state sequence is projected from the acoustic feature space to the hidden vector space corresponding to the speech emotion recognition model using a projection mapping model to obtain an acoustic embedding sequence; the acoustic embedding sequence and the speech recognition command are input into the speech emotion recognition model, and the speech signal is classified and recognized using the speech emotion recognition model to obtain a classification and recognition result.

[0025] This application constructs a projection mapping model to achieve projection mapping between an acoustic encoder and a speech perception recognition model. The projection mapping model precisely maps the acoustic hidden state sequence obtained by the acoustic encoder from the acoustic feature space to the hidden vector space of the speech perception recognition model, generating an acoustic embedding sequence aligned with the text embedding sequence. This acoustic embedding sequence and the text embedding sequence are then input into the speech perception recognition model for inference. This achieves the goal of fusing non-textual acoustic embedding sequences and speech recognition commands in a unified high-dimensional semantic space, enabling effective interaction and controllable output of multimodal information. This improves the accuracy of classification and recognition results, thereby solving the technical problem of poor accuracy in speech recognition using related technologies. Attached Figure Description

[0026] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0027] Figure 1 This is a flowchart of a speech emotion recognition method according to an embodiment of this application;

[0028] Figure 2 This is a schematic diagram of a voice emotion recognition system according to an embodiment of this application;

[0029] Figure 3 This is a schematic diagram of an optional voice emotion recognition process according to an embodiment of this application;

[0030] Figure 4 This is a schematic diagram of an optional voice emotion recognition method according to an embodiment of this application;

[0031] Figure 5 This is a schematic diagram of an optional voice emotion recognition device according to an embodiment of this application;

[0032] Figure 6 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] According to an embodiment of this application, a method embodiment for voice emotion recognition is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0036] This embodiment provides a voice emotion recognition method. Figure 1 This is a flowchart of a speech emotion recognition method according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps:

[0037] Step S102: In response to receiving a voice recognition command, acquire the voice signal corresponding to the voice recognition command.

[0038] The aforementioned voice recognition commands refer to the commands or signals issued by the user through a voice interaction interface to trigger voice processing tasks. Voice recognition commands can contain semantic content, such as querying the weather or playing music, and control information for the recognition task, such as specifying the range of emotion categories to be recognized or output format requirements.

[0039] Speech recognition commands can guide large language models to perform sentiment inference, transforming sentiment classification into an open and controllable generation task. This allows for flexible adaptation to the sentiment analysis needs of different application scenarios, avoiding the problem of insufficient generalization caused by fixed classification heads. The diversity of speech recognition commands enhances the ability of large language models to understand complex contexts, while ensuring the interpretability and controllability of sentiment recognition results, making human-computer interaction more natural and intelligent.

[0040] The aforementioned speech signal refers to the physical quantity of sound waves generated by the vibration of a sound source, propagating through a medium, and received by the human ear or microphone. Speech signals are the raw data source for emotion recognition and can be represented as a pressure fluctuation waveform that changes over time. Speech signals possess a certain signal-to-noise ratio to facilitate the extraction of effective features. Speech signals can reflect the speaker's identity, the content of their speech, and emotional cues such as intonation, stress, speech rate, pitch, and energy. Speech signals may include various time-varying structures such as steady segments, bursts, and pauses, providing fundamental material for subsequent acoustic feature extraction.

[0041] Voice emotions may be hidden in the acoustic features of the speech signal, such as the high energy and rapid speech rate when angry, and the low tone and slow rhythm when sad. Therefore, acquiring high-quality speech signals is a prerequisite for ensuring the accuracy of emotion recognition. Speech signals can be acquired through microphone arrays or mono devices, and after noise reduction and standardization, they are sent to an acoustic encoder to transform physical sound waves into a computable feature sequence, realizing the mapping from sound waves to digital signal space.

[0042] In one optional embodiment, a beamforming and voice activity detection collaborative acquisition mechanism based on a microphone array is used to obtain the voice signal. An array structure consisting of multiple microphones deployed at different locations on the smart terminal is used to collect the sound field signal in the environment in real time. Using an adaptive beamforming algorithm, the weighting coefficients of each microphone channel are dynamically adjusted based on the direction estimation of the user's sound source, forming a highly directional main beam, thereby effectively suppressing background noise and reverberation interference from non-user directions. Before acquiring the voice signal, a voice activity detection module is first run. By analyzing the signal's short-time energy, zero-crossing rate, and spectral entropy, it is determined whether a valid human voice signal exists in the current environment. Once the starting point of voice activity is detected, beamforming processing is activated to extract a clear human voice segment. Subsequently, combined with the triggering logic of voice recognition commands, such as detecting a specific wake word or command prefix, after confirming the user's intent, the voice data within that time window is locked as the voice signal.

[0043] In another alternative embodiment, the acquisition of the voice signal is offloaded to the local processor of the terminal device, such as a smart speaker, in-vehicle control unit, or robot main control board. Using the audio driver interface provided by the operating system, the digital audio stream, converted from analog signals to digital signals, is continuously captured at a fixed sampling rate (e.g., 16kHz) and bit depth (e.g., 16bit). To reduce transmission latency, the data undergoes preliminary preprocessing locally, including gain control, echo reduction, and noise reduction. After parsing the user's voice recognition command, a complete voice segment corresponding to the command is extracted from the circular buffer. This segment includes the silence segment before the command, the command itself, and any possible subsequent continuation of the speech. This streaming processing mechanism ensures the real-time acquisition of the voice signal. Local processing not only reduces latency caused by network transmission but also avoids the security risks that may arise from uploading audio data to the cloud. Simultaneously, local preprocessing can dynamically adjust the acquisition gain according to the ambient noise level, ensuring that the amplitude of the recorded voice signal is appropriate and avoiding clipping distortion or excessively low signal-to-noise ratio, thereby providing a stable and standardized input signal for the acoustic encoder.

[0044] In another optional embodiment, the terminal device acquires a continuous audio stream via a microphone and uploads it to a cloud server after real-time compression and encoding. On the cloud server, a speech recognition engine transcribes the received audio stream in real time and, combined with contextual semantic analysis, identifies the user's speech recognition command. Once the speech recognition command is successfully parsed, the cloud server returns a timestamp or sequence identifier corresponding to the command. Based on the returned timestamp, the terminal device accurately extracts the corresponding speech signal from the locally cached raw audio data. By utilizing the computing resources of the cloud server for high-precision speech recognition and command parsing, even if the speech recognition command contains complex semantics or multi-turn dialogue context, the user's intent can be accurately understood. Simultaneously, a time synchronization mechanism ensures the correspondence between the acquired speech signal and the command content.

[0045] Step S104: Use an acoustic encoder to extract acoustic features from the speech signal to obtain the acoustic hidden state sequence of the speech signal.

[0046] The aforementioned acoustic encoder is a core module for deep learning representations of high-level semantic and emotional information from acoustic features. Through temporal modeling, the acoustic encoder captures dynamic features in speech signals that change over time. It maps low-dimensional acoustic features to a high-dimensional hidden state space, extracting abstract features related to emotion, thereby abstracting local prosody and global contextual information from the speech signal. Deep learning automatically discovers acoustic patterns that are discriminative for emotion classification, such as subtle changes in formants or the trajectory of the fundamental frequency, providing rich feature representations for subsequent emotion projection and classification, thus improving representational ability and generalization performance.

[0047] The acoustic hidden state sequence described above is a high-dimensional vector sequence output by the acoustic encoder after deep temporal modeling, representing the deep semantic and emotional representation of the speech signal at a specific point in time. The acoustic hidden state sequence can be a series of continuous feature vectors, with each vector corresponding to a time step or frame of the speech.

[0048] Acoustic hidden state sequences reside in the acoustic feature space, possessing high dimensionality and density, and encapsulating emotional cues such as prosody, energy, and intonation in speech. These sequences can be high-dimensional tensors in the real number domain, such as 512-dimensional or higher-dimensional vectors. Acoustic hidden state sequences preserve subtle emotional fluctuations in speech, such as tremolo, pauses, and stress, improving the accuracy and subtlety of emotion recognition.

[0049] In one alternative embodiment, the acoustic encoder is constructed based on a combination of convolutional neural networks and bidirectional long short-term memory networks. The acoustic encoder utilizes one-dimensional or multi-dimensional convolutional layers to extract local features from the input speech signal, capturing local patterns in the speech signal, such as phoneme transitions, energy abrupt changes, or formant variations, through a sliding window. Simultaneously, pooling layers reduce the sequence dimensionality and enhance the translation invariance of features. Subsequently, the convolutionally processed feature sequence is input into the bidirectional long short-term memory network, which processes temporal information from left to right and from right to left, respectively. This allows for the simultaneous use of past and future contextual information to model the state at the current time step, preventing gradient vanishing and ensuring the effective transmission of emotional cues in long speech sequences. The final output acoustic hidden state sequence combines the sensitivity of local convolutional features with the temporal memory capability of recurrent networks, comprehensively reflecting the dynamic changes in emotion within the speech.

[0050] In another alternative embodiment, the acoustic encoder is constructed based on causal convolution and a deep residual network. To address the vanishing gradient problem in deep networks and accelerate convergence, the model introduces residual connections, directly adding the input to the output, enabling the network to learn identity mappings or subtle feature changes more effectively. Simultaneously, by introducing dilated convolutions with different dilation coefficients, the receptive field is expanded exponentially without increasing the number of parameters, allowing the acoustic encoder to capture long-range temporal dependencies, such as the intonation trend of an entire sentence. The resulting acoustic hidden state sequence preserves the temporal resolution of the original speech while extracting highly abstract emotional features through deep nonlinear transformations.

[0051] In another alternative embodiment, a general acoustic encoder is pre-trained using unlabeled speech data through self-supervised tasks such as masked acoustic modeling and contrastive predictive coding. During the pre-training phase, the acoustic encoder learns to reconstruct masked acoustic features or distinguish between positive and negative sample pairs, thereby learning a general speech representation, including low-level, emotion-sensitive features such as pitch, energy, and speech rate. When extracting the acoustic hidden state sequence, the pre-trained acoustic encoder is frozen or fine-tuned, and a specific speech signal is input. Since the pre-training process has already enabled the acoustic encoder to grasp the speech structure information, the hidden state sequence output by the acoustic encoder carries higher-order semantic information related to emotion. This reduces the reliance on labeled data, especially in few-shot or zero-shot emotion recognition scenarios, where the pre-trained encoder can extract more robust acoustic hidden state sequences.

[0052] Step S106: The acoustic hidden state sequence is projected from the acoustic feature space to the hidden vector space corresponding to the speech perception recognition model using a projection mapping model to obtain the acoustic embedding sequence.

[0053] The acoustic embedding sequence is aligned with the text embedding sequence corresponding to the speech recognition command.

[0054] The aforementioned projection mapping model refers to the component used to connect the acoustic encoder and the speech emotion recognition model. By constructing a learnable nonlinear transformation space, the projection mapping model reduces the heterogeneity of feature representations across different modalities. Through mathematical mapping operations, it transforms low-level acoustic features belonging to the acoustic domain into semantic-level embedded representations that the speech emotion recognition model can directly understand and process. The projection mapping model performs independent, parameter-consistent dimensional expansion or compression on the acoustic hidden states at each time step, allowing previously isolated acoustic features to be seamlessly embedded into the speech emotion recognition model. This enables the speech emotion recognition model to leverage semantic understanding capabilities to handle speech emotion recognition tasks, improving scalability and the accuracy and controllability of emotion recognition.

[0055] The aforementioned acoustic embedding sequence is a vector sequence obtained by projecting and mapping the acoustic hidden state sequence. The acoustic embedding sequence is the acoustic representation in the hidden space of the speech emotion recognition model. Through the acoustic embedding sequence, acoustic information can be directly processed by the speech emotion recognition model, thus participating in inference. This achieves modal alignment between acoustic features and linguistic features, enabling large language models to utilize instruction understanding and generation capabilities to process speech emotions. By concatenating acoustic embeddings with spoken instructions, the result of emotion recognition can be flexibly adjusted according to the spoken instructions, such as outputting specific emotion tags or descriptions. This alignment mechanism not only improves the accuracy of recognition but also enhances the controllability and interpretability of the speech emotion recognition model, making emotion recognition a command-based generation process, thereby improving the naturalness and intelligence of human-computer interaction.

[0056] The acoustic feature space mentioned above refers to the mathematical space containing the output vectors of the acoustic encoder. The acoustic feature space can be a high-dimensional Euclidean space, where each point represents a specific acoustic state. The dimension of the acoustic feature space is determined by the final output layer of the acoustic encoder, such as 512 dimensions, thus storing and representing the deep acoustic features of speech. The acoustic feature space can be a combination of various acoustic parameters, such as fundamental frequency, energy, and spectral envelope, abstracted in the hidden layers, providing a target space for projection mapping.

[0057] Because the acoustic feature space and the hidden space of the speech emotion recognition model differ in semantic distribution and dimensionality, directly inputting them into the speech emotion recognition model will prevent the model from effectively utilizing acoustic information. Therefore, it is necessary to define a clear acoustic feature space as the source space for projection. By performing feature transformation and alignment in this space, the effectiveness of acoustic information in the speech emotion recognition model can be improved, ensuring that emotional cues are not distorted or lost. The construction of the acoustic feature space is a key step in improving the multimodal fusion effect.

[0058] The aforementioned hidden vector space refers to the continuous vector space within the speech emotion recognition model used to represent text and input features. The hidden vector space can be a high-dimensional semantic space, where each vector represents a word, phrase, or concept. The dimensionality of the hidden vector space is relatively large, such as 2048 dimensions or higher. The hidden vector space can be used to capture semantic information and contextual relationships of language, such as through word embedding, positional encoding, and context embedding after multi-layer transformation processing. The hidden vector space provides a target mapping space for acoustic embedding. Through projection mapping, the acoustic hidden states are transformed into the hidden vector space, enabling the speech emotion recognition model to understand acoustic information in a text-like manner. This spatial alignment allows the speech emotion recognition model to utilize its pre-trained linguistic knowledge to assist in emotion recognition, such as understanding subtle differences in emotional vocabulary or emotional tendencies in context. The introduction of the hidden vector space transforms emotion recognition into a generative task, improving the flexibility and generalization ability of the speech emotion recognition model. Therefore, the speech recognition method in this embodiment can handle unseen instructions and complex emotional expressions, achieving more intelligent emotional interaction.

[0059] In one optional embodiment, since the speech emotion recognition model has a high number of parameters, a low-rank decomposition strategy is employed to decompose the projection mapping model into the product of two low-rank matrices, i.e., matrix W equals matrix A multiplied by matrix B, where A and B are both trainable low-rank matrices, while the original weight matrix remains frozen or used only as an initialization reference. When processing the acoustic hidden state sequence, the projection mapping model updates the parameters of these two low-rank matrices, mapping the acoustic features to the hidden vector space through matrix multiplication. By updating only a small number of parameters, catastrophic forgetting can be effectively prevented, ensuring that the speech emotion recognition model's language understanding and reasoning abilities are not negatively affected when adapting to acoustic feature inputs. Furthermore, the low-rank structure helps the projection mapping model learn more generalizable mapping relationships, resulting in generated acoustic embedding sequences that are not only dimensionally aligned but also semantically closer to the natural language features of the speech emotion recognition model, thereby improving the accuracy and stability of emotion recognition.

[0060] In another alternative embodiment, an adaptive projection mapping method based on cross-modal contrastive learning is employed to obtain the acoustic embedding sequence. A contrastive learning mechanism is introduced to reduce the distance between the acoustic embedding sequence and the corresponding text embedding sequence in the feature space during the projection process, while maximizing the distance between irrelevant samples. The projection mapping model is trained as a differentiable transformation function, with the objective function including reconstruction loss and contrastive loss. During training, the projection mapping model forces the acoustic hidden state to converge towards the linguistic semantic space by maximizing the similarity of positive sample pairs (i.e., the acoustic embedding and text embedding corresponding to the same speech segment). This adaptive projection method can dynamically adjust the mapping parameters to compensate for the inherent distributional differences between the acoustic feature space and the speech emotion recognition model space. In this way, the projected acoustic embedding sequence is not only dimensionally aligned but also semantically consistent with the internal representation of the speech emotion recognition model, enabling the speech emotion recognition model to more accurately understand subtle emotional nuances in speech, such as sarcasm or implied sadness in intonation, thereby improving the semantic understanding ability and robustness of emotion recognition.

[0061] Step S108: Input the acoustic embedding sequence and speech recognition command into the speech emotion recognition model, and use the speech emotion recognition model to classify and recognize the speech signal to obtain the classification and recognition result.

[0062] The aforementioned speech emotion recognition model refers to a large, adapted language model. This model can be a generative language model based on a transducer architecture. After fine-tuning to accept acoustic embedding input and perform emotion classification tasks, the speech emotion recognition model can generate emotion category labels based on the input acoustic embedding sequence and speech recognition instructions.

[0063] Voice emotion recognition models can transform multimodal inputs into emotion recognition results. Through semantic understanding and command compliance capabilities, they make emotion recognition more flexible and controllable. These models can output emotion descriptions in different formats or with varying granularities based on commands, ranging from basic emotions to more nuanced psychological states. Voice emotion recognition models improve the accuracy and robustness of emotion recognition while reducing training costs, enabling emotion recognition systems to be easily transferred to different application scenarios, such as in-vehicle systems, smart homes, and companion robots, thus enhancing the user experience.

[0064] In one optional embodiment, the voice emotion recognition model performs emotion classification based on thought chain reasoning. Voice recognition commands guide the model to analyze acoustic embedding sequences and output classification results. Before generating the final emotion category label, the model first generates a natural language reasoning description explaining the basis for the emotion category label judgment. This decomposes the complex emotion recognition task into multiple sub-steps, thereby improving the interpretability and accuracy of the decision. The acoustic embedding sequence, as a conditional input, influences the attention distribution of the voice emotion recognition model, causing it to focus on emotion-related acoustic features when generating reasoning text. The final emotion category label serves as the end marker of the reasoning process. This thought chain reasoning approach simulates the logical process of human emotion judgment—first perceiving features, then analyzing semantics, and finally making a comprehensive judgment—thus reducing misclassifications caused by "illusion" or chance factors. Especially when dealing with subtle emotional distinctions (such as "disappointment" versus "sadness," "excitement" versus "anger"), thought chain reasoning provides clear discriminative criteria, enhancing the performance of the voice emotion recognition model in complex emotional scenarios.

[0065] In another optional embodiment, a classification and recognition method based on confidence calibration and multi-class probability output is used to obtain the classification and recognition results. When classifying speech signals using a speech emotion recognition model, not only is the final emotion category label output, but also the probability distribution or confidence score of each candidate emotion category is output. The speech emotion recognition model maps the hidden states to the logical space of emotion category words and then normalizes them into probability values. To further improve the reliability of classification, a confidence calibration mechanism is introduced, such as using temperature scaling, to adjust the probability distribution output by the speech emotion recognition model, so that the confidence score more realistically reflects the correctness of the prediction. In addition, by monitoring the confidence distribution, the quality of the acoustic embedding sequence or the design of the text instructions can be adjusted in reverse, forming a closed-loop mechanism.

[0066] In another alternative embodiment, the speech emotion recognition model achieves classification based on a multimodal fusion gating mechanism. A learnable gating mechanism is introduced within the speech emotion recognition model to dynamically adjust the contribution weights of the acoustic embedding sequence and the text instruction embedding sequence corresponding to the speech recognition command to the final emotion classification result. Specifically, before each processing layer of the speech emotion recognition model, a gating vector is computed, which is determined based on the interaction strength between the current input acoustic features and text features. For example, when the text content has strong emotional connotations, the gating mechanism may reduce the weight of the acoustic embedding, relying mainly on text semantics for judgment; conversely, when the text content is neutral, the gating mechanism will increase the weight of the acoustic embedding, focusing on acoustic cues such as intonation and energy. This dynamic weight allocation mechanism allows the speech emotion recognition model to adaptively balance the importance of multimodal information, preventing one modality from dominating or interfering with the information of another modality. The gating mechanism is implemented by a small neural network, with input being a concatenation or intermediate representation of the acoustic embedding and the text command, and outputting weight values ​​between 0 and 1. The final emotion classification result is generated from the weighted fusion of the multimodal representations. This improves the robustness of the speech emotion recognition model to noise and modal loss. For example, when background noise is high and causes acoustic feature distortion, the speech emotion recognition model can automatically reduce its reliance on acoustic features; conversely, when there is significant text ambiguity, it can increase its reliance on acoustic features. Through this intelligent multimodal fusion strategy, the speech emotion recognition model can maintain stable high performance in various complex environments, achieving multimodal collaborative emotion understanding.

[0067] In this embodiment, in response to receiving a speech recognition command, the speech signal corresponding to the speech recognition command is acquired; acoustic features are extracted from the speech signal using an acoustic encoder to obtain an acoustic hidden state sequence of the speech signal; the acoustic hidden state sequence is projected from the acoustic feature space to the hidden vector space corresponding to the speech emotion recognition model using a projection mapping model to obtain an acoustic embedding sequence; the acoustic embedding sequence and the speech recognition command are input into the speech emotion recognition model, and the speech signal is classified and recognized using the speech emotion recognition model to obtain a classification and recognition result.

[0068] This application constructs a projection mapping model to achieve projection mapping between an acoustic encoder and a speech perception recognition model. The projection mapping model precisely maps the acoustic hidden state sequence obtained by the acoustic encoder from the acoustic feature space to the hidden vector space of the speech perception recognition model, generating an acoustic embedding sequence aligned with the text embedding sequence. This acoustic embedding sequence and the text embedding sequence are then input into the speech perception recognition model for inference. This achieves the goal of fusing non-textual acoustic embedding sequences and speech recognition commands in a unified high-dimensional semantic space, enabling effective interaction and controllable output of multimodal information. This improves the accuracy of classification and recognition results, thereby solving the technical problem of poor accuracy in speech recognition using related technologies.

[0069] Optionally, before using an acoustic encoder to extract acoustic features from the speech signal to obtain the acoustic hidden state sequence of the speech signal, the above method further includes: using a preprocessing unit to extract features from the speech signal to obtain an acoustic feature sequence.

[0070] The aforementioned preprocessing unit is used to standardize and extract features from the original speech signal, such as performing resampling, normalization, short-time Fourier transform, and filter bank operations. The preprocessing unit processes the raw audio waveform data, converting the unstructured speech signal into a structured sequence of acoustic features. This reduces the impact of noise and irrelevant information, allowing the acoustic encoder to focus on extracting emotion-related features. For example, it captures frequency distribution through Mel-spectrum analysis, reduces sequence length through low frame rate processing, improves computational efficiency, and better simulates the human ear's perception of sound, thereby enhancing the effectiveness of emotion recognition.

[0071] In one alternative embodiment, in order to reduce the data heterogeneity caused by differences in recording devices, sampling rates and channels, and to ensure that the input data has a uniform time resolution and amplitude range, thereby providing a stable and consistent signal basis for subsequent feature extraction, a preprocessing unit can be used to transform the original waveform into a structured acoustic feature sequence through operations such as resampling, normalization, and frequency transformation, thereby reducing the impact of noise interference and hardware deviations.

[0072] The acoustic encoder includes a single-modal feedforward temporal memory coding unit, a memory-enhanced self-attention temporal modeling unit, and a projection and causal convolution enhancement unit. The acoustic encoder extracts acoustic features from the speech signal to obtain the acoustic hidden state sequence of the speech signal. This includes: using the single-modal feedforward temporal memory coding unit to perform temporal enhancement on the acoustic feature sequence corresponding to the speech signal to obtain a context feature sequence; using the memory-enhanced self-attention temporal modeling unit to perform attention processing on the context feature sequence to obtain an attention feature sequence; and using the projection and causal convolution enhancement unit to project and enhance the attention feature sequence to obtain the acoustic hidden state sequence.

[0073] The aforementioned single-modal feedforward temporal memory coding unit is a submodule of the acoustic encoder. The single-modal feedforward temporal memory coding unit can be a feedforward neural network incorporating finite temporal memory. It receives an acoustic feature sequence as input, extracts preliminary representations through linear transformations and activation functions, and utilizes information from historical time steps for local context enhancement. Before entering the complex attention mechanism, it establishes local linear dependencies between speech frames. Due to the strong local continuity of speech signals, the single-modal feedforward temporal memory coding unit can capture dynamic changes and local prosodic patterns within phonemes. This lightweight modeling approach is computationally efficient, providing richly localized initial representations for subsequent self-attention mechanisms. By introducing temporal memory, the single-modal feedforward temporal memory coding unit can better capture short-term emotional cues in speech, such as sudden energy changes or intonation shifts, thereby improving the sensitivity of emotion recognition.

[0074] The aforementioned memory-enhanced self-attention temporal modeling unit is a core module in the acoustic encoder. This unit can be a hybrid architecture combining global self-attention and local temporal memory. By modeling the global context and enhancing local dynamic information through temporal memory branches, it captures long-distance dependencies and subtle local changes in speech, preserving the dynamic changes between local time steps and avoiding information smoothing. This hybrid modeling approach enables the speech emotion recognition model to comprehensively understand the emotional expression of speech, grasping both the overall intonation and focusing on local emotional outbursts, thus improving the accuracy and robustness of emotion recognition.

[0075] The aforementioned projection and causal convolution enhancement unit is a module within the acoustic encoder. It adjusts the feature dimension through linear transformation and enhances local temporal features through causal convolution, mapping attention features to the acoustic hidden state space and highlighting short-term energy and prosodic variations, thus ensuring temporal causality. Linear projection ensures alignment between the output dimension and the hidden space of the speech emotion recognition model. Through local enhancement and dimensionality transformation, the projection and causal convolution enhancement unit ensures that the acoustic hidden state sequence contains rich acoustic information while also meeting the input requirements of the speech emotion recognition model.

[0076] In one alternative embodiment, in order to extract the acoustic hidden state, the acoustic feature sequence corresponding to the speech signal can first be temporally enhanced using a single-modal feedforward temporal memory coding unit. For example, by introducing a sliding context window of finite length or a memory convolution mechanism, the local temporal dependencies between speech frames can be captured without the need for a recurrent network, and short-term acoustic dynamic cues such as phoneme transitions and local energy changes can be extracted. This achieves efficient temporal modeling with low computational complexity and parameter quantity, providing a contextual feature sequence containing prosody and other information for subsequent deep feature extraction.

[0077] Then, the memory-enhanced self-attention temporal modeling unit is used to perform attention processing on the context feature sequence, extracting the global emotional tone and key intonation transitions across the entire sentence or even paragraph, generating attention feature sequences that have both a global perspective and retain local details, improving the ability to capture long-distance emotional cues and local emotional outbursts in speech, achieving complementary fusion of local and global temporal information, and enhancing the emotional discrimination power of acoustic representation.

[0078] Therefore, by using projection and causal convolution enhancement units to project and enhance the attention feature sequence, the risk of future information leakage that may be caused by the self-attention mechanism is reduced, and the causal consistency during inference is ensured. At the same time, it highlights the fine-grained acoustic cues for emotion recognition, such as short-term energy changes and pause patterns in speech. The attention feature sequence is then subjected to nonlinear transformation and feature enhancement, and finally a fixed-dimensional acoustic hidden state sequence is obtained. This completes the transformation from high-level semantic attention features to low-level emotion discrimination features, ensuring the causal integrity of temporal information. As a result, the generated acoustic hidden state sequence has both rich semantic and emotional connotations and retains accurate temporal alignment information.

[0079] Optionally, a single-modal feedforward temporal memory coding unit is used to temporally enhance the acoustic feature sequence to obtain a context feature sequence. This includes: for the acoustic features at any time step in the acoustic feature sequence, performing a linear transformation and activation function processing on the acoustic features at any time step to obtain an intermediate feature at any time step; using the historical intermediate features of historical time steps, enhancing the intermediate features at any time step to obtain an enhanced feature at any time step, where the historical time step is the time step preceding any time step; and fusing the intermediate feature at any time step with the enhanced feature at any time step to obtain the context feature at any time step.

[0080] In one optional embodiment, for the acoustic features at any time step in the acoustic feature sequence, the original high-dimensional acoustic features are mapped to a higher-dimensional or more suitable feature space, and a nonlinear activation function is introduced to enhance expressive power, transforming the underlying acoustic signal into a feature representation with a certain level of abstraction, extracting the independent acoustic properties within that time step, and capturing local patterns and nonlinear relationships in the acoustic features through the nonlinear mapping of the fully connected layer, obtaining the intermediate features of each time step, and providing input for the fusion of subsequent time-series information.

[0081] Then, the single-modal feedforward temporal memory coding unit utilizes historical intermediate features from previous time steps to enhance the intermediate features of the current time step. This enhancement is achieved through mechanisms such as sliding windows or memory convolution, aggregating historical contextual information prior to the current moment to construct a memory vector containing past speech states. This vector captures short-term dynamic changes in the speech signal, such as intonation continuity, energy gradation, and phoneme transitions. By calculating the features from historical time steps and the intermediate features from the current time step, an enhanced feature reflecting local temporal dependencies is generated. This effectively utilizes information from past frames to assist feature discrimination in the current frame, enhancing sensitivity to speech prosody and local emotional cues.

[0082] Finally, the single-modal feedforward temporal memory encoding unit fuses the intermediate features of any time step with the enhanced features of that time step. This combines the independent acoustic features of the current frame with the enhanced features that have fused with past temporal information, such as through residual connections or concatenation. This preserves the integrity of the original acoustic features while injecting temporal contextual information, generating a contextual feature that contains both local details and global context awareness. Through feature fusion, the intermediate features of any time step and the temporal enhanced features are integrated into a unified contextual feature, achieving effective complementarity between local features and temporal information. This results in a final contextual feature sequence that not only retains the rich details of the original acoustic signal but also possesses the ability to perceive dynamic changes in speech temporal sequence. This provides high-quality, semantically rich input for subsequent, more complex self-attention temporal modeling, thereby improving the speech emotion recognition model's understanding and robustness in complex speech scenarios.

[0083] Optionally, by utilizing the historical intermediate features of historical time steps, the intermediate features of any time step can be enhanced to obtain the enhanced features of any time step. This includes: performing a weighted summation of the intermediate features of any time step and the historical intermediate features to obtain the enhanced features of any time step.

[0084] In one alternative embodiment, to capture the continuity of intonation, the gradual trend of energy changes, and the transitional features of phonemes, the intermediate features extracted at the current time step can be fused with multiple historical intermediate features accumulated over several past time steps. For example, a fixed-length historical window can be predefined, and several historical intermediate features from past time steps can be selected within this window. Each historical intermediate feature is assigned a learnable weight coefficient, reflecting the degree of influence of different historical moments on the emotional state at the current moment. These historical intermediate features are then multiplied by their corresponding weights and accumulated to form a memory enhancement vector. Finally, this memory enhancement vector is added to or concatenated with the intermediate features of the current time step to obtain the enhanced features of the current time step. This allows the strong temporal dependence in speech signals to be utilized, and the limitations of single-frame features in expressing dynamic changes in emotion can be overcome by explicitly introducing historical state information. By constructing a lightweight and efficient local temporal modeling mechanism, the speech emotion recognition model can perceive the local context of the speech sequence. Through weighted summation, it can adaptively focus on those moments that have historical reference value for the current emotion judgment, thereby generating an enhanced feature sequence containing rich temporal context information. This improves the acoustic encoder's ability to capture subtle changes in speech prosody and emotion. By preserving local temporal dependencies, the accuracy and stability of the speech emotion recognition system are improved.

[0085] Optionally, the memory-enhanced self-attention temporal modeling unit includes a global enhancement unit, a temporal memory unit, and a fusion unit. The memory-enhanced self-attention temporal modeling unit performs attention processing on the context feature sequence to obtain an attention feature sequence, including: performing global attention processing on the context feature sequence using the global enhancement unit to obtain a global feature sequence of the context feature sequence; performing temporal enhancement processing on the value matrix of the context feature sequence using the temporal memory unit to obtain a local feature sequence of the context feature sequence, wherein the memory parameters of the temporal memory unit are different from those of the unimodal feedforward temporal memory encoding unit; and fusing the global feature sequence of the context feature sequence and the local feature sequence of the context feature sequence using the fusion unit to obtain the attention feature sequence.

[0086] In one optional embodiment, the global enhancement unit utilizes the global receptive field characteristic of the self-attention mechanism to calculate the correlation between all time steps in the sequence, thereby capturing semantic and emotional dependencies spanning the entire speech segment or even long distances. This breaks the limitations of local windows, focusing on contextual information that is temporally distant but crucial to the current emotion judgment, such as the emotional tone of the entire sentence or emotional transitions that echo each other. This results in a global feature sequence containing global contextual information, achieving a macro-level emotion understanding capability without ignoring the overall context due to local noise or short-sightedness, providing a globally consistent representational basis for emotion recognition. For example, by calculating the dot product similarity of the query matrix, key matrix, and value matrix, an attention weight distribution is generated, resulting in a global feature sequence containing global contextual information.

[0087] Then, a temporal memory unit is used to perform temporal enhancement processing on the value matrix of the context feature sequence. A lightweight, time-based sliding memory mechanism is introduced into the value vector path of the self-attention mechanism to capture continuous dynamic changes in the local temporal dimension, thereby compensating for the shortcomings of the standard self-attention mechanism in capturing high-frequency local details (such as speech rate changes, instantaneous pauses, and subtle intonation tremors). By introducing a set of learnable temporal memory weights, the value vectors of the current time step and the adjacent historical time steps are weighted and aggregated to generate a local feature sequence that has undergone local temporal smoothing and enhancement. This strengthens the ability to perceive the micro-temporal structure in the speech signal, thereby capturing the instantaneous fluctuations and prosodic details of emotional expression, complementing the global features.

[0088] Finally, the global feature sequence obtained through global attention processing is integrated with the local feature sequence enhanced by temporal memory. This process preserves both long-distance emotional logic connections and local acoustic dynamic details, such as through residual connections, splicing, or weighted summation, generating a final attention feature sequence that combines global perspective and local precision. This achieves deep fusion of multi-scale temporal features, enabling the generated attention feature sequence to comprehensively reflect multi-level information in the speech signal, from macro-level emotional tone to micro-level emotional fluctuations, thus improving the emotional discrimination power and robustness of the acoustic hidden state.

[0089] Optionally, the projection and causal convolution enhancement unit includes a linear projection layer and a causal one-dimensional convolution layer; using the projection and causal convolution enhancement unit, the attention feature sequence is projected and enhanced to obtain the acoustic hidden state sequence, including: using the linear projection layer to perform a linear transformation on the attention feature sequence to obtain the hidden dimension sequence of the speech signal; using the causal one-dimensional convolution layer to perform convolution processing on the hidden dimension sequence to obtain the acoustic hidden state sequence.

[0090] The linear projection layer described above is part of the projection and causal convolution enhancement unit. The linear projection layer can be a fully connected layer used to perform dimensionality mapping, transforming high-dimensional or low-dimensional attention features into the hidden dimensions required by the acoustic encoder.

[0091] Since acoustic encoders may output features of different dimensions, while speech emotion recognition models require input of specific dimensions, linear projection layers achieve flexible transformation of the feature space through learnable parameters, ensuring that acoustic information is not distorted during projection. At the same time, it provides a suitable input format for subsequent causal convolution. Linear projection is the basis for dimension alignment, enabling multimodal information to be fused in a unified vector space, thereby improving the overall efficiency and accuracy of speech emotion recognition models.

[0092] The aforementioned causal one-dimensional convolutional layer is the core component of the projection and causal convolution enhancement unit. The receptive field of the causal one-dimensional convolutional layer is limited to the current and past time steps, enhancing the dynamic connections between speech frames through local convolution.

[0093] Because causal convolution expands the receptive field through sparse connections (dilation coefficients) while maintaining causality, speech emotion recognition models can capture long-range local trends, such as slow changes in intonation, without introducing future information. This characteristic is crucial for real-time emotion recognition, ensuring the real-time performance and accuracy of the speech emotion recognition model's output. Causal one-dimensional convolutional layers enhance local temporal features, highlighting short-term energy changes and prosodic patterns in speech, providing important acoustic cues for emotion recognition.

[0094] In one optional embodiment, a linear projection layer is used to linearly transform the attention feature sequence. For example, through a learnable fully connected layer matrix multiplication operation, the dimension of the feature vector output by the self-attention mechanism is mapped to the dimension of the hidden state preset by the acoustic encoder. This solves the problem of inconsistency between the self-attention output dimension and the subsequent acoustic hidden state dimension, achieving unified alignment of the feature space. This compresses or expands the attention features containing global and local temporal information to the target dimension, completing the format conversion of the hidden features. This lays a foundation for dimension consistency for subsequent projection mapping to the hidden vector space of the speech perception recognition model, enhancing the expressive power of the features.

[0095] Then, a causal one-dimensional convolutional layer is used to convolve the hidden dimension sequence. A causally constrained one-dimensional convolutional kernel slides along the time axis, using information from the current time step and previous time steps to update the features at the current time step. This ensures no future information leakage is introduced, maintaining the causal consistency of temporal processing. Simultaneously, the convolutional operation captures feature change patterns within local time windows. Through the sliding computation of multiple one-dimensional convolutional kernels, local feature enhancement and nonlinear transformation are performed on the hidden dimension sequence, highlighting fine-grained acoustic cues such as short-term energy changes, local prosodic fluctuations, and pause patterns in the speech signal. This further strengthens the temporal local features of the acoustic hidden state sequence, resulting in an acoustic hidden state that not only possesses global semantic alignment capabilities but also retains precise local acoustic dynamic details. This improves the emotional discriminative power of the acoustic features, ensuring that the acoustic embedding sequence input to the speech emotion recognition model is semantically aligned, temporally complete, and free from future information contamination, thus enhancing the accuracy and interpretability of the speech emotion recognition system.

[0096] Optionally, the acoustic hidden state at any time step can be determined by a causal one-dimensional convolutional layer using the hidden dimension of any time step and the hidden dimension of the time step preceding any time step.

[0097] In one optional embodiment, the acoustic hidden state at any time step is determined by a causal one-dimensional convolutional layer using the hidden dimension of any time step and the hidden dimension of any time step prior to that time step. This clarifies the computational constraint mechanism of causal convolution in the time dimension, meaning the output feature of the current time step depends on the input feature at the current moment and all historical input features earlier than the current moment, while excluding information from future moments. This ensures the causality of speech signal processing and prevents the speech perception recognition model from using future information to predict the current state during inference or training. This ensures that the speech perception recognition model can meet the needs of real-time speech stream processing in actual deployment and conforms to the linear temporal order logic of human sound perception. By setting a mask or a special convolutional kernel structure, the convolution operation can be made unidirectional on the time axis, making the acoustic hidden state vector of the current time step an accumulation and refined expression of historical acoustic context information. By establishing directional constraints for the internal temporal modeling of the acoustic encoder, the causal consistency of the generated acoustic hidden state sequence in the temporal structure is ensured. This not only avoids the risk of overfitting caused by information leakage, but also makes the emotional features learned by the speech emotion recognition model more realistically reflect the evolution of speech over time.

[0098] Optionally, the projection mapping model is a multilayer perceptron; the acoustic hidden state sequence is projected from the acoustic feature space to the hidden vector space corresponding to the speech perception recognition model using the projection mapping model to obtain the acoustic embedding sequence, including: using the multilayer perceptron, the acoustic hidden state sequence is processed by linear processing, activation processing, regularization processing and linear processing respectively to obtain the acoustic embedding sequence, wherein the multilayer perceptron processes each time step of the acoustic hidden state sequence with the same network parameters, and the time steps of the acoustic hidden state sequence and the acoustic embedding sequence are the same.

[0099] The aforementioned multilayer perceptron is a neural network composed of multiple fully connected layers and nonlinear activation functions. The multilayer perceptron maps acoustic features to the hidden space of the speech emotion recognition model through nonlinear transformations, achieving alignment between the acoustic space and the language space. By sharing parameters, the multilayer perceptron performs the same mapping at each time step, maintaining the sequence structure. Simultaneously, the nonlinear activations introduce complex feature transformation capabilities, allowing acoustic information to exist in a higher-level semantic form within the speech emotion recognition model. Shared parameters ensure the consistency of the mapping and avoid deviations between time steps.

[0100] In one optional embodiment, a mapping structure is constructed comprising two fully connected layers and intermediate nonlinear transformations and regularization operations. For example, the acoustic hidden states are mapped to an intermediate high-dimensional space, then a nonlinear factor is introduced through an activation function, followed by regularization to prevent overfitting, and finally, linear processing is used to map the features to the hidden vector dimension of the speech emotion recognition model. This achieves a high-dimensional nonlinear mapping from the feature space specific to the acoustic encoder to the general semantic space of the speech emotion recognition model, ensuring that acoustic features can be effectively understood and processed by the speech emotion recognition model. By converting the original acoustic-dimensional hidden states into a vector form consistent with the word embedding dimension of the speech emotion recognition model, and simultaneously capturing the complex nonlinear relationships between acoustic and linguistic features through nonlinear transformations, a bridge for cross-modal feature transformation is established. This allows previously incompatible acoustic representations to find their corresponding geometric positions in the semantic space of the speech emotion recognition model, thereby allowing the speech emotion recognition model to directly utilize these acoustic vectors for subsequent emotion inference and generation.

[0101] Furthermore, the multilayer perceptron processes the acoustic hidden state sequence using the same network parameters at each time step, and the acoustic hidden state sequence and the acoustic embedding sequence have the same time steps. Through the "weight sharing" characteristic of the multilayer perceptron—that is, the weight matrix and bias vector in the network are fixed at all time steps, and the length of the input sequence remains consistent before and after projection—without temporal downsampling or upsampling operations, the temporal sequence structure of the speech signal is maintained, avoiding the loss or misalignment of temporal information introduced during projection. This ensures that each acoustic time step can be independently and consistently mapped to the hidden vector space. Moreover, by sharing parameters, computational complexity and memory usage are reduced, while ensuring that the acoustic features at each time step have the same transformation logic in the language space. This ensures the temporal alignment of the acoustic embedding sequence and the text embedding sequence, enabling the speech emotion recognition model to accurately correspond to the acoustic emotion information and speech recognition command at each time step, maintaining the structure of the temporal features of speech emotion.

[0102] Optionally, the preprocessing unit includes a time-frequency transformation unit and a filter; the preprocessing unit performs initial feature extraction on the speech signal to obtain an acoustic feature sequence, including: using the time-frequency transformation unit to perform Fourier transform on the speech signal to obtain the spectral features of the speech signal; and using the filter to compress the spectral features according to the time sequence to obtain the acoustic feature sequence.

[0103] In one optional embodiment, a time-frequency transformation unit is used to perform a Fourier transform on the speech signal, decomposing the original waveform signal in the time domain into energy distributions of different frequency components, thereby generating a spectrum. This transforms a non-stationary time-domain signal that is difficult to analyze directly into a time-frequency representation with clear physical meaning in the frequency domain, extracting acoustic information such as the fundamental frequency, formant frequencies, and energy distribution from the speech signal. By mapping the signal amplitudes on the time and frequency axes to complex or real matrices, spectral features reflecting the spectral structure of the speech signal are formed, revealing the structural characteristics of the speech signal in the frequency domain.

[0104] Then, filters can be used to compress the spectral features according to the time series, and weighted summation can be performed on the linear spectral features to simulate the nonlinear perception characteristics of human ears regarding frequency. That is, the human ear is more sensitive to changes in low frequencies and relatively insensitive to changes in high frequencies, thereby reducing the dimensionality of the spectral features, removing redundant information, and extracting acoustic features that conform to human auditory perception characteristics. For example, by converting the spectrum of the linear frequency axis into a log-Mel spectrum, and smoothing spectral details by merging adjacent frequencies, the main energy envelope is retained, thereby realizing the transformation from physical acoustic features to perceptual acoustic features. This not only reduces the feature dimensionality and thus reduces computational complexity, but also enhances the robustness of features to noise, making the extracted acoustic feature sequence more consistent with the human perception mechanism of speech emotion, providing high-quality, low-redundancy input for deep feature extraction of acoustic encoders.

[0105] The acoustic embedding sequence and speech recognition command are input into the speech emotion recognition model. The speech signal is then classified and recognized using the speech emotion recognition model to obtain the classification and recognition results. This includes: performing word embedding processing on the speech recognition command to obtain the text embedding sequence corresponding to the speech recognition command; concatenating the acoustic embedding sequence and the text embedding sequence to obtain the recognition sequence; and inputting the recognition sequence into the speech emotion recognition model to obtain the classification and recognition results.

[0106] In one optional embodiment, the speech recognition instructions are processed by word embedding, such as by looking up a pre-trained word embedding table or a finely tuned embedding layer, to convert discrete text instruction character sequences (such as "Please identify the emotion category of this speech") into continuous, dense text embedding sequences. This transforms the speech recognition instructions into numerical inputs that the speech emotion recognition model can understand and process, making instructions with similar semantics or syntax close in distance in the vector space. By mapping symbolic text information to the same high-dimensional semantic space as the acoustic embedding sequence, a mapping relationship is established between the speech recognition instructions and the internal representation of the speech emotion recognition model. This allows the speech emotion recognition model to clearly understand the goal, constraints, and output format requirements of the current task, thereby guiding the speech emotion recognition model to perform emotion reasoning under specific instruction constraints.

[0107] Then, the acoustic embedding sequence and the text embedding sequence are concatenated. For example, in the time dimension, the acoustic embedding sequence representing the acoustic emotion of the speech is connected with the text embedding sequence representing the task instruction to form a unified long sequence, constructing a complete context. This allows the speech emotion recognition model to simultaneously see the emotional cues to be processed and the specific task instructions to be executed. Through sequence concatenation, the features of the two different modalities are fused into a single data stream, forming the input of the speech emotion recognition model. This enables the speech emotion recognition model to focus on specific parts of the acoustic sequence based on the speech recognition instructions, thereby realizing an instruction-driven emotion recognition mechanism. This makes the classification and recognition results no longer a fixed, single output, but rather allows for flexible adjustment of the candidate category range or output format according to the speech recognition instructions.

[0108] In other words, the recognized sequence is input into the speech emotion recognition model, which then performs deep semantic understanding and reasoning. Finally, the corresponding emotion category label or description is generated through the output layer. By leveraging the context modeling and instruction-following capabilities of the speech emotion recognition model, combined with the emotional information contained in the acoustic embedding sequence, accurate classification and recognition results that meet the instruction requirements are generated. This completes the closed loop from multimodal input to final emotion decision-making, which not only improves the accuracy of emotion recognition (especially for semantically neutral but emotionally rich speech), but also achieves flexible and controllable output format for emotion recognition.

[0109] Optionally, the above method further includes: acquiring training data, wherein the training data includes a speech training signal, a classification label corresponding to the speech training signal, and a text training instruction; obtaining a training embedding sequence based on the speech training signal and the text training instruction; and training an initial model based on the training embedding sequence and the classification label to obtain a speech emotion recognition model.

[0110] In one optional embodiment, training data is collected, including the original speech waveform, the corresponding manually labeled emotion categories (such as discrete labels or continuous dimension ratings such as happy, sad, angry, etc.), and text training instructions (such as "Please judge the emotion of this speech, the options are neutral or angry") to guide the output of the speech emotion recognition model. The purpose is to provide the speech emotion recognition model with the "input-output" mapping relationship required for supervised learning, to ensure that the speech emotion recognition model can learn the association between acoustic features and specific emotion labels and instruction constraints, and to provide a source of samples for the training of the speech emotion recognition model.

[0111] Then, the speech training signals and text training instructions are uniformly converted into homogeneous vector representations acceptable to the initial model, constructing a training embedding sequence for updating the initial model parameters. This training embedding sequence is then input into the initial model, and the model parameters are updated using the classification labels and backpropagation algorithm to reduce prediction errors. This enables the speech emotion recognition model to learn to accurately predict the corresponding emotion category based on the training embedding sequence.

[0112] Optionally, the initial model is trained based on the training embedding sequence and classification labels to obtain a speech emotion recognition model, including: inputting the training embedding sequence into the initial model to obtain an initial classification result; constructing a loss function based on the initial classification result and classification labels; and adjusting the parameters of the initial model based on the loss function to obtain the speech emotion recognition model.

[0113] In one alternative embodiment, the training embedding sequence is input into the initial model to generate the initial model's sentiment prediction result for the current input sample, i.e., the initial classification result, which transforms the implicit vector representation into explicit class probabilities, providing predicted values ​​for subsequent loss calculations.

[0114] Then, based on the initial classification results and labels, a loss function is constructed. For example, the cross-entropy loss function is used to measure the difference between the probability distribution of the initial model output and the true label distribution. The loss is calculated for the output positions corresponding to the target sentiment category, while the input positions corresponding to the acoustic embedding and text instructions are masked and not included in gradient calculation. This quantifies the degree of error in the initial model's predictions, providing a clear direction for parameter updates. The degree of classification error is transformed into a differentiable numerical scalar, allowing the initial model to find parameter adjustment directions to reduce error through gradient descent. By reducing the difference between the predictions and the true labels, the initial model is guided to gradually correct its misunderstanding of acoustic sentiment features, thereby improving recognition accuracy.

[0115] Furthermore, the parameters of the initial model are adjusted based on the loss function, shifting them towards reducing the loss. This eventually converges to a set of parameter values ​​that accurately reflect the relationship between speech emotion and labels, thus obtaining the speech emotion recognition model. Additionally, the parameters of the projection mapping module and the low-rank adaptation module in the speech emotion recognition model can be updated, while the parameters of the acoustic encoder and the backbone network of the speech emotion recognition model are frozen. This achieves efficient and low-cost fine-tuning, utilizing the semantic understanding and instruction-following capabilities of the pre-trained speech emotion recognition model while adapting to a specific speech emotion recognition task with a small number of trainable parameters, avoiding the computational overhead and catastrophic forgetting risk associated with full-scale fine-tuning.

[0116] According to an embodiment of this application, an embodiment of a speech recognition system is also provided.

[0117] This embodiment provides a speech recognition system. Figure 2 This is a schematic diagram of a speech recognition system according to an embodiment of this application, such as... Figure 2 As shown, the system includes:

[0118] The voice sampling module 20 is used to obtain the voice signal corresponding to the received voice recognition command.

[0119] The acoustic coding module 22 is used to extract acoustic features from the speech signal using an acoustic encoder to obtain the acoustic hidden state sequence of the speech signal.

[0120] The projection mapping module 24 is used to project the acoustic hidden state sequence from the acoustic feature space to the hidden vector space corresponding to the speech recognition model using the projection mapping model, so as to obtain the acoustic embedding sequence, wherein the acoustic embedding sequence is aligned with the text embedding sequence corresponding to the speech recognition command.

[0121] The emotion recognition output module 26 is used to input the acoustic embedding sequence and speech recognition command into the speech emotion recognition model, and use the speech emotion recognition model to classify and recognize the speech signal to obtain the classification and recognition result.

[0122] This system achieves accuracy and controllability in speech emotion recognition by constructing a closed-loop processing chain of acoustic feature extraction, cross-modal space alignment, and multimodal command fusion. The acoustic coding module extracts deep temporal features, preserving acoustic characteristics such as intonation and energy. The projection mapping module linearly maps and projects the heterogeneous acoustic hidden states onto the isomorphic hidden vector space of the speech emotion recognition model. The emotion recognition output module inputs the projected acoustic embedding and speech recognition commands into the speech emotion recognition model. This model not only understands the emotional prosody in speech but also adheres to the category constraints and format requirements of the speech recognition commands, thus accurately recognizing subtle emotional changes in complex scenarios and flexibly adapting to different command requirements.

[0123] The technical solution proposed in this application is described below with reference to an optional embodiment. This application proposes a speech emotion recognition method based on acoustic projection adaptation. This method can realize speech emotion recognition, acoustic signal processing, cross-modal representation learning, and large language model adaptation. The method encodes the hidden state of speech using an acoustic encoder and uses projection mapping to match the acoustically encoded speech to vectors in the hidden space of a large language model, thereby realizing speech emotion categories. It can be applied to various human-machine interfaces capable of voice interaction, such as companion robots, service robots, in-vehicle smart cockpits, and smart speakers, to achieve speech emotion detection and emotion tag generation on different smart terminals.

[0124] Currently, speech emotion recognition uses acoustic features to feed into a fixed classifier, and the output is usually several discrete emotion labels. However, the model often lacks adaptability to task expression, candidate categories, and output formats, making it difficult to adapt the recognition results to different situations. Methods based on deep acoustic encoders can extract emotional information carried by various signals such as prosody, energy, speech rate, pitch changes, and pauses in speech, but the output acoustic hidden state space is inconsistent with the hidden state space of large language models, and cannot be directly used as input to large language models. Therefore, large language models cannot effectively use the emotional information in the acoustic hidden state. Some solutions convert speech to text and then perform text sentiment analysis, but this approach relies too heavily on semantic content and ignores non-textual emotional information such as intonation, energy, speed, stress, and pauses. When the content spoken by the user is relatively neutral, but the tone indicates that the speaker is depressed, angry, excited, or anxious, the accuracy of speech emotion labeling is poor. In addition, voice emotion recognition systems often use pooling layers or fixed classification heads to simplify variable-length speech sequences into fixed-length features before emotion classification. This ignores different emotions in different time segments of the speech and results in the loss of acoustic temporal information in the speech.

[0125] This embodiment addresses the challenges of directly integrating acoustic encoder outputs into large language models, the inconsistency between acoustic hidden states and the hidden vector space of large language models, insufficient controllability of fixed classification head output methods, and limited stability of emotion recognition in complex speech scenarios. For example, the acoustic hidden representations emitted by existing acoustic encoders cannot be directly used in large language models. Without a suitable projection adaptation structure, large language models cannot effectively utilize the emotional information carried in the acoustic hidden representations. Conversely, emotion recognition models use fixed classification heads to classify emotions, failing to incorporate instructions from natural language representation tasks to flexibly constrain candidate emotion categories and result formats.

[0126] Therefore, this embodiment proposes a new speech emotion recognition scheme, which includes an acoustic encoder, a projection mapping network, and a large language model working together to extract deep emotional features from speech. The projection mapping network maps the acoustic hidden state of speech onto the hidden vector space of the large language model, and then concatenates it with the embedded text task instructions. This enables the large language model to complete stable and controllable emotion classification in the speech acoustic information under the guidance of text instructions, thereby accurately and controllably completing speech emotion recognition.

[0127] This embodiment also provides a speech emotion recognition system based on acoustic projection adaptation. The system includes: a speech acquisition and preprocessing module, an acoustic coding module, a projection mapping module, an acoustic coding embedding and text command mixing module, and an emotion recognition output module.

[0128] The speech sampling and preprocessing module completes speech preprocessing such as user speech sampling, speech signal resampling, channel normalization, spectrum transformation, and acoustic feature compression, thereby acquiring the user's speech signal and performing feature transformation on the speech signal to obtain an acoustic feature sequence suitable for the acoustic encoder.

[0129] The acoustic coding module implements deep temporal modeling of the preprocessed acoustic feature sequence to obtain an acoustic hidden state sequence carrying pitch, intensity, speech rate, phonology, and emotional information.

[0130] The projection mapping module projects the hidden state of the acoustic encoder from the acoustic feature space to the hidden vector space of the large language model, so that the acoustic features become a continuous embedding sequence and are fed into the large language model.

[0131] The acoustic coding embedding and text instruction hybrid module concatenates the acoustic projection embedding sequence received by the large language model with the text embedding sequence corresponding to the emotion recognition text instruction to obtain the final joint embedding sequence input by the large language model.

[0132] The emotion recognition output module realizes the output of emotion recognition, and predicts the emotion category corresponding to the user's voice through a large language model.

[0133] The aforementioned speech emotion recognition system effectively connects the speech acoustic encoder with a large language model, enabling the large language model to predict emotion categories using speech acoustic features under the constraints of text task instructions. This avoids problems such as fixed output methods relying on fixed classification heads, difficulty in integrating large language models with speech acoustic features, and difficulty in controlling emotion classification performance.

[0134] like Figure 3 As shown, the system implementation process is as follows: After the original audio undergoes Mel-spectrum extraction and frame stacking downsampling, it is processed by an acoustic encoder and a mapper. Projected acoustic embedding and recognition instructions are then embedded to achieve emotion recognition. Specifically, the acoustic encoder, after processing by a single-modal feedforward temporal memory encoding unit, a memory-enhanced self-attention temporal modeling unit, and a projection and causal convolution enhancement unit, obtains the acoustic hidden state sequence. The mapper includes a first linear layer, an activation layer, a regularization layer, and a second linear layer.

[0135] The projection mapping module addresses the inconsistency between the output dimension of the acoustic encoder and the hidden dimension of the large language model. It inputs the speech acoustic hidden states as a continuous embedding sequence into the large language model, aligning the acoustic hidden feature space with the hidden vector space of the large language model across different modalities. By combining the projected acoustic embeddings with natural language task instructions, the model is constrained by the range of candidate emotion categories and the output format, enhancing the controllability of speech emotion recognition. Furthermore, the projection mapping module projects data from each acoustic time step, ensuring that each time step receives the same mapping. The projection process does not affect the acoustic time series structure or change the feature dimension, further preserving speech emotion information across different time periods. By combining the acoustic representation capabilities of the pre-trained acoustic encoder with the instruction understanding capabilities of the large language model, the stability, accuracy, and scalability of speech emotion recognition in complex speech scenarios are improved. Moreover, a lightweight large language model is utilized, combined with low-rank adaptation for low-rank training, primarily training the projection mapping module and the adaptation part. This reduces training costs while improving the system's deployability and transferability.

[0136] This embodiment also provides a speech emotion recognition method based on acoustic projection adaptation, such as... Figure 4 As shown, an alternative speech emotion recognition method is presented. This method includes speech signal acquisition and normalization preprocessing, deep feature extraction from an acoustic encoder, projection mapping of acoustic hidden states, and emotion category generation based on a large language model.

[0137] In the first step, the speech signal is acquired by a microphone or other speech acquisition device. After the speech signal is acquired, it is first processed to unify the audio channels. If the input speech is multi-channel, it is first converted from multi-channel audio to mono audio. The speech signal is then resampled to ensure that the sampling rate of speech from different sources is consistent at 16kHz, thus ensuring that speech from different sources has the same input speech signal format.

[0138] After unifying the sampling rate, the speech signal is subjected to Short-Time Fourier Transform (SFT) and Mel filter bank feature extraction. In this implementation, the SFT uses a Fast Fourier Transform (FFT) with 512 points, a window length of 400 points, and a frame shift of 160 points, and extracts 80-dimensional Log-Mel (log-Mel spectrum) filter bank features. To avoid instability in logarithmic values ​​due to excessively small values, a lower bound is first applied to the energy of the Mel spectrum before taking the logarithm, thus obtaining a stable Log-Mel acoustic feature sequence.

[0139] To simultaneously preserve local contextual information of consecutive speech frames, reduce sequence length, and better utilize subsequent coding efficiency, this system further performs LFR (low frame-rate processing) frame stacking and downsampling on consecutive Log-Mel (log-Mel spectrum) feature frames. Specifically, the LFR frame stacks 7 consecutive Log-Mel (log-Mel spectrum) feature frames and then downsamples them with a step size of 6 frames. Assuming that each Log-Mel (log-Mel spectrum) feature is 80-dimensional, after stacking the frames, the acoustic feature dimension at each time step is 560. The compressed acoustic feature sequence is thus represented by the following formula:

[0140] ;

[0141] in, This is the compressed acoustic feature sequence. Let T be the set of real numbers, where T is the number of acoustic time steps obtained by stacking and downsampling LFR frames, and the feature dimension of each time step is 560.

[0142] In the second step, the acoustic feature sequence is input into the acoustic coding module to perform deep temporal modeling on the input sequence, extracting pitch changes, energy changes, speech rate changes, pause features, prosodic patterns, and emotional expression cues contained in different time segments of speech.

[0143] Specifically, the acoustic coding module includes a single-modal feedforward temporal memory coding unit, a memory-enhanced self-attention temporal modeling unit, and a projection and causal convolution enhancement unit.

[0144] A single-modal feedforward temporal memory encoding unit is used to extract preliminary representation features from the input acoustic features, incorporating a finite-length temporal memory model into the feedforward network. This temporal memory is not an externally readable and writable memory matrix, but rather a one-dimensional memory convolution or sliding context summation along the linear direction. For the input features at time step t... The coding unit first obtains an intermediate representation through a linear transformation and activation function. Then, information from several historical time steps at time step t is aggregated to obtain enhanced memory features. In the causal coding scenario of this embodiment, the current time step and historical time steps are used, while future time steps are not used to avoid future information leakage. Therefore, the single-modal feedforward temporal memory coding unit can model the contextual information between speech frames without applying a recurrent network structure.

[0145] Enhanced memory features are represented by the following formula:

[0146]

[0147] Among them, enhanced memory features , Indicates the length of sequential memory. This represents the learnable weight matrix or convolutional weights corresponding to the k-th historical time step. This is the intermediate representation at the (tk)th time step. It can be set between 3 and 15, for example, it can be 8.

[0148] Then, the current representation With enhanced memory features Fusion is represented by the following formula:

[0149] ;

[0150] in, Let be the contextual features at time step t.

[0151] Memory-enhanced self-attention temporal modeling units are used to further model long-distance dependencies between different time segments in a speech sequence. For the input sequence X, the query matrix is ​​calculated respectively. Key matrix Sum matrix The query matrix, key matrix, and value matrix are represented as follows:

[0152] ;

[0153] in, Let represent the learnable linear projection weight matrices, respectively.

[0154] And calculate the correlation between points in the scaled dot product at different time steps, such as using the following formula:

[0155] ;

[0156] Among them, the query matrix Key matrix Sum matrix ; For matrix multiplication, calculate the dot product similarity between the query vector and the key vector; Represents the dimension of the key vector. This represents an attention mask.

[0157] The memory-enhanced self-attention temporal modeling unit introduces a temporal memory branch in addition to the self-attention branch, performing finite context modeling on the value vector V along the time direction to obtain a new value vector. :

[0158] ;

[0159] in, The memory capacity of the self-attention unit is represented by t, which represents the time step. This represents the learnable temporal memory weights. It can be set between 3 and 31, for example, it can be 15.

[0160] This memory-enhanced self-attention temporal modeling unit can capture global temporal correlations using self-attention mechanisms, and retain dynamic change information in local continuous frames through temporal memory branches, thereby enhancing the ability to model changes in phonological prosody and emotional expression.

[0161] Projection and causal convolutional augmentation units are used to perform feature dimension transformation and local temporal augmentation on the output of the memory-enhanced self-attention temporal modeling unit.

[0162] The projection and causal convolutional enhancement unit can consist of a linear projection layer and a causal one-dimensional convolutional layer. The linear projection layer is used to transform the input features to the hidden dimension of the acoustic encoder output, while the causal one-dimensional convolutional layer can be used to enhance the dynamic features of local continuous speech frames along the time direction.

[0163] For time step t, causal convolution involves computation of the current time step and previous time steps, as represented by the following formula:

[0164]

[0165] in, Causal convolution output features at time step t Indicates the kernel size. Indicates the coefficient of thermal expansion. Represents the convolution kernel parameters. This represents the input features after linear projection; This represents the input feature vector corresponding to the offset r in dilated causal convolution. The number of causal convolution layers can be set to 1 to 3, and the kernel size can be set to 3 to 7, for example, 3. The stride is set to 1. The dilation coefficient can be set to 1, 2, or 4. Through the causal convolution enhancement unit, the acoustic coding module can further highlight emotion-related acoustic cues in speech, such as short-term energy changes, local prosodic changes, and pause changes.

[0166] After processing by the acoustic coding module described above, the original speech is transformed into a 512-dimensional acoustic hidden state sequence with temporal context and emotional expression information.

[0167] Acoustic hidden state sequence It can be represented as:

[0168] ;

[0169] Each acoustic hidden state satisfies: The overall acoustic hidden state sequence satisfies: T represents the number of compressed acoustic time steps.

[0170] In the third step, since the dimension of the acoustic hidden state output by the acoustic encoder is not necessarily the same as the dimension of the hidden vector of the large language model, the acoustic hidden state cannot be directly embedded as the input of the large language model. Therefore, projection mapping is used to map the acoustic hidden state to the hidden vector space of the large language model. This achieves alignment between the output space of the acoustic encoder and the hidden vector space of the large language model, and the speech acoustic information is represented as a continuous embedding sequence and input to the large language model.

[0171] The projection mapping part is a shared multilayer perceptron. The projection mapping module maps the acoustic hidden state sequence using the same network parameters at each time step, without increasing the number of time steps in the acoustic sequence. It maps each acoustic hidden state from 512 dimensions to the hidden dimension of the large language model. For example, when the hidden dimension of the large language model is 2048, the projection mapping part consists of a first linear layer, GELU (Gaussian Error Linear Unit) activation, Dropout (random deactivation) regularization, and a second linear layer, i.e.: .

[0172] For any acoustic hidden state The projection process can be represented as:

[0173]

[0174] in, and These represent the weights and biases of the first linear layer, respectively. and These represent the weights and biases of the second linear layer, respectively.

[0175] After projection mapping, the projected acoustic embedding sequence is obtained. Represented as: ,in, , .

[0176] In the fourth step, the text task instructions are input into the word embedding layer of the large language model to obtain the text instruction embedding sequence. Subsequently, the acoustic embedding sequence output by the projection mapping module is... Embedded sequences of text instructions The sequences are concatenated along the time dimension to form a joint input embedding sequence. :

[0177] ;

[0178] Where Z represents the input embedding sequence of a large language model that simultaneously contains speech acoustic emotion information and text task instructions and constraints of the large language model.

[0179] Then, the joint input embedding sequence is fed into the large language model. The large language model autoregressively outputs the corresponding target sentiment category word based on the emotional acoustic information of the projected acoustic embedding, the candidate categories provided by the text task instructions, and the constraints of the large language model.

[0180] During the training phase, a supervised language model is trained. Each training data set includes a speech data path, a corresponding sentiment category label, and a corresponding text task instruction. The speech data is processed by the speech acquisition and preprocessing module, the acoustic coding module, and the projection mapping module to obtain the projected acoustic embedding training sequence. The text task instructions are processed through the word embedding layer of the large language model to obtain the text task instruction embedding training sequence, and the two are concatenated to obtain the joint training embedding sequence.

[0181] The acoustic coding module retains its pre-trained parameters and does not perform reverse updates. The backbone parameters of the large language model remain unchanged and are not updated backwards. The parameters of the projection mapping module are learnable and are updated backwards. The low-rank adaptation parameters of the large language model backbone are learnable and are updated backwards. Therefore, during training, only the parameters of the projection mapping module and the low-rank adaptation parameters are updated, reducing training costs and improving training ease of use.

[0182] The low-rank adaptation module operates on the attention item layer and feedforward layer of the large language model. Specifically, the low-rank adaptation module can be configured at the query item, value item, and item output of the attention layer, the gate item of the feedforward layer, and the up and down projections. The rank r of the low-rank adaptation layer can be set to 16, the scaling factor α can be set to 32, and the dropout of the adaptation layer can be set to 0.1. These values ​​can be adjusted according to the size of the large language model and the size of the training data.

[0183] The training loss uses the cross-entropy loss of an autoregressive language model. Labels corresponding to the projected acoustic embedding and text task instructions are masked and not included in the loss calculation. The loss is calculated for the target sentiment category answer. Let the target answer sequence be (Y=[ , ,…, If ]), then the training objective can be expressed as:

[0184] ;

[0185] Where Z represents the joint input embedding sequence, ( The ) represents the generated words preceding the nth word in the target answer, where N is the length of the target sentiment answer sequence, i.e., the number of words in the answer. The nth real word in the target sentiment answer sequence.

[0186] For example, the target answer might be a sentiment word among candidate answers such as neutral, happy, sad, or angry. The model is trained using this loss function to learn the speech acoustic embedding and text task instructions to obtain the correct sentiment category for the target answer.

[0187] In summary, this embodiment constructs an acoustic encoder consisting of a feedforward temporal memory encoding unit, a memory-enhanced self-attention temporal modeling unit, and a projection and causal convolution enhancement unit. Through multi-layer temporal modeling, it progressively extracts prosody, energy, speech rate, pitch, and long-term dependency features from speech, achieving deep acoustic representation learning for emotion recognition. Combining temporal memory modeling with self-attention modeling utilizes both the feedforward temporal memory structure to capture local temporal contextual information and the memory-enhanced self-attention mechanism to establish long-distance temporal associations, thereby improving the expressive power of speech emotion information. Furthermore, a shared projection mapping network is used to map the hidden state sequence output by the acoustic encoder to the hidden space of a large language model, achieving unified alignment between acoustic and linguistic representations, allowing speech acoustic information to be directly integrated into the large language model in the form of embedded sequences. The projected acoustic embeddings and text task instructions are input into the large language model. While preserving the speech emotion expression information, the contextual modeling and instruction constraint capabilities of the large language model are used to complete emotion category recognition, improving the controllability and generalization ability of the recognition results.

[0188] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0189] According to an embodiment of this application, a device embodiment for a speech recognition apparatus is provided. It should be noted that the apparatus can be used to perform the speech recognition method described above.

[0190] This embodiment provides a voice recognition device. Figure 5 This is a schematic diagram of a voice recognition device according to an embodiment of this application, such as... Figure 5 As shown, the device includes:

[0191] The acquisition module 50 is used to acquire the voice signal corresponding to the received voice recognition command.

[0192] Extraction module 52 is used to extract acoustic features from speech signals using an acoustic encoder to obtain the acoustic hidden state sequence of the speech signals.

[0193] Projection module 54 is used to project the acoustic hidden state sequence from the acoustic feature space to the hidden vector space corresponding to the speech recognition model using a projection mapping model, so as to obtain an acoustic embedding sequence, wherein the acoustic embedding sequence is aligned with the text embedding sequence corresponding to the speech recognition command.

[0194] The recognition module 56 is used to input the acoustic embedding sequence and speech recognition instructions into the speech perception recognition model, and use the speech perception recognition model to classify and recognize the speech signal to obtain the classification and recognition results.

[0195] Optionally, the extraction module is further configured to: extract features from the speech signal using the preprocessing unit to obtain an acoustic feature sequence; the acoustic encoder includes a single-modal feedforward temporal memory coding unit, a memory-enhanced self-attention temporal modeling unit, and a projection and causal convolution enhancement unit; the extraction module is further configured to: use the single-modal feedforward temporal memory coding unit to perform temporal enhancement on the acoustic feature sequence corresponding to the speech signal to obtain a context feature sequence; use the memory-enhanced self-attention temporal modeling unit to perform attention processing on the context feature sequence to obtain an attention feature sequence; and use the projection and causal convolution enhancement unit to perform projection enhancement on the attention feature sequence to obtain an acoustic hidden state sequence.

[0196] Optionally, the extraction module is also used to: perform linear transformation and activation function processing on the acoustic features at any time step in the acoustic feature sequence to obtain the intermediate features at any time step; enhance the intermediate features at any time step using the historical intermediate features of historical time steps to obtain the enhanced features at any time step, where the historical time step is the time step before any time step; and fuse the intermediate features at any time step with the enhanced features at any time step to obtain the context features at any time step.

[0197] Optionally, the extraction module is also used to: perform weighted summation of the intermediate features at any time step and the historical intermediate features to obtain the enhanced features at any time step.

[0198] Optionally, the memory-enhanced self-attention temporal modeling unit includes a global enhancement unit, a temporal memory unit, and a fusion unit; the extraction module is further configured to: use the global enhancement unit to perform global attention processing on the context feature sequence to obtain the global feature sequence of the context feature sequence; use the temporal memory unit to perform temporal enhancement processing on the value matrix of the context feature sequence to obtain the local feature sequence of the context feature sequence, wherein the memory parameters of the temporal memory unit are different from those of the unimodal feedforward temporal memory encoding unit; and use the fusion unit to fuse the global feature sequence of the context feature sequence and the local feature sequence of the context feature sequence to obtain the attention feature sequence.

[0199] Optionally, the projection and causal convolution enhancement unit includes a linear projection layer and a causal one-dimensional convolution layer; the extraction module is also used to: use the linear projection layer to perform a linear transformation on the attention feature sequence to obtain the hidden dimension sequence of the speech signal; and use the causal one-dimensional convolution layer to perform convolution processing on the hidden dimension sequence to obtain the acoustic hidden state sequence.

[0200] Optionally, the acoustic hidden state at any time step can be determined by a causal one-dimensional convolutional layer using the hidden dimension of any time step and the hidden dimension of the time step preceding any time step.

[0201] Optionally, the projection mapping model is a multilayer perceptron; the projection module is also used to: use the multilayer perceptron to process the acoustic hidden state sequence through linear processing, activation processing, regularization processing and linear processing respectively to obtain an acoustic embedding sequence, wherein the multilayer perceptron processes each time step of the acoustic hidden state sequence with the same network parameters, and the time steps of the acoustic hidden state sequence and the acoustic embedding sequence are the same.

[0202] Optionally, the preprocessing unit includes a time-frequency transformation unit and a filter; the extraction module is further configured to: use the time-frequency transformation unit to perform Fourier transform on the speech signal to obtain the spectral features of the speech signal; and use the filter to compress the spectral features according to the time sequence to obtain an acoustic feature sequence.

[0203] Optionally, the recognition module is also used to: perform word embedding processing on the speech recognition command to obtain the text embedding sequence corresponding to the speech recognition command; concatenate the acoustic embedding sequence and the text embedding sequence to obtain the recognition sequence; and input the recognition sequence into the speech perception recognition model to obtain the classification recognition result.

[0204] Optionally, the recognition module is further configured to: acquire training data, wherein the training data includes a speech training signal, a classification label corresponding to the speech training signal, and a text training instruction; obtain a training embedding sequence based on the speech training signal and the text training instruction; and train the initial model based on the training embedding sequence and the classification label to obtain a speech emotion recognition model.

[0205] Optionally, the recognition module is also used to: input the training embedding sequence into the initial model to obtain the initial classification result; construct a loss function based on the initial classification result and classification label; and adjust the parameters of the initial model based on the loss function to obtain the speech emotion recognition model.

[0206] Embodiments of this application also provide a vehicle, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods described in various embodiments of this application when it runs.

[0207] This application also provides an electronic device 60, please refer to... Figure 6 It includes a memory 610 and a processor 620, wherein the memory is used to store computer programs; and the processor is used to execute the programs stored in the memory to implement the methods in the various embodiments of this application.

[0208] Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0209] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0210] Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.

[0211] Embodiments of this application also provide a computer program that, when executed by a processor, implements the methods described in the various embodiments of this application.

[0212] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0213] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0214] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0215] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0216] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0217] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A speech recognition method, characterized in that, include: In response to receiving a voice recognition command, the voice signal corresponding to the voice recognition command is acquired; Acoustic features are extracted from the speech signal using an acoustic encoder to obtain the acoustic hidden state sequence of the speech signal; The acoustic hidden state sequence is projected from the acoustic feature space to the hidden vector space corresponding to the speech recognition model using a projection mapping model to obtain an acoustic embedding sequence, wherein the acoustic embedding sequence is aligned with the text embedding sequence corresponding to the speech recognition instruction. The acoustic embedding sequence and the speech recognition command are input into the speech emotion recognition model, and the speech signal is classified and recognized using the speech emotion recognition model to obtain the classification and recognition result.

2. The method according to claim 1, characterized in that, Before extracting acoustic features from the speech signal using an acoustic encoder to obtain the acoustic hidden state sequence of the speech signal, the method further includes: The speech signal is used to extract features using a preprocessing unit to obtain an acoustic feature sequence; The acoustic encoder includes a single-modal feedforward temporal memory coding unit, a memory-enhanced self-attention temporal modeling unit, and a projection and causal convolution enhancement unit; the acoustic encoder is used to extract acoustic features from the speech signal to obtain the acoustic hidden state sequence of the speech signal, including: Using the single-modal feedforward temporal memory coding unit, the acoustic feature sequence corresponding to the speech signal is temporally enhanced to obtain the context feature sequence; The memory-enhanced self-attention temporal modeling unit is used to perform attention processing on the context feature sequence to obtain an attention feature sequence. The attention feature sequence is projected and enhanced using the projection and causal convolution enhancement unit to obtain the acoustic hidden state sequence.

3. The method according to claim 2, characterized in that, Using the single-modal feedforward temporal memory coding unit, the acoustic feature sequence is temporally enhanced to obtain a context feature sequence, including: For the acoustic features at any time step in the acoustic feature sequence, the acoustic features at any time step are subjected to linear transformation and activation function processing to obtain the intermediate features at any time step. By utilizing the historical intermediate features of historical time steps, the intermediate features of any time step are enhanced to obtain the enhanced features of any time step, wherein the historical time step is the time step located before the arbitrary time step. By fusing the intermediate features of any time step with the enhanced features of any time step, the contextual features of any time step can be obtained.

4. The method according to claim 3, characterized in that, By utilizing the historical intermediate features of historical time steps, the intermediate features of any time step are enhanced to obtain the enhanced features of any time step, including: The enhanced features of any time step are obtained by weighted summation of the intermediate features at any time step and the historical intermediate features.

5. The method according to claim 2, characterized in that, The memory enhancement self-attention temporal modeling unit includes a global enhancement unit, a temporal memory unit, and a fusion unit; The memory-enhanced self-attention temporal modeling unit is used to perform attention processing on the context feature sequence to obtain an attention feature sequence, including: The global enhancement unit is used to perform global attention processing on the context feature sequence to obtain the global feature sequence of the context feature sequence; Using the temporal memory unit, temporal enhancement processing is performed on the value matrix of the context feature sequence to obtain the local feature sequence of the context feature sequence, wherein the memory parameters of the temporal memory unit are different from the memory parameters of the single-modal feedforward temporal memory encoding unit; The fusion unit is used to fuse the global feature sequence and the local feature sequence of the context feature sequence to obtain the attention feature sequence.

6. The method according to claim 2, characterized in that, The projection and causal convolution enhancement unit includes a linear projection layer and a causal one-dimensional convolution layer; The attention feature sequence is projected and enhanced using the projection and causal convolution enhancement unit to obtain the acoustic hidden state sequence, including: The attention feature sequence is linearly transformed using the linear projection layer to obtain the hidden dimension sequence of the speech signal; The hidden dimension sequence is convolved using the causal one-dimensional convolutional layer to obtain the acoustic hidden state sequence.

7. The method according to claim 6, characterized in that, The acoustic hidden state at any time step is determined by the causal one-dimensional convolutional layer using the hidden dimension of any time step and the hidden dimension of the time step preceding any time step.

8. The method according to claim 1, characterized in that, The projection mapping model is a multilayer perceptron; the acoustic hidden state sequence is projected from the acoustic feature space to the hidden vector space corresponding to the speech perception recognition model using the projection mapping model to obtain the acoustic embedding sequence, including: Using the multilayer perceptron, the acoustic hidden state sequence is processed through linear processing, activation processing, regularization processing, and linear processing to obtain the acoustic embedding sequence. The multilayer perceptron processes each time step of the acoustic hidden state sequence using the same network parameters, and the time steps of the acoustic hidden state sequence and the acoustic embedding sequence are the same.

9. The method according to claim 2, characterized in that, The preprocessing unit includes a time-frequency conversion unit and a filter; The preprocessing unit performs initial feature extraction on the speech signal to obtain an acoustic feature sequence, including: The speech signal is subjected to Fourier transform using the time-frequency transformation unit to obtain the spectral characteristics of the speech signal; The spectral features are compressed according to the time series using the filter to obtain the acoustic feature sequence.

10. The method according to claim 1, characterized in that, The acoustic embedding sequence and the speech recognition command are input into a speech emotion recognition model. The speech emotion recognition model is then used to classify and recognize the speech signal to obtain the classification and recognition results, including: The speech recognition command is subjected to word embedding processing to obtain the text embedding sequence corresponding to the speech recognition command; The acoustic embedding sequence and the text embedding sequence are concatenated to obtain the recognition sequence; The recognition sequence is input into the speech emotion recognition model to obtain the classification recognition result.

11. The method according to claim 10, characterized in that, The method further includes: Acquire training data, wherein the training data includes a speech training signal, a classification label corresponding to the speech training signal, and a text training instruction; Based on the speech training signal and the text training instruction, a training embedding sequence is obtained; The initial model is trained based on the training embedding sequence and classification labels to obtain the speech emotion recognition model.

12. The method according to claim 11, characterized in that, The initial model is trained based on the training embedding sequence and classification labels to obtain the speech emotion recognition model, including: The trained embedding sequence is input into the initial model to obtain the initial classification result; Based on the initial classification result and the classification label, a loss function is constructed; The parameters of the initial model are adjusted based on the loss function to obtain the speech emotion recognition model.

13. A speech recognition system, characterized in that, include: A voice sampling module is used to obtain the voice signal corresponding to the received voice recognition command in response to the voice recognition command. An acoustic coding module is used to extract acoustic features from the speech signal using an acoustic encoder to obtain an acoustic hidden state sequence of the speech signal; The projection mapping module is used to project the acoustic hidden state sequence from the acoustic feature space to the hidden vector space corresponding to the speech recognition model using a projection mapping model to obtain an acoustic embedding sequence, wherein the acoustic embedding sequence is aligned with the text embedding sequence corresponding to the speech recognition instruction. The emotion recognition output module is used to input the acoustic embedding sequence and the speech recognition command into the speech emotion recognition model, and use the speech emotion recognition model to classify and recognize the speech signal to obtain the classification and recognition result.

14. A vehicle, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the method according to any one of claims 1 to 12.