A voice emotion time-frequency dual-flow interaction algorithm for an emotional companion robot
Patent Information
- Application Number
- CN202611257383.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]在情感陪伴机器人的真实工作场景下,现有技术普遍存在的不足之处具体体现在:现有方法在前端通常采用单一窗长的梅尔频谱作为输入,较长窗长有利于刻画准稳态语音段的精细频率结构而不利于刻画瞬变段的精细时间结构、较短窗长则恰好相反,单一窗长的设计无法在同一段语音中同时获得较优的频率分辨率和时间分辨率,致使在情感承载频带的辨析与情感边界过渡的捕捉之间难以兼顾
[0072]本发明设计了双尺度梅尔特征张量,S1步骤通过对预加重和分帧后的语音信号同时计算长窗梅尔频谱和短窗梅尔频谱,再依据由S15步骤所计算的归一化短时能量映射得到的长窗权重对两路梅尔频谱按通道分别加权后沿通道维堆叠为双尺度梅尔特征张量,使后端编码器在每一帧位置上都依据当前帧的能量特性自动平衡对频率分辨率与时间分辨率的依赖比例。基于上述按帧自适应平衡的特征前端,本发明的方法在多语种、多类别数情感识别任务上对识别准确率有较为稳定的提升,且所述提升在不同语种的公开情感语料库上具有跨语种的一致性。
Smart Images

Figure CN122821998A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech emotion recognition and human-computer interaction technology, specifically involving a speech emotion time-frequency dual-stream interaction algorithm for emotional companion robots. Background Technology
[0002] Emotional companion robots are intelligent devices that provide continuous companionship to users with an anthropomorphic appearance and interactive methods. Their target audience includes elderly people living alone, children, and individuals with emotional support needs. Voice is the most natural interaction medium between the user and the emotional companion robot. The robot needs to identify the user's current emotional state from their continuous speech and then adjust its responses—voice content, speed, tone, and body language—to provide responses tailored to the user's current psychological state during extended companionship. Therefore, voice emotion recognition is one of the key technologies of emotional companion robots.
[0003] Existing speech emotion recognition methods, when applied to continuous companionship interaction scenarios for emotional companion robots, suffer from technical challenges such as the inability to simultaneously achieve optimal time-frequency resolution at the front end, the lack of hierarchical interaction in dual-stream coding, the absence of prior phonological input at the front end, and the difficulty in balancing stability and timeliness in the decoding stage. These issues make it difficult to meet the comprehensive requirements of emotional companion robots for recognition accuracy, training efficiency, category balance, and response stability. The technical path of speech emotion recognition generally includes several stages: front-end preprocessing of the speech signal, feature extraction, feature encoding, and emotion classification. In the front-end preprocessing and feature extraction stages, the most common approach is to use short-time Fourier transform, Mel filter banks, and logarithmic energy mapping on the pre-emphasized and segmented speech signal to obtain the Mel spectrum as input features. In the feature encoding stage, the industry has developed hybrid encoding methods based on convolutional neural networks and recurrent neural networks, methods based on dual-stream parallel coding followed by end-point fusion, and methods based on multi-stream convolution combined with multi-head attention. In the sentiment classification process, the classification layer typically uses a fully connected layer in conjunction with a normalized exponential function to obtain a window-by-window sentiment category probability vector, which is then smoothed by window-by-window hard decision or exponential weighting with fixed coefficients to obtain the output sentiment category.
[0004] In real-world scenarios involving emotional companion robots, existing technologies generally suffer from the following shortcomings: Current methods typically use a single-window-length Mel spectrum as input. While a longer window length is advantageous for depicting the fine frequency structure of quasi-steady-state speech segments, it is less effective for depicting the fine temporal structure of transient segments; conversely, a shorter window length has the opposite effect. This single-window-length design cannot simultaneously achieve optimal frequency and temporal resolution within the same speech segment, making it difficult to balance the identification of emotionally carrying frequency bands with the capture of emotional boundary transitions. Even when existing methods employ a parallel encoding structure for time and frequency streams in the dual-stream encoding stage, the two streams are typically simply spliced or fused at the end of the encoding process. There is no information exchange between the time and frequency streams during encoding, making it difficult to establish deep time-frequency joint representations at a shallow level. This is particularly challenging under conditions of limited training samples and small sample sizes, making it difficult to learn emotion-related frequency band patterns. Existing methods lack prior phonetic information on the physiological structure of human pronunciation at the front end. The model can only learn the distribution of emotion-related frequency bands autonomously from a massive Mel frequency band, resulting in a highly discrete distribution of recognition accuracy for various emotions on a limited emotion corpus. It also lacks sufficient ability to recognize emotion categories that rely on formant shifts and high-frequency energy of voiceless fricatives. In the decoding stage, existing methods often employ window-by-window hard decision or exponentially weighted smoothing with a fixed smoothing coefficient. They cannot automatically adjust the reliance on historical information based on the confidence level of the current observation. This leads to frequent window-by-window jitter when the user's emotion is stable and subject to occasional misidentification, and delayed response due to excessively large smoothing coefficients when the user's emotion truly shifts. Consequently, the robot's response strategy frequently switches or reacts slowly, affecting the consistency and naturalness of long-term companionship interactions.
[0005] Based on the above problems, it is of great significance to design a voice-emotion time-frequency dual-stream interaction algorithm for emotional companion robots, taking into account the application characteristics of emotional companion robots. Summary of the Invention
[0006] To address the problems existing in the background technology, this invention provides a voice-emotion time-frequency dual-stream interaction algorithm for emotional companion robots, comprising the following steps:
[0007] S1: The user voice signal collected by the emotional companion robot is pre-emphasized and framed, and the long-window Mel spectrum and short-window Mel spectrum are weighted and stacked along the channel dimension based on short-time energy to obtain the dual-scale Mel feature tensor.
[0008] S2: Global pooling is performed on the dual-scale Mel feature tensor along the time dimension and frequency band dimension to obtain the frequency dimension feature sequence and the time dimension feature sequence; after applying the frequency band weighting value based on the formant prior to the frequency dimension feature sequence, it is input into the frequency stream encoder, and the time dimension feature sequence is input into the time stream encoder for parallel encoding;
[0009] S3: Perform bidirectional time-frequency interaction attention between every two adjacent coding layers of the time-stream encoder and the frequency-stream encoder, and use the interaction results as input to the next coding layer in the form of residual connections and layer normalization;
[0010] S4: The final outputs of the time-stream encoder and the frequency-stream encoder are concatenated after global pooling and input into the classification layer to obtain the emotion category probability vector of the current speech window;
[0011] S5: Perform exponentially weighted smoothing based on prediction confidence on the sequence of emotion category probability vectors of continuous speech windows to obtain smoothed emotion category probability vectors, output emotion category and smoothed confidence;
[0012] S6: The processor generates speech synthesis control commands and body movement control commands based on the output emotion category and smoothed confidence. The speech synthesis control commands are transmitted to the speech synthesis module of the emotional companion robot to output a speech response, and the body movement control commands are transmitted to the body actuator of the emotional companion robot to generate a body response.
[0013] Furthermore, S1 includes the following specific steps:
[0014] S11: Perform pre-emphasis filtering on the collected user voice signal;
[0015] S12: Set a preset long window, a preset short window, and a preset frame shift for the pre-emphasized speech signal. The window length of the long window is not less than twice the window length of the short window. The long window and the short window are aligned with the same frame center and generate a frame sequence using the same frame shift. Hamming windows are added to both the long window and the short window.
[0016] S13: For each frame position The long window is used to apply the pre-emphasized speech signal to the frame position. A short-time Fourier transform is performed on the frame segment containing the given location, and then a long-window logarithmic Mel feature map is obtained through a preset Mel filter bank. The long-window logarithmic Mel feature map is then processed along the frame index. After normalizing the mean and variance, the long-window Mel spectrum is obtained. ; For the long window Mel spectrum in the th The Mel frequency band, the first The normalized log-Mel eigenvalues of the frame; For Mel band indexing; For frame indexing;
[0017] S14: For each frame position The short window is used to apply the pre-emphasized speech signal at the frame position. A short-time Fourier transform is performed on the frame segment containing the given location, and then a short-window logarithmic Mel feature map is obtained through the Mel filter bank. The short-window logarithmic Mel feature map is then linearly interpolated along the time dimension to be compared with the long-window Mel spectrum obtained in S13. The frame index is aligned, and the short-window log-Mel feature map after interpolation alignment is aligned along the frame index. Perform mean-variance normalization to obtain the short-window Mel spectrum. ; For the short-window Mel spectrum in the 1st The Mel frequency band, the first The normalized log-Mel eigenvalues of the frame;
[0018] S15: For each frame position Calculate its normalized short-time energy using the following formula. and long window weight :
[0019] ;
[0020] ;
[0021] In the formula: For the first Normalized short-time energy of a frame; For the first The short-time energy of a frame is equal to the sum of the squares of the amplitudes of all sampling points within that frame; This is the operation of taking the median of the elements in a set; Total number of frames; To prevent division by zero by default positive decimals; For the first Frame long window weight, value range ; For natural index calculations; This is the steepness coefficient; Energy switching threshold; For frame indexing;
[0022] S16: The following formula is used to... and stated Stacked along the channel dimension as a dual-scale Mel feature tensor :
[0023] , ;
[0024] In the formula: The two-scale Mel feature tensor has the following shape: 2 represents the number of channels; The total number of Mel bands; Total number of frames; for In the 0th channel, the The Mel frequency band, the first The element value at the frame position; for In the first channel, the The Mel frequency band, the first The element value at the frame position; For Mel band indexing; For frame index.
[0025] Furthermore, step S2 includes the following specific steps:
[0026] S21: Calculate the prior weighted initial value of the resonance peak using the following formula. :
[0027] ;
[0028] In the formula: For the first The prior weighted initial values of the resonant peaks in each Mel band, with a range of values. ; For index The operation to find the maximum value; For natural index calculations; For the first The center frequency of the Mel band; For the first A priori central frequency; For the first A priori bandwidth; For Mel band indexing; For prior central index;
[0029] S22: Take 6 sets of the aforementioned prior center frequencies and prior bandwidths. Take numbers 1 to 6 in sequence, which correspond to the fundamental frequency region of human voice, the first formant region, the second formant region, the third formant region, the fourth formant region, and the high frequency region of unvoiced fricatives, respectively.
[0030] S23: Obtain a length equal to the total number of Melbands Band weighted parameter vector The These are fixed parameter vectors obtained through pre-training; for The Middle One element; The total number of Mel bands; For Mel band indexing;
[0031] S24: Generate the prior weighting value of the resonance peak using the following formula. :
[0032] ;
[0033] In the formula: For the first Prior weighting values of the resonant peaks in the Mel frequency band, with a range of values. ; The sigmoid function is expressed as follows: ; For natural index calculations; for Input variables; For Mel band indexing;
[0034] S25: Convert the dual-scale Mel feature tensor Global average pooling is performed along the time dimension to obtain the frequency dimension feature sequence. , The shape is 2 represents the number of channels; [The sentence is incomplete and requires more context to translate accurately.] Global average pooling is performed along the frequency band dimension to obtain the time-dimensional feature sequence. , The shape is ;in For the total number of Mel bands, Total number of frames;
[0035] S26: The frequency dimension feature sequence At each Mel band position The value multiplied by the The weighted frequency-dimensional feature sequence is obtained. , and Same shape;
[0036] S27: The weighted frequency-dimensional feature sequence... Input the first coding layer of the frequency stream encoder to obtain the output features of the first coding layer of the frequency stream. The time-dimensional feature sequence The first coding layer of the time-stream encoder is input to obtain the output features of the first coding layer of the time-stream encoder. .
[0037] Furthermore, step S3 includes the following specific steps:
[0038] S31: The time-stream encoder and the frequency-stream encoder each include... One coding layer, The input is an integer not less than 2; each coding layer of the time stream sequentially includes a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, and a bidirectional gated recurrent unit layer, ensuring that the time dimension length of the feature sequence input to each coding layer of the time stream remains constant before and after processing by that coding layer. Each coding layer of the frequency stream sequentially includes a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, and a multi-head self-attention layer, ensuring that the frequency band dimension length of the feature sequence input to each coding layer of the frequency stream remains constant before and after processing by that coding layer. ; in the Coding layer and the first A time-frequency interaction attention module is set up between the encoding layers. Take 1 to 1 in sequence ;in For the coding layer index, The total number of frames. The total number of Mel bands;
[0039] S32: Configure 6 linear projection matrices in each of the time-frequency interaction attention modules. , , , , , The six linear projection matrices are all fixed parameters obtained through pre-training; the time flow... Coding layer output The dimension is OK The column, the frequency stream number Coding layer output The dimension is OK Column; the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK Column; among which For the time flow channel dimension, For frequency flow channel dimension, The attention dimension after projection. , and All are preset positive integers; For encoding layer index;
[0040] S33: In the Coding layer and the first In the time-frequency interaction attention module between coding layers, the interaction features injected into the time stream by frequency information are calculated according to the following formula. :
[0041] ;
[0042] In the formula: For the first At the encoding layer, frequency information is injected into the interactive feature matrix of the time stream, with dimensions of [missing information]. OK List; Normalized exponentiation is performed row-by-row; For encoding layer index;
[0043] S34: In the time-frequency interaction attention module, the interaction features injected into the frequency stream by time information are calculated according to the following formula. :
[0044] ;
[0045] In the formula: For the first The interaction feature matrix at the encoding layer, into which temporal information is injected into the frequency stream, has a dimension of [missing information]. OK List;
[0046] S35: Generate the time stream according to the following formulas respectively. Input of the coding layer and frequency flow Input of the coding layer :
[0047] , ;
[0048] In the formula: For the time flow The input feature matrix of the encoding layer has the same dimension. ; For frequency flow number The input feature matrix of the encoding layer has the same dimension. ; For layer normalization operations; For encoding layer index;
[0049] S36: The Input the time stream Encoding layer, to obtain the time stream Coding layer output ; will the Input the frequency stream The coding layer yields the frequency stream. Coding layer output ;
[0050] According to the methods described in S33 to S36, Take 1 to 1 in sequence Iterative execution yields the final output of the time stream. and frequency stream final layer output .
[0051] Furthermore, step S4 includes the following specific steps:
[0052] S41: Output the last layer of the time stream Global average pooling is performed along the time dimension to obtain the global feature vector of the time flow. , The length is equal to the dimension of the time stream channel. ;
[0053] S42: Output the final layer of the frequency stream. Global average pooling is performed along the frequency band dimension to obtain the global eigenvector of the frequency flow. , The length is equal to the frequency flow channel dimension ;
[0054] S43: The above and stated Concatenate along the channel dimension to form a joint feature vector , will the The current speech window is output after passing through a fully connected layer and a Softmax layer. sentiment category probability vector The weight parameters of the fully connected layer are fixed parameters obtained through pre-training. The length is equal to the total number of preset emotion categories. The probability vector; The total number of preset emotion categories; This is the window timing index.
[0055] Furthermore, S5 includes the following specific steps:
[0056] S51: Calculate the current voice window using the following formula Prediction confidence :
[0057] ;
[0058] In the formula: For the current voice window The prediction confidence level, and its range. ; This is an operation that takes the maximum value of each element of a vector. For the current voice window The probability vector of the sentiment category; For window timing index;
[0059] S52: The following formula is used to... Mapped to the current voice window Adaptive smoothing coefficient :
[0060] ;
[0061] In the formula: For the current voice window The adaptive smoothing coefficient; This is the lower bound of the smoothing coefficient; This is the upper bound of the smoothing coefficient;
[0062] S53: When At that time, make the current voice window Smoothed sentiment category probability vector ;when When the value is greater than 0, the current voice window is calculated using the following formula. Smoothed sentiment category probability vector :
[0063] ;
[0064] In the formula: For the current voice window The smoothed sentiment category probability vector; Previous voice window The smoothed sentiment category probability vector; For window timing index;
[0065] S54: Regarding the above Take the index of the maximum value as the current voice window. The output sentiment category is taken from the... The maximum value is used as the current voice window. Smoothed post-confidence .
[0066] Furthermore, step S6 includes the following specific steps:
[0067] S61: Combine the output sentiment category and the smoothed post-confidence obtained in S5. The information is transmitted to the empathic response module of the emotional companion robot; the empathic response module includes a pre-stored response strategy library; the response strategy library is a mapping table between emotion categories and response strategies, and each response strategy includes at least a voice text template, a prosodic parameter set, and a body movement number;
[0068] S62: The empathic response module determines a stable emotional state according to the following rules: if from the... The first voice window to the second A total of 1 voice window The output emotion categories of each consecutive speech window are the same, and the smoothed post-confidence of each speech window is... All are not less than the preset output threshold. ,in Take in sequence to Then, the same output sentiment category is recorded as a stable sentiment state. ; The preset number of consecutive windows is an integer between 3 and 10; For the first Smoothed post-confidence of each voice window; To stabilize emotional state; Index for the voice window;
[0069] S63: If the above Not less than the Then the empathic response module will respond according to the current voice window. The processor retrieves the corresponding response strategy from the response strategy library for the output emotion category, and generates the speech synthesis control command and the body movement control command based on the retrieved response strategy.
[0070] S64: If the above Smaller than the And there are recorded stable emotional states. Then the empathic response module maintains the The corresponding response strategy; if the Smaller than the Furthermore, no such stable emotional state has been recorded. Then the empathic response module adopts a preset neutral response strategy, that is, the response strategy corresponding to the neutral emotion category in the response strategy library.
[0071] The beneficial effects achieved by this invention are as follows:
[0072] This invention designs a dual-scale Mel feature tensor. Step S1 calculates both the long-window and short-window Mel spectra of the pre-emphasized and framed speech signals simultaneously. Then, based on the long-window weights obtained from the normalized short-time energy mapping calculated in step S15, the two Mel spectra are weighted by channel and stacked along the channel dimension to form a dual-scale Mel feature tensor. This allows the back-end encoder to automatically balance the dependence ratio of frequency resolution and temporal resolution at each frame position according to the energy characteristics of the current frame. Based on the above-mentioned frame-adaptive balanced feature front-end, the method of this invention provides a relatively stable improvement in recognition accuracy for multilingual and multi-category sentiment recognition tasks, and the improvement is consistent across different language public sentiment corpora.
[0073] This invention features bidirectional time-frequency interactive attention. Step S3 performs bidirectional time-frequency interactive attention between every two adjacent coding layers of the time-stream encoder and the frequency-stream encoder. Step S33 calculates the interactive features injected into the time stream by frequency information, and step S34 calculates the interactive features injected into the frequency stream by time information. Step S35 then uses the interactive results as the input to the next coding layer through residual connections and layer normalization. This allows the features of the two streams to guide each other in a layered and iterative manner during forward propagation. The joint time-frequency representation can be initially formed at a shallow layer and further refined at a deeper layer. Therefore, the method of this invention can achieve a higher recognition accuracy relatively quickly under the same training cost. Compared with existing schemes that perform splicing or end-to-end fusion at the end of the coding, it improves both training convergence speed and final recognition accuracy.
[0074] This invention designs a formant prior weighting mechanism. In step S2, the formant prior weighting initial values are constructed using the Gaussian envelope form described in step S21, with the fundamental frequency region of human voice, the first to fourth formant regions, and the high-frequency region of unvoiced fricatives as prior centers. Then, in step S24, the initial values are finely tuned by a learnable frequency band by frequency using a trained frequency band weighting parameter vector and a sigmoid function to obtain the formant prior weighting values. In step S26, the frequency dimension feature sequence is weighted by frequency band according to the weighting values and then fed into the frequency stream encoder. The frequency stream is guided at the input end to focus on the frequency band carrying emotional cues. The model has a relatively balanced distribution of recognition accuracy for various emotions, and it has a significant improvement in the recognition ability of happy emotions that depend on formant shift and fear emotions that depend on high-frequency energy enhancement.
[0075] This invention designs a confidence-adaptive exponential weighted smoothing decoding and a stable emotional state determination and response control mechanism. Step S5 establishes a linearly increasing relationship between the smoothing coefficient and the current predicted confidence level, following step S52. Step S53 then recursively obtains the smoothed emotional category probability vector, ensuring that high-confidence observations dominate the output while low-confidence observations rely more on historical smoothing results. Step S6 further introduces an E_st determination rule based on stable consistency across several consecutive windows, generating the robot's speech synthesis control commands and body movement control commands accordingly. This suppresses window-by-window jitter caused by occasional misidentification during periods of stable user emotion and maintains a relatively timely response when the user's emotion truly shifts. The robot's response strategy switching frequency is reasonable, ensuring the continuity and naturalness of responses during long-term interaction, thus balancing the comprehensive requirements of recognition stability and response timeliness. Attached Figure Description
[0076] Figure 1 This is a comparison chart of the User Assistance (UA) metrics of Examples 1, 2, and 3 with Comparative Examples 1, 2, and 3 on three public datasets: IEMOCAP, CASIA, and EMO-DB. The chart is a grouped bar chart.
[0077] Figure 2 This is a comparison chart of the training convergence curves of Examples 1, 2, and 3 with Comparative Examples 1, 2, and 3 on the IEMOCAP dataset. The chart is a multi-curve line graph.
[0078] Figure 3 This is a comparison chart of the accuracy of each of the six emotion categories in Example 1, Example 2, Example 3 and Comparative Examples 1, 2, and 3 on the CASIA dataset. The chart is a two-dimensional heatmap.
[0079] Figure 4 The figures are comparison charts of Example 1, Example 2, Example 3 and Comparative Example 1, Comparative Example 2, Comparative Example 3 in terms of the stability of emotion prediction. Among them, (a) is a ladder line graph comparing the emotion category prediction trajectory of Example 1 and Comparative Example 1, Comparative Example 2, Comparative Example 3 on continuous speech, and (b) is a bar chart comparing the number of emotion switching per minute of Example 1, Example 2, Example 3 and Comparative Example 1, Comparative Example 2, Comparative Example 3.
[0080] Figure 5 This is a flowchart of a voice-emotion time-frequency dual-stream interaction algorithm for emotional companion robots according to the present invention. Detailed Implementation
[0081] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. In addition, the forms of the various structures described in the following embodiments are merely illustrative. The present invention is not limited to the structures described in the following embodiments. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0082] This invention relates to a voice emotion recognition and response control processing flow running on an emotional companion robot. The process is executed by the processor inside the emotional companion robot and relies on fixed model parameters obtained during the offline training phase to complete online inference. The training phase is not included in the execution flow of the method described in this specification. Parameters such as the frequency band weighting parameter vector, linear projection matrix, and fully connected layer weights used by the algorithm are frozen after training and are only read and called during inference.
[0083] Reference Figure 5 The algorithm of this invention consists of six main steps, S1 to S6. S1 extracts dual-scale Mel features from the collected speech; S2 decouples these features along both time and frequency directions and applies formant prior weighting; S3 performs bidirectional time-frequency interactive attention in layers between two parallel encoders; S4 aggregates the output of the final encoding layer into a window-level emotion category probability vector; S5 performs confidence-adaptive smoothing decoding on the probability vectors of consecutive windows; and S6 converts the decoding result into voice and body response control commands for the robot. The emotional companion robot includes at least a microphone, a processor, a memory, a speech synthesis module, and a body execution mechanism. The microphone, speech synthesis module, and body execution mechanism are communicatively connected to the processor. The memory stores a computer program, a response strategy library, and an empathic response program module. The processor is configured to run the computer program to execute steps S1 to S6.
[0084] The dual-scale Mel feature extraction process described in S1 is explained. S1 consists of six specific steps, from S11 to S16. The goal is to simultaneously acquire the long-window Mel spectrum and the short-window Mel spectrum of the same speech signal, and then adaptively weight them frame by frame and stack them along the channel dimension to form a dual-scale Mel feature tensor, so as to alleviate the inherent trade-off that time resolution and frequency resolution cannot be optimized at the same time.
[0085] S11 performs pre-emphasis filtering on the acquired user speech signal, with the pre-emphasis coefficient preferably between 0.95 and 0.97. The pre-emphasis aims to counteract the effects of lip radiation and glottal excitation on spectral flatness using a first-order high-pass finite impulse response filter, resulting in a more stable energy distribution in the mid-to-high frequency region of the subsequent Mel spectrum. The first-order high-pass finite impulse response filter can be implemented by directly calling the corresponding function in the AudioToolbox or librosa open-source library.
[0086] S12 sets a preset long window, a preset short window, and a preset frame shift for the pre-emphasized speech signal. The length of the long window is not less than twice the length of the short window. The long window is used to characterize the fine harmonic structure of the quasi-steady-state speech segment, and the short window is used to characterize the fine temporal structure of the transient segment. The long window and the short window are aligned with the same frame center and generate a frame sequence using the same frame shift, thereby ensuring that the frame is indexed within the same frame. The long and short windows observe the same speech segment center, differing only in observation duration, thus avoiding temporal misalignment of the subsequent two feature paths. Both the long and short windows are fitted with Hamming windows, a well-known windowing function that replaces rectangular truncation with a gradually varying envelope, reducing spectral leakage caused by truncation.
[0087] S13 for each frame position The long window is used to apply the pre-emphasized speech signal to the frame position. The frame segment containing the given location undergoes a short-time Fourier transform (SFT), and then a long-window logarithmic Mel feature map is obtained through a pre-defined Mel filter bank. The SFT, Mel filter bank, and logarithmic energy mapping are all well-known methods in the field of speech signal processing, and the Mel scale satisfies a well-known form. ,in The frequency is linear, measured in Hz. Then, the long-window log-Mel feature map is processed along the frame index. Perform mean-variance normalization, i.e., for the same Mel frequency band The mean and standard deviation are calculated along all frame indices. Then, the log-Mel eigenvalues of each frame are shifted by the mean and scaled by the standard deviation to obtain the long-window Mel spectrum. This normalization ensures that each Mel frequency band maintains a zero-mean, unit-variance distribution under different recording conditions, facilitating subsequent weighting and splicing processes. For the long window Mel spectrum in the th The Mel frequency band, the first The normalized log-Mel eigenvalue of the frame is dimensionless; For Mel band indexing; For frame index.
[0088] S14 for each frame position The short window is used to apply the pre-emphasized speech signal at the frame position. A short-time Fourier transform is performed on the frame segment containing the given location, and then the short-window logarithmic Mel feature map is obtained through the Mel filter bank. Since the short window is shorter than the long window, the frame positions corresponding to the short window are more densely packed for the same total duration. Therefore, the short-window logarithmic Mel feature map is linearly interpolated along the time dimension to be compared with the long-window Mel spectrum obtained in S13. Frame index alignment is achieved through linear interpolation, which uses a weighted average of adjacent frames—a well-known resampling method. Then, the short-window log-Mel feature map after interpolation alignment is processed along the frame index. Perform mean and variance normalization, using the same method as in S13, to obtain the short-window Mel spectrum. . For the short-window Mel spectrum in the 1st The Mel frequency band, the first The normalized log-Mel eigenvalue of the frame is dimensionless.
[0089] S15 positions each frame Calculate its normalized short-time energy using the following formula. , In the formula, For the first The normalized short-time energy of a frame is dimensionless. For the first The short-time energy of a frame is equal to the sum of the squares of the amplitudes of all sampling points within that frame; This is the operation of taking the median of the elements in a set; Total number of frames; To prevent division by zero of preset positive decimals, Preferred selection to The values between; For frame indexing. Divide by the median of the normalized short-time energy of the entire speech segment to make Maintaining stable distribution characteristics at different recording loudness levels avoids introducing additional bias due to microphone sensitivity or ambient loudness when directly mapping absolute energy; Its function is to prevent numerical overflow when the entire speech is in a near-silent state and the median is close to zero.
[0090] Then, according to the following formula, the... Mapped to the first Frame long window weight , In the formula, For the first The long window weight of the frame, with a value range of 1. ; For natural index calculations; The steepness coefficient is dimensionless. The larger relatively The steeper the transition, The preferred value is between 2 and 6; The meaning is the same as the above formula; The energy switching threshold is dimensionless. The preferred value is one that is near the median of the normalized short-time energy distribution of the entire speech segment; For frame index. This mapping is mathematically known as the sigmoid function pair. The result of this is that it smoothly transforms the difference between frame energy and the threshold into... Continuous weights within the interval, corresponding to high-energy frames Approaching 1, low-energy frames corresponding to Approaching 0 avoids introducing discontinuities between adjacent frames due to hard handover.
[0091] S16 is to use the following formula to... and stated Stacked along the channel dimension as a dual-scale Mel feature tensor :
[0092] , ;
[0093] In the formula, The two-scale Mel feature tensor has the following shape: 2 represents the number of channels; The total number of Mel bands; Total number of frames; for In the 0th channel, the The Mel frequency band, the first The element value at the frame position; Same meaning as S15; Same meaning as S13; for In the first channel, the The Mel frequency band, the first The element value at the frame position; Same meaning as S14; For Mel band indexing; For frame indexing. When the frame is located in the high-energy steady-state segment. Approaching 1, When channel 0 is close to the original long window feature and channel 1 is close to 0, the frame is primarily frequency-resolution; when the frame is in a low-energy transient segment... Approaching 0 The 0th channel is close to 0, and the 1st channel is close to the original short window feature. At this point, the frame is mainly time-resolution. The two channels exist side by side along the channel dimension, allowing the subsequent encoder to utilize features from both resolution perspectives simultaneously.
[0094] The following explains the time-frequency decoupling pooling and formant prior weighting process described in S2. S2 consists of seven specific steps, from S21 to S27, and its goal is to transform the dual-scale Mel feature tensor obtained in S1 into... The sequences are pooled along two complementary directions to form time-dimensional feature sequences and frequency-dimensional feature sequences, respectively. Prior weights based on the statistical regularity of human speech formants are applied to the frequency-dimensional feature sequences, and then fed into the time-stream encoder and frequency-stream encoder, respectively.
[0095] S21 calculates the prior weighted initial value of the resonance peak using the following formula. , ;
[0096] In the formula, For the first The prior weighted initial values of the resonant peaks in the Mel frequency band, with a range of values of [value range missing]. ; For index The operation to find the maximum value; For natural index calculations; For the first The center frequency of each mech band, in Hz; For the first A priori center frequency, in Hz; For the first A priori bandwidth, in Hz; For Mel band indexing; This is a priori central index. The formula uses several... Centered on, with Using a Gaussian bell-shaped function with standard deviation as the component, the point-by-point maximum value of all components at each frequency position is taken as the envelope. Falling on a certain When nearby Approaching 1, when Stay away from any one hour Approaching 0, thus injecting phonetic priors into the input of the frequency stream.
[0097] S22 selects 6 sets of the aforementioned prior center frequencies and the aforementioned prior bandwidths. The numbers 1 to 6 are assigned sequentially, corresponding to the fundamental frequency region, the first formant region, the second formant region, the third formant region, the fourth formant region, and the high-frequency region of voiceless fricatives, respectively. Formants are the inherent resonant peaks of the vocal tract cavity, carrying vowel recognition information and conveying emotional cues through their frequency shifts and bandwidth variations under different emotional states. The fundamental frequency is related to the vocal cord vibration frequency and is closely associated with emotional attributes such as pitch and tension. The high-frequency region of voiceless fricatives is related to emotional perception cues such as fricatives and aerophones. These six a priori regions collectively cover the main frequency domain carrying capacity of human vocal emotions. A value between 100 and 300 is preferred. A value between 300 and 900 is preferred. The preferred value is between 900 and 1900. A value between 1900 and 2800 is preferred. A value between 2800 and 4500 is preferred. The preferred values are between 4500 and 8000, all in Hz. A value between 50 and 300 is preferred. A value between 200 and 600 is preferred. A value between 400 and 900 is preferred. A value between 500 and 1200 is preferred. A value between 500 and 1200 is preferred. The preferred values are between 800 and 2000, all in Hz. These values are based on the statistical distribution of formant frequencies in mixed male and female adult speech patterns. Those skilled in the art may adjust these values appropriately for children, the elderly, or people speaking different languages.
[0098] S23 obtains a length equal to the total number of Melbands. Band weighted parameter vector The It is a fixed parameter vector obtained through pre-training, whose elements are obtained during the training phase by gradient backpropagation and frozen after training. for The Middle One element; The total number of Mel bands; For Mel band indexing. (Introduction) The goal is to retain a certain degree of data-driven adjustment freedom on the basis of a strict phonetics prior envelope, so that the model can be refined and adapted to specific training corpora under the constraints of fixed priors.
[0099] S24 generates the prior weighting value of the resonance peak using the following formula. : ;
[0100] In the formula, For the first The prior weighted values of the resonant peaks in the Mel frequency band, with values ranging from... ; Same meaning as S21; The sigmoid function is expressed as follows: , is a well-known nonlinear mapping function; For natural index calculations; for Input variables; Same meaning as S23; For Mel band indexing. (Introduction) The purpose of the function is to... The range of values is compressed from any real number that may appear after training to... interval, making Strictly fall into In between, maintain consistency with the semantics of "weighted" to avoid When a negative value or a large number is encountered during training, make Loss of weighting.
[0101] S25 converts the dual-scale Mel feature tensor Global average pooling is performed along the time dimension to obtain the frequency dimension feature sequence. , The shape is 2 represents the number of channels; this pooling averages the values of each Mel band across the entire window into a single scalar, thereby extracting a feature vector reflecting the overall structure of the spectrum. Then, the aforementioned... Global average pooling is performed along the frequency band dimension to obtain the time-dimensional feature sequence. , The shape is 2 represents the number of channels; this pooling averages the values of all Mel bands at each time step into a single scalar, thereby extracting the feature vectors that reflect rhythmic changes. For the total number of Mel bands, The total number of frames. Pooling along different dimensions completely decouples the processing objects in the subsequent encoder in terms of dimension, making it easier to focus on inter-band structure and inter-frame dynamics respectively.
[0102] S26 will use the frequency dimension feature sequence At each Mel band position The value multiplied by the The weighted frequency-dimensional feature sequence is obtained. , and The shapes are identical. The multiplication is element-wise, acting only on the frequency band dimension and not changing the channel dimension. After this weighting, the feature values of the Mel band falling in the formant and deafness fricative regions receive a larger weight, while the feature values of the Mel band falling outside the formant regions receive a smaller weight, allowing the subsequent frequency stream encoder to focus more attention on the emotion-related frequency bands.
[0103] S27 will use the weighted frequency dimension feature sequence Input the first coding layer of the frequency stream encoder to obtain the output features of the first coding layer of the frequency stream. The time-dimensional feature sequence The first coding layer of the time-stream encoder is input to obtain the output features of the first coding layer of the time-stream encoder. The specific structures of each coding layer of the frequency stream encoder and the time stream encoder are uniformly described in S31. The two features extracted along different dimensions are processed by two structurally dual encoders. Compared with the approach of processing mixed features by a single encoder, the two coding paths can focus on the extraction of inter-band harmonic structures and the extraction of inter-frame prosodic variations, respectively, and then be fused by bidirectional time-frequency interactive attention in S3.
[0104] The following describes the bidirectional time-frequency interaction attention process described in S3. S3 consists of six specific steps, from S31 to S36, with the goal of performing a bidirectional time-frequency interaction between every two adjacent coding layers of the two parallel encoders, so that the time stream and frequency stream guide each other iteratively in a layered manner during the forward propagation process.
[0105] S31 defines that the time-stream encoder and the frequency-stream encoder each include... One coding layer, The integer is not less than 2. Each coding layer of the temporal stream sequentially includes a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, and a bidirectional gated recurrent unit layer. The one-dimensional convolutional layer extracts local temporal patterns along the time dimension, the batch normalization layer accelerates convergence and alleviates internal covariate shifts, the ReLU activation layer introduces nonlinearity, and the bidirectional gated recurrent unit layer captures long-range dependencies along the time dimension. Each coding layer of the temporal stream ensures that the time dimension length of the feature sequence input to the coding layer remains constant before and after processing by that coding layer. This can be achieved by setting equal-length padding for the one-dimensional convolutional layer and taking the full sequence output for the bidirectional gated recurrent unit layer. Each coding layer of the frequency stream sequentially includes a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, and a multi-head self-attention layer. The multi-head self-attention layer is used to establish global correlations across all Mel frequency bands, complementing the local harmonic structures captured by the one-dimensional convolutional layer. Each coding layer of the frequency stream ensures that the frequency band dimension length of the feature sequence input to the coding layer remains constant before and after processing by that coding layer. This can be achieved by setting equal-length padding for one-dimensional convolutional layers and taking the full sequence output for multi-head self-attention layers. In the... Coding layer and the first A time-frequency interaction attention module is set up between the encoding layers. Take 1 to 1 in sequence . For the coding layer index, The total number of frames. This represents the total number of Mel frequency bands.
[0106] S32 configures 6 linear projection matrices in each of the time-frequency interaction attention modules. , , , , , The six linear projection matrices are all fixed parameters obtained through pre-training. The time stream... Coding layer output The dimension is OK The column, the frequency stream number Coding layer output The dimension is OK The column. The aforementioned The dimension is OK The column, the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK List. For the time flow channel dimension, For frequency flow channel dimension, The attention dimension after projection. , and All are preset positive integers. Preferred selection and The value near the geometric mean; This serves as the index for the encoding layer. The above dimensional settings ensure that subsequent attention operations are strictly self-consistent along the matrix dimension, and also... and The outputs fall on the same side. and The same channel space facilitates the subsequent addition of residuals.
[0107] S33 in Coding layer and the first In the time-frequency interaction attention module between coding layers, the interaction features injected into the time stream by frequency information are calculated according to the following formula. :
[0108] ;
[0109] In the formula, For the first At the encoding layer, frequency information is injected into the interactive feature matrix of the time stream, with dimensions of [missing information]. OK List; The normalized exponentiation operation is performed row-wise, and is a well-known probability normalization function; , , , and The meaning is the same as S32; Same meaning as S32; This is an index for the encoding layer. This formula, based on the standard form, places the query source in the time stream and the key-value source in the frequency stream, thus extending same-stream attention to cross-stream attention. Divide by Its function is to suppress the inner product amplitude from changing. Increasing the softmax saturation caused by the increase makes the gradient more stable during the training phase.
[0110] S34 In the time-frequency interaction attention module, the interaction features injected into the frequency stream by time information are calculated according to the following formula. ,
[0111] ;
[0112] In the formula, For the first The interaction feature matrix at the encoding layer, into which temporal information is injected into the frequency stream, has a dimension of [missing information]. OK List; , and The meaning of S32 is the same; the meanings of the other symbols are the same as in S33. This formula is directionally dual to formula S33. The frequency stream provides the query, and the time stream provides the key and value. The relevant information in the time dimension is weighted according to the position in the frequency dimension and then fed back to the frequency stream. S33 and S34 are combined to form the bidirectional time-frequency interactive attention proposed in this method. Compared with schemes that only perform splicing at the end of the encoder or only guide in a single direction, the mutual guidance relationship of time-frequency features can be repeatedly established between every two adjacent coding layers, so that the features of the two streams are coupled layer by layer during the forward propagation process, thereby forming a joint time-frequency representation earlier.
[0113] S35 generates the time stream according to the following formula: Input of the coding layer and frequency flow Input of the coding layer ,
[0114] , ;
[0115] In the formula, For the time flow The input feature matrix of the encoding layer has the same dimension. ; For frequency flow number The input feature matrix of the encoding layer has the same dimension. ; Layer normalization is a well-known operation that calculates the mean and variance of each element in a row and then shifts the mean and scales the result by the standard deviation. , , and The meaning is the same as S33 and S34; This serves as the index for the coding layers. The purpose of residual summation is to preserve the original information pathways of each coding layer and avoid gradient degradation in deep networks; the purpose of layer normalization is to stabilize the feature distribution between different coding layers and different samples.
[0116] S36 will say Input the time stream The encoding layer obtains the time stream. Coding layer output ; will the Input the frequency stream The coding layer yields the frequency stream. Coding layer output According to the methods described in S33 to S36... Take 1 to 1 in sequence The process is iterative, with each iteration performing a bidirectional time-frequency interaction between two adjacent coding layers. After each iteration, the output of the final time-flow layer is obtained. and frequency stream final layer output Both serve as inputs to S4.
[0117] The following describes the window-level feature aggregation and sentiment classification process described in S4. S4 consists of three specific steps, S41 to S43. The goal is to aggregate the positional feature matrices of the time stream and frequency stream obtained in S3 into window-level global vectors, and then concatenate them and input them into the classification layer to obtain the sentiment category probability vector of the current window.
[0118] S41 outputs the final layer of the time stream. Global average pooling is performed along the time dimension to obtain the global feature vector of the time flow. , The length is equal to the dimension of the time stream channel. The global average pooling then... along The weighted average of all positions in a direction is a well-known way to compress a time series signal into a fixed-length feature vector.
[0119] S42 outputs the final layer of the frequency stream. Global average pooling is performed along the frequency band dimension to obtain the global eigenvector of the frequency flow. , The length is equal to the frequency flow channel dimension The specific steps differ from the pooling direction of S41, but the operation form is the same, corresponding to the aggregation methods of the two streams respectively.
[0120] S43 will say and stated Concatenate along the channel dimension to form a joint feature vector , will the The current speech window is output after passing through a fully connected layer and a Softmax layer. sentiment category probability vector The weight parameters of the fully connected layer are fixed parameters obtained through pre-training. The Softmax layer is a known function that normalizes a real-valued vector into a non-negative probability vector whose elements sum to 1. The length is equal to the total number of preset emotion categories. The probability vector; To preset the total number of emotion categories, The preferred integer is between 4 and 8; This is a window time-series index. The set of emotion categories can be set according to the selected training dataset. Commonly used emotion categories in this field include calm, happy, sad, angry, surprised, fearful, disgusted, neutral, etc.
[0121] The following describes the confidence-adaptive smoothing decoding process described in S5. S5 consists of four specific steps, S51 to S54, aiming to transform the sequence of sentiment category probability vectors generated window by window into a stable output sentiment category and smoothed confidence, avoiding label jitter caused by hard decision-making in window by window. The design basis for jitter suppression is that human emotional states usually exhibit gradual characteristics in natural interactions, and the probability fluctuations in window by window are mostly caused by noise, speech transition segments, or atypical pronunciations. These should be smoothed through historical information rather than directly reflected in the output.
[0122] S51 calculates the current voice window using the following formula. Prediction confidence , ;
[0123] In the formula, For the current voice window The prediction confidence level, with a value range of ; This is an operation that takes the maximum value of each element of a vector. For the current voice window The probability vector of the sentiment category; This is the window time series index. The maximum probability value, as the confidence level, reflects the model's degree of confidence in the current decision and is a well-known confidence assessment method in classification models.
[0124] S52 will use the following formula to... Mapped to the current voice window Adaptive smoothing coefficient ,
[0125] ;
[0126] In the formula, For the current voice window The adaptive smoothing coefficient; The smoothing coefficient is a lower bound, dimensionless. The preferred value is between 0.2 and 0.4; The upper bound of the smoothing coefficient is dimensionless. A value between 0.8 and 0.95 is preferred. The meaning is the same as S51. This linear mapping makes Follow Monotonically increasing, when the confidence level of the current observation is high. Approaching The smoothing results are biased towards the current observation, especially when the confidence level of the current observation is low. Approaching The smoothing results are biased towards historical results, thereby automatically adjusting the dependence ratio on current observations and historical information at different confidence levels.
[0127] S53 when At that time, make the current voice window Smoothed sentiment category probability vector , as the initial condition for the recursion; when When the value is greater than 0, the current voice window is calculated using the following formula. Smoothed sentiment category probability vector ,
[0128] ;
[0129] In the formula, For the current voice window The smoothed sentiment category probability vector; Same meaning as S52; Same meaning as S51; Previous voice window The smoothed sentiment category probability vector; This is the window timing index. Mathematically, this recursion represents time-varying coefficients. The dominant exponentially weighted moving average is an adaptive extension of classical smoothing filtering. Compared to exponentially weighted smoothing with fixed coefficients, this method automatically adjusts the level of confidence in the current observation as the confidence level changes. This can alleviate the excessive contamination of historical results by low-confidence observations, and also mitigate the problem of high-confidence observations being over-smoothed and losing their timeliness.
[0130] S54 regarding the above Take the index of the maximum value as the current voice window. The output sentiment category is taken from the... The maximum value is used as the current voice window. Smoothed post-confidence The output emotion category corresponds to a certain category in the emotion category set described in S43. It reflects the credibility of the smoothed decision and serves as an input for robot empathic response control in S6.
[0131] The following describes the robot empathic response control process described in S6. S6 consists of four specific steps, from S61 to S64. The goal is to convert the output of S5 into the robot's voice and body response control commands, and to ensure the continuity of the response strategy when the user's voice is briefly distracted or in a transitional phase through a stable emotional state determination mechanism.
[0132] S61 combines the output sentiment category and the smoothed post-confidence obtained in S5. The information is transmitted to the empathic response module of the emotional companion robot. This module includes a pre-stored response strategy library. The response strategy library is a mapping table between emotion categories and response strategies. Each response strategy includes at least a speech text template, a prosodic parameter set, and a body movement number. The speech text template drives the speech synthesis module to generate synthesized speech of the corresponding response text. The prosodic parameter set includes parameters such as speech rate, fundamental frequency center, fundamental frequency range, and intensity, giving the synthesized speech prosodic features that match the target emotion. The body movement number points to a specific action script in a pre-stored action library, used to drive the limb actuator to perform the corresponding limb response. The response strategy library is built offline during the design phase, and those skilled in the art can design it based on a pre-set emotional response corpus and a pre-set robot action library.
[0133] The empathic response module described in S62 determines stable emotional states according to the following rules: if from the... The first voice window to the second A total of 1 voice window The output emotion categories of each consecutive speech window are the same, and the smoothed post-confidence of each speech window is... All are not less than the preset output threshold. ,in Take in sequence to Then, the same output sentiment category is recorded as a stable sentiment state. .in A value between 0.4 and 0.6 is preferred. It is a dimensionless quantity; To preset the number of consecutive windows, The preferred integer is between 3 and 10; For the first The smoothed post-confidence of each speech window is calculated in the same way as S54. To stabilize emotional state; This serves as an index for the voice window. The purpose of introducing stable emotional states is to distinguish between occasional high-confidence decisions and sustained high-confidence decisions, thereby avoiding frequent switching of response strategies when there are brief misrecognitions in the user's voice. and The value of represents a trade-off between stability and response speed. The larger the value, the more conservative the switching and the slower but more stable the response. Conversely, the smaller the value, the more aggressive the switching and the faster the response, but the more prone to jitter.
[0134] S63 If the above Not less than the Then the empathic response module will respond according to the current voice window. The processor retrieves the corresponding response strategy from the response strategy library for the output emotion category, and generates the speech synthesis control command and the body movement control command based on the retrieved response strategy. The speech synthesis control command includes a speech text template and a prosodic parameter set, and is transmitted to the speech synthesis module to output a speech response; the body movement control command includes a body movement number, and is transmitted to the body actuator to generate a body response.
[0135] S64 If the above Smaller than the And there are recorded stable emotional states. Then the empathic response module maintains the The corresponding response strategy maintains the consistency of the robot's response even when the user's voice is briefly distracted or in a transitional phase; if the Smaller than the Furthermore, no such stable emotional state has been recorded. The empathic response module then adopts a preset neutral response strategy, which is the response strategy in the response strategy library corresponding to the neutral emotion category, so that the robot can still give a reasonable response even before it has established an understanding of the user's emotional state.
[0136] Steps S1 to S6 together constitute the complete method of this invention for completing a voice emotion recognition and response control on an emotional companion robot. The entire process is implemented collaboratively at the hardware level by a microphone, processor, memory, speech synthesis module, and limb actuators. The microphone is used to collect user voice signals and can be a single microphone or a microphone array, preferably configured with a sampling rate between 8kHz and 48kHz. The processor executes all steps S1 to S6 and can be implemented using an embedded system-on-a-chip, embedded graphics processor, or dedicated neural network accelerator. The memory stores the computer program, response strategy library, and empathic response program module, and can be implemented using flash memory, an embedded multimedia card, or a solid-state drive. The speech synthesis module outputs voice responses according to speech synthesis control commands and can be implemented by a separate speech synthesis chip or a speech synthesis program running within the processor. The limb actuators generate limb responses according to limb movement control commands and can include head servos, expression displays, arm joint actuators, or whole-body joint actuators. The microphone, speech synthesis module, and limb actuators are each communicatively connected to the processor, and the communication method can be implemented using known bus interfaces such as I2S, I2C, UART, or SPI. Based on the above methods and hardware configurations, those skilled in the art can implement the specific technical solutions of this invention without any creative effort.
[0137] Example 1 provides a specific application of a voice-emotion time-frequency dual-stream interaction method for elderly family companionship scenarios. The emotional companion robot integrates a 16kHz single microphone, an embedded neural network accelerator, solid-state memory, a speech synthesis chip, a head servo motor, and an expression display screen. During daily conversations with the user, the robot continuously executes steps S1 to S6 of this invention, generating voice and body responses according to the output emotion category and smoothed post-confidence.
[0138] In step S1, pre-emphasis filtering is performed on the acquired speech according to step S11, with a pre-emphasis coefficient of 0.97; in step S12, the long window is set to 50ms, the short window to 15ms, and the frame shift to 10ms, with the long window and the short window aligned with the same frame center and a Hamming window added; in step S13, the normalized long window Mel spectrum is obtained. Press S14 to obtain the aligned and normalized short-window Mel spectrum. Calculate the normalized short-time energy according to S15. and long window weight ,in Take 4. Take 1, Pick The shape generated according to S16 is Dual-scale Mel feature tensor Mel frequency band total Take 80.
[0139] In S2, prior weighted initial values of resonance peaks are constructed according to the Gaussian envelope form as described in S21. According to S22, take 6 sets of prior center frequencies and prior bandwidths. to 200, 500, 1500, 2500, 3700, and 6000 Hz were selected respectively. to Select 150, 400, 600, 800, 800, and 1500 Hz respectively; obtain the pre-trained fixed parameter vector according to S23. ; according to S24 Functional constraints generate prior weighted values of resonance peaks Obtain the frequency dimension feature sequence according to S25. and time dimension feature sequence Generate a weighted frequency-dimensional feature sequence according to S26. ; Press S27 and They are fed into the first encoding layer of the frequency stream encoder and the time stream encoder, respectively.
[0140] The total number of coding layers is configured according to S31 in S3. Taking option 2, each coding layer of the temporal stream sequentially includes a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, and a bidirectional gated recurrent unit layer; each coding layer of the frequency stream sequentially includes a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, and a multi-head self-attention layer; the temporal stream channel dimensions are configured according to S32. Take 128, frequency flow channel dimension Take 128, attention projection dimension Take 64; calculate the interaction characteristics of the time stream injected with frequency information according to S33. ; Calculate the interaction characteristics of the frequency stream injected with time information according to S34. ; Generate normalized residuals according to S35 and The final layer output is obtained by iterating according to S36. and .
[0141] In S4, the outputs of the two streams are aggregated according to S41 to S43, and then passed through a fully connected layer and a Softmax layer to obtain the current speech window. sentiment category probability vector Preset total number of emotion categories The answer is 4, which corresponds to the four categories of calm, happiness, sadness, and anger.
[0142] In S5, the prediction confidence level is calculated according to S51. Configured according to S52 Take 0.3, Taking 0.9 yields the adaptive smoothing coefficient. The smoothed sentiment category probability vector is obtained by recursion according to S53. Press S54 to output the current window's sentiment category and smoothed confidence score. .
[0143] In S6, according to S61, the output emotion category and the... Send to the empathic response module; configure according to S62 Take 0.5, A score of 5 is used to determine a stable emotional state. S63 and S64 generate voice synthesis control commands and body movement control commands, which are then transmitted to the voice synthesis chip and head servo motor respectively to output soothing voice and body responses that match the user's current emotional state.
[0144] Example 2 provides a home companionship application method for applications with low signal-to-noise ratios and high robustness requirements. The emotional companion robot employs a 4-microphone array, embedded graphics processor, solid-state memory, independent speech synthesis chip, and multi-degree-of-freedom head-arm joints. The difference from Example 1 is that the pre-emphasis coefficient in S11 is 0.96, and in S15... Take 2.5, Pick Make long window weight The transition between different energy levels is relatively smooth, in S22 to Take 180, 450, 1400, 2400, 3500, and 5500 Hz respectively. to 100, 300, 500, 700, 700, and 1200 Hz were selected respectively to approximate the formant distribution of adult Chinese pronunciation. The total number of coding layers in S31 To increase the number of layers for mutual guidance between the time-frequency dual streams, in S32... Take 128, Take 128, Take 64, the total number of preset emotion categories in S43 Take 6, in S52 Take 0.2, Take 0.85, in S62 Take 0.45, Take 7. The remaining steps are the same as in Example 1.
[0145] Example 3 provides a lightweight application method for low-computing-power embedded platforms. The emotional companion robot uses a low-power dual-core processor, on-chip flash memory, a speaker, and an expression display screen. The difference from Example 1 is that the pre-emphasis coefficient in S11 is 0.95, and in S15... Take 5.5, Pick S22 to 250, 600, 1700, 2700, 4000, and 6500 Hz were selected respectively. to The total number of coding layers in S31 is calculated using Hz values of 200, 500, 700, 900, 1000, and 1700. Take 2, in S32 Take 64, Take 64, To reduce the number of network parameters and inference latency, 32 is chosen in S43. Take 7, in S52 Take 0.35, Take 0.92, in S62 Take 0.55, Choosing option 3 allows the response strategy to switch quickly at a higher confidence level to maintain the speed of interactive response. The remaining steps are the same as in Example 1.
[0146] Comparative Example 1 presents a baseline sentiment recognition method based on a single-window Mel spectrum and a CNN+BiLSTM hybrid network. The hardware platform and data acquisition method are consistent with Example 1. The differences are: in S1, Mel spectra are generated only with a single 25ms window length, and the dual-window weighted stacking process described in S12 to S16 of Example 1 is not performed; in S2, formant prior weighting is not applied, and the Mel spectrum is directly fed into the two-dimensional convolutional feature extraction network; in S3, a time-frequency interaction attention module between the time-stream and frequency-stream encoders is not set; in S5, only window-by-window hard decision is performed without confidence-adaptive smoothing decoding; and in S6, stable sentiment states are not distinguished. The sentiment category output in each window directly triggers the response strategy retrieval.
[0147] Comparative Example 2 presents a speech emotion recognition method based on time-frequency dual-stream parallel coding and end-point concatenation fusion. The difference from Example 1 is that in S3, no time-frequency interactive attention module is set between the time-stream encoder and the frequency-stream encoder; the time-stream and frequency-stream independently complete the forward propagation of all coding layers; only in S4 are the outputs of the two streams' final layers concatenated along the channel dimension and input into the classification layer; formant prior weighting is not applied in S2; and in S5, confidence-adaptive smoothing decoding is replaced with exponentially weighted smoothing with a fixed smoothing coefficient of 0.5. The remaining steps are the same as in Example 1.
[0148] Comparative Example 3 presents a speech emotion recognition method based on multi-stream convolution and multi-head attention fusion at the end of the encoding process. The differences from Example 1 are: in S2, a multi-branch fully convolutional network based on AlexNet is used to process the Mel spectrum, and no formant prior weighting is applied; in S3, no inter-layer time-frequency interaction attention module is set, and multi-stream features are fused only once at the end of the encoding process using multi-head attention; in S5, confidence-free adaptive exponential weighted smoothing is performed, with a smoothing coefficient set to a fixed value of 0.7. The remaining steps are the same as in Example 1.
[0149] Experiment 1 employs six methods—Examples 1, 2, and 3, and Comparative Examples 1, 2, and 3—to evaluate the unweighted average accuracy (UA) on three public datasets: the IEMOCAP English sentiment corpus, the CASIA Chinese sentiment corpus, and the EMO-DB Berlin German sentiment corpus, using 5-fold cross-validation. UA is defined as the arithmetic mean of the accuracy rates for each category on the test set. During evaluation, each method automatically adjusts the total number of preset sentiment categories described in S43 of this invention based on the number of categories in the selected dataset. .
[0150] The experimental results of this example are as follows: Figure 1 As shown. Figure 1 The horizontal axis represents the names of the three datasets, and the vertical axis represents the User-Agent (UA) metric. Under each dataset, the UA values for six different methods are displayed side-by-side. From... Figure 1 It can be observed that the UA of the three embodiments is significantly higher than that of the three comparative embodiments on all three datasets. The embodiments show a 5.3 percentage point improvement over the strongest comparative embodiment in IEMOCAP, a 4.4 percentage point improvement in CASIA, and a 3.5 percentage point improvement in EMO-DB. Moreover, the UA differences between embodiments are small, and the increasing UA gradient between comparative embodiments reflects the evolution relationship between the baseline method, the two-stream end fusion method, and the multi-stream end fusion method. This indicates that steps S1 to S6 of the present invention consistently provide recognition accuracy higher than that of existing typical schemes on public datasets of different languages. Comparative embodiment 1 uses only the Mel spectrum of a single window length, and comparative embodiments 2 and 3, although adopting a two-stream coding structure, still do not perform frame-by-frame adaptive balancing of the long-window quasi-steady-state frequency resolution and the short-window transient temporal resolution at the front end. In contrast, step S1 of the present invention uses the long-window weight based on short-time energy defined by formula S15. By weighting and stacking the long-window Mel spectrum and the short-window Mel spectrum separately along the trailing edge channels, the back-end encoder automatically selects the most suitable time-frequency resolution based on the signal characteristics at each frame position, thus achieving a consistent accuracy improvement in scenarios involving multiple languages and multiple sentiment categories.
[0151] Experiment Example 2: This experiment uses six methods, namely Example 1, Example 2, Example 3, and Comparative Example 1, Comparative Example 2, and Comparative Example 3. All methods are trained on the training set of the IEMOCAP dataset with the same batch size and Adam optimizer configuration. The User Abilities (UA) on the validation set are recorded every few iterations.
[0152] The experimental results of this example are as follows: Figure 2 As shown. Figure 2 The horizontal axis represents the number of training iterations, and the vertical axis represents the IEMOCAP validation set UA metric. Each line represents the training convergence curve of one method. From Figure 2It can be observed that the line graphs of the three embodiments consistently lie above the line graphs of the three comparative examples during the iteration process. Furthermore, the rate of increase in UA of the embodiments is significantly higher than that of the comparative examples during the 10th to 20th iteration interval. The embodiments reach a UA level close to the final value around the 30th iteration, while the comparative examples require more iterations to approach their respective convergence upper limits. The final UA values of the three embodiments are 79.6, 78.2, and 77.5, respectively, while the final UA values of the three comparative examples are 70.5, 72.8, and 74.3, respectively. This illustrates the necessity of the bidirectional time-frequency interactive attention module described in step S3 of this invention. Comparative Examples 2 and 3 both employ an end-fusion approach, where the time stream and frequency stream do not interact during encoding, and the two streams are passively spliced together only in the final stage. Such structures struggle to learn deep time-frequency joint representations during training. In contrast, step S3 of this invention performs interactive attention between every two adjacent encoding layers, injecting frequency information into the time stream as defined by formula S33 and time information into the frequency stream as defined by formula S34. The residual connection and layer normalization described in formula S35 serve as inputs to the next encoding layer. This allows the features of the two streams to iteratively guide each other during forward propagation, enabling the time-frequency joint representation to be initially formed at a shallow level and further refined at a deeper level. Therefore, with the same training cost, higher recognition accuracy can be achieved earlier.
[0153] Experiment 3: This experiment uses six methods, namely Example 1, Example 2, Example 3, and Comparative Examples 1, 2, and 3. After uniformly adapting the classification layer to six emotion categories on the CASIA Chinese Emotion Corpus, the results are evaluated to obtain the class-by-class recognition accuracy of each method for the six emotions: anger, happiness, sadness, calmness, fear, and surprise. The unit is %.
[0154] Experimental results are as follows Figure 3 As shown. Figure 3 The horizontal axis represents the six emotion categories, and the vertical axis represents the six method names. The color of each cell in the heatmap represents the recognition accuracy of the corresponding method in the corresponding emotion category; the warmer the color, the higher the accuracy. Figure 3It can be observed that the overall color of the three rows containing the three embodiments is warmer than that of the three rows containing the three comparative examples. Furthermore, the difference between the embodiments and the comparative examples is most pronounced in the categories of happiness and fear, reaching approximately 6.5 to 8.0 percentage points and 6.5 to 9.0 percentage points respectively. The accuracy distribution of the three embodiments for the six emotion categories is more balanced, with the difference between the highest and lowest accuracy rates between 5 and 6 percentage points. In contrast, the accuracy distribution of the three comparative examples for the six emotion categories is more dispersed, with the difference between the highest and lowest accuracy rates exceeding 10 percentage points, indicating that S2 of the present invention… The importance of the formant prior weighting mechanism described in the steps is highlighted. The spectral characteristics of happiness are mainly reflected in the upward shift of the fundamental frequency and the displacement of the second formant F2. The spectral characteristics of fear are mainly reflected in the instability of the fundamental frequency and the enhancement of high-frequency energy at the fourth formant F4. In contrast, due to the lack of formant prior information at the front end, the model struggles to autonomously learn specific frequency band patterns related to emotions from a massive amount of Mel frequency bands under limited training sample conditions. In contrast, step S21 of this invention constructs a priori centers based on the fundamental frequency region, the first to fourth formant regions, and the high-frequency region of the deafening fricative in the form of a Gaussian envelope. And generated through the sigmoid constraint described in S24 Then, the frequency dimension feature sequence is processed by S26 according to... By using frequency band weighting, the frequency stream is guided to focus on emotion-related frequency bands at the input end, thus achieving a relatively balanced recognition of various emotions and a high overall accuracy.
[0155] Experiment 4 consists of two complementary parts. Part 1 uses four methods—Example 1, Comparative Examples 1, 2, and 3—to output the predicted emotion category for each window in a continuous family conversation recording with 100 voice windows of the same duration. The actual emotion label sequence of the conversation recording is: windows 1-25 are calm, 26-55 are happy, 56-80 are sad, and 81-100 are calm. The three actual emotion switches occur between windows 25-26, 55-56, and 80-81, respectively. Part 2 uses six methods—Example 1, 2, 3, Comparative Examples 1, 2, and 3—to evaluate the number of emotion category switches per minute in a self-built 10-minute continuous family companionship conversation recording. The number of switches is defined as the number of times the output emotion category of adjacent voice windows changes divided by the total duration. In Part 2, the switching frequency of Example 1 was 4.2 times / min, Example 2 was 3.8 times / min, Example 3 was 4.5 times / min, Comparative Example 1 was 14.5 times / min, Comparative Example 2 was 8.8 times / min, and Comparative Example 3 was 9.6 times / min.
[0156] The experimental results of this example are as follows: Figure 4 As shown, Figure 4In (a), the horizontal axis represents the speech window index, the vertical axis represents the predicted sentiment category, and the five step lines represent the real sentiment label, the prediction sequence of Example 1, the prediction sequence of Comparative Example 1, the prediction sequence of Comparative Example 2, and the prediction sequence of Comparative Example 3, respectively. Figure 4 In (b), the horizontal axis represents the method name, and the vertical axis represents the number of sentiment category switches per minute. From... Figure 4 In Figure (a), it can be observed that the predicted sequence of Example 1 lags only by 1 to 2 windows at each sentiment category transition boundary and exhibits no window-by-window jitter within each stable sentiment segment, closely matching the step-like shape of the true sentiment label. The predicted sequence of Comparative Example 1 shows frequent window-by-window jitter in all stable segments, significantly deviating from the true label. Although the predicted sequences of Comparative Example 2 and Comparative Example 3 have relatively reduced jitter due to the introduction of a fixed smoothing coefficient, their responses at sentiment switching boundaries are significantly delayed, with Comparative Example 2 lagging by approximately 5 windows and Comparative Example 3 lagging by approximately 7 windows. Figure 4 In (b), it can be observed that the number of emotion switching times per minute in the three embodiments is in the low range of 3.8 to 4.5 times / min, while the number of emotion switching times per minute in the three comparative examples is significantly higher than this range, with comparative example 1 even reaching 14.5 times / min. This shows that the confidence-adaptive smoothing decoding described in step S5 and the stable emotion state determination mechanism described in step S6 of the present invention are effective. Comparative example 1 makes a hard decision on the emotion probability of each window, while comparative examples 2 and 3 use a fixed smoothing coefficient and cannot automatically adjust the dependence ratio on historical information according to the current observation confidence. Therefore, they are prone to jitter when the user's emotion is stable and is subject to occasional misidentification disturbances, and the response is delayed when the user's emotion actually switches due to the large fixed smoothing coefficient. However, step S5 of the present invention uses the smoothing coefficient according to the formula in S52. Compared with the current prediction confidence level Establishing a linearly increasing relationship allows high-confidence observations to dominate the output, while low-confidence observations rely more on historical smoothing results, in conjunction with the continuous-based approach described in step S6. Each window is stable and consistent The judgment rules can suppress occasional misidentifications during periods of stable user emotions and maintain timely responses during periods of shifting user emotions.
[0157] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A voice-emotion time-frequency dual-stream interaction algorithm for emotional companion robots, characterized in that, Includes the following steps: S1: The user voice signal collected by the emotional companion robot is pre-emphasized and framed, and the long-window Mel spectrum and short-window Mel spectrum are weighted and stacked along the channel dimension based on short-time energy to obtain the dual-scale Mel feature tensor. S2: Global pooling is performed on the dual-scale Mel feature tensor along the time dimension and frequency band dimension to obtain the frequency dimension feature sequence and the time dimension feature sequence; after applying the frequency band weighting value based on the formant prior to the frequency dimension feature sequence, it is input into the frequency stream encoder, and the time dimension feature sequence is input into the time stream encoder for parallel encoding; S3: Perform bidirectional time-frequency interaction attention between every two adjacent coding layers of the time-stream encoder and the frequency-stream encoder, and use the interaction results as input to the next coding layer in the form of residual connections and layer normalization; S4: The final outputs of the time-stream encoder and the frequency-stream encoder are concatenated after global pooling and input into the classification layer to obtain the emotion category probability vector of the current speech window; S5: Perform exponentially weighted smoothing based on prediction confidence on the sequence of emotion category probability vectors of continuous speech windows to obtain smoothed emotion category probability vectors, output emotion category and smoothed confidence; S6: The processor generates speech synthesis control commands and body movement control commands based on the output emotion category and smoothed confidence. The speech synthesis control commands are transmitted to the speech synthesis module of the emotional companion robot to output a speech response, and the body movement control commands are transmitted to the body actuator of the emotional companion robot to generate a body response.
2. The algorithm according to claim 1, characterized in that, S1 includes the following specific steps: S11: Perform pre-emphasis filtering on the collected user voice signal; S12: Set a preset long window, a preset short window, and a preset frame shift for the pre-emphasized speech signal. The window length of the long window is not less than twice the window length of the short window. The long window and the short window are aligned with the same frame center and generate a frame sequence using the same frame shift. Hamming windows are added to both the long window and the short window. S13: For each frame position The long window is used to apply the pre-emphasized speech signal to the frame position. A short-time Fourier transform is performed on the frame segment containing the given location, and then a long-window logarithmic Mel feature map is obtained through a preset Mel filter bank. The long-window logarithmic Mel feature map is then processed along the frame index. After normalizing the mean and variance, the long-window Mel spectrum is obtained. ; For the long window Mel spectrum in the th The Mel frequency band, the first The normalized log-Mel eigenvalues of the frame; For Mel band indexing; For frame indexing; S14: For each frame position The short window is used to apply the pre-emphasized speech signal at the frame position. A short-time Fourier transform is performed on the frame segment containing the given location, and then a short-window logarithmic Mel feature map is obtained through the Mel filter bank. The short-window logarithmic Mel feature map is then linearly interpolated along the time dimension to be compared with the long-window Mel spectrum obtained in S13. The frame index is aligned, and the short-window log-Mel feature map after interpolation alignment is aligned along the frame index. Perform mean-variance normalization to obtain the short-window Mel spectrum. ; For the short-window Mel spectrum in the 1st The Mel frequency band, the first The normalized log-Mel eigenvalues of the frame; S15: For each frame position Calculate its normalized short-time energy using the following formula. and long window weight : ; ; In the formula: For the first Normalized short-time energy of a frame; For the first The short-time energy of a frame is equal to the sum of the squares of the amplitudes of all sampling points within that frame; This is the operation of taking the median of the elements in a set; Total number of frames; To prevent division by zero by default positive decimals; For the first Frame long window weight, value range ; For natural index calculations; This is the steepness coefficient; Energy switching threshold; For frame indexing; S16: The following formula is used to... and stated Stacked along the channel dimension as a dual-scale Mel feature tensor : , ; In the formula: The two-scale Mel feature tensor has the following shape: 2 represents the number of channels; The total number of Mel bands; Total number of frames; for In the 0th channel, the The Mel frequency band, the first The element value at the frame position; for In the first channel, the The Mel frequency band, the first The element value at the frame position; For Mel band indexing; For frame index.
3. The algorithm according to claim 1, characterized in that, S2 includes the following specific steps: S21: Calculate the prior weighted initial value of the resonance peak using the following formula. : ; In the formula: For the first The prior weighted initial values of the resonant peaks in each Mel band, with a range of values. ; For index The operation to find the maximum value; For natural index calculations; For the first The center frequency of the Mel band; For the first A priori central frequency; For the first A priori bandwidth; For Mel band indexing; For prior central index; S22: Take 6 sets of the aforementioned prior center frequencies and prior bandwidths. Take numbers 1 to 6 in sequence, which correspond to the fundamental frequency region of human voice, the first formant region, the second formant region, the third formant region, the fourth formant region, and the high frequency region of unvoiced fricatives, respectively. S23: Obtain a length equal to the total number of Melbands Band weighted parameter vector The These are fixed parameter vectors obtained through pre-training; for The Middle One element; The total number of Mel bands; For Mel band indexing; S24: Generate the prior weighting value of the resonance peak using the following formula. : ; In the formula: For the first Prior weighting values of the resonant peaks in the Mel frequency band, with a range of values. ; The sigmoid function is expressed as follows: ; For natural index calculations; for Input variables; For Mel band indexing; S25: Convert the dual-scale Mel feature tensor Global average pooling is performed along the time dimension to obtain the frequency dimension feature sequence. , The shape is 2 represents the number of channels; [The sentence is incomplete and requires more context to translate accurately.] Global average pooling is performed along the frequency band dimension to obtain the time-dimensional feature sequence. , The shape is ;in For the total number of Mel bands, Total number of frames; S26: The frequency dimension feature sequence At each Mel band position The value multiplied by the The weighted frequency-dimensional feature sequence is obtained. , and Same shape; S27: The weighted frequency-dimensional feature sequence... Input the first coding layer of the frequency stream encoder to obtain the output features of the first coding layer of the frequency stream. The time-dimensional feature sequence The first coding layer of the time-stream encoder is input to obtain the output features of the first coding layer of the time-stream encoder. .
4. The algorithm according to claim 1, characterized in that, S3 includes the following specific steps: S31: The time-stream encoder and the frequency-stream encoder each include... One coding layer, The input is an integer not less than 2; each coding layer of the time stream sequentially includes a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, and a bidirectional gated recurrent unit layer, ensuring that the time dimension length of the feature sequence input to each coding layer of the time stream remains constant before and after processing by that coding layer. Each coding layer of the frequency stream sequentially includes a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, and a multi-head self-attention layer, ensuring that the frequency band dimension length of the feature sequence input to each coding layer of the frequency stream remains constant before and after processing by that coding layer. ; in the Coding layer and the first A time-frequency interaction attention module is set up between the encoding layers. Take 1 to 1 in sequence ;in For the coding layer index, The total number of frames. The total number of Mel bands; S32: Configure 6 linear projection matrices in each of the time-frequency interaction attention modules. , , , , , The six linear projection matrices are all fixed parameters obtained through pre-training; the time flow... Coding layer output The dimension is OK The column, the frequency stream number Coding layer output The dimension is OK Column; the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK The column, the The dimension is OK Column; among which For the time flow channel dimension, For frequency flow channel dimension, The attention dimension after projection. , and All are preset positive integers; For encoding layer index; S33: In the Coding layer and the first In the time-frequency interaction attention module between coding layers, the interaction features injected into the time stream by frequency information are calculated according to the following formula. : ; In the formula: For the first At the encoding layer, frequency information is injected into the interactive feature matrix of the time stream, with dimensions of [missing information]. OK List; Normalized exponentiation is performed row-by-row; For encoding layer index; S34: In the time-frequency interaction attention module, the interaction features injected into the frequency stream by time information are calculated according to the following formula. : ; In the formula: For the first The interaction feature matrix at the encoding layer, into which temporal information is injected into the frequency stream, has a dimension of [missing information]. OK List; S35: Generate the time stream according to the following formulas respectively. Input of the coding layer and frequency flow Input of the coding layer : , ; In the formula: For the time flow The input feature matrix of the encoding layer has the same dimension. ; For frequency flow number The input feature matrix of the encoding layer has the same dimension. ; For layer normalization operations; For encoding layer index; S36: The Input the time stream Encoding layer, to obtain the time stream Coding layer output ; will the Input the frequency stream The coding layer yields the frequency stream. Coding layer output ; According to the methods described in S33 to S36, Take 1 to 1 in sequence Iterative execution yields the final output of the time stream. and frequency stream final layer output .
5. The algorithm according to claim 1, characterized in that, S4 includes the following specific steps: S41: Output the last layer of the time stream Global average pooling is performed along the time dimension to obtain the global feature vector of the time flow. , The length is equal to the dimension of the time stream channel. ; S42: Output the final layer of the frequency stream. Global average pooling is performed along the frequency band dimension to obtain the global eigenvector of the frequency flow. , The length is equal to the frequency flow channel dimension ; S43: The above and stated Concatenate along the channel dimension to form a joint feature vector , will the The current speech window is output after passing through a fully connected layer and a Softmax layer. sentiment category probability vector The weight parameters of the fully connected layer are fixed parameters obtained through pre-training. The length is equal to the total number of preset emotion categories. The probability vector; The total number of preset emotion categories; This is the window timing index.
6. The algorithm according to claim 1, characterized in that, S5 includes the following specific steps: S51: Calculate the current voice window using the following formula Prediction confidence : ; In the formula: For the current voice window The prediction confidence level, range of values ; This is an operation that takes the maximum value of each element of a vector. For the current voice window The probability vector of the sentiment category; For window timing index; S52: The following formula is used to... Mapped to the current voice window Adaptive smoothing coefficient : ; In the formula: For the current voice window The adaptive smoothing coefficient; This is the lower bound of the smoothing coefficient; This is the upper bound of the smoothing coefficient; S53: When At that time, make the current voice window Smoothed sentiment category probability vector ;when When the value is greater than 0, the current voice window is calculated using the following formula. Smoothed sentiment category probability vector : ; In the formula: For the current voice window The smoothed sentiment category probability vector; Previous voice window The smoothed sentiment category probability vector; For window timing index; S54: Regarding the above Take the index of the maximum value as the current voice window. The output sentiment category is taken from the... The maximum value is used as the current voice window. Smoothed post-confidence .
7. The algorithm according to claim 1, characterized in that, S6 includes the following specific steps: S61: Combine the output sentiment category and the smoothed post-confidence obtained in S5. The information is transmitted to the empathic response module of the emotional companion robot. The empathic response module includes a pre-stored response strategy library; the response strategy library is a mapping table between emotion categories and response strategies, and each response strategy includes at least a voice text template, a prosodic parameter set, and a body movement number; S62: The empathic response module determines a stable emotional state according to the following rules: if from the... The first voice window to the second A total of 1 voice window The output emotion categories of each consecutive speech window are the same, and the smoothed post-confidence of each speech window is... All are not less than the preset output threshold. ,in Take in sequence to Then, the same output sentiment category is recorded as a stable sentiment state. ; The preset number of consecutive windows is an integer between 3 and 10; For the first Smoothed post-confidence of each voice window; To stabilize emotional state; Index for the voice window; S63: If the above Not less than the Then the empathic response module will respond according to the current voice window. The processor retrieves the corresponding response strategy from the response strategy library for the output emotion category, and generates the speech synthesis control command and the body movement control command based on the retrieved response strategy. S64: If the above Smaller than the And there are recorded stable emotional states. Then the empathic response module maintains the The corresponding response strategy; if the Smaller than the Furthermore, no such stable emotional state has been recorded. Then the empathic response module adopts a preset neutral response strategy, that is, the response strategy corresponding to the neutral emotion category in the response strategy library.