Emotional speech recognition method of intelligent health care robot based on deep network
By combining the Flamingo algorithm with multi-dimensional feature modeling and hyperparameter optimization of recurrent neural networks, the problems of poor generalization ability and static interaction strategies of existing emotional speech recognition systems in health care scenarios are solved, achieving a high-accuracy and natural human-computer interaction experience, and improving the stability and adaptability of emotion recognition.
Patent Information
- Application Number
- CN202510946378.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing emotional speech recognition systems have poor generalization capabilities in health care scenarios, and find it difficult to accurately model and discern delicate and slowly changing speech emotions. They also lack weighted recognition strategies for multiple emotion categories, making it difficult for recognition results to drive actual interactive behaviors. Traditional hyperparameter tuning is inefficient, and robot interaction strategies cannot be dynamically adjusted, making it difficult to meet the requirements of high accuracy, high real-time performance, strong adaptability, and natural interactivity.
Combining the Flamingo algorithm, recurrent neural network and emotional speech feature extraction technology, an emotional speech recognition model is constructed through multi-dimensional feature modeling and intelligent hyperparameter optimization. Deep recurrent neural network is used to extract high-order emotional dynamic features. Combined with the Flamingo algorithm, global search and local fine-grained optimization are performed to dynamically adjust voice interaction strategies and behavioral responses.
It significantly improves the accuracy of emotion recognition and model stability, achieves a highly adaptable and natural human-computer interaction experience, and enhances the humanization level of health care services and the quality of system interaction.
Smart Images

Figure CN120673746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of emotional speech recognition technology, and in particular to an emotional speech recognition method for an intelligent health care robot based on a deep network. Background Art
[0002] Smart healthcare technology, a new model integrating artificial intelligence and health services, is becoming a crucial support tool in areas such as elderly care, chronic disease management, and emotional counseling. Voice-based human-computer interaction systems, due to their natural and accessible nature, have been widely adopted in scenarios such as smart elderly care robots, voice-assisted systems, and health Q&A platforms. In these applications, the interaction between users and healthcare devices requires not only understanding of the language content but also, more importantly, recognizing the emotional information contained in the speech to achieve emotional perception and personalized responses, thereby improving the quality of interaction and user experience.
[0003] As one of the key technologies for achieving the goal of "human-machine empathy," emotional speech recognition aims to identify the user's emotional state by analyzing the emotional features in speech signals. Currently, emotional speech recognition systems mostly use an acoustic feature modeling approach to extract and combine features such as short-term energy, intonation, speech rate, and resonance peaks to construct an emotion classifier, thereby achieving emotion recognition. However, such methods have poor generalization capabilities in natural language scenarios, especially in health care scenarios. Speech emotions are often delicate, slowly changing, and highly context-dependent, making it difficult for traditional methods to accurately model and identify them.
[0004] In recent years, with the widespread application of deep learning technology, long short-term memory networks and gated recurrent unit recurrent neural network structures have been introduced into the field of emotional speech recognition, effectively improving the ability to model speech temporal features. These can capture long-range semantic dependencies and extract emotional dynamic features with better contextual understanding capabilities. However, in actual deployment, the performance of deep models is highly dependent on the reasonable configuration of hyperparameters, such as learning rate, network depth, regularization coefficient, etc. In existing technologies, grid search, random search or manual parameter adjustment are often used to adjust hyperparameters, which are inefficient and difficult to adapt to complex and changing health and wellness interaction scenarios. In addition, existing research often focuses only on improving recognition accuracy and lacks integrated modeling of human-computer interaction strategies that are ultimately deployed on health and wellness robots, making it difficult for recognition results to naturally drive actual interactive behaviors.
[0005] At the same time, most current emotional speech recognition systems remain at the laboratory testing stage, lacking systematic design for real-world healthcare environments. For example, user emotion recognition is often performed in isolation, failing to integrate feedback with the robot's actual execution system, making it difficult to achieve a true "recognition-response-adjustment" closed-loop control. Furthermore, many existing systems fail to fully consider the complexity and imbalance of emotion categories and lack weighted recognition strategies for multiple emotion categories, leading to confusion and misidentification when faced with segmented emotional states.
[0006] In terms of optimization algorithms, although some studies have attempted to introduce swarm intelligence optimization algorithms to improve the efficiency of model parameter adjustment, such as particle swarm optimization and genetic algorithms, these algorithms are still prone to falling into local optimality in high-dimensional search space, and the convergence speed is unstable. Moreover, most of them have not been improved in combination with the characteristics of emotion recognition tasks, resulting in limited effect of the optimization results on improving the recognition accuracy of the actual system.
[0007] In terms of interactive response mechanisms, traditional healthcare robots often rely on fixed rules-driven behavioral decision-making models. These robots are unable to dynamically adjust their interaction strategies based on different users and emotional states, making it difficult to meet personalized, context-aware, and other intelligent healthcare interaction needs. Existing interaction logic lacks modeling for the linkage between emotion recognition results, voice intonation modulation, and nonverbal behavioral feedback, limiting the robot's ability to empathize.
[0008] In summary, the existing emotional speech recognition and health care human-computer interaction systems still have many shortcomings in feature modeling, model optimization, real-time deployment, strategy control, etc., and it is difficult to fully meet the comprehensive requirements of high accuracy, high real-time performance, strong adaptability and natural interactivity in smart health care scenarios.
[0009] Therefore, how to provide an emotional speech recognition method for an intelligent health care robot based on a deep network is an urgent problem that technicians in this field need to solve. Summary of the Invention
[0010] One purpose of the present invention is to propose an emotional speech recognition method for an intelligent health care robot based on a deep network. The present invention combines the Flamingo algorithm, recurrent neural network and emotional speech feature extraction technology, and effectively captures the emotional fluctuations and temporal variation characteristics in the user's voice through multi-dimensional feature modeling, emotional dynamic construction and intelligent hyperparameter optimization. The system uses a deep recurrent neural network to extract high-order emotional dynamic features, and combines the Flamingo algorithm to perform global search and local fine-grained optimization of hyperparameters to dynamically improve the accuracy of emotion recognition and model stability. The emotional speech feature extraction process fully integrates spectral features, time domain features and emotional label mapping relationships to ensure that the model has good discrimination and generalization capabilities for different emotional states. Ultimately, the intelligent health care robot intelligently adjusts the voice interaction strategy and behavioral response based on the emotional category identified in real time, achieving a highly adaptable, natural, and emotionally integrated human-computer interaction experience, significantly improving the humanization level of health care services and the quality of system interaction.
[0011] The emotional speech recognition method of the intelligent health care robot based on a deep network according to an embodiment of the present invention includes the following steps:
[0012] S1. Collect and pre-process the voice data of users in the health care environment to construct an emotional voice feature vector sequence;
[0013] S2. Use a recurrent neural network to build an emotional speech recognition model, set the input layer to receive the emotional speech feature vector at each time step, and use long short-term memory units to extract the emotional dynamic features within the emotional speech feature vector sequence;
[0014] S3. Based on the emotional speech recognition model, set the hyperparameter set to be optimized, use the Flamingo algorithm to encode the hyperparameter set into the position vector of the flamingo individual, and initialize the flamingo population;
[0015] S4. Update the position vector of each flamingo based on its foraging behavior mechanism, retrain the emotional speech recognition model using the updated hyperparameter set, and calculate the emotion recognition accuracy on the validation set as the fitness value.
[0016] S5. Select the flamingo individual with the best emotion recognition fitness value as the current optimal hyperparameter set, update the position vector of the flamingo population, and determine the optimal hyperparameter set of the emotional speech recognition model;
[0017] S6. Based on the optimal hyperparameter set, the emotional speech recognition model is finally trained to determine the final emotional speech recognition model parameters;
[0018] S7. Deploy the final emotional speech recognition model to the smart healthcare robot, and adjust the smart healthcare robot's voice interaction strategy and behavioral response based on the identified emotion category.
[0019] Optionally, the speech data specifically includes original speech waveform, speech content text, speech acoustic features, emotion annotations, speaker information and background noise features.
[0020] Optionally, the hyperparameters to be optimized specifically include a learning rate, a number of hidden layer units, a regularization coefficient, and a dropout rate.
[0021] Optionally, the S1 specifically includes:
[0022] S11, performing pre-emphasis processing on the collected voice data, subtracting the voice signal value of the previous sampling point from the voice signal value of the current sampling point and multiplying it by a pre-emphasis coefficient, wherein the pre-emphasis coefficient is between 0.9 and 1.0;
[0023] S12, performing frame processing on the pre-emphasized voice data, setting the number of sampling points of each frame to a fixed value, setting the overlapping part between frames to a fixed number of sampling points, and extracting the voice segments of each frame;
[0024] S13, performing weighted processing on each frame segment to smooth the boundary of the transition frame segment signal;
[0025] S14, extracting frequency domain features from each frame segment to generate data representation reflecting signal energy distribution and change trend;
[0026] S15. Extracting multiple different types of feature indicators based on the frequency domain feature data, including features representing frequency distribution, pitch information, and signal change rate;
[0027] S16, combining the feature indicators extracted from each frame segment into an emotional speech feature vector in order;
[0028] S17. Arrange all single-frame emotional speech feature vectors in chronological order to construct an emotional speech feature vector sequence.
[0029] Optionally, the S2 specifically includes:
[0030] S21, determining the input data type of emotion recognition, selecting the emotional speech feature vector sequence as input data, and defining an emotional speech feature vector corresponding to each time step;
[0031] S22. Based on a recurrent neural network, an emotional speech recognition model is constructed, and an input layer is set. The input layer sequentially receives the features of each time step in the emotional speech feature vector sequence;
[0032] S23, configuring a long short-term memory unit after the input layer to process the emotional speech feature vector sequence and extract the potential emotion-related dynamic features within the time series;
[0033] S24. A hidden layer is set after the long short-term memory unit to further abstract and enhance the extracted emotional dynamic features. The hidden layer output features are generated using a weighted summation mechanism:
[0034]
[0035] Among them, h agg is the emotional dynamic feature, T is the total number of time steps, t is the sequence number of the time step currently being processed, and h t is the hidden state feature vector output by the long short-term memory unit at time step t, ω t is the weighting coefficient of time step t;
[0036] S25. Setting a fully connected layer after the hidden layer to map the extracted emotional dynamic feature information to the corresponding emotional category prediction space;
[0037] S26. Set the output layer to normalize the prediction results output by the fully connected layer to generate the predicted probability distribution of each emotion category.
[0038] Optionally, the S3 specifically includes:
[0039] S31. Setting a set of hyperparameters to be optimized for the emotional speech recognition model, and representing the position vector of each flamingo individual as a set of specific hyperparameter combinations;
[0040] S32. Randomly initialize the positions of several flamingo individuals in the hyperparameter search space, where each flamingo individual position corresponds to a set of hyperparameter combinations, to form an initial flamingo population position set;
[0041] S33. Setting an initial velocity vector for each flamingo individual in the flamingo population, where the initial velocity vector is the movement direction and change amplitude of the flamingo individual in the hyperparameter search space, to generate an initial velocity set;
[0042] S34, binding the position vector of each flamingo individual to the hyperparameter combination of the emotional speech recognition model to generate an initial flamingo population hyperparameter distribution;
[0043] S35. The initial flamingo population hyperparameter distribution is the initial state of different individuals in the search space. An initial velocity vector is assigned to each flamingo individual. The velocity vector is the direction and amplitude of movement in the hyperparameter space.
[0044] Optionally, the S4 specifically includes:
[0045] S41. Based on the flamingo foraging behavior model, the new position vector of each flamingo is calculated, and local search and global jumping behaviors are simulated. The flamingo moves closer to the optimal solution in the hyperparameter search space.
[0046] S42, using the updated position vector as a new hyperparameter combination of the emotional speech recognition model, and retraining the new emotional speech recognition model, where the training input is the emotional speech feature vector sequence;
[0047] S43, using the trained emotional speech recognition model to predict the validation set samples, and outputting the emotional category probability distribution of each speech sample;
[0048] S44. For each flamingo individual, calculate the emotion recognition accuracy on the validation set as the fitness value:
[0049]
[0050] Among them, f(p i ) is the fitness value of the i-th flamingo individual, c is the emotion category number, C is the total number of emotion categories, w c is the importance weight of the c-th category emotion, N c is the number of samples in the validation set whose true labels belong to the cth class, k is the recognition result score of each sample, When the predicted emotion category is consistent with the true emotion category, the value is 1, otherwise the value is 0. is the emotional dynamic feature vector extracted from the emotional speech recognition model of the kth verification sample, is the modulus of the emotional dynamic feature vector, α and β are non-negative weight coefficients, i is the i-th flamingo individual, p i is the position vector of the i-th flamingo individual.
[0051] Optionally, the S5 specifically includes:
[0052] S51. In each round of iteration, all individuals in the flamingo population are compared according to their fitness values, and the individual with the highest current recognition accuracy is selected as the optimal individual of the current iteration round;
[0053] S52. Update the positions of the remaining flamingo individuals, and calculate based on the position of the current individual, the position of the optimal individual, and the position of the random individual:
[0054]
[0055] Among them, f(p i ) is the fitness value of the i-th flamingo individual, is the new position vector of the i-th flamingo individual at the t+1th iteration, p i is the position vector of the i-th flamingo individual, r1 and r2 are random factors between 0 and 1, i is the i-th flamingo individual, is the current position vector of the i-th flamingo individual, is the position of a flamingo individual randomly selected from the current population, P (t) is the set of position vectors of all flamingo individuals at the tth iteration, j is a flamingo individual randomly selected from the entire flamingo population, and argmax is the position vector of the hyperparameter combination that maximizes the fitness value of emotion recognition in the flamingo population;
[0056] S53. Generate a new hyperparameter combination based on the updated position vector of each flamingo individual, apply the new hyperparameter combination to the trained emotional speech recognition model, calculate the recognition accuracy and the extracted emotional dynamic feature strength for each emotional category in the validation set, accumulate the sum of the weighted accuracy and weighted feature strength of all emotional categories, and use the weighted average as the fitness value of the flamingo individual in the current iteration;
[0057] S54: Determine whether the preset maximum number of rounds 200 is reached. If not, return to step S51 to continue iteration;
[0058] S55. When the preset maximum number of rounds of 200 is reached, the individual with the highest fitness value is selected from the current flamingo population, and the corresponding position vector is the optimal hyperparameter combination of the final emotional speech recognition model.
[0059] Optionally, the S6 specifically includes:
[0060] S61. Initializing the training configuration of the final emotional speech recognition model based on the optimal hyperparameter set, including the learning rate, the number of hidden units, the regularization coefficient, and the dropout rate parameters;
[0061] S62, inputting the emotional speech feature vector sequence in the training set into the final emotional speech recognition model;
[0062] S63, using long short-term memory units to sequentially process the emotional speech feature vectors, extracting the emotional dynamic features within the time series, performing weighted aggregation on the hidden state features of all time steps, and generating an emotional dynamic feature aggregation vector for each speech sample in the training set;
[0063] S64, inputting the emotion dynamic feature aggregation vector into the fully connected layer and the output layer, and after weighted summation and bias adjustment, using a normalization function mapping to generate the emotion category probability distribution corresponding to each speech sample in the training set;
[0064] S65. Repeat S61-S64 using the emotional speech feature vector sequence in the validation set to obtain the emotional category probability distribution corresponding to each sample in the validation set;
[0065] S66. Calculate the latest emotion recognition accuracy on the training set and the validation set respectively. The latest emotion recognition accuracy is the ratio of the number of emotion category samples correctly predicted by the emotion speech recognition model to the total number of samples. A correctly identified sample is scored as one point, and an incorrectly identified sample is scored as zero. The sum is accumulated and divided by the total number of samples to obtain the final emotion recognition accuracy.
[0066] S67. Based on the final emotion recognition accuracy of the training set and the validation set, determine all training parameters of the final emotion speech recognition model.
[0067] Optionally, the S7 specifically includes:
[0068] S71. Deploy the final emotional speech recognition model to the local deep neural network inference module of the smart healthcare robot, load the optimal hyperparameter set, and establish a data channel connection with the speech acquisition module;
[0069] S72. During the operation of the smart health care robot, collect user voice signals in real time, perform framing, pre-emphasis, and endpoint detection on the voice data, and extract a time series emotional voice feature vector sequence;
[0070] S73, inputting the emotional speech feature vector sequence into the final emotional speech recognition model, extracting the hidden feature representation of each time frame through a recurrent neural network, and aggregating the hidden features of all time frames according to a weighted rule to generate an overall real-time emotional dynamic feature;
[0071] S74: Input the real-time emotional dynamic features into the fully connected layer and the output layer to generate a probability distribution result of multiple emotion categories corresponding to the current user's voice, and select the emotion category with the largest probability value as the current emotion recognition output label;
[0072] S75. Based on the identified emotion category label, call the preset voice interaction strategy library and behavior control module in the smart healthcare robot, and select the voice feedback method, tone and speed adjustment scheme, and non-verbal behavior response corresponding to the emotion category label in the strategy library.
[0073] The beneficial effects of the present invention are:
[0074] This paper introduces emotional speech feature extraction technology, combined with deep recurrent neural network modeling, to effectively improve the accuracy of recognizing the emotional state of users' speech in healthcare environments. Compared to traditional acoustic feature-driven recognition methods, this paper optimizes the feature system for the emotional dimension, capable of capturing more subtle and complex emotional changes in speech signals. It demonstrates higher accuracy and stronger noise immunity, especially in fine-grained emotion classification tasks.
[0075] This paper innovatively employs the Flamingo intelligent optimization algorithm to perform global search and local fine-tuning of the emotion recognition model's hyperparameters, overcoming the inefficiency and tendency to get stuck in local optima associated with traditional grid and random search methods. By dynamically updating the hyperparameter space by simulating the foraging behavior of flamingos, the system can rapidly converge to a high-performance region in a complex parameter space, significantly reducing tuning time and improving training efficiency. Furthermore, it maintains stable and robust recognition performance despite the uneven distribution of samples across different emotion categories.
[0076] At the deployment level, this invention seamlessly integrates a deep network-based emotional speech recognition model into a smart healthcare robot system, establishing a complete chain from real-time user speech acquisition, emotional feature extraction, deep reasoning, to dynamic mapping of interaction strategies. Based on real-time emotion recognition results, the robot system automatically adjusts speech interaction content, intonation and speed parameters, and action feedback, forming a dynamic and adaptive human-machine emotional interaction mechanism. This significantly enhances the robot's ability to perceive the user's emotional state and the naturalness of interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0078] Figure 1 This is a flowchart of the emotional speech recognition method for the intelligent health care robot based on deep network proposed by the present invention;
[0079] Figure 2 This is a schematic diagram of the emotional speech recognition method for the intelligent health care robot based on deep network proposed in the present invention;
[0080] Figure 3 This is the data flow diagram of the emotional speech recognition method for the intelligent health care robot based on deep network proposed in this invention. DETAILED DESCRIPTION
[0081] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0082] refer to Figure 1-3 , an emotional speech recognition method for an intelligent health care robot based on a deep network includes the following steps:
[0083] S1. Collect and pre-process the voice data of users in the health care environment to construct an emotional voice feature vector sequence;
[0084] S2. Use a recurrent neural network to build an emotional speech recognition model, set the input layer to receive the emotional speech feature vector at each time step, and use long short-term memory units to extract the emotional dynamic features within the emotional speech feature vector sequence;
[0085] S3. Based on the emotional speech recognition model, set the hyperparameter set to be optimized, use the Flamingo algorithm to encode the hyperparameter set into the position vector of the flamingo individual, and initialize the flamingo population;
[0086] S4. Update the position vector of each flamingo based on its foraging behavior mechanism, retrain the emotional speech recognition model using the updated hyperparameter set, and calculate the emotion recognition accuracy on the validation set as the fitness value.
[0087] S5. Select the flamingo individual with the best emotion recognition fitness value as the current optimal hyperparameter set, update the position vector of the flamingo population, and determine the optimal hyperparameter set of the emotional speech recognition model;
[0088] S6. Based on the optimal hyperparameter set, the emotional speech recognition model is finally trained to determine the final emotional speech recognition model parameters;
[0089] S7. Deploy the final emotional speech recognition model to the smart healthcare robot, and adjust the smart healthcare robot's voice interaction strategy and behavioral response based on the identified emotion category.
[0090] This paper leverages the foraging behavior update mechanism of the Flamingo algorithm to globally optimize the hyperparameters of the emotional speech recognition model, improving its sensitivity to emotional changes and recognition accuracy in a healthcare setting. By modeling the dynamic evolution of a Flamingo population and incorporating emotion recognition fitness values to guide the hyperparameter search, it avoids being trapped in local optima. Ultimately, this method is deployed on a smart healthcare robot to implement personalized voice interaction strategy adjustments based on emotion categories, enhancing the naturalness and adaptability of human-computer interaction.
[0091] In this embodiment, the speech data specifically includes original speech waveform, speech content text, speech acoustic features, emotion annotation, speaker information and background noise features.
[0092] This method constructs a multidimensional speech data space by collecting raw speech waveforms, speech content text, acoustic features, emotion annotations, speaker information, and background noise characteristics to comprehensively characterize user speech. By introducing multimodal feature joint analysis, emotion recognition relies not only on a single acoustic feature but also combines text semantics, emotion annotations, and environmental noise conditions to achieve more accurate and robust emotion understanding. Furthermore, speaker information is leveraged to improve the system's adaptability to individual emotional differences.
[0093] In this embodiment, the hyperparameters to be optimized specifically include learning rate, number of hidden layer units, regularization coefficient and dropout rate.
[0094] The present invention builds a comprehensive hyperparameter tuning space by incorporating the learning rate, number of hidden layer units, regularization coefficient, and dropout rate into the hyperparameter optimization range. Through multi-dimensional hyperparameter joint search, taking into account training speed, capacity control, and overfitting suppression, the generalization ability and stability of the emotional speech recognition model are effectively improved. At the same time, combined with a dynamically adaptive hyperparameter adjustment mechanism, the emotional speech recognition model can maintain a high recognition accuracy and robustness under different health care environment noise levels and emotional change amplitudes.
[0095] In this embodiment, S1 specifically includes:
[0096] S11, performing pre-emphasis processing on the collected voice data, subtracting the voice signal value of the previous sampling point from the voice signal value of the current sampling point and multiplying it by a pre-emphasis coefficient, wherein the pre-emphasis coefficient is between 0.9 and 1.0;
[0097] S12, performing frame processing on the pre-emphasized voice data, setting the number of sampling points of each frame to a fixed value, setting the overlapping part between frames to a fixed number of sampling points, and extracting the voice segments of each frame;
[0098] S13, performing weighted processing on each frame segment to smooth the boundary of the transition frame segment signal;
[0099] S14, extracting frequency domain features from each frame segment to generate data representation reflecting signal energy distribution and change trend;
[0100] S15. Extracting multiple different types of feature indicators based on the frequency domain feature data, including features representing frequency distribution, pitch information, and signal change rate;
[0101] S16, combining the feature indicators extracted from each frame segment into an emotional speech feature vector in order;
[0102] S17. Arrange all single-frame emotional speech feature vectors in chronological order to construct an emotional speech feature vector sequence.
[0103] This method fully exploits the fine-grained dynamic variation characteristics of speech signals by sequentially performing pre-emphasis, framing, weighting, and frequency domain feature extraction on the collected speech data. By extracting multidimensional features representing frequency distribution, pitch information, and signal change rate, a sequence of emotional speech feature vectors is constructed. This ensures the richness of feature expression and enhances the temporal coherence of emotional variation patterns, helping subsequent emotional speech recognition models accurately model speech emotion dynamics and improve recognition accuracy and robustness.
[0104] In this embodiment, S2 specifically includes:
[0105] S21, determining the input data type of emotion recognition, selecting the emotional speech feature vector sequence as input data, and defining an emotional speech feature vector corresponding to each time step;
[0106] S22. Based on a recurrent neural network, an emotional speech recognition model is constructed, and an input layer is set. The input layer sequentially receives the features of each time step in the emotional speech feature vector sequence;
[0107] S23, configuring a long short-term memory unit after the input layer to process the emotional speech feature vector sequence and extract the potential emotion-related dynamic features within the time series;
[0108] S24. A hidden layer is set after the long short-term memory unit to further abstract and enhance the extracted emotional dynamic features. The hidden layer output features are generated using a weighted summation mechanism:
[0109]
[0110] Among them, h agg is the emotional dynamic feature, T is the total number of time steps, t is the sequence number of the time step currently being processed, and h t is the hidden state feature vector output by the long short-term memory unit at time step t, ω t is the weighting coefficient of time step t;
[0111] S25. Setting a fully connected layer after the hidden layer to map the extracted emotional dynamic feature information to the corresponding emotional category prediction space;
[0112] S26. Set the output layer to normalize the prediction results output by the fully connected layer to generate the predicted probability distribution of each emotion category.
[0113] This paper constructs a recurrent neural network model that uses a sequence of emotional speech feature vectors as input. It uses long-short-term memory units to extract dynamic emotional features from time series and enhances key information through a weighted mechanism, improving feature expression. Hidden layers and fully connected layers are used to effectively map emotional features to emotional categories, ultimately generating a normalized predicted probability distribution for emotional categories. This improves the recognition accuracy and classification stability of complex speech emotional changes.
[0114] In this embodiment, S3 specifically includes:
[0115] S31. Setting a set of hyperparameters to be optimized for the emotional speech recognition model, and representing the position vector of each flamingo individual as a set of specific hyperparameter combinations;
[0116] S32. Randomly initialize the positions of several flamingo individuals in the hyperparameter search space, where each flamingo individual position corresponds to a set of hyperparameter combinations, to form an initial flamingo population position set;
[0117] S33. Setting an initial velocity vector for each flamingo individual in the flamingo population, where the initial velocity vector is the movement direction and change amplitude of the flamingo individual in the hyperparameter search space, to generate an initial velocity set;
[0118] S34, binding the position vector of each flamingo individual to the hyperparameter combination of the emotional speech recognition model to generate an initial flamingo population hyperparameter distribution;
[0119] S35. The initial flamingo population hyperparameter distribution is the initial state of different individuals in the search space. An initial velocity vector is assigned to each flamingo individual. The velocity vector is the direction and amplitude of movement in the hyperparameter space.
[0120] This method maps the set of hyperparameters to be optimized for the emotional speech recognition model to the position vectors of individual flamingos. This is combined with the hyperparameter search space to initialize the position and velocity of the flamingo population, forming a comprehensive initial hyperparameter distribution. By setting independent movement directions and amplitudes for each individual, the search process is diversified and exploratory, enhancing global search capabilities. This provides a good initial foundation for subsequent dynamic adjustment and evolutionary optimization, improving the efficiency of hyperparameter optimization and ultimately performance.
[0121] In this embodiment, the S4 specifically includes:
[0122] S41. Based on the flamingo foraging behavior model, the new position vector of each flamingo is calculated, and local search and global jumping behaviors are simulated. The flamingo moves closer to the optimal solution in the hyperparameter search space.
[0123] S42, using the updated position vector as a new hyperparameter combination of the emotional speech recognition model, and retraining the new emotional speech recognition model, where the training input is the emotional speech feature vector sequence;
[0124] S43, using the trained emotional speech recognition model to predict the validation set samples, and outputting the emotional category probability distribution of each speech sample;
[0125] S44. For each flamingo individual, calculate the emotion recognition accuracy on the validation set as the fitness value:
[0126]
[0127] Among them, f(p i) is the fitness value of the i-th flamingo individual, c is the emotion category number, C is the total number of emotion categories, w c is the importance weight of the c-th category emotion, N c is the number of samples in the validation set whose true labels belong to the cth class, k is the recognition result score of each sample, When the predicted emotion category is consistent with the true emotion category, the value is 1, otherwise the value is 0. is the emotional dynamic feature vector extracted from the emotional speech recognition model of the kth verification sample, is the modulus of the emotional dynamic feature vector, α and β are non-negative weight coefficients, i is the i-th flamingo individual, p i is the position vector of the i-th flamingo individual.
[0128] This method, based on modeling flamingo foraging behavior, simulates the local search and global jump dynamics of individuals within a hyperparameter search space, driving the continuous evolution of hyperparameter combinations toward the optimal solution. The emotional speech recognition model is retrained using the updated hyperparameters, and the fitness value is calculated based on the emotion recognition accuracy of the validation set to guide optimization. This method balances the exploratory and convergent nature of the hyperparameter search, effectively improving the performance and stability of the emotional speech recognition model in environments with diverse speech characteristics.
[0129] In this embodiment, the S5 specifically includes:
[0130] S51. In each round of iteration, all individuals in the flamingo population are compared according to their fitness values, and the individual with the highest current recognition accuracy is selected as the optimal individual of the current iteration round;
[0131] S52. Update the positions of the remaining flamingo individuals, and calculate based on the position of the current individual, the position of the optimal individual, and the position of the random individual:
[0132]
[0133] Among them, f(p i ) is the fitness value of the i-th flamingo individual, is the new position vector of the i-th flamingo individual at the t+1th iteration, p i is the position vector of the i-th flamingo individual, r1 and r2 are random factors between 0 and 1, i is the i-th flamingo individual, is the current position vector of the i-th flamingo individual, is the position of a flamingo individual randomly selected from the current population, P (t)is the set of position vectors of all flamingo individuals at the tth iteration, j is a flamingo individual randomly selected from the entire flamingo population, and arg max is the position vector of the hyperparameter combination that maximizes the fitness value of emotion recognition in the flamingo population;
[0134] S53. Generate a new hyperparameter combination based on the updated position vector of each flamingo individual, apply the new hyperparameter combination to the trained emotional speech recognition model, calculate the recognition accuracy and the extracted emotional dynamic feature strength for each emotional category in the validation set, accumulate the sum of the weighted accuracy and weighted feature strength of all emotional categories, and use the weighted average as the fitness value of the flamingo individual in the current iteration;
[0135] S54: Determine whether the preset maximum number of rounds 200 is reached. If not, return to step S51 to continue iteration;
[0136] S55. When the preset maximum number of rounds of 200 is reached, the individual with the highest fitness value is selected from the current flamingo population, and the corresponding position vector is the optimal hyperparameter combination of the final emotional speech recognition model.
[0137] This method selects the flamingo individual with the highest recognition accuracy as a guide in each iteration, dynamically updating the population distribution by combining individual positions, optimal positions, and random individual positions. This achieves multi-directional guidance for hyperparameter search and local fine-grained optimization. Individual fitness is comprehensively assessed based on weighted emotion recognition accuracy and feature strength, improving the optimization process's ability to detect subtle emotional changes. Ultimately, through iterative convergence, the globally optimal hyperparameter combination is obtained, significantly improving the performance stability and generalization of the emotional speech recognition model.
[0138] In this embodiment, S6 specifically includes:
[0139] S61. Initializing the training configuration of the final emotional speech recognition model based on the optimal hyperparameter set, including the learning rate, the number of hidden units, the regularization coefficient, and the dropout rate parameters;
[0140] S62, inputting the emotional speech feature vector sequence in the training set into the final emotional speech recognition model;
[0141] S63, using long short-term memory units to sequentially process the emotional speech feature vectors, extracting the emotional dynamic features within the time series, performing weighted aggregation on the hidden state features of all time steps, and generating an emotional dynamic feature aggregation vector for each speech sample in the training set;
[0142] S64, inputting the emotion dynamic feature aggregation vector into the fully connected layer and the output layer, and after weighted summation and bias adjustment, using a normalization function mapping to generate the emotion category probability distribution corresponding to each speech sample in the training set;
[0143] S65. Repeat S61-S64 using the emotional speech feature vector sequence in the validation set to obtain the emotional category probability distribution corresponding to each sample in the validation set;
[0144] S66. Calculate the latest emotion recognition accuracy on the training set and the validation set respectively. The latest emotion recognition accuracy is the ratio of the number of emotion category samples correctly predicted by the emotion speech recognition model to the total number of samples. A correctly identified sample is scored as one point, and an incorrectly identified sample is scored as zero. The sum is accumulated and divided by the total number of samples to obtain the final emotion recognition accuracy.
[0145] S67. Based on the final emotion recognition accuracy of the training set and the validation set, determine all training parameters of the final emotion speech recognition model.
[0146] Based on the optimal set of hyperparameters obtained through optimization, the present invention configures the final emotional speech recognition model training process, sequentially extracts emotional dynamic features and generates emotional category probability distributions, ensuring that the final emotional speech recognition model achieves optimal recognition results on both the training set and the validation set. By statistically analyzing the final emotional recognition accuracy, the performance of the final emotional speech recognition model is quantified to avoid overfitting and underfitting problems. This method takes into account both feature extraction quality and classification effect, ultimately determining the training parameters and significantly improving the accuracy, stability, and generalization adaptability of the final emotional speech recognition model.
[0147] In this embodiment, the S7 specifically includes:
[0148] S71. Deploy the final emotional speech recognition model to the local deep neural network inference module of the smart healthcare robot, load the optimal hyperparameter set, and establish a data channel connection with the speech acquisition module;
[0149] S72. During the operation of the smart health care robot, collect user voice signals in real time, perform framing, pre-emphasis, and endpoint detection on the voice data, and extract a time series emotional voice feature vector sequence;
[0150] S73, inputting the emotional speech feature vector sequence into the final emotional speech recognition model, extracting the hidden feature representation of each time frame through a recurrent neural network, and aggregating the hidden features of all time frames according to a weighted rule to generate an overall real-time emotional dynamic feature;
[0151] S74: Input the real-time emotional dynamic features into the fully connected layer and the output layer to generate a probability distribution result of multiple emotion categories corresponding to the current user's voice, and select the emotion category with the largest probability value as the current emotion recognition output label;
[0152] S75. Based on the identified emotion category label, call the preset voice interaction strategy library and behavior control module in the smart healthcare robot, and select the voice feedback method, tone and speed adjustment scheme, and non-verbal behavior response corresponding to the emotion category label in the strategy library.
[0153] This method deploys the final emotional speech recognition model to the local inference module of the intelligent healthcare robot, enabling real-time acquisition and processing of user voice signals. A recurrent neural network extracts and aggregates real-time dynamic emotional features, generates a probability distribution for multiple emotion categories, and selects the current emotion label based on the maximum probability. Combined with the emotion category labels, the voice interaction strategy library and behavior control module are dynamically invoked to flexibly adjust voice feedback and non-verbal behavioral responses, significantly improving the naturalness of human-machine interaction, emotional adaptability, and user experience perception of the intelligent healthcare robot.
[0154] Example 1:
[0155] To verify the feasibility of this invention, it was applied to a large-scale smart healthcare center located in Changping District, Beijing. This center primarily serves the elderly and some patients in rehabilitation, with an average daily active user base of approximately 500. The center was originally equipped with traditional voice-interactive robots capable of performing basic voice Q&A, reminders, and entertainment announcements. However, these robots were unable to effectively identify the user's emotional state, resulting in poor results in nursing intervention, emotional comfort, and companionship communication. These robots exhibited low emotion recognition, stiff responses, and insufficient interaction satisfaction.
[0156] To address this situation, the present invention deploys an intelligent health care robot emotion speech recognition system based on a deep recurrent neural network and the Flamingo optimization algorithm. First, voice acquisition terminals are installed in different functional areas of the health care center. Highly sensitive microphones collect real-time speech data from users during natural communication. The collected speech data undergoes preprocessing through noise removal, endpoint detection, and emotion feature extraction to form an emotion speech feature vector sequence, which is then fed into the emotion speech recognition model in real time.
[0157] To address this situation, the present invention deploys an intelligent health care robot emotion speech recognition system based on a deep recurrent neural network and the Flamingo optimization algorithm. First, voice acquisition terminals are installed in different functional areas of the health care center. Highly sensitive microphones collect real-time speech data from users during natural communication. The collected speech data undergoes preprocessing through noise removal, endpoint detection, and emotion feature extraction to form an emotion speech feature vector sequence, which is then fed into the emotion speech recognition model in real time.
[0158] In actual operation, the health care robot automatically selects different interaction strategies by identifying the emotional state of the user's voice, such as happiness, sadness, anxiety, and indifference. For example, when the system recognizes the presence of sadness in the user's voice, the robot not only uses a softer tone to communicate, but also actively pushes light music, provides psychological comfort prompts, and notifies caregivers to intervene in a timely manner. If the user is identified as emotionally excited or excited, the robot adjusts the speech speed of the voice interaction, recommends exercise rehabilitation programs, or organizes invitations to social events. This entire process does not require the user to actively input emotions; the system is able to perceive and respond to emotional changes through deep learning and intelligent reasoning.
[0159] To evaluate the effectiveness of this invention, we randomly sampled user interaction records from different time periods daily over 30 days of actual operation, compiling a total of 900 sets of emotion recognition and interaction results. By comparing manual annotation with machine output, we calculated metrics such as the system's emotion recognition accuracy, response latency, and user satisfaction. During the experiment, the health and wellness center also conducted a user satisfaction survey, collecting 752 valid responses.
[0160] Table 1 Comparison of the optimization effects of emotional speech recognition of smart health care robots based on deep networks
[0161]
[0162] Table 1 shows that after deploying this invention, the overall emotion recognition accuracy of the health care robot increased from 68.3% in the original system to 88.9%, response latency decreased from an average of 3.2 seconds to 1.4 seconds, and overall user satisfaction increased from 78.6 to 92.1. Furthermore, in the identification and intervention of specific emotion categories, the post-intervention user emotion return rate reached 82.4%, significantly higher than the 59.7% in the control group. These data clearly demonstrate that this invention can significantly enhance the emotional perception and interactive intelligence of smart health care robots.
[0163] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. The emotional speech recognition method of the intelligent health care robot based on deep network is characterized by: The steps include: S1. Collect and pre-process the voice data of users in the health care environment to construct an emotional voice feature vector sequence; S2. Use a recurrent neural network to build an emotional speech recognition model, set the input layer to receive the emotional speech feature vector at each time step, and use long short-term memory units to extract the emotional dynamic features within the emotional speech feature vector sequence; S3. Based on the emotional speech recognition model, set the hyperparameter set to be optimized, use the Flamingo algorithm to encode the hyperparameter set into the position vector of the flamingo individual, and initialize the flamingo population; S4. Update the position vector of each flamingo based on its foraging behavior mechanism, retrain the emotional speech recognition model using the updated hyperparameter set, and calculate the emotion recognition accuracy on the validation set as the fitness value. S5. Select the flamingo individual with the best emotion recognition fitness value as the current optimal hyperparameter set, update the position vector of the flamingo population, and determine the optimal hyperparameter set of the emotional speech recognition model; S6. Based on the optimal hyperparameter set, the emotional speech recognition model is finally trained to determine the final emotional speech recognition model parameters; S7. Deploy the final emotional speech recognition model to the smart healthcare robot, and adjust the smart healthcare robot's voice interaction strategy and behavioral response based on the identified emotion category.
2. The emotional speech recognition method of the intelligent health care robot based on deep network according to claim 1 is characterized in that: The speech data specifically includes original speech waveform, speech content text, speech acoustic features, emotion annotations, speaker information and background noise features.
3. The emotional speech recognition method of the intelligent health care robot based on deep network according to claim 1 is characterized in that: The hyperparameters to be optimized specifically include learning rate, number of hidden layer units, regularization coefficient and dropout rate.
4. The emotional speech recognition method of the intelligent health care robot based on deep network according to claim 1 is characterized in that: Said S1 specifically includes: S11, performing pre-emphasis processing on the collected voice data, subtracting the voice signal value of the previous sampling point from the voice signal value of the current sampling point and multiplying it by a pre-emphasis coefficient, wherein the pre-emphasis coefficient is between 0.9 and 1.0; S12, performing frame processing on the pre-emphasized voice data, setting the number of sampling points of each frame to a fixed value, setting the overlapping part between frames to a fixed number of sampling points, and extracting the voice segments of each frame; S13, performing weighted processing on each frame segment to smooth the boundary of the transition frame segment signal; S14, extracting frequency domain features from each frame segment to generate data representation reflecting signal energy distribution and change trend; S15. Extracting multiple different types of feature indicators based on the frequency domain feature data, including features representing frequency distribution, pitch information, and signal change rate; S16, combining the feature indicators extracted from each frame segment into an emotional speech feature vector in order; S17. Arrange all single-frame emotional speech feature vectors in chronological order to construct an emotional speech feature vector sequence.
5. The emotional speech recognition method of the intelligent health care robot based on deep network according to claim 1 is characterized in that: The S2 specifically includes: S21, determining the input data type of emotion recognition, selecting the emotional speech feature vector sequence as input data, and defining an emotional speech feature vector corresponding to each time step; S22. Based on a recurrent neural network, an emotional speech recognition model is constructed, and an input layer is set. The input layer sequentially receives the features of each time step in the emotional speech feature vector sequence; S23, configuring a long short-term memory unit after the input layer to process the emotional speech feature vector sequence and extract the potential emotion-related dynamic features within the time series; S24. A hidden layer is set after the long short-term memory unit to further abstract and enhance the extracted emotional dynamic features. The hidden layer output features are generated using a weighted summation mechanism: Among them, h agg is the emotional dynamic feature, T is the total number of time steps, t is the sequence number of the time step currently being processed, and h t is the hidden state feature vector output by the long short-term memory unit at time step t, ω t is the weighting coefficient of time step t; S25. Setting a fully connected layer after the hidden layer to map the extracted emotional dynamic feature information to the corresponding emotional category prediction space; S26. Set the output layer to normalize the prediction results output by the fully connected layer to generate the predicted probability distribution of each emotion category.
6. The emotional speech recognition method of the intelligent health care robot based on deep network according to claim 1 is characterized in that: The S3 specifically includes: S31. Setting a set of hyperparameters to be optimized for the emotional speech recognition model, and representing the position vector of each flamingo individual as a set of specific hyperparameter combinations; S32. Randomly initialize the positions of several flamingo individuals in the hyperparameter search space, where each flamingo individual position corresponds to a set of hyperparameter combinations, to form an initial flamingo population position set; S33. Setting an initial velocity vector for each flamingo individual in the flamingo population, where the initial velocity vector is the movement direction and change amplitude of the flamingo individual in the hyperparameter search space, to generate an initial velocity set; S34, binding the position vector of each flamingo individual to the hyperparameter combination of the emotional speech recognition model to generate an initial flamingo population hyperparameter distribution; S35. The initial flamingo population hyperparameter distribution is the initial state of different individuals in the search space. An initial velocity vector is assigned to each flamingo individual. The velocity vector is the direction and amplitude of movement in the hyperparameter space.
7. The emotional speech recognition method of the intelligent health care robot based on deep network according to claim 1 is characterized in that: The S4 specifically includes: S41. Based on the flamingo foraging behavior model, the new position vector of each flamingo is calculated, and local search and global jumping behaviors are simulated. The flamingo moves closer to the optimal solution in the hyperparameter search space. S42, using the updated position vector as a new hyperparameter combination of the emotional speech recognition model, and retraining the new emotional speech recognition model, where the training input is the emotional speech feature vector sequence; S43, using the trained emotional speech recognition model to predict the validation set samples, and outputting the emotional category probability distribution of each speech sample; S44. For each flamingo individual, calculate the emotion recognition accuracy on the validation set as the fitness value: Among them, f(p i ) is the fitness value of the i-th flamingo individual, c is the emotion category number, C is the total number of emotion categories, w c is the importance weight of the c-th category emotion, N c is the number of samples in the validation set whose true labels belong to the cth class, k is the recognition result score of each sample, When the predicted emotion category is consistent with the true emotion category, the value is 1, otherwise the value is 0. is the emotional dynamic feature vector extracted from the emotional speech recognition model of the kth verification sample, is the modulus of the emotional dynamic feature vector, α and β are non-negative weight coefficients, i is the i-th flamingo individual, p i is the position vector of the i-th flamingo individual.
8. The emotional speech recognition method of the intelligent health care robot based on deep network according to claim 1 is characterized in that: The S5 specifically includes: S51. In each round of iteration, all individuals in the flamingo population are compared according to their fitness values, and the individual with the highest current recognition accuracy is selected as the optimal individual of the current iteration round; S52. Update the positions of the remaining flamingo individuals, and calculate based on the position of the current individual, the position of the optimal individual, and the position of the random individual: Among them, f(p i ) is the fitness value of the i-th flamingo individual, is the new position vector of the i-th flamingo individual at the t+1th iteration, p i is the position vector of the i-th flamingo individual, r1 and r2 are random factors between 0 and 1, i is the i-th flamingo individual, is the current position vector of the i-th flamingo individual, is the position of a flamingo individual randomly selected from the current population, P (t) is the set of position vectors of all flamingo individuals at the tth iteration, j is a flamingo individual randomly selected from the entire flamingo population, and argmax is the position vector of the hyperparameter combination that maximizes the fitness value of emotion recognition in the flamingo population; S53. Generate a new hyperparameter combination based on the updated position vector of each flamingo individual, apply the new hyperparameter combination to the trained emotional speech recognition model, calculate the recognition accuracy and the extracted emotional dynamic feature strength for each emotional category in the validation set, accumulate the sum of the weighted accuracy and weighted feature strength of all emotional categories, and use the weighted average as the fitness value of the flamingo individual in the current iteration; S54: Determine whether the preset maximum number of rounds 200 is reached. If not, return to step S51 to continue iteration; S55. When the preset maximum number of rounds of 200 is reached, the individual with the highest fitness value is selected from the current flamingo population, and the corresponding position vector is the optimal hyperparameter combination of the final emotional speech recognition model.
9. The emotional speech recognition method of the intelligent health care robot based on deep network according to claim 1 is characterized in that: The S6 specifically includes: S61. Initializing the training configuration of the final emotional speech recognition model based on the optimal hyperparameter set, including the learning rate, the number of hidden units, the regularization coefficient, and the dropout rate parameters; S62, inputting the emotional speech feature vector sequence in the training set into the final emotional speech recognition model; S63, using long short-term memory units to sequentially process the emotional speech feature vectors, extracting the emotional dynamic features within the time series, performing weighted aggregation on the hidden state features of all time steps, and generating an emotional dynamic feature aggregation vector for each speech sample in the training set; S64, inputting the emotion dynamic feature aggregation vector into the fully connected layer and the output layer, and after weighted summation and bias adjustment, using a normalization function mapping to generate the emotion category probability distribution corresponding to each speech sample in the training set; S65. Repeat S61-S64 using the emotional speech feature vector sequence in the validation set to obtain the emotional category probability distribution corresponding to each sample in the validation set; S66. Calculate the latest emotion recognition accuracy on the training set and the validation set respectively. The latest emotion recognition accuracy is the ratio of the number of emotion category samples correctly predicted by the emotion speech recognition model to the total number of samples. A correctly identified sample is scored as one point, and an incorrectly identified sample is scored as zero. The sum is accumulated and divided by the total number of samples to obtain the final emotion recognition accuracy. S67. Based on the final emotion recognition accuracy of the training set and the validation set, determine all training parameters of the final emotion speech recognition model.
10. The emotional speech recognition method of the intelligent health care robot based on deep network according to claim 1 is characterized in that: The S7 specifically includes: S71. Deploy the final emotional speech recognition model to the local deep neural network inference module of the smart healthcare robot, load the optimal hyperparameter set, and establish a data channel connection with the speech acquisition module; S72. During the operation of the smart health care robot, collect user voice signals in real time, perform framing, pre-emphasis, and endpoint detection on the voice data, and extract a time series emotional voice feature vector sequence; S73, inputting the emotional speech feature vector sequence into the final emotional speech recognition model, extracting the hidden feature representation of each time frame through a recurrent neural network, and aggregating the hidden features of all time frames according to a weighted rule to generate an overall real-time emotional dynamic feature; S74: Input the real-time emotional dynamic features into the fully connected layer and the output layer to generate a probability distribution result of multiple emotion categories corresponding to the current user's voice, and select the emotion category with the largest probability value as the current emotion recognition output label; S75. Based on the identified emotion category label, call the preset voice interaction strategy library and behavior control module in the smart healthcare robot, and select the voice feedback method, tone and speed adjustment scheme, and non-verbal behavior response corresponding to the emotion category label in the strategy library.