Emotion resonance feedback method and device based on facial expression recognition and medium
Through facial expression recognition technology, combining convolutional neural networks and long and short-term memory networks, accurately identifying the user's emotions and matching animation sequence frames, the problem of lack of depth of emotional interaction in the existing technology is solved, and more natural and accurate emotional feedback is achieved.
Patent Information
- Application Number
- CN202510071225.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
Existing facial expression recognition technologies rely mostly on simple emotion classification, ignoring the diversity of emotional intensity and emotional change processes, resulting in robots' performance in emotional interactions being more mechanical and lacking emotional depth.
By obtaining the user's facial image information and preprocessing, the convolutional neural network in the emotion recognition model is used to extract facial features, and the emotional state is analyzed in combination with the long and short-term memory network in the emotion classification model, accurately identify the user's emotions type, and use the emotion retrieval algorithm to match the corresponding animation sequence frames from the animation database and present it to the robot's face in real time.
It achieves more accurate and natural emotional recognition and feedback, enhances the interaction and emotional resonance between the robot and the user, and improves the immersion and interaction quality of the user experience.
Smart Images

Figure CN120014682A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of emotional resonance feedback, and in particular to an emotional resonance feedback method, device and medium based on facial expression recognition. Background Art
[0002] Facial expressions, as a non-verbal communication method that directly reflects emotions, have been widely used in the field of affective computing. Through facial expression recognition technology, robots can capture and analyze the user's facial emotional changes in real time, thereby providing targeted emotional feedback. However, existing facial expression recognition technologies mostly rely on simple emotion classification, while ignoring the diversity of emotion intensity and emotion change processes, which makes the robot's performance in emotional interaction appear to be more mechanical and lack of emotional depth. Summary of the invention
[0003] In order to provide a robot interaction experience with greater emotional depth, the present application provides an emotional resonance feedback method, device and medium based on facial expression recognition.
[0004] The above-mentioned invention objective of the present application is achieved through the following technical solutions: An emotional resonance feedback method based on facial expression recognition, the emotional resonance feedback method based on facial expression recognition comprising: Acquire facial image information of a user, and perform image preprocessing on the facial image information to obtain preprocessed facial image information; Inputting the preprocessed facial image information into an emotion recognition model, wherein the emotion recognition model extracts facial features from the preprocessed facial image information to obtain facial feature information; The facial feature information is input into an emotion classification model, and the emotion classification model obtains an emotion type result by performing an emotion state analysis on the facial feature information; According to the emotion type result, an animation sequence frame matching the emotion feature result is retrieved from an animation database through an emotion retrieval algorithm to obtain a selected animation sequence frame; The selected animation sequence frames are presented to the robot face in real time.
[0005] By adopting the above technical solution, the user's facial image information is obtained and preprocessed to ensure high-quality input for subsequent emotion recognition. Then, the convolutional neural network in the emotion recognition model is used to extract facial features and accurately capture the user's emotional state. Next, the emotional state of the facial feature information is analyzed through the emotion classification model to accurately identify the user's emotion type, and the emotion retrieval algorithm is used to match the corresponding animation sequence frames from the animation database. Finally, the selected animation sequence frames are presented to the robot's face in real time, intuitively showing emotional feedback and enhancing the sense of interaction with the user.
[0006] In a preferred example, the present application may be further configured as follows: the pre-processed facial image information is input into an emotion recognition model, and the emotion recognition model extracts facial features from the pre-processed facial image information to obtain facial feature information, including: The preprocessed facial image information is analyzed by a convolutional neural network algorithm in the emotion recognition model to obtain the facial feature information.
[0007] By adopting the above technical solution and using the convolutional neural network algorithm, facial features can be efficiently identified from complex facial images, and the user's emotional state can be automatically analyzed and extracted. The accuracy and processing speed of emotion recognition are improved, allowing the robot to perceive and respond to the user's emotions in real time and accurately. In addition, the advantages of convolutional neural networks in image recognition make facial image processing more stable, and can maintain good recognition effects even in complex environments such as lighting changes and angle deviations, thereby improving the naturalness and intelligence of the interactive experience.
[0008] In a preferred example, the present application may be further configured as follows: the facial feature information is input into an emotion classification model, and the emotion classification model performs an emotion state analysis on the facial feature information to obtain an emotion type result, including: The emotional state of the facial feature information is analyzed by the long short-term memory network algorithm in the emotion classification and intensity prediction model to identify the emotion type of the user and obtain the emotion type result.
[0009] By adopting the above technical solution and using the long short-term memory network (LSTM) to analyze the emotional state of facial feature information, the accuracy and real-time performance of emotion recognition can be significantly improved. LSTM can process and learn long-term dependencies, especially when facing time series data of user emotional changes, it can better capture the dynamic characteristics of emotional fluctuations. This not only improves the accuracy of emotion type recognition, but also enables more accurate and timely responses to changes in different emotions.
[0010] In a preferred example, the present application can be further configured as follows: according to the emotion type result, through the emotion retrieval algorithm, the animation sequence frame matching the emotion feature result is retrieved from the animation database to obtain the selected animation sequence frame, including: Mapping the emotion type result to a preset emotion label in the animation database to obtain an emotion label corresponding to the emotion type result, wherein the emotion label matches the emotion state corresponding to each animation sequence frame; Based on the emotion tag, similarity calculation is performed with the emotion tags of all animation sequence frames in the animation database by using an emotion retrieval algorithm to obtain a similarity score; The similarity scores are sorted from high to low to obtain a sorting result, and according to the sorting result, the animation sequence frame with the highest similarity score is selected as the selected animation sequence frame.
[0011] By adopting the above technical solution, the emotion label is matched with the emotional state of each animation sequence frame in the animation database, making the animation selection more targeted and accurate. Secondly, by calculating the similarity score and sorting through the emotion retrieval algorithm, it can be ensured that the selected animation sequence frame best matches the user's current emotional state, thereby improving the naturalness and authenticity of the robot's emotional expression. At the same time, the robot can dynamically adjust the expression method and present the most suitable animation in real time according to the user's emotional changes, which enhances the interactive immersion and emotional resonance effect and improves the user experience.
[0012] In a preferred example, the present application may be further configured as follows: the similarity calculation is performed based on the emotion tag and the emotion tags of all animation sequence frames in the animation database by using an emotion retrieval algorithm to obtain a similarity score, including: The sentiment retrieval algorithm obtains the similarity score through the following cosine similarity formula: Among them, the S i is the similarity score of the i-th animation sequence frame, the w k is the weight coefficient of the kth emotion dimension, and L e (k) is the value of the input emotion label on the kth emotion dimension, the L ani (k) is the value of the emotion label of the kth animation sequence frame, the ‖L e (k) ‖ is the vector modulus length of the emotion label in the kth dimension, and the ‖L ani (k) ‖ is the vector modulus of the emotion label of the animation sequence frame in the kth dimension.
[0013] By adopting the above technical solution, it is possible to accurately match the emotional state of the emotional label with the animation sequence frame, thereby achieving more accurate and personalized emotional feedback. The use of the cosine similarity formula ensures the quantitative comparison of the emotional state of the emotional label and the animation sequence frame, so that the animation selection is based on the similarity of emotional features rather than a simple matching rule. By calculating the weight coefficient of the emotional dimension, the adaptability of the model is further improved, allowing the robot to dynamically select the most appropriate animation sequence frame according to the differences in different emotional dimensions. This similarity-based retrieval method enables the robot to respond to user emotions more flexibly and accurately, thereby achieving a more natural and emotional interaction, and improving the immersion and interaction quality of the user experience.
[0014] In a preferred example, the present application may be further configured as follows: presenting the selected animation sequence frames to the robot face in real time, further comprising: The selected animation sequence frames are presented to the robot face in real time, and voice feedback corresponding to the selected animation sequence frames is generated through speech synthesis technology, and the voice feedback is played synchronously with the selected animation sequence frames; based on the voice feedback and the selected animation sequence frames, the robot is controlled by an action control algorithm to perform an emotional feedback action that matches the emotion type result.
[0015] By adopting the above technical solutions, a more natural and emotional interaction between the robot and the user can be achieved. By synchronously playing the animation sequence frames and voice feedback, the robot can express emotions more vividly and realistically, improving the accuracy of emotional expression and the immersiveness of interaction. At the same time, the addition of motion control algorithms allows the robot to convey emotions not only through visual and auditory feedback, but also through body movements to enhance the layering and realism of emotional expression. This multimodal emotional feedback method not only makes the robot more expressive, but also can make timely responses according to the user's emotional state, further enhancing the user's emotional experience and interaction quality.
[0016] In a preferred example, the present application can be further configured as follows: the emotional resonance feedback method based on facial expression recognition further includes: Acquire user interaction information, and construct an emotion supplement model based on the user interaction information, wherein the emotion supplement model analyzes the preprocessed facial image information to obtain new emotion features; According to the new emotion feature, the emotion recognition model is optimized and adjusted to obtain an optimized emotion recognition model; based on the output emotion feature result of the optimized emotion recognition model, the output emotion feature result is input into the emotion classification model for verification and update, and the adaptability of the new emotion feature to the emotion classification model is analyzed; if the new emotion feature matches the classification rule of the emotion classification model abnormally, the classification rule of the emotion classification model is adjusted using the emotion feature result to generate an optimized emotion classification model; According to the emotion type results output by the optimized emotion classification model, the matching situation between the emotion type results and the animation sequence frames in the animation database is analyzed, and based on the matching situation, the animation database is updated to add new animation sequence frames corresponding to the new emotion features.
[0017] By adopting the above technical solutions, real-time adaptive optimization of robot emotion recognition and feedback can be achieved. First, the emotion supplement model continuously improves emotion recognition and classification through user interaction information, ensuring that the robot can accurately capture and understand the user's emotional changes. Secondly, with the continuous update of emotional features, the optimized emotion recognition and classification model can improve the robot's emotional response accuracy, making it more natural and realistic when interacting with users. Finally, based on the emotion type results, the animation sequence frames and animation database are updated, so that the robot can provide richer feedback content that is more in line with the user's emotions, thereby improving the user experience and emotional connection.
[0018] The second object of the invention is achieved by the following technical solutions: An emotional resonance feedback device based on facial expression recognition, the emotional resonance feedback device based on facial expression recognition comprising: A facial image acquisition and preprocessing module is used to obtain facial image information of a user, perform image preprocessing on the facial image information, and obtain preprocessed facial image information; An emotion recognition and feature extraction module, used to input the preprocessed facial image information into an emotion recognition model, and the emotion recognition model obtains facial feature information by extracting facial features from the preprocessed facial image information; an emotion classification and analysis module, used to input the facial feature information into an emotion classification model, and the emotion classification model obtains an emotion type result by performing an emotion state analysis on the facial feature information; An animation retrieval and matching module, used to retrieve animation sequence frames matching the emotion feature results from an animation database through an emotion retrieval algorithm according to the emotion type results, and obtain the selected animation sequence frames; The animation presentation module is used to present the selected animation sequence frames to the robot face in real time.
[0019] By adopting the above technical solution, the user's facial image information is obtained and preprocessed to ensure high-quality input for subsequent emotion recognition. Then, the convolutional neural network in the emotion recognition model is used to extract facial features and accurately capture the user's emotional state. Next, the emotional state of the facial feature information is analyzed through the emotion classification model to accurately identify the user's emotion type, and the emotion retrieval algorithm is used to match the corresponding animation sequence frames from the animation database. Finally, the selected animation sequence frames are presented to the robot's face in real time, intuitively showing emotional feedback and enhancing the sense of interaction with the user.
[0020] The third objective of the present application is achieved through the following technical solutions: A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned emotional resonance feedback method based on facial expression recognition when executing the computer program.
[0021] The fourth objective of the present application is achieved through the following technical solutions: A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned emotional resonance feedback method based on facial expression recognition.
[0022] In summary, the present application includes at least one of the following beneficial technical effects: 1. Obtain the user's facial image information and perform preprocessing to ensure high-quality input for subsequent emotion recognition. Then, use the convolutional neural network in the emotion recognition model to extract facial features and accurately capture the user's emotional state. Next, use the emotion classification model to analyze the emotional state of the facial feature information, accurately identify the user's emotion type, and use the emotion retrieval algorithm to match the corresponding animation sequence frames from the animation database. Finally, the selected animation sequence frames are presented to the robot's face in real time, intuitively showing emotional feedback and enhancing the sense of interaction with the user; 2. It can accurately match the emotional state of the emotional labels with the animation sequence frames, thereby achieving more accurate and personalized emotional feedback. The use of the cosine similarity formula ensures the quantitative comparison of the emotional state of the emotional labels and the animation sequence frames, so that the animation selection is based on the similarity of emotional features rather than simple matching rules. By calculating the weight coefficients of the emotional dimensions, the adaptability of the model is further improved, allowing the robot to dynamically select the most appropriate animation sequence frames based on the differences in different emotional dimensions. This similarity-based retrieval method enables the robot to respond to user emotions more flexibly and accurately, thereby achieving more natural and emotional interactions, and improving the immersion and interaction quality of the user experience; 3. It can realize real-time adaptive optimization of robot emotion recognition and feedback. First, the emotion supplement model continuously improves emotion recognition and classification through user interaction information, ensuring that the robot can accurately capture and understand the user's emotional changes. Secondly, with the continuous update of emotion characteristics, the optimized emotion recognition and classification model can improve the robot's emotional response accuracy, making it more natural and real when interacting with users. Finally, based on the emotion type results, the animation sequence frames and animation database are updated, so that the robot can provide richer feedback content that is more in line with the user's emotions, thereby improving the user experience and emotional connection. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a flow chart of an emotional resonance feedback method based on facial expression recognition in one embodiment of the present application; Figure 2 It is a flowchart for implementing step S20 in the emotional resonance feedback method based on facial expression recognition in an embodiment of the present application; Figure 3 It is a flowchart for implementing step S30 in the emotional resonance feedback method based on facial expression recognition in one embodiment of the present application; Figure 4 It is a flowchart for implementing step S40 in the emotional resonance feedback method based on facial expression recognition in an embodiment of the present application; Figure 5 It is a flowchart for implementing step S402 in the emotional resonance feedback method based on facial expression recognition in one embodiment of the present application; Figure 6 It is a flowchart for implementing step S50 in the emotional resonance feedback method based on facial expression recognition in one embodiment of the present application; Figure 7 This is a flowchart for implementing the step S50 in the emotional resonance feedback method based on facial expression recognition in one embodiment of the present application; Figure 8 It is a principle block diagram of an emotional resonance feedback device based on facial expression recognition in one embodiment of the present application; Fig. 9 It is a schematic diagram of a device in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The present application is further described in detail below in conjunction with the accompanying drawings.
[0025] In one embodiment, if Figure 1 As shown, the present application discloses an emotional resonance feedback method based on facial expression recognition, which specifically includes the following steps: S10: Acquire the user's facial image information, perform image preprocessing on the facial image information, and obtain preprocessed facial image information.
[0026] Specifically, the user's facial image information is obtained by using a high-definition camera or a depth camera. The image information contains various details of the user's face, including facial expressions, skin color, lighting, eyes, mouth and other areas. During preprocessing, the acquired image is first denoised, and a Gaussian filter algorithm or a median filter algorithm is used to remove noise in the image to ensure the clarity of the image. Subsequently, the image size and lighting are adjusted so that the contrast and brightness of the facial expressions meet the model input requirements. Common methods such as histogram equalization ensure that useful facial feature information can be extracted under different lighting conditions. Next, facial alignment technology is applied, and facial key point detection in the Dlib library is used to locate important feature points such as the user's eyes, nose, and mouth, and geometric transformation is performed to unify the facial image to a standard alignment angle to ensure that the facial image input into the emotion recognition model has a unified standard.
[0027] Furthermore, the Dlib library is used to extract feature points from the user's facial image and then analyze the facial expression.
[0028] S20: Inputting the preprocessed facial image information into the emotion recognition model, the emotion recognition model extracts facial features from the preprocessed facial image information to obtain facial feature information.
[0029] Specifically, the preprocessed facial image information is input into the emotion recognition model, which uses a deep convolutional neural network (CNN) algorithm to extract features from facial images. First, the emotion recognition model extracts local features from the input image through multiple convolutional layers. For example, by using convolution kernels of different sizes, the model can capture various facial detail features from the image, such as the position, shape, and dynamic changes of the eyes, mouth, and eyebrows. After the convolution layer, a pooling layer (such as maximum pooling) is usually used to reduce the spatial dimension and computational complexity of the image while retaining the most significant facial features. After a series of convolution and pooling processes, the key structural features in the facial image are extracted to generate a set of feature maps containing key facial areas. These feature maps not only include facial contours, the shapes of eyes and mouth, but also capture subtle changes in facial muscle movements, such as smiling, frowning, and eyes opening or closing.
[0030] S30: Input the facial feature information into the emotion classification model. The emotion classification model analyzes the emotional state of the facial feature information to obtain an emotion type result.
[0031] Specifically, facial feature information is passed to the deep neural network in the emotion classification model. The model will gradually process the input feature information through multiple hidden layers, and each layer will learn different levels of features, such as basic facial muscle movements (such as eyebrow raising, mouth corners up, etc.), and the combination of these movements. For each expression, the network will capture the spatial relationship and change pattern between feature points, especially the facial landmarks related to emotions, such as the shape and position changes of eyes and mouth. The model will classify emotions based on the temporal characteristics of facial expressions. Taking the long short-term memory network (LSTM) as an example, LSTM can learn and remember the relationship between the previous and next time steps through memory units, which can effectively process the dynamic characteristics of facial expressions that change over time. LSTM will capture the trend and evolution of emotional expression based on the facial features of each frame and the information of the previous and next moments. For example, when the facial expression gradually becomes nervous and angry, the model will identify this emotional change trend through LSTM, and then judge the emotion type as anger. Finally, the emotion classification model outputs the emotion type result based on the facial feature analysis results, which usually includes several common emotion categories, such as happiness, anger, sadness, surprise, fear, etc.
[0032] S40: According to the emotion type result, an emotion retrieval algorithm is used to retrieve animation sequence frames matching the emotion feature result from an animation database to obtain selected animation sequence frames.
[0033] Specifically, the emotion type results are mapped to the preset emotion label space, for example, joy corresponds to the joy label, and anger corresponds to the anger label. Then, through the emotion retrieval algorithm, the current emotion label is compared with the emotion labels of all animation sequence frames stored in the database, and the similarity between them is calculated. Several animation sequence frames with the highest similarity to the current emotion label are selected as candidate animation frames. Finally, the most suitable animation sequence frame is selected based on the similarity sorting.
[0034] S50: Presenting the selected animation sequence frames to the robot face in real time.
[0035] Specifically, the selected animation sequence frames are displayed in real time on the robot's display screen through a rendering engine. Usually, the display screen is an OLED or LCD screen that can display high-definition animation content. The animation frames of the robot's face are presented at a fast refresh rate (for example, 60 frames per second) to ensure the smoothness and naturalness of the animation. The animation frame content includes changes in facial expressions, such as smiling, frowning, blinking and other actions.
[0036] In one embodiment, if Figure 2 As shown, in step S20, the pre-processed facial image information is input into the emotion recognition model, and the emotion recognition model extracts facial features from the pre-processed facial image information to obtain facial feature information, including: S201: Analyze the preprocessed facial image information through the convolutional neural network algorithm in the emotion recognition model to obtain facial feature information.
[0037] Specifically, the emotion recognition model uses a convolutional neural network (CNN) to process the input pre-processed facial image. At this stage, CNN will extract low-level features of the facial image, such as edges, textures, contours, etc. through a series of convolutional layers; then, the model will downsample the low-level features through a pooling operation. Convolutional neural networks can adaptively learn facial features under different expressions, such as the raising of eyebrows and the raising of the corners of the mouth. For example, when the rising amplitude of eyebrows and the smile of the corners of the mouth are recognized, the model will extract this information as facial features, reflecting the emotional state of joy.
[0038] In one embodiment, if Figure 3 As shown, in step S30, the facial feature information is input into the emotion classification model, and the emotion classification model obtains the emotion type result by analyzing the emotional state of the facial feature information, including: S301: Analyze the emotional state of facial feature information through the long short-term memory network algorithm in the emotion classification and intensity prediction model, identify the user's emotion type, and obtain the emotion type result.
[0039] Specifically, the emotion classification and intensity prediction model uses the long short-term memory network (LSTM) algorithm to analyze the input facial feature information. The LSTM algorithm first receives the facial feature information obtained from the emotion recognition model, including expression features, facial action units (such as raised corners of the mouth, raised eyebrows, etc.) and changes in the time dimension. The LSTM algorithm effectively captures the dynamic changes of facial expressions in the time series by transmitting feature information layer by layer, and then infers the emotional state at the current moment. For example, when the input sequence contains the feature of a facial smile accompanied by changes in slightly closed eyes, the LSTM model will analyze the emotional trends in the historical data and predict that the user's current emotion is joy or happiness. Finally, based on the output results of the LSTM, the emotional state is classified into different emotional types, such as pleasure, anger, sadness, etc., and a corresponding output result is generated for each emotional type, indicating the emotional state of the current user.
[0040] In one embodiment, if Figure 4 As shown, in step S40, based on the emotion type result, the emotion retrieval algorithm is used to retrieve the animation sequence frames matching the emotion feature result from the animation database to obtain the selected animation sequence frames, including: S401: Mapping the emotion type result to a preset emotion label in the animation database to obtain an emotion label corresponding to the emotion type result, wherein the emotion label matches the emotion state corresponding to each animation sequence frame.
[0041] Specifically, it is first necessary to convert the emotion type results obtained from the emotion classification model into a format that matches the emotion tags in the animation database. To this end, the emotion type results output by the emotion classification model will be mapped to a preset set of emotion tags, which usually include happiness, anger, sadness, surprise, fear, etc. Each tag corresponds to a specific emotional state, and then indicates the animation sequence frames associated with this emotion. Through this mapping, a direct association can be established between the emotion type and the animation sequence frames stored in the database. For example, if the emotion type result is anger, the result will be mapped to the emotion tag related to anger in the database, and then all animation sequence frames related to anger will be filtered out.
[0042] S402: Based on the emotion tag, similarity calculation is performed with the emotion tags of all animation sequence frames in the animation database by using an emotion retrieval algorithm to obtain a similarity score.
[0043] Specifically, a similarity algorithm is used to match the emotion labels and calculate the similarity between the emotion labels. The similarity algorithm obtains the similarity score of each animation sequence frame by comparing the similarity between the label of the emotion type result and the label of the animation sequence frame. For example, if the emotion type is anger and the emotion label of an animation sequence frame is anger, the similarity score is 1; if the emotion label of the animation sequence frame is happy, the similarity score is lower, which may be 0.3 or lower.
[0044] S403: Sort the similarity scores from high to low to obtain a sorting result, and according to the sorting result, select the animation sequence frame with the highest similarity score as the selected animation sequence frame.
[0045] Specifically, the similarity scores of all animation sequence frames will be sorted from high to low to obtain a list sorted by matching degree. According to the sorting results, the animation sequence frame with the highest similarity score is selected as the final selected animation sequence frame. This process ensures that the selected animation sequence frame is the most accurate expression of the user's current emotions, thereby providing an animation performance that is more in line with emotional feedback. For example, after calculating the similarity scores of all animation sequence frames, assuming that animation A has a score of 0.95, animation B has a score of 0.85, and animation C has a score of 0.78, after sorting, animation A has the highest score and will be selected as the selected animation sequence frame.
[0046] In one embodiment, if Figure 5 As shown, in step S402, based on the emotion tag, the emotion retrieval algorithm is used to calculate the similarity with the emotion tags of all animation sequence frames in the animation database to obtain a similarity score, including: S4021: The sentiment retrieval algorithm obtains the similarity score using the following cosine similarity formula: Among them, S i is the similarity score of the i-th animation sequence frame, w k is the weight coefficient of the kth emotion dimension, L e (k) is the value of the input emotion label on the kth emotion dimension, L ani (k) is the value of the emotion label of the kth animation sequence frame, ‖L e (k) ‖ is the vector modulus of the emotion label in the kth dimension, ‖L ani (k) ‖ is the vector modulus of the emotion label of the animation sequence frame in the kth dimension.
[0047] Specifically, the dot product of the emotion label and the animation sequence frame emotion label in each dimension k is calculated, L e (k)and L ani (k) are the values of the input emotion label and animation sequence frame on the kth emotion dimension. Then, the vector modulus length of each emotion label and animation sequence frame emotion label on this dimension is calculated, which are ‖L e (k) ‖ and ‖L ani (k) ‖, reflects the size or strength of each label in that dimension. Next, the similarity on each sentiment dimension is calculated using the cosine similarity formula The calculation result indicates the similarity between the emotion label and the animation sequence frame in this dimension. The closer the value is to 1, the more similar they are. For each dimension k, the similarity value is multiplied by the weight coefficient w. k , to reflect the impact of different emotional dimensions on the overall emotional matching. Different emotional dimensions may have different importance to the overall emotional expression. For example, some emotional dimensions (such as anger) may be more important to the user's emotional feedback, so they will be given a higher weight. Finally, the weighted similarity values of all emotional dimensions will be accumulated to obtain the final similarity score S for each animation sequence frame. i .
[0048] In one embodiment, if Figure 6 As shown, in step S50, the selected animation sequence frames are presented to the robot face in real time, and the process also includes: S501: Presenting the selected animation sequence frame to the robot face in real time, and generating voice feedback corresponding to the selected animation sequence frame through speech synthesis technology, and the voice feedback is played synchronously with the selected animation sequence frame.
[0049] Specifically, the selected animation sequence frames will be loaded and presented to the robot face in real time. This process usually relies on an image display module or an LCD screen to ensure that the selected animation can be presented clearly and smoothly. Voice feedback synchronized with this animation sequence frame will be generated through speech synthesis technology. Speech synthesis technology can generate appropriate voice feedback according to the emotional type of the animation content. For example, if the animation expresses anger, the voice feedback will be a powerful or angry voice. If the animation expresses happiness, the voice feedback will be a gentle and pleasant voice. Speech synthesis technology can generate natural speech through a deep learning model to ensure that the speech and animation are emotionally consistent. The generated voice feedback will be synchronized with the playback of the animation, that is, during the playback of the animation, the voice feedback will always be synchronized with it to enhance the effect of emotional resonance, so that the user can perceive the robot's emotions through both vision and hearing. For example, if the selected animation is an animation sequence frame of anger, the robot's face will show an angry facial expression, and the voice feedback may be a high-pitched and hurried voice. The voice playback is synchronized with the facial animation to enhance the realism of the robot's emotions.
[0050] S502: According to the voice feedback and the selected animation sequence frames, the robot is controlled by a motion control algorithm to execute an emotional feedback action that matches the emotion type result.
[0051] Specifically, the robot uses the action control algorithm to perform the corresponding emotional feedback action according to the selected animation sequence frames and voice feedback. The role of the action control algorithm is to convert the visual content and voice content of the animation into the specific behavior of the robot. For example, the robot may wave its arms or make other body movements while showing an angry expression to strengthen the emotional feedback of anger. The algorithm will select the appropriate action mode according to the emotion type and the content of the voice feedback, and control the robot's moving parts (such as joints, head, arms, etc.) to perform coordinated actions. For example, if the emotion type result is anger, the action control algorithm may instruct the robot to make angry movements such as clenching fists, shaking arms or turning its head quickly. If the emotion type is happy, the robot may make movements such as waving and nodding to express happiness.
[0052] In one embodiment, if Figure 7 As shown, after step S50, the emotional resonance feedback based on facial expression recognition further includes: S60: Obtain user interaction information, build an emotion supplement model based on the user interaction information, and use the emotion supplement model to analyze the preprocessed facial image information to obtain new emotion features.
[0053] Specifically, the emotion supplement model extracts the emotion patterns and changing trends of user feedback by analyzing the user's interaction information and emotional feedback. The process first identifies the emotions expressed by the user during a specific interaction process, such as through various interaction methods such as voice, facial expressions or body movements, to capture the user's emotional response. Then, based on these interaction data and feedback, the model uses deep learning algorithms to discover possible emotional features and compares these new emotional features with existing emotional categories. Through these comparisons, the model can supplement new emotional features.
[0054] In addition, the emotion supplement model can also dynamically adjust the parameters of the emotion recognition model and the emotion classification model according to the new emotion data. S70: Optimizing and adjusting the emotion recognition model according to the new emotion features to obtain an optimized emotion recognition model.
[0055] Specifically, the new emotion features provided by the emotion supplement model will be used as input data and fed back to the emotion recognition model. The emotion recognition model identifies possible deviations or deficiencies by comparing these analysis results with the existing model output. Based on these feedbacks, the emotion recognition model will perform self-optimization adjustments. This process includes adjusting algorithm parameters, updating neural network weights, introducing new emotion labels or emotion features, adjusting the data distribution of the training set, etc. For example, if the analysis results show that the emotion recognition model has a low recognition accuracy in a specific emotion category (such as anxiety), the learning of this emotion category will be strengthened based on these feedbacks to optimize the model's recognition ability.
[0056] S80: Based on the output emotion feature results of the optimized emotion recognition model, the output emotion feature results are input into the emotion classification model for verification and updating, and the adaptability of the new emotion features to the emotion classification model is analyzed. If the new emotion features do not match the classification rules of the emotion classification model abnormally, the emotion feature results are used to adjust the classification rules of the emotion classification model to generate an optimized emotion classification model.
[0057] Specifically, the emotion feature results output by the emotion recognition model after optimization will be passed to the emotion classification model for verification. The emotion classification model verifies the adaptability of the emotion recognition model in practical applications by analyzing these new emotion features and checks whether these features conform to the existing classification rules. For example, if the emotion recognition model finds that the new anxiety feature does not conform to the existing rules, it will identify that the feature has a low degree of match with the classification rules. In order to improve the classification accuracy, the emotion classification model will adjust the classification rules based on the feedback of this match degree, further refine the classification boundaries, and enable the model to correctly handle new emotion features. The adjusted classification rules will form an optimized emotion classification model, which can maintain high accuracy and stability when facing new emotion features.
[0058] S90: According to the emotion type results output by the optimized emotion classification model, the matching between the emotion type results and the animation sequence frames in the animation database is analyzed, and based on the matching, the animation database is updated to add new animation sequence frames corresponding to the new emotion features.
[0059] Specifically, the emotion type results output by the optimized emotion classification model will be used to match the animation sequence frames in the animation database. First, based on the emotion type results output by the emotion classification model, the animation database will be searched for existing animation sequence frames that match the emotion type. If the existing animation sequence frames cannot perfectly match the emotion type, this matching problem will be identified, and new animation sequence frames will be generated based on the new emotion feature requirements. The new animation sequence frames will be added as updates to the database to ensure that the database can cover all emotion types and corresponding animation performances. For example, if the newly added emotion type is anxiety, and there is a lack of animation matching the anxiety emotion in the database, a new animation sequence frame will be generated based on this emotion feature, and the animation can show anxious facial expressions and body movements.
[0060] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0061] In one embodiment, an emotional resonance feedback device based on facial expression recognition is provided, and the emotional resonance feedback device based on facial expression recognition corresponds one-to-one to the emotional resonance feedback method based on facial expression recognition in the above embodiment. Figure 8 As shown, the emotional resonance feedback device based on facial expression recognition includes a facial image acquisition and preprocessing module, an emotion recognition and feature extraction module, an emotion classification and analysis module, an animation retrieval and matching module, and an animation presentation module. The functional modules are described in detail as follows: The facial image acquisition and preprocessing module is used to obtain the facial image information of the user, perform image preprocessing on the facial image information, and obtain the preprocessed facial image information; The emotion recognition and feature extraction module is used to input the preprocessed facial image information into the emotion recognition model, and the emotion recognition model obtains facial feature information by extracting facial features from the preprocessed facial image information; The emotion classification and analysis module is used to input the facial feature information into the emotion classification model. The emotion classification model analyzes the emotional state of the facial feature information to obtain the emotion type result; The animation retrieval and matching module is used to retrieve the animation sequence frames matching the emotion feature results from the animation database according to the emotion type results through the emotion retrieval algorithm to obtain the selected animation sequence frames; The animation presentation module is used to present the selected animation sequence frames to the robot face in real time.
[0062] Optionally, the emotion recognition and feature extraction module includes: The facial feature extraction submodule is used to analyze the preprocessed facial image information through the convolutional neural network algorithm in the emotion recognition model to obtain facial feature information.
[0063] Optionally, the animation retrieval and matching module includes: The emotional state analysis submodule is used to analyze the emotional state of facial feature information through the long short-term memory network algorithm in the emotion classification and intensity prediction model, identify the user's emotional type, and obtain the emotional type result.
[0064] Optionally, the animation retrieval and matching module includes: The emotion label mapping submodule is used to map the emotion type results to the preset emotion labels in the animation database to obtain the emotion labels corresponding to the emotion type results, and the emotion labels are matched with the emotion states corresponding to each animation sequence frame; The emotion matching retrieval submodule is used to calculate the similarity between the emotion tags and the emotion tags of all animation sequence frames in the animation database through the emotion retrieval algorithm to obtain a similarity score; The animation sorting and selection submodule is used to sort the similarity scores from high to low to obtain a sorting result. According to the sorting result, the animation sequence frame with the highest similarity score is selected as the selected animation sequence frame.
[0065] Optionally, the sentiment matching retrieval submodule includes: The similarity scoring submodule is used in the sentiment retrieval algorithm to obtain the similarity score through the following cosine similarity formula: Among them, S i is the similarity score of the i-th animation sequence frame, w k is the weight coefficient of the kth emotion dimension, L e (k) is the value of the input emotion label on the kth emotion dimension, L ani (k) is the value of the emotion label of the kth animation sequence frame, ‖L e (k) ‖ is the vector modulus of the emotion label in the kth dimension, ‖L ani (k) ‖ is the vector modulus of the emotion label of the animation sequence frame in the kth dimension.
[0066] Optionally, the animation presentation module further includes: The animation presentation and voice synchronization submodule is used to present the selected animation sequence frames to the robot face in real time, and generate voice feedback corresponding to the selected animation sequence frames through speech synthesis technology, and the voice feedback is played synchronously with the selected animation sequence frames; The emotion feedback action control submodule is used to control the robot to perform emotion feedback actions that match the emotion type results through the action control algorithm based on the voice feedback and the selected animation sequence frames.
[0067] Optionally, the animation rendering module also includes: The emotion supplement model building module is used to obtain user interaction information and build an emotion supplement model based on the user interaction information. The emotion supplement model analyzes the preprocessed facial image information to obtain new emotion features. The emotion recognition model optimization module is used to optimize and adjust the emotion recognition model according to the new emotion characteristics to obtain an optimized emotion recognition model; The emotion classification model update module is used to input the output emotion feature results of the optimized emotion recognition model into the emotion classification model for verification and update, analyze the adaptability of the new emotion features to the emotion classification model, and if the new emotion features do not match the classification rules of the emotion classification model, adjust the classification rules of the emotion classification model using the emotion feature results to generate an optimized emotion classification model; The animation database update module is used to analyze the matching between the emotion type results output by the optimized emotion classification model and the animation sequence frames in the animation database, and based on the matching, update the animation database and add new animation sequence frames corresponding to the new emotion features.
[0068] For the specific definition of the emotional resonance feedback device based on facial expression recognition, please refer to the definition of the emotional resonance feedback method based on facial expression recognition above, which will not be repeated here. Each module in the above-mentioned emotional resonance feedback device based on facial expression recognition can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0069] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig. 9As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used for an animation database. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an emotional resonance feedback method based on facial expression recognition is implemented.
[0070] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: Acquire the user's facial image information, perform image preprocessing on the facial image information, and obtain preprocessed facial image information; input the preprocessed facial image information into the emotion recognition model, and the emotion recognition model obtains facial feature information by extracting facial features from the preprocessed facial image information; The facial feature information is input into the emotion classification model, and the emotion classification model obtains the emotion type result by analyzing the emotional state of the facial feature information; According to the emotion type result, the animation sequence frames matching the emotion feature result are retrieved from the animation database through the emotion retrieval algorithm to obtain the selected animation sequence frames; The selected animation sequence frames are presented to the robot face in real time.
[0071] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented: Acquire the user's facial image information, perform image preprocessing on the facial image information, and obtain preprocessed facial image information; input the preprocessed facial image information into the emotion recognition model, and the emotion recognition model obtains facial feature information by extracting facial features from the preprocessed facial image information; The facial feature information is input into the emotion classification model, and the emotion classification model obtains the emotion type result by analyzing the emotional state of the facial feature information; According to the emotion type result, the animation sequence frames matching the emotion feature result are retrieved from the animation database through the emotion retrieval algorithm to obtain the selected animation sequence frames; The selected animation sequence frames are presented to the robot face in real time.
[0072] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0073] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0074] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An emotional resonance feedback method based on facial expression recognition, characterized in that: The emotional resonance feedback method based on facial expression recognition includes: Acquire facial image information of a user, and perform image preprocessing on the facial image information to obtain preprocessed facial image information; Inputting the preprocessed facial image information into an emotion recognition model, wherein the emotion recognition model extracts facial features from the preprocessed facial image information to obtain facial feature information; The facial feature information is input into an emotion classification model, and the emotion classification model obtains an emotion type result by performing an emotion state analysis on the facial feature information; According to the emotion type result, an animation sequence frame matching the emotion feature result is retrieved from an animation database through an emotion retrieval algorithm to obtain a selected animation sequence frame; The selected animation sequence frames are presented to the robot face in real time.
2. The emotional resonance feedback method based on facial expression recognition according to claim 1, characterized in that: The step of inputting the preprocessed facial image information into an emotion recognition model, wherein the emotion recognition model extracts facial features from the preprocessed facial image information to obtain facial feature information, comprises: The preprocessed facial image information is analyzed by a convolutional neural network algorithm in the emotion recognition model to obtain the facial feature information.
3. The emotional resonance feedback method based on facial expression recognition according to claim 1, characterized in that: The step of inputting the facial feature information into an emotion classification model, wherein the emotion classification model performs an emotion state analysis on the facial feature information to obtain an emotion type result, includes: The emotional state of the facial feature information is analyzed by the long short-term memory network algorithm in the emotion classification and intensity prediction model to identify the emotion type of the user and obtain the emotion type result.
4. The emotional resonance feedback method based on facial expression recognition according to claim 1, characterized in that: The step of retrieving, according to the emotion type result, an animation sequence frame matching the emotion feature result from an animation database through an emotion retrieval algorithm to obtain the selected animation sequence frame comprises: Mapping the emotion type result to a preset emotion label in the animation database to obtain an emotion label corresponding to the emotion type result, wherein the emotion label matches the emotion state corresponding to each animation sequence frame; Based on the emotion tag, similarity calculation is performed with the emotion tags of all animation sequence frames in the animation database by using an emotion retrieval algorithm to obtain a similarity score; The similarity scores are sorted from high to low to obtain a sorting result, and according to the sorting result, the animation sequence frame with the highest similarity score is selected as the selected animation sequence frame.
5. The emotional resonance feedback method based on facial expression recognition according to claim 4, characterized in that: The step of calculating similarity between the emotion tags and the emotion tags of all animation sequence frames in the animation database by using an emotion retrieval algorithm according to the emotion tags to obtain a similarity score includes: The sentiment retrieval algorithm obtains the similarity score through the following cosine similarity formula: Among them, the S i is the similarity score of the i-th animation sequence frame, the w k is the weight coefficient of the kth emotion dimension, and L e (k) is the value of the input emotion label on the kth emotion dimension, the L ani (k) is the value of the emotion label of the kth animation sequence frame, the ‖L e (k) ‖ is the vector modulus length of the emotion label in the kth dimension, and the ‖L ani (k) ‖ is the vector modulus of the emotion label of the animation sequence frame in the kth dimension.
6. The emotional resonance feedback method based on facial expression recognition according to claim 1, characterized in that: The step of presenting the selected animation sequence frames to the robot face in real time also includes: Presenting the selected animation sequence frames to the robot face in real time, and generating voice feedback corresponding to the selected animation sequence frames through speech synthesis technology, wherein the voice feedback is played synchronously with the selected animation sequence frames; According to the voice feedback and the selected animation sequence frames, the robot is controlled by an action control algorithm to perform an emotional feedback action that matches the emotion type result.
7. The emotional resonance feedback method based on facial expression recognition according to claim 1, characterized in that: The emotional resonance feedback method based on facial expression recognition also includes: Acquire user interaction information, and construct an emotion supplement model based on the user interaction information, wherein the emotion supplement model analyzes the preprocessed facial image information to obtain new emotion features; According to the new emotion feature, the emotion recognition model is optimized and adjusted to obtain an optimized emotion recognition model; Based on the output emotion feature results of the optimized emotion recognition model, the output emotion feature results are input into the emotion classification model for verification and updating, and the adaptability of the new emotion features to the emotion classification model is analyzed. If the new emotion features do not match the classification rules of the emotion classification model, the classification rules of the emotion classification model are adjusted using the emotion feature results to generate an optimized emotion classification model; According to the emotion type results output by the optimized emotion classification model, the matching situation between the emotion type results and the animation sequence frames in the animation database is analyzed, and based on the matching situation, the animation database is updated to add new animation sequence frames corresponding to the new emotion features.
8. An emotional resonance feedback device based on facial expression recognition, characterized in that: The emotional resonance feedback device based on facial expression recognition comprises: A facial image acquisition and preprocessing module is used to obtain facial image information of a user, perform image preprocessing on the facial image information, and obtain preprocessed facial image information; An emotion recognition and feature extraction module, used to input the preprocessed facial image information into an emotion recognition model, and the emotion recognition model obtains facial feature information by extracting facial features from the preprocessed facial image information; an emotion classification and analysis module, used to input the facial feature information into an emotion classification model, and the emotion classification model obtains an emotion type result by performing an emotion state analysis on the facial feature information; An animation retrieval and matching module, used to retrieve animation sequence frames matching the emotion feature results from an animation database through an emotion retrieval algorithm according to the emotion type results, and obtain the selected animation sequence frames; The animation presentation module is used to present the selected animation sequence frames to the robot face in real time.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the emotional resonance feedback method based on facial expression recognition as claimed in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the emotional resonance feedback method based on facial expression recognition as claimed in any one of claims 1 to 7 are implemented.