Health care accompanying robot based on multi-modal facial feature fusion

By using multimodal facial feature fusion technology to acquire and process video streams of facial expressions of elderly users, the problem of misjudgment in the identification of emotional states of elderly users by health and wellness companion robots has been solved, and more accurate emotional understanding and interactive response have been achieved.

CN120726685BActive Publication Date: 2025-11-21ZHEJIANG FUBAO INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511188236.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-21
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing elderly care companion robots struggle to accurately identify the emotional state of elderly users, especially when there is a conflict between macro and micro expressions, leading to misjudgments and impacting the user experience.

Method used

By acquiring video streams of facial expressions from elderly individuals, and performing preprocessing, multimodal facial feature fusion is used, including facial expression spatial feature extraction, saliency enhancement, and hidden encoding, to generate more accurate emotion vectors to guide robot actions.

Benefits of technology

It achieves accurate and unambiguous recognition of the emotional state of elderly users, enhancing the realism and effectiveness of the elderly care companion robot's interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726685B_ABST
    Figure CN120726685B_ABST
Patent Text Reader

Abstract

The application relates to the field of health care robots, and specifically discloses a health care accompanying robot based on multi-modal facial feature fusion, which acquires facial expression video streams of the elderly, carries out fine pretreatment, captures subtle features of each region of the face, intelligently identifies the subtle features of each region of the face, and gives higher weights to facial micro-expressions. Then, the time dynamics of the captured expressions are combined with a time attention module to further focus on the key moments and modes of expression changes, so as to generate more accurate final emotion vectors. Finally, appropriate action instructions are generated for the robot based on the emotion vectors, so as to overcome the dependence of traditional methods on single macro-expression and the neglect of contradictory signals, and realize accurate and unambiguous understanding and response to the emotional state of the elderly user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of health and wellness robots, and more specifically, to a health and wellness companion robot based on multimodal facial feature fusion. Background Technology

[0002] With the increasingly significant trend of global population aging, elderly care companion robots, as an emerging intelligent service device, have shown great potential in improving the quality of life for the elderly and alleviating the social pressure of elderly care. However, for elderly care companion robots to truly fulfill their companionship role, accurately understanding the true emotional state of elderly users is crucial. When expressing emotions, the elderly often exhibit expressions that do not match their true feelings due to various reasons (such as not wanting to trouble others, habitual concealment, etc.). For example, even if they feel lost or sad, they may habitually smile. This ambiguity and contradiction in emotional expression poses a serious challenge to traditional emotion recognition technology, which may lead to inappropriate interactive responses from the robot, thereby damaging the user experience or even having negative impacts.

[0003] Existing elderly care companion robots or emotion recognition systems mostly rely on the recognition of macro-expressions (such as a raised corner of the mouth or a furrowed brow) or are limited to the analysis of single-modal facial features. However, such single or superficial emotion recognition methods struggle to effectively capture deeper and more nuanced emotional cues such as micro-expressions, subtle changes in the eyes, eyebrow movements, and blinking frequency. When there is a conflict between macro-expressions and micro-expressions (for example, a happy macro-expression in the mouth, while the eyes or eyebrows express neutral or sad micro-expressions), systems relying solely on macro-expressions often selectively ignore these contradictory signals, leading to misjudgments of the user's true emotions. Such misjudgments not only fail to meet the deeper emotional needs of elderly users but may also exacerbate their negative emotions due to inappropriate interactions, severely limiting the practical application effectiveness of elderly care companion robots.

[0004] Therefore, how to effectively integrate features from different regions and levels of the face, especially to resolve the conflict between macro-expressions and micro-expressions, so as to achieve accurate and unambiguous recognition of the emotional state of elderly users, is a key technical problem that urgently needs to be solved in the field of elderly care companion robots. Summary of the Invention

[0005] This application is made in order to solve the above-mentioned technical problems.

[0006] According to one aspect of this application, a multimodal facial feature fusion-based elderly care companion robot is provided, comprising: an elderly person facial expression video stream acquisition module for acquiring a facial expression video stream of a target elderly person captured by a camera of the elderly care companion robot; a facial expression preprocessing module for preprocessing the facial expression video stream to obtain a facial expression frame sequence; a facial expression spatial feature extraction module for extracting facial expression spatial features from each facial expression frame in the facial expression frame sequence to obtain a facial expression spatial feature map sequence; a facial expression spatial saliency module for inputting the facial expression spatial feature map sequence into a spatial attention module to obtain a facial expression spatial saliency feature vector sequence; a facial expression hiding encoding module for performing temporal dynamic encoding on the facial expression spatial saliency feature vector sequence to obtain a facial expression hiding state encoding vector sequence; an emotion vector generation module for inputting the facial expression hiding state encoding vector sequence into a temporal attention module to obtain a final emotion vector; and an instruction generation module for generating robot action instructions based on the final emotion vector.

[0007] Compared with existing technologies, this application provides a multimodal facial feature fusion-based elderly care companion robot, aiming to solve the problems of ambiguity and contradiction in the emotional expression of the elderly and the inaccuracy of existing systems in the background technology. This solution acquires video streams of facial expressions of elderly individuals and performs refined preprocessing, then uses a facial expression spatial feature extraction module to capture subtle features of various facial regions. Crucially, the facial expression spatial saliency module intelligently identifies and assigns higher weight to micro-expressions (such as subtle changes in the eyes and eyebrows), effectively distinguishing genuine emotions even when they conflict with macro-expressions of the mouth (such as forced smiles). Subsequently, a facial expression hiding encoding module captures the temporal dynamics of expressions, and combined with a temporal attention module, further focuses on key moments and patterns of expression changes, thereby generating a more accurate final emotion vector. Finally, based on this emotion vector, the instruction generation module can generate appropriate action commands for the robot, overcoming the reliance on single macro-expressions and the neglect of contradictory signals in traditional methods, achieving accurate and unambiguous understanding and response to the emotional state of elderly users. Attached Figure Description

[0008] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0009] Figure 1This is a block diagram of a health and wellness companion robot based on multimodal facial feature fusion according to an embodiment of this application.

[0010] Figure 2 This is a schematic diagram of data flow for a health and wellness companion robot based on multimodal facial feature fusion according to an embodiment of this application.

[0011] Figure 3 This is a block diagram of the facial expression preprocessing module in a health and wellness companion robot based on multimodal facial feature fusion according to an embodiment of this application.

[0012] Figure 4 This is a block diagram of the facial expression spatial saliency module in a health and wellness companion robot based on multimodal facial feature fusion, according to an embodiment of this application.

[0013] Figure 5 This is a block diagram of the facial expression hiding coding module in a health and wellness companion robot based on multimodal facial feature fusion according to an embodiment of this application. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0016] This application is made in order to overcome the aforementioned problems mentioned in the background art. Figure 1 This is a block diagram of a health and wellness companion robot based on multimodal facial feature fusion according to an embodiment of this application. Figure 2 This is a schematic diagram illustrating the data flow of a health and wellness companion robot based on multimodal facial feature fusion according to an embodiment of this application. Specifically, as shown... Figure 1 and Figure 2As shown, the elderly care companion robot 100 based on multimodal facial feature fusion according to an embodiment of this application includes: an elderly person facial expression video stream acquisition module 110, used to acquire a facial expression video stream of a target elderly person captured by the camera of the elderly care companion robot; a facial expression preprocessing module 120, used to preprocess the facial expression video stream to obtain a facial expression frame sequence; a facial expression spatial feature extraction module 130, used to extract facial expression spatial features from each facial expression frame in the facial expression frame sequence to obtain a facial expression spatial feature map sequence; a facial expression spatial saliency module 140, used to input the facial expression spatial feature map sequence into a spatial attention module to obtain a facial expression spatial saliency feature vector sequence; a facial expression hiding encoding module 150, used to perform temporal dynamic encoding on the facial expression spatial saliency feature vector sequence to obtain a facial expression hiding state encoding vector sequence; an emotion vector generation module 160, used to input the facial expression hiding state encoding vector sequence into a temporal attention module to obtain a final emotion vector; and an instruction generation module 170, used to generate robot action instructions based on the final emotion vector. In particular, this solution can transform complex emotional results derived from single-modal depth analysis into highly contextualized robotic behavior that integrates auditory, kinesthetic, and visual senses. This shift from single-modal perception to multimodal expression makes the robot's response no longer monotonous and rigid, but rather closer to natural human interaction, thereby greatly enhancing the realism and effectiveness of health and wellness companionship.

[0017] Specifically, the elderly person's facial expression video stream acquisition module 110 is used to acquire the facial expression video stream of the target elderly person captured by the camera of the elderly care companion robot. It is understandable that accurately capturing and understanding the user's emotional state is crucial for effective companionship during the interaction between the elderly care companion robot and the elderly user. However, elderly people often exhibit facial expressions that do not match their true feelings when facing the robot, for various reasons such as not wanting to trouble their children or out of habit. For example, even if they feel lost or sad, they may habitually smile. This ambiguity and contradiction in emotional expression makes it difficult for the robot to judge the user's true emotions based solely on simple dialogue or preset programs. In order to deeply analyze these subtle and complex facial expression changes, distinguish between forced smiles and genuine joy, and further identify potential anxiety or sadness, the elderly care companion robot must first have the ability to acquire the user's facial expression information. Therefore, acquiring the facial expression video stream of the target elderly person captured by the camera of the elderly care companion robot is the foundation for all subsequent emotion recognition and intelligent interaction, providing the robot with raw data input for perceiving the user's emotions.

[0018] In one specific implementation, the process of the elderly person's facial expression video stream acquisition module 110 is as follows: The elderly care companion robot integrates a high-resolution visual sensor, i.e., a camera. This camera is strategically mounted on the robot's head or chest and has a certain degree of mechanical freedom, enabling it to pitch or rotate horizontally according to the interaction scenario and the position of the target elderly person, to ensure that the facial area of ​​the target elderly person can be captured stably and clearly.

[0019] When the robot initiates companion mode or receives a user interaction command, the camera module is activated and enters working mode. The camera first performs a series of self-check procedures to confirm that its optical components, such as the lens, are unobstructed, the sensor chip is functioning properly, and the communication link with the main processing unit is unobstructed. Then, the camera begins continuously acquiring image data within its field of view at a preset frame rate and resolution. For example, a frame rate of 30 frames per second and a resolution of 1920x1080 pixels can be set to ensure smooth video streaming and rich image detail. To adapt to dynamic changes in indoor lighting conditions, the camera's built-in image signal processor (ISP) runs automatic exposure and white balance algorithms in real time. The automatic exposure algorithm dynamically adjusts the sensor's exposure time, gain, and aperture size based on the current scene's brightness information to ensure that the acquired image has appropriate brightness, avoiding overexposure or underexposure in facial areas that could lead to loss of detail. The white balance algorithm adjusts the gain of the three primary colors (red, green, and blue) according to the scene's color temperature to ensure accurate reproduction of skin tone and facial expressions.

[0020] The acquired raw image data is transmitted to the robot's main processing unit in digital signal form via a high-speed data interface, such as USB 3.0 or MIPI CSI. During transmission, to optimize bandwidth usage and facilitate subsequent storage and processing, the raw image data undergoes real-time video compression encoding, for example, using H.264 or H.265 encoding formats. Upon receiving the encoded data stream, the main processing unit decapsulates and decodes it using a corresponding decoder, thereby restoring a continuous, uncompressed sequence of image frames. These continuous image frame sequences constitute the aforementioned facial expression video stream.

[0021] Specifically, the facial expression preprocessing module 120 is used to preprocess the facial expression video stream to obtain a facial expression frame sequence. Correspondingly, the raw video stream data typically contains a large amount of redundant information, such as background clutter, lighting variations, and consecutive frames that may be irrelevant to facial expression analysis. Directly performing deep analysis on this unprocessed raw data not only incurs a huge computational burden and reduces processing efficiency but may also introduce noise, affecting the accuracy of subsequent feature extraction. Furthermore, deep learning models for facial expression recognition have strict requirements on the format, size, and pixel distribution of the input data. To improve processing efficiency, reduce noise interference, and provide high-quality, standardized input data for subsequent feature extraction and emotion recognition, the acquired facial expression video stream needs to be preprocessed. This series of preprocessing operations aims to accurately locate and extract key facial regions from the continuous video stream and convert them into a uniform format, thereby laying a solid foundation for subsequent deep analysis.

[0022] In particular, in a specific implementation, Figure 3 This is a block diagram of the facial expression preprocessing module in a health and wellness companion robot based on multimodal facial feature fusion, according to an embodiment of this application. Figure 3 As shown, the facial expression preprocessing module 120 includes: a facial expression sampling unit 121, used to sample the facial expression video stream at a fixed frequency to obtain a facial expression sampling frame sequence; a face region detection unit 122, used to perform face detection on each facial expression sampling frame in the facial expression sampling frame sequence to obtain a face region image sequence; and a face region processing unit 123, used to perform grayscale conversion, size normalization, and standardization on each face region image in the face region image sequence to obtain the facial expression frame sequence.

[0023] In this embodiment, the facial expression preprocessing module 120 is implemented as follows: First, after the elderly care companion robot acquires the facial expression video stream, preliminary processing is required to improve the efficiency and accuracy of subsequent analysis. Since the original video has a high frame rate, such as 30 frames per second, and contains a large amount of redundant information, direct processing would increase computational burden, introduce noise, and affect feature extraction and emotion recognition. To preserve key facial expression dynamics and reduce data volume, a fixed-frequency sampling strategy should be adopted to reduce redundancy and computational pressure, while providing efficient and concise input for subsequent refined analysis. The facial expression sampling unit 121 receives the facial expression video stream from the camera acquisition module. This video stream consists of a series of continuous image frames and has a fixed original frame rate, such as 30 frames per second. This unit internally includes a frame counter and a sampling interval calculation logic. When processing the video stream, the frame counter increments from the starting frame of the video stream. Simultaneously, based on the preset fixed sampling frequency and the frame rate of the original video stream, the interval between samples is calculated. For example, if the original video stream has a frame rate of 30 frames per second (fps) and the preset fixed sampling frequency is 5 fps, the sampling interval will be calculated as 30 fps ÷ 5 fps = 6 frames. This means that the facial expression sampling unit will extract one frame from the original video stream every 6 frames. When the frame counter reaches an integer multiple of the sampling interval (e.g., frame 1, frame 7, frame 13, etc.), the current frame is selected and added to the facial expression sampling frame sequence. Specifically, this fixed frequency setting is determined through empirical evaluation and optimization. For facial expression recognition tasks, a sampling frequency of 5 fps to 10 fps is considered a reasonable range, effectively capturing facial expression dynamics while significantly reducing data volume. For example, in practical deployment, this fixed frequency can be set to 8 fps based on the specific interaction scenario of the robot, such as daily conversations with the elderly, observation of emotional fluctuations, and available computing resources.

[0024] Next, considering that after the elderly care companion robot acquires the sampled facial expression frame sequence, although these image frames have reduced redundancy, they still contain background, body, and other regions unrelated to facial expression analysis. To ensure that subsequent facial expression spatial feature extraction can focus on the most core facial information, avoid interference from background noise, and improve processing efficiency and accuracy, it is necessary to accurately identify and crop the face region from each sampled frame. This provides a clean and focused input for refined facial expression analysis, enabling subsequent feature extraction modules to more effectively capture subtle facial muscle movements and expression changes, thus laying the foundation for accurate emotion recognition. The face region detection unit 122 receives the facial expression sampled frame sequence from the facial expression sampling unit and performs face detection processing. Face detection utilizes a pre-trained deep learning model, such as a multi-task cascaded convolutional network model. This model consists of three cascaded convolutional neural networks: P-Net, R-Net, and O-Net. P-Net is a lightweight convolutional neural network whose primary task is to rapidly scan an image pyramid (i.e., scaled versions of images at different sizes) to generate a large number of coarse face candidate boxes and their corresponding confidence scores. These candidate boxes are numerous and highly overlapping. R-Net receives the candidate boxes output by P-Net, normalizes their size, and then feeds them into a deeper convolutional neural network. R-Net's task is to further classify these coarse candidate boxes (determining whether they are faces) and perform bounding box regression (refining the position and size of the candidate boxes), while filtering out most non-face regions, thus obtaining fewer and more accurate candidate boxes. Finally, O-Net receives the refined candidate boxes from R-Net, normalizes them again, and feeds them into the deepest convolutional neural network. O-Net performs the final face classification, bounding box regression, and additionally predicts five facial landmarks (such as the centers of the left and right eyes, the tip of the nose, and the corners of the mouth). This model is trained on a public dataset containing a large number of face images and their corresponding bounding boxes and landmark annotations, and its internal weights and bias parameters are continuously adjusted through backpropagation and an optimizer. When a facial expression sample frame is input into the trained face detection model, it first generates preliminary candidate boxes using P-Net. These candidate boxes are then refined by R-Net and O-Net, ultimately outputting the bounding box coordinates (e.g., top-left x, y coordinates, width, and height) of possible faces in the image, along with their corresponding confidence scores. For example, if a sample frame containing the face of an elderly person is input, the model might output a bounding box with coordinates (100, 50, 200, 250), indicating that the face region starts at pixel (100, 50) in the image, has a width of 200 pixels, and a height of 250 pixels, along with a confidence score of 0.98.In some cases, the model may detect multiple faces, such as other faces appearing in the background, or the model generating multiple overlapping bounding boxes for the same face. To ensure that only the face of the target elderly object is extracted, the face region detection unit uses the non-maximum suppression (NMS) algorithm to eliminate overlapping bounding boxes and selects the most suitable face according to a preset strategy. For example, it selects the bounding box of the face with the highest confidence score; or, if a face was successfully detected in the previous frame, it selects the bounding box closest to the face position in the previous frame to maintain the continuity of target tracking. For example, a confidence threshold, such as 0.9, can be set, and only bounding boxes with a confidence score higher than this threshold will be considered. If there are still multiple faces that meet the criteria, the face with the largest area is selected as the target face. Once the bounding box of the target elderly object's face is determined, the face region detection unit will accurately crop the corresponding face region image from the original sampling frame based on this coordinate information. These cropped face region images together form a face region image sequence.

[0025] Finally, after the facial region image sequence is extracted, problems such as color redundancy, inconsistent sizes, and uneven pixel distribution still exist. Directly inputting these into the deep model will lead to low efficiency, slow convergence, and decreased accuracy. To improve the accuracy and efficiency of subsequent expression recognition, the images need to be grayscaled, size normalized, and standardized to unify the data format, optimize pixel distribution, and ensure high quality and consistency of model input, providing a reliable foundation for deep feature extraction. The facial region processing unit 123 first performs grayscale processing. A commonly used grayscale method is the weighted average method based on the human eye's sensitivity to different colors. Specifically, for each pixel in the image, its original red (R), green (G), and blue (B) components are linearly combined according to preset weights to calculate a grayscale value. For example, the following formula can be used: Grayscale value = 0.299*R + 0.587*G + 0.114*B. This formula reflects the characteristic that the human eye is most sensitive to green light, followed by red, and least sensitive to blue light. After grayscale conversion, each color image in the face region image sequence is converted to a corresponding grayscale image, forming a grayscale face region image sequence. Next, size normalization is performed. Each grayscale face region image is scaled or enlarged to a preset uniform size. For example, the target size can be set to 96x96 pixels. During size adjustment, interpolation algorithms are used to fill or calculate the values ​​of new pixels to maintain image sharpness and detail. Commonly used interpolation algorithms include bilinear interpolation or bicubic interpolation, which can smoothly transition based on the grayscale values ​​of surrounding pixels, avoiding obvious jagged edges or distortion during scaling. After size normalization, all face region images will have the same width and height, forming a size-normalized face region image sequence. Finally, standardization is performed. The pixel values ​​of the size-normalized images remain within the integer range of 0 to 255. To optimize the training effect and convergence speed of deep learning models, these pixel values ​​need to be standardized to distribute them within a range more suitable for model learning. A commonly used standardization method is Z-score standardization. This method transforms pixel values ​​into a distribution with a mean of 0 and a standard deviation of 1 by subtracting the mean of the pixel values ​​and dividing by their standard deviation. Specifically, for each image after size normalization, the average gray value (μ) and standard deviation (σ) of all its pixels are first calculated. These mean and standard deviation are global statistics pre-calculated based on the large image dataset used to train the facial expression recognition model to ensure consistency in standardization. Then, each pixel value (P) in the image is transformed according to the formula P=(P-μ) / σ to obtain the standardized pixel value P'. After these three consecutive processing steps—grayscale conversion, size normalization, and standardization—the sequence of face region images is finally converted into a sequence of facial expression frames with a uniform format and standard distribution.

[0026] Specifically, the facial expression spatial feature extraction module 130 is used to extract facial expression spatial features from each facial expression frame in the facial expression frame sequence to obtain a facial expression spatial feature map sequence. It should be understood that after grayscale conversion, size normalization, and standardization of the face region image, a series of facial expression frames with a uniform format and standard distribution are obtained. However, these frames are still raw pixel data. Although preprocessed, the complex facial expression information they contain, such as fine lines around the eyes, the curvature of the corners of the mouth, and subtle raising or lowering of the eyebrows, cannot be directly identified and quantified simply through pixel values. To extract more abstract, discriminative, and sensitive deep features from these pixel data, thereby effectively distinguishing subtle emotional differences such as forced smiles and genuine joy in the elderly, a powerful feature learning mechanism is needed. Therefore, extracting facial expression spatial features from each facial expression frame in the facial expression frame sequence to obtain a facial expression spatial feature map sequence is a crucial step in achieving accurate emotion recognition. It transforms the raw image data into a high-dimensional, semantically rich feature representation, laying the foundation for subsequent emotion analysis.

[0027] Specifically, in one embodiment, the facial expression spatial feature extraction module 130 is used to: input each facial expression frame in the facial expression frame sequence into a trained 2D convolutional neural network model to obtain the facial expression spatial feature map sequence. It is worth mentioning that the 2D convolutional neural network can automatically learn and extract multi-level, multi-scale spatial features from raw pixel data, eliminating the need for manually designed complex feature extractors and greatly simplifying feature engineering. Its local receptive field and weight-sharing mechanism make it robust to translation, scaling, and rotation in images, effectively capturing subtle texture and shape changes in different facial regions (such as eyes, eyebrows, and mouth). Furthermore, by stacking multiple layers of convolution and pooling operations, the model can learn abstract features from low-level edges to high-level semantic concepts. These features are crucial for distinguishing subtle emotional differences in the elderly, providing a more discriminative representation for subsequent emotion recognition.

[0028] In this embodiment, the facial expression spatial feature extraction module 130 is implemented as follows: The trained 2D convolutional neural network model employs a deep learning architecture, such as a modified VGG network or a ResNet network, designed to automatically learn and extract multi-level visual features from images. The model consists of multiple convolutional layers, activation function layers, pooling layers, and possibly batch normalization layers stacked together. Specifically, the architecture of this 2D convolutional neural network model may include: an input layer that receives preprocessed facial expression frames; convolutional layers, each containing multiple learnable convolutional kernels (or filters). These kernels slide across the input image or the feature map of the previous layer, performing convolution operations to detect local patterns such as edges, textures, corners, etc. As the network depth increases, the convolutional layers can learn increasingly abstract and complex facial features, such as the shape of the eyes, the contour of the mouth, the direction of the eyebrows, etc.; and activation function layers, applying non-linear activation functions, such as rectified linear units (ReLU), after each convolutional layer. The ReLU function introduces nonlinearity, enabling the network to learn and represent more complex nonlinear relationships, thereby enhancing the model's expressive power. Pooling layers, such as max pooling or average pooling, are used to downsample the feature maps, reducing their spatial size and thus lowering computational cost. This also makes the extracted features more robust to small translations and deformations of the input image. Specifically, this 2D convolutional neural network model is trained on a large-scale facial expression dataset, such as millions of facial images with emotional labels (e.g., happy, sad, angry). During training, the model continuously adjusts its internal convolutional kernel weights and bias parameters through backpropagation and an optimizer. These trained weights and bias parameters enable the model to automatically extract effective spatial features from facial images.

[0029] In practical applications, when each facial expression frame in the sequence is input into this trained 2D convolutional neural network model, it sequentially passes through the model's convolutional layers, activation function layers, and pooling layers. The model does not execute the final classification layer; instead, it extracts the feature map output from the penultimate convolutional layer or the last pooling layer. This output is a multi-channel tensor. Each value in this tensor represents the activation intensity of a specific region in the input image along a specific feature dimension, and together they constitute the facial expression space feature map of that facial expression frame. By repeating this process for each facial expression frame in the sequence, a final sequence of facial expression space feature maps containing the feature maps of all frames is obtained.

[0030] Specifically, the facial expression spatial saliency module 140 is used to input the facial expression spatial feature map sequence into the spatial attention module to obtain a facial expression spatial saliency feature vector sequence. Correspondingly, after the facial expression spatial feature extraction module successfully extracts the high-dimensional facial expression spatial feature map sequence from the facial expression frame, although these feature maps contain rich spatial information, not all regions are equally important for emotion recognition. For example, when distinguishing between a forced smile and genuine joy in an elderly person, micro-expressions around the eyes and subtle movements of the eyebrows may reflect true emotions more effectively than a broad expression like a raised corner of the mouth. Traditional feature extraction methods often treat all spatial locations in the feature map equally, which may cause the model to fail to effectively focus on the most discriminative regions when processing complex or contradictory expressions, thus affecting the accuracy of emotion recognition. In order to enable the model to dynamically focus on the regions in facial images that contribute more to emotion recognition, improve the ability to perceive subtle expressions, and effectively resolve the conflict between macro expressions and micro expressions, this application inputs the facial expression space feature map sequence into the spatial attention module to obtain a facial expression space saliency feature vector sequence to highlight key information and suppress irrelevant or interfering information, thereby generating a more discriminative feature representation.

[0031] In particular, in a specific implementation, Figure 4 This is a block diagram of the facial expression spatial saliency module in a health and wellness companion robot based on multimodal facial feature fusion, according to an embodiment of this application. Figure 4 As shown, the facial expression space saliency module 140 includes: a facial expression space attention generation unit 141, used to input the facial expression space feature map into the space attention module to obtain a facial expression space attention weight tensor; a facial expression space attention weighting unit 142, used to calculate the element-wise product between the facial expression space feature map and the facial expression space attention weight tensor to obtain a facial expression space saliency feature map; and a facial expression space feature pooling unit 143, used to perform global mean pooling on the facial expression space saliency feature map to obtain the facial expression space saliency feature vector.

[0032] In this embodiment, the facial expression space saliency module 140 is implemented as follows: First, a facial expression space attention generation unit 141 is used. This unit inputs the facial expression space feature map, for example, an HxWxC tensor, where H and W are spatial dimensions and C is the number of channels, into a specially designed spatial attention module. This spatial attention module consists of a series of convolutional layers and activation functions, and its purpose is to learn and generate a facial expression space attention weight tensor (e.g., an HxWx1 tensor) with the same spatial dimensions as the input feature map. Specifically, the spatial attention module can first perform global average pooling and global max pooling on the input feature map along the channel dimension to obtain two HxWx1 feature maps, representing the average and maximum information along the channel dimension, respectively. Then, these two feature maps are concatenated along the channel dimension to form an HxWx2 tensor. Next, this concatenated tensor is processed through a convolutional layer, for example, a 7x7 convolutional kernel with 1 output channel, and then through a sigmoid activation function. The Sigmoid function compresses the output values ​​to between 0 and 1, and these values ​​constitute the facial expression spatial attention weight tensor, where each value represents the importance of the corresponding spatial location. Specifically, the parameters in this spatial attention module are obtained through end-to-end learning.

[0033] Secondly, there is the facial expression space attention weighting unit 142. This unit broadcasts the facial expression space attention weight tensor along the channel dimension (if the attention weight tensor is single-channel, it is copied C times to match the number of channels in the feature map), and then performs element-wise multiplication with the original facial expression space feature map. This multiplication operation achieves spatial weighting of the original feature map: regions with high attention weights have their feature values ​​preserved or amplified; regions with low attention weights have their feature values ​​suppressed or reduced. For example, if the attention weight for the eye region is 0.9 and the attention weight for the background region is 0.1, then the feature information of the eye region will be enhanced in the saliency feature map, while the feature information of the background region will be weakened. In this way, the facial expression space saliency feature map highlights the facial region features that are more important for emotion recognition.

[0034] Finally, there is the facial expression space feature pooling unit 143. Global mean pooling is a common method for converting a two-dimensional feature map into a one-dimensional feature vector. It achieves this by calculating the average of all spatial locations in each channel of the salient feature map. For example, if the size of the facial expression space salient feature map is HxWxC, then global mean pooling will calculate the average of each C channel, ultimately resulting in a 1x1xC vector, i.e., a C-dimensional facial expression space salient feature vector. This feature vector is a generalized representation of the entire salient feature map; it aggregates all spatial information after attention weighting and compresses it into a fixed-length vector, facilitating processing by the subsequent temporal dynamic coding module. By performing the above operation on each facial expression space feature map in the sequence, they are all transformed into a compact feature vector rich in key spatial information, collectively forming the sequence of facial expression space salient feature vectors.

[0035] Specifically, the facial expression hiding encoding module 150 is used to perform temporal dynamic encoding on the spatially salient feature vector sequence of facial expressions to obtain a facial expression hiding state encoding vector sequence. It should be understood that human emotional expression is not an isolated, static moment, but a dynamic, evolving process. For example, sadness may be a gradually deepening emotion, while surprise is a brief outburst. Relying solely on the feature vector of a single frame to judge emotion ignores the continuity of expression over time, contextual information, and dynamic change patterns, leading to an inability to accurately capture emotional fluctuations, duration, and subtle emotional shifts. This is especially true for complex emotions such as forced smiles that may exist in the elderly, whose authenticity often needs to be judged by observing the dynamic process of the expression. In order to capture the dynamic changes of facial expressions over time, understand the evolution of emotions, and utilize historical information to assist in the judgment of current emotions, thereby achieving more accurate and context-aware recognition of genuine emotions, this application obtains a facial expression hiding state encoding vector sequence by performing temporal dynamic encoding on the spatially salient feature vector sequence of facial expressions.

[0036] In particular, in a specific implementation, Figure 5 This is a block diagram of the facial expression hiding coding module in a health and wellness companion robot based on multimodal facial feature fusion, according to an embodiment of this application. Figure 5 As shown, the facial expression hiding encoding module 150 includes: a facial expression space feature calibration unit 151, used to perform temporal collaborative dynamic attention calibration on the facial expression space saliency feature vector sequence to obtain a facial expression space saliency calibration feature vector sequence; and a hidden state encoding unit 152, used to input the facial expression space saliency calibration feature vector sequence into a gated recurrent network to obtain the facial expression hidden state encoding vector sequence.

[0037] In this embodiment, the facial expression hiding encoding module 150 is implemented as follows: Specifically, in the facial expression spatial saliency module, the facial expression spatial feature map sequence is converted into a facial expression spatial saliency feature vector sequence through spatial attention and global mean pooling. However, in the process of compressing the two-dimensional feature map into a one-dimensional vector, global mean pooling inevitably smooths out or even loses some key local spatial attention information, resulting in the final feature vector possibly failing to fully and faithfully express the image semantics enhanced by the spatial attention module in the original feature map. The degree of this information loss may not be consistent at different time steps of the sequence, thus introducing temporal imbalance. To compensate for this loss and ensure high-fidelity preservation of the temporal dimension of spatial attention information during feature vectorization, a temporal collaborative calibration mechanism needs to be established to correct the generated feature vector sequence, making it more consistent with the information expression of the original feature map sequence.

[0038] Based on this, in a specific embodiment, the facial expression space feature calibration unit 151 is used to: first, calculate the standard deviation of the difference between each corresponding facial expression space feature map and facial expression space saliency feature vector in the facial expression space feature map sequence and the facial expression space saliency feature vector sequence to obtain the facial expression difference standard deviation sequence, that is: ;in, It refers to the individual facial expression spatial feature maps in the facial expression spatial feature map sequence. It refers to each facial expression space salient feature vector in the sequence of facial expression space salient feature vectors. This indicates the calculation of the standard deviation. To take the absolute value, This represents the standard deviation of each facial expression difference in the facial expression difference standard deviation sequence. It can be understood that in the expression analysis scenario of a companion robot for elderly care, the facial expression spatial feature map retains rich, spatially attention-enhanced details of various facial regions (such as the corners of the eyes and mouth), while the facial expression spatial saliency feature vector is a generalized representation of this information after global pooling. Calculating the standard deviation of each feature set and finding their difference aims to quantify the difference in hierarchical saliency balance during the dimensionality reduction process from a two-dimensional feature map to a one-dimensional feature vector. This difference standard deviation can be considered a metric reflecting the degree of information loss or distortion caused by the uneven distribution of attention in the spatial dimension during global mean pooling. Performing this calculation for each frame in the sequence yields a facial expression difference standard deviation sequence, which dynamically reveals the quality of information retention at each time point throughout the entire expression evolution process.

[0039] Next, the facial expression difference standard deviation sequence is reset and activated to obtain the facial expression activation threshold sequence, i.e.: ;in, and Let represent the maximum and mean values ​​in the standard deviation sequence of facial expression differences, respectively. This refers to the activation thresholds of each facial expression in the facial expression activation threshold sequence. However, simply obtaining the facial expression activation threshold sequence is insufficient for direct correction, as its values ​​may contain noise or inconsistencies in scale. To generate a more stable and globally context-aware correction factor, the sequence needs to be reactivated. This step uses a dynamically weighted formula that combines the standard deviation of the difference at the current time step, the maximum standard deviation of the difference across the entire sequence, and the mean standard deviation. The aim is to maximize response compatibility alignment globally; that is, if the standard deviation of the difference at a certain time point is large, indicating significant information loss, its corresponding activation threshold will be closer to the maximum standard deviation of the difference in the sequence, thus generating a stronger correction signal; conversely, it will tend towards the mean, resulting in a gentler correction. This process generates a facial expression activation threshold sequence, where each threshold considers both the current frame and the dynamic calibration weights of the global sequence.

[0040] Finally, based on the facial expression activation threshold sequence, the facial expression space saliency feature vector sequence is subjected to hard constraint correction to obtain the facial expression space saliency calibration feature vector sequence, namely: ;in, For calculation The reciprocal, It is a dot product by position. This refers to each facial expression space saliency calibration feature vector in the facial expression space saliency calibration feature vector sequence. In other words, after obtaining a globally calibrated facial expression activation threshold sequence that accurately reflects the degree of information loss, the final step is to apply correction. This step multiplies each facial expression space saliency feature vector element-wise with the inverse of its corresponding activation threshold. This is a hard constraint correction, the effect of which is that for feature vectors that suffer greater information loss during pooling (i.e., higher facial expression activation thresholds), their values ​​are suppressed accordingly; while for feature vectors that retain better information (i.e., lower facial expression activation thresholds), their influence is preserved or even amplified. This operation constrains the activation of saliency bias from the perspective of the temporal distribution of saliency balance differences, optimizing the correlation between the temporal distribution of saliency in the facial expression space feature map sequence and the facial expression space saliency feature vector sequence, thereby establishing a collaborative attention mechanism between the two, ultimately resulting in a calibrated facial expression space saliency calibration feature vector sequence that more faithfully reflects the original spatial attention information.

[0041] It is worth mentioning that gated recurrent networks (GRUs) are a special type of recurrent neural network structure. By introducing a gating mechanism, they control the flow and retention of information within a sequence, effectively solving the gradient vanishing or exploding problems that traditional recurrent neural networks often encounter when processing long sequences, thus enabling them to better capture long-term dependencies. Therefore, the hidden state encoding unit 152 is implemented as follows: A common gated recurrent network structure is the Gated Recurrent Unit (GRU). The GRU model mainly consists of two gates: an update gate and a reset gate. The update gate determines how much of the hidden state information from the previous time step is retained in the current time step, and how much of the input information from the current time step is adopted. The reset gate determines how the hidden state from the previous time step is combined with the current input. Through these gating mechanisms, the GRU can selectively remember or forget historical information, thereby effectively learning the temporal dependencies in sequence data. Specifically, this gated recurrent network model is trained on a large-scale dynamic facial expression video dataset. During training, the model continuously adjusts its internal weights and bias parameters through backpropagation algorithm and optimizer. These trained weights and bias parameters enable the model to automatically extract temporal dynamic features from facial expression sequences.

[0042] In practical applications, when a sequence of facial expression spatial saliency calibration feature vectors (e.g., a sequence containing T vectors, each with dimension D, i.e., TxD) is input into this trained gated recurrent network, the network processes these vectors one by one according to time steps. At each time step t, the network receives the input feature vector at the current time step (i.e., the t-th facial expression spatial saliency calibration feature vector in the sequence) and the hidden state at the previous time step. Through the calculation of update and reset gates, the network generates a new hidden state. This new hidden state not only contains the facial expression features at the current time step but also integrates the facial expression information and temporal dynamic patterns from all previous time steps. For example, if the sequence length is 10, the network will iteratively calculate 10 hidden states starting from the first vector. Finally, the network outputs a sequence of facial expression hidden state encoding vectors with the same length as the input sequence, where each vector encodes all relevant facial expression dynamic information from the beginning of the sequence to the current time step.

[0043] Specifically, the emotion vector generation module 160 is used to input the facial expression hidden state encoding vector sequence into the time attention module to obtain the final emotion vector. It is understood that not all hidden states at every time step in the entire facial expression hidden state encoding vector sequence have equal importance to the final emotion judgment. For example, when judging whether an elderly person is truly sad, a micro-expression or eye contact at a certain moment may be more decisive than expressions at other times; or in a forced smile, a brief slip may be the key to revealing the true emotion. Simply averaging all hidden states or taking the last hidden state as the final emotion representation may dilute key information or ignore important historical context. In order to dynamically identify and focus on the time steps in the sequence that contribute most to the final emotion judgment, thereby more accurately capturing subtle changes in emotion and key clues, this application inputs the facial expression hidden state encoding vector sequence into the time attention module to obtain the final emotion vector. By assigning different weights to the hidden states at different time steps, it highlights information at key time points and suppresses irrelevant or interfering information, thereby generating a more discriminative emotion representation.

[0044] In particular, in a certain specific embodiment, the emotion vector generation module 160 includes: a time weight vector generation unit 161, used to input the facial expression hidden state encoding vector sequence into a time attention module to obtain a time weight vector; and a time weight vector weighting unit 162, used to calculate the weighted sum of the facial expression hidden state encoding vector sequence based on the time weight of each position in the time weight vector to obtain the final emotion vector.

[0045] In this embodiment, the emotion vector generation module 160 is implemented as follows: First, there is a time weight vector generation unit 161. It should be understood that a time attention module is a mechanism that dynamically assigns a weight to each element in the input sequence (here, each hidden state encoding vector) based on the context of the input sequence, to represent its importance to the final output. Specifically, in a particular embodiment, the time weight vector generation unit 161 is used to: calculate the time weight of each facial expression hidden state encoding vector in the facial expression hidden state encoding vector sequence using the following formula: ;in, This is the weight matrix. For bias vectors, It is the hidden state encoding vector of each facial expression in the sequence of hidden state encoding vectors. As a reference vector for time weighting, It is a vector transpose. It is the value of an exponential function with the natural constant e as its base. yes Normalization function, yes The corresponding time weights. Specifically, for each vector in the sequence... It first undergoes a linear transformation Here It is a learnable weight matrix. It is a learnable bias vector. These weights and bias parameters are automatically learned during model training through backpropagation and the optimizer. Its goal is to enable the model to accurately identify the time steps that are most important for sentiment judgment. This is a learnable time-weighted reference vector, a parameter automatically learned during model training through backpropagation and the optimizer. This vector can be viewed as a query vector, representing the feature patterns the model wants to focus on in the current task. It is calculated... After linear transformation The inner product between them can measure the similarity or relevance between each facial expression hidden state encoding vector and the temporal weight metric reference vector, thus obtaining a scalar value representing the importance score of the hidden state. It is a normalization function that transforms the importance scores of all time steps into a probability distribution, ensuring that all time weights... The sum is 1. Thus, This represents the contribution ratio of the i-th facial expression hidden state encoding vector to the final emotion vector. Thus, the temporal weight vector generation unit generates a corresponding temporal weight for each hidden state encoding vector in the sequence, and these weights together form the temporal weight vector.

[0046] Next is the temporal weight vector weighting unit 162. Each facial expression hidden state encoding vector is scalar multiplied by its corresponding temporal weight, and then all weighted vectors are summed. The final emotion vector is... This weighted summation operation effectively aggregates information from all hidden states in the sequence, but assigns higher weights to more important time steps. For example, if a micro-expression at a certain time step reveals the true emotion, its corresponding hidden state vector will receive higher weights, thus occupying a larger proportion in the final emotion vector. In this way, the final emotion vector is a more discriminative emotion representation that integrates all temporal information of the sequence and highlights the contributions of key time points.

[0047] Specifically, the instruction generation module 170 is used to generate robot action instructions based on the final emotion vector. That is, after the elderly care companion robot generates a final emotion vector that accurately reflects the elderly user's true emotions through multimodal facial feature fusion, this high-dimensional numerical representation itself cannot directly drive the robot's behavior. The core value of the elderly care companion robot lies in its ability to provide appropriate, personalized, and empathetic companionship and services based on the user's true emotional state. If the robot merely stays at the level of emotion recognition and cannot translate this perception into concrete and meaningful interactive actions, its companionship effect will be greatly reduced, and it may even exacerbate the elderly user's negative emotions due to inappropriate reactions. To enable the robot to truly achieve intelligent and adaptive interaction, the deep understanding of the elderly user's emotions is transformed into actual robot actions that can influence the user experience, thereby providing timely and effective psychological comfort or positive guidance. Furthermore, based on the final emotion vector, robot action instructions are generated. This step transforms abstract emotional perception into concrete physical actions, enabling the robot to truly become an emotional supporter and life assistant for the elderly user.

[0048] Specifically, in one particular embodiment, the instruction generation module 170 is configured to: input the final emotion vector into a fully connected layer classifier to obtain an emotion recognition result; input the emotion recognition module into a behavior tree-based decision module to obtain a response strategy; and generate the robot action instructions based on the response strategy.

[0049] In this embodiment, the instruction generation module 170 is implemented as follows: First, the final emotion vector is input into a fully connected layer classifier to obtain the emotion recognition result. The final emotion vector is a high-dimensional numerical representation containing emotional information after complex processing and attention weighting. To convert it into human-understandable emotion categories, the vector is input into a pre-trained fully connected layer classifier. This classifier consists of one or more fully connected layers (also called dense layers), each taking the output of the previous layer as input and processing it through linear transformations and non-linear activation functions such as ReLU. The last layer is a fully connected layer whose output dimension matches the number of predefined emotion categories, such as happiness, sadness, anger, surprise, calmness, forced smile, etc., followed by a Softmax activation function. The Softmax function converts the output value into a probability distribution representing the likelihood that the final emotion vector belongs to each emotion category. The classifier is trained on a large-scale, diverse emotion-labeled dataset, and its internal weights and bias parameters are adjusted through backpropagation and an optimizer (such as the Adam optimizer). For example, if the probability distribution of the classifier output after the final emotion vector is input is: happy 0.1, sad 0.05, angry 0.02, surprised 0.03, calm 0.05, forced smile 0.75, then the emotion recognition result is forced smile, because it has the highest probability.

[0050] Secondly, the emotion recognition result is input into a behavior tree-based decision module to obtain a response strategy. For example, a forced smile is used as input and fed into a behavior tree-based decision module. Those skilled in the art will know that a behavior tree is a hierarchical, modular control architecture widely used in robotics and game AI to manage complex behavioral logic. It consists of different types of nodes, including: control flow nodes, such as sequence nodes (executing child nodes sequentially until one fails or all succeed), selector nodes (trying child nodes sequentially until one succeeds or all fail), and parallel nodes (executing all child nodes simultaneously). Conditional nodes are used to check whether specific conditions are met (e.g., whether the user is alone, or whether it is nighttime). Behavioral nodes execute specific actions or tasks. The decision module pre-designs a series of response strategies for different emotions and situations. When an emotion recognition result is received, the behavior tree traverses from its root node. The root node is a selector node that attempts to activate its subordinate branches based on the current emotion recognition result. For example, if the emotion recognition result is forced smile, the selector node will find and activate a branch specifically for handling the forced smile emotion. This branch is a sequence node that defines a series of subtasks or sub-behaviors that need to be executed sequentially. Within this forced smile sequence branch, the first child node might be a condition node, such as checking whether the user continues to exhibit a forced smile. This condition node queries the robot's internal state or perception module to determine whether the forced smile emotion has persisted over a period of time or whether a preset duration threshold has been reached. If the condition node evaluates successfully (i.e., the user does indeed continue to exhibit a forced smile), the sequence node will continue to execute its next child node. If the evaluation fails, the current sequence node will fail, and the behavior tree may backtrack and try other branches (although branches for specific emotions are designed to terminate directly when the condition is not met, or have backup strategies). Once the condition node successfully passes, the sequence node will activate subsequent behavior nodes in sequence. For example, the next behavior node might be to initiate an empathetic dialogue. When this behavior node is activated, it sends instructions to the robot's speech synthesis module to prepare for a preset dialogue expressing empathy. Next, the sequence node activates a behavior node that adjusts the robot's posture. This node sends commands to the robot's motion control module, causing the robot to make gestures that express concern, such as slightly tilting its head or leaning forward slightly. Finally, the sequence node activates a behavior node that provides comfort, which may trigger the robot to speak comforting words or play soothing music. This specific path, activated from the root node to the final behavior node, and the higher-level action plans represented by all the behavior nodes along this path, together constitute the response strategy to the current emotion.For example, for the emotion recognition result of forced smile, if the conditions are consistently met, the generated response strategy may be an ordered list of actions including initiating empathic dialogue, adjusting the robot's posture, and providing comfort.

[0051] Finally, based on the response strategy, the robot's action instructions are generated. The response strategy is a high-level action plan that needs to be translated into low-level action instructions that the robot can understand and execute. For example, a response strategy for the emotion of "forcing a smile" might be an ordered list of actions including initiating empathic dialogue, adjusting the robot's posture, and providing comfort. These high-level action plans need to be translated into low-level action instructions that the robot can understand and execute. The instruction generation module contains a predefined mapping mechanism or conversion rule set. It iterates through each high-level action plan in the response strategy and parses it into a series of specific, executable low-level robot action instructions. For example, when the instruction generation module receives the high-level action plan of initiating empathic dialogue, it generates a specific speech synthesis instruction, such as `speak("You seem to have something on your mind, would you like to talk to me?")`, based on a preset dialogue script or context template. This instruction contains the text content to be synthesized and is sent to the robot's speech synthesis module for processing. Next, for the action plan of adjusting the robot's posture, the instruction generation module parses it into a series of fine-grained motion control instructions. For example, to express concern, it might generate a `move_head(pitch=-5)` command, instructing the robot's head to tilt downwards by 5 degrees; and simultaneously generate a `move_torso(lean_forward=0.1)` command, instructing the robot's torso to tilt forward by 0.1 meters. These parameter values ​​are pre-set based on the robot's kinematic characteristics and the posture requirements for expressing concern, ensuring the robot can move in a natural and expressive manner. These movement commands are sent to the robot's motion control module. The action plan of providing comfort might be converted into combined commands to achieve a multimodal comforting effect. For example, the command generation module might generate a speech synthesis command `speak` ("I'm here with you"), and simultaneously generate a display module command `display_emoticon(sad_face)` to synchronously display a sad or concerned expression on the robot's screen, enhancing the richness of emotional expression. These commands are sent to the speech synthesis module and the display module, respectively. These robot movement commands together constitute the final sequence of robot movement commands. These instructions are then sent to the robot's execution layer, driving its various functional modules (such as the speech synthesis module, motion control module, and display module) to work together, thereby enabling intelligent responses to the emotions of elderly users.

[0052] In summary, a multimodal facial feature fusion-based elderly care companion robot 100 based on embodiments of this application is presented, aiming to address the problems of ambiguity and contradiction in emotional expression among the elderly, as well as the inaccuracy of existing systems in the prior art. This solution acquires video streams of facial expressions from elderly individuals and performs refined preprocessing, then utilizes a facial expression spatial feature extraction module to capture subtle features of various facial regions. Crucially, the facial expression spatial saliency module intelligently identifies and assigns higher weight to micro-expressions (such as subtle changes in the eyes and eyebrows), effectively distinguishing genuine emotions even when they conflict with macro-expressions of the mouth (such as forced smiles). Subsequently, a facial expression hiding encoding module captures the temporal dynamics of expressions, and combined with a temporal attention module, further focuses on key moments and patterns of expression changes, thereby generating a more accurate final emotion vector. Finally, based on this emotion vector, a command generation module can generate appropriate action commands for the robot, overcoming the reliance on single macro-expressions and the neglect of contradictory signals in traditional methods, achieving accurate and unambiguous understanding and response to the emotional state of elderly users.

[0053] As described above, the health and wellness companion robot 100 based on multimodal facial feature fusion according to the embodiments of this application can be implemented in various wireless terminals, such as servers with health and wellness companion algorithms based on multimodal facial feature fusion. In one possible implementation, the health and wellness companion robot 100 based on multimodal facial feature fusion according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or hardware module. For example, the health and wellness companion robot 100 based on multimodal facial feature fusion can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the health and wellness companion robot 100 based on multimodal facial feature fusion can also be one of many hardware modules of the wireless terminal.

[0054] Alternatively, in another example, the health care companion robot 100 based on multimodal facial feature fusion and the wireless terminal can also be separate devices, and the health care companion robot 100 based on multimodal facial feature fusion can be connected to the wireless terminal via wired and / or wireless networks, and transmit interactive information in accordance with the agreed data format.

[0055] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive. Furthermore, it is not limited to the disclosed implementations, and many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations.

Claims

1. A health and wellness companion robot based on facial feature fusion, characterized in that, include: The elderly person's facial expression video stream acquisition module is used to acquire the facial expression video stream of the target elderly person captured by the camera of the elderly care companion robot; A facial expression preprocessing module is used to preprocess the facial expression video stream to obtain a facial expression frame sequence; a facial expression spatial feature extraction module is used to extract facial expression spatial features from each facial expression frame in the facial expression frame sequence to obtain a facial expression spatial feature map sequence; a facial expression spatial saliency module is used to input the facial expression spatial feature map sequence into a spatial attention module to obtain a facial expression spatial saliency feature vector sequence. The facial expression hiding encoding module is used to perform temporal dynamic encoding on the facial expression space saliency feature vector sequence to obtain the facial expression hiding state encoding vector sequence; the emotion vector generation module is used to input the facial expression hiding state encoding vector sequence into the temporal attention module to obtain the final emotion vector. The instruction generation module is used to generate robot action instructions based on the final emotion vector; The facial expression hiding encoding module includes: a facial expression space feature calibration unit, used to perform temporal collaborative dynamic attention calibration on the facial expression space saliency feature vector sequence to obtain a facial expression space saliency calibration feature vector sequence; and a hidden state encoding unit, used to input the facial expression space saliency calibration feature vector sequence into a gated recurrent network to obtain the facial expression hidden state encoding vector sequence.

2. The health and wellness companion robot based on facial feature fusion according to claim 1, characterized in that, The facial expression preprocessing module includes: a facial expression sampling unit, used to sample the facial expression video stream at a fixed frequency to obtain a facial expression sampling frame sequence; a face region detection unit, used to perform face detection on each facial expression sampling frame in the facial expression sampling frame sequence to obtain a face region image sequence; and a face region processing unit, used to perform grayscale conversion, size normalization, and standardization on each face region image in the face region image sequence to obtain the facial expression frame sequence.

3. The health and wellness companion robot based on facial feature fusion according to claim 2, characterized in that, The facial expression spatial feature extraction module is used to: input each facial expression frame in the facial expression frame sequence into a trained 2D convolutional neural network model to obtain the facial expression spatial feature map sequence.

4. The health and wellness companion robot based on facial feature fusion according to claim 3, characterized in that, The facial expression space saliency module includes: a facial expression space attention generation unit, used to input the facial expression space feature map into the spatial attention module to obtain a facial expression space attention weight tensor; a facial expression space attention weighting unit, used to calculate the element-wise product between the facial expression space feature map and the facial expression space attention weight tensor to obtain a facial expression space saliency feature map; and a facial expression space feature pooling unit, used to perform global mean pooling on the facial expression space saliency feature map to obtain the facial expression space saliency feature vector.

5. The health and wellness companion robot based on facial feature fusion according to claim 1, characterized in that, The facial expression space feature calibration unit is used to: calculate the standard deviation of the difference between each group of facial expression space feature maps and facial expression space saliency feature vectors in the facial expression space feature map sequence and the facial expression space saliency feature vector sequence to obtain the facial expression difference standard deviation sequence; The facial expression difference standard deviation sequence is reset and activated to obtain a facial expression activation threshold sequence; based on the facial expression activation threshold sequence, the facial expression spatial saliency feature vector sequence is hard-constrained and corrected to obtain the facial expression spatial saliency calibration feature vector sequence.

6. The health and wellness companion robot based on facial feature fusion according to claim 1, characterized in that, The emotion vector generation module includes: a time weight vector generation unit, used to input the facial expression hidden state encoding vector sequence into the time attention module to obtain a time weight vector; and a time weight vector weighting unit, used to calculate the weighted sum of the facial expression hidden state encoding vector sequence based on the time weight of each position in the time weight vector to obtain the final emotion vector.

7. The health and wellness companion robot based on facial feature fusion according to claim 6, characterized in that, The time weight vector generation unit is used to: calculate the time weight of each facial expression hidden state encoding vector in the facial expression hidden state encoding vector sequence using the following formula: ;in, This is the weight matrix. For bias vectors, It is the hidden state encoding vector of each facial expression in the sequence of hidden state encoding vectors. As a reference vector for time weighting, It is a vector transpose. It is the value of an exponential function with the natural constant e as its base. yes Normalization function, yes The corresponding time weight.

8. The health and wellness companion robot based on facial feature fusion according to claim 1, characterized in that, The instruction generation module is used to: input the final emotion vector into a fully connected layer classifier to obtain an emotion recognition result; and input the emotion recognition module into a behavior tree-based decision module to obtain a response strategy. Based on the response strategy, the robot action commands are generated.

Citation Information

Patent Citations

  • Old people supporting and accompanying service robot system and old people supporting and accompanying method

    CN107856039A