Emotion accompanying humanoid robot system based on multi-modal interaction and control method thereof
By using a multimodal humanoid robot system to extract user state feature vectors through data acquisition and recognition models, personalized dialogue and massage control are generated, solving the problem of low accuracy in single-modal recognition and achieving more intelligent user emotion recognition and personalized care.
Patent Information
- Application Number
- CN202511550785.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-03-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, humanoid companion robots mainly rely on a single modality for emotion recognition, resulting in low recognition accuracy, insufficient intelligence in response, difficulty in achieving personalized companionship and response, and limiting the effectiveness of intelligent companionship.
The emotional companion robot system, which adopts multimodal interaction, acquires physiological parameters and behavioral data through a data acquisition unit, extracts user state feature vectors using a user state recognition model, and generates personalized outputs through a generative dialogue model and a massage control algorithm.
It improves the accuracy and intelligence of user emotion recognition, enabling more personalized companionship and responses that are closer to the user's psychological state, thus enhancing the robot's intelligent companionship effect.
Smart Images

Figure CN121697000A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of humanoid robots, and in particular to a sentiment accompanying humanoid robot system based on multi-modal interaction and a control method thereof. BACKGROUND
[0002] A humanoid accompanying robot is an intelligent device with a humanoid appearance and interaction capability, which is widely used in medical rehabilitation fields (postoperative psychological counseling, cognitive training), smart elderly care scenarios (emotional accompanying of lonely old people, chronic disease management), special education industries (behavior intervention for autistic children), and family services (child growth accompanying, pet emotional interaction), etc. The core technologies thereof mainly include speech recognition and synthesis, natural language processing (NLP), computer vision, human-computer interaction, multi-modal perception and emotion recognition, etc. Among them, the speech recognition and synthesis technology enables the robot to interact with the user through voice, improving the natural communication experience; the natural language processing technology enables the robot to understand the user's language intention and respond intelligently; the computer vision enables the robot to recognize faces, postures, actions, etc., to assist in understanding user behavior; the human-computer interaction technology combines voice, vision and motion control to build a more humanized interaction experience.
[0003] Although current humanoid accompanying robots have made progress in voice and visual interaction, they still lack the ability to deeply integrate multi-modal data, and cannot accurately identify complex emotions of users. Most of the existing technologies rely only on a single mode, such as voice, for emotion recognition, resulting in low recognition accuracy, insufficient intelligent response, and difficulty in achieving personalized accompanying and response based on user emotions, which limits the intelligent accompanying effect of the robot. SUMMARY
[0004] The present application provides a sentiment accompanying humanoid robot system based on multi-modal interaction and a control method thereof, which at least solves the problem that most of the existing technologies rely only on a single mode, resulting in low recognition accuracy, insufficient intelligent response, difficulty in achieving personalized accompanying and response based on user emotions, and limiting the intelligent accompanying effect of the robot.
[0005] In one aspect, a sentiment accompanying humanoid robot system based on multi-modal interaction is provided, which is connected to a humanoid robot device, and includes a data acquisition unit, a data processing unit and a control unit: The data acquisition unit is configured to acquire physiological parameters and / or behavior data of a target user; The data processing unit is configured to obtain a user state feature vector through a user state recognition model according to the physiological parameters and / or behavior data; The control unit is configured to read the real-time user state feature vector and input the user state feature vector into a generative dialogue model and / or a massage control algorithm to obtain an output for the request of the target user in response to the request of the target user and / or monitoring that the target user triggers a preset condition.
[0006] Optionally, the data acquisition unit is configured to obtain the physiological parameters and / or behavior data of the target user by at least one of the following three ways: acquiring the physiological parameters and / or behavior data of the target user through sensors arranged on the humanoid robot device; acquiring the physiological parameters and / or behavior data of the target user through sensors worn by the target user and in communication with the humanoid robot device; acquiring the physiological parameters and / or behavior data of the target user through sensors worn by the target user and in communication with the Internet.
[0007] Optionally, the user state recognition model is configured to: extract local features of the physiological parameters and / or behavior data of the target user through a pre-model to obtain time series feature data; analyze the time series feature data through an LSTM model to obtain a user state feature vector, the user state feature vector including a user psychological state feature and a user psychological state change trend feature; the pre-model and the LSTM model belong to the user state recognition model.
[0008] Optionally, the LSTM model is configured to include: at least 2 LSTM layers for extracting time-dependent features of the time series feature data; a fully connected layer for outputting a final user psychological state feature vector; an activation function for representing probability or state intensity; the output of the LSTM model is configured to include: a binary classification scalar for representing the user psychological state feature obtained by the activation function; a binary classification scalar for the user psychological state change trend feature obtained by the activation function.
[0009] Optionally, the reading of the real-time user state feature vector and the input of the user state feature vector into a generative dialogue model and / or a massage control algorithm to obtain an output for the request of the target user in response to the request of the target user and / or monitoring that the target user triggers a preset condition includes: reading the real-time user state feature vector in response to the dialogue request of the target user; generating a state prompt according to the real-time user state feature vector; inputting the conversation request of the target user and the state prompt into a generative dialogue model to obtain an output for the conversation request of the target user.
[0010] Optionally, the generating a state prompt according to the real-time user state feature vector comprises: obtaining a prompt mapping corresponding to the target user according to the change of the user state in the historical dialogue process of the target user and the generative dialogue model; generating a state prompt according to the real-time user state feature vector and the prompt mapping.
[0011] Optionally, the reading a real-time user state feature vector in response to the request of the target user and / or monitoring that the target user triggers a preset condition, and inputting the user state feature vector into a generative dialogue model and / or a massage control algorithm to obtain an output for the request of the target user comprises: reading a real-time user state feature vector in response to the massage request of the target user and / or monitoring that the target user triggers a preset condition; obtaining a massage target according to the real-time user state feature vector; controlling a massage device arranged on a humanoid robot device through a massage control algorithm according to the massage target.
[0012] Optionally, the preset condition comprises at least one of the following conditions: monitoring that a physiological parameter of the target user is abnormal; monitoring that behavior data of the target user is abnormal; monitoring that a user psychological state feature of the target user is abnormal; monitoring that a user psychological state change trend feature of the target user is abnormal.
[0013] In still another aspect, a control method of an emotional companion humanoid robot based on multi-modal interaction comprises: obtaining physiological parameters and / or behavior data of a target user; obtaining a user state feature vector through a user state recognition model according to the physiological parameters and / or behavior data; reading a real-time user state feature vector in response to the request of the target user and / or monitoring that the target user triggers a preset condition, and inputting the user state feature vector into a generative dialogue model and / or a massage control algorithm to obtain an output for the request of the target user.
[0014] In another aspect, a computer device includes a memory and a processor, the memory having stored therein a computer program, and the processor executes the computer program to implement the method described above.
[0015] In another aspect, a computer storage medium has stored thereon a computer program, and a processor executes the computer program to implement the method described above.
[0016] Compared with the prior art, the present application has the following advantages and beneficial effects: The present application is a kind of based on multi-modal interaction emotional companion humanoid robot system and control method thereof, the system connects humanoid robot equipment, the system includes data acquisition unit, data processing unit and control unit: the data acquisition unit is configured to, obtain the physiological parameters and / or behavior data of target user;The data processing unit is configured to, according to the physiological parameters and / or behavior data, obtain user state feature vector through user state identification model;The control unit is configured to, in response to the request of the target user and / or monitoring the target user triggers preset condition, read real-time user state feature vector, and input the user state feature vector into generative dialogue model and / or massage control algorithm to obtain the output for the request of the target user.At least solved the problem that the prior art mostly depends on single mode, resulting in low recognition accuracy, not enough intelligent reaction, difficult to realize personalized companion and response based on user emotion, limiting the intelligent companion effect of robot. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced below.In all the drawings, similar elements or parts are generally identified by similar reference numerals.The elements or parts in the drawings are not necessarily drawn according to the actual scale.
[0018] Figure 1 The flowchart of the control method of the emotional companion humanoid robot based on multi-modal interaction in the present application is shown in the figure. Figure 2 The structural diagram of the computer device in the present application is shown in the figure.
[0019] In the figure, the marks are:101-processor, 102-communication bus, 103-network interface, 104-user interface, 105-memory.
[0020] The implementation of the present application, functional characteristics and advantages will be further described with reference to the drawings. DETAILED DESCRIPTION
[0021] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0023] Example 1 like Figure 1 As shown, an emotional companionship humanoid robot system based on multimodal interaction is disclosed. The system is connected to a humanoid robot device and includes a data acquisition unit, a data processing unit, and a control unit. The data acquisition unit is configured to acquire the physiological parameters and / or behavioral data of the target user; The data processing unit is configured to obtain a user state feature vector based on physiological parameters and / or behavioral data through a user state recognition model; The control unit is configured to, in response to a request from a target user and / or to detect that the target user has triggered a preset condition, read the real-time user state feature vector and input the user state feature vector into a generative dialogue model and / or a massage control algorithm to obtain an output for the target user's request.
[0024] Optionally, the humanoid robot device is configured as follows: a humanoid robot with a biomimetic appearance and flexible tactile feedback, wherein the humanoid robot uses a skin-like skin, such as silicone material with a Shore hardness of 30A-40A and a tensile strength of 1.5MPa-2.0MPa, and has a deformable facial expression mechanism capable of presenting a variety of basic expression combinations. The humanoid robot device is also equipped with multimodal sensors and a UWB environmental perception unit, wherein the multimodal sensors integrate cameras, piezoelectric tactile sensors, impedance spectroscopy blood glucose monitoring sensors, and millimeter-wave vital signs sensors, etc. The humanoid robot device is also equipped with a network communication module.
[0025] Optional physiological parameters include one or more of the following data: heart rate, blood pressure, body temperature, respiratory rate, blood oxygen saturation, and blood glucose.
[0026] Optionally, behavioral data may include one or more of the following: voice data, activity data, dietary data, sleep data, social interaction data, and environmental interaction data.
[0027] Optionally, the user state recognition model can adopt the following architectures: multi-channel neural network, multi-modal fusion Transformer model, multi-modal convolutional neural network, and deep fusion autoencoder; When using a multi-channel neural network, a separate neural network branch is used for each data modality to process each input data. The outputs of each branch are fused in subsequent layers, usually by concatenation. At the same time, self-attention mechanisms or weighted summation can be used to handle the relative importance of different modalities. When using the multimodal fusion Transformer model, each modality's data is processed using an independent Transformer encoder. Each encoder outputs a contextual representation of a specific modality. The feature information of each modality is fused through a cross-modal self-attention mechanism, thereby enabling information exchange between multiple modalities. When using a multimodal convolutional neural network, for each modality of data, a convolutional layer is used to extract low-level features, and then a pooling layer is used to reduce the dimensionality. Next, a long short-term memory network or a gated recurrent unit is applied to the features of each modality for temporal modeling. The outputs of each modality are fused by concatenation or weighted averaging to generate a comprehensive feature representation. The fused features are then fed into a fully connected layer to output a psychological state feature vector and a trend vector.
[0028] When using a deep fusion autoencoder, each modality of data has an independent encoder, and a deep neural network is used to extract the low-dimensional feature representation of each modality; the feature representations of each modality are fused together through weighted summation or concatenation operations to form a unified high-dimensional representation, and the fused high-dimensional representation is restored into a psychological state feature vector and a change trend vector through a decoder.
[0029] Optionally, the user state feature vector can be input into the generative dialogue model and / or massage control algorithm to obtain output tailored to the target user's request, including: In response to the target user's dialogue request, read the real-time user state feature vector; The target user's dialogue request and real-time user state feature vector are input into the generative dialogue model to obtain the output for the target user's dialogue request.
[0030] Optionally, the user state feature vector can be input into the generative dialogue model and / or massage control algorithm to obtain output tailored to the target user's request, including: In response to a massage request from a target user and / or upon detecting that the target user has triggered a preset condition, the system reads the real-time user state feature vector. Based on real-time user state feature vectors, a massage control algorithm is used to control the massage device installed on the humanoid robot.
[0031] The above method employs a pre-defined user state recognition model to identify multimodal data and obtain user state feature vectors. These feature vectors are then used as input to a generative dialogue model and / or a massage control model, enabling the output content or control commands of these models to better provide companionship to the user. This addresses the problem that most existing technologies rely solely on a single modality, resulting in low recognition accuracy, insufficient intelligent response, and difficulty in achieving personalized companionship and responses based on user emotions, thus limiting the effectiveness of intelligent companionship robots.
[0032] Example 2 This embodiment, based on Embodiment 1, provides an emotional companionship humanoid robot system based on multimodal interaction. The system connects to a humanoid robot device and includes a data acquisition unit, a data processing unit, and a control unit. The data acquisition unit is configured to acquire at least one physiological parameter and / or behavioral data of the target user.
[0033] Optionally, the data acquisition unit is configured to acquire the target user's physiological parameters and / or behavioral data through at least one of the following three methods: Physiological parameters and / or behavioral data of the target user are acquired by sensors installed on the humanoid robot device; Physiological parameters and / or behavioral data of the target user are acquired through sensors worn by the target user and communicating with the humanoid robot device; Physiological parameters and / or behavioral data of the target user are acquired through sensors worn by the target user and communicating with the Internet.
[0034] Optionally, the data acquisition unit can acquire the target user's physiological parameters and / or behavioral data from sensors on the humanoid robot device via a data bus, or acquire the target user's physiological parameters and / or behavioral data from other devices via near-field communication or network communication. Other devices can be dedicated extension devices, or wearable devices or smart devices such as smartwatches with physiological detection functions. Specifically, camera data of target users and their interactions with the environment can be obtained by installing cameras on humanoid robot devices; Blood oxygen saturation data and blood glucose data can be obtained by using an impedance spectral blood glucose monitoring sensor installed on a humanoid robot device; Body temperature data, respiratory rate data, etc. can be obtained by millimeter-wave vital sign sensors installed on humanoid robot devices; Optionally, heart rate data, blood pressure data, etc., can be obtained by communicating with the smartwatch worn by the target user.
[0035] The data processing unit is configured to obtain a user state feature vector based on physiological parameters and / or behavioral data through a user state recognition model.
[0036] Optionally, the user state recognition model is configured as follows: Temporal feature data is obtained by extracting local features of the target user's physiological parameters and / or behavioral data through a pre-model; The user state feature vector is obtained by analyzing the time series feature data through the LSTM model. The user state feature vector includes user psychological state features and user psychological state change trend features. The pre-model and LSTM model belong to the user state recognition model.
[0037] Optionally, the LSTM model is configured to include: At least two LSTM layers are used to extract time-dependent features from time-series data; Fully connected layer used to output the final user mental state feature vector; Activation functions used to represent probabilities or state strengths; The output of the LSTM model is configured to include: A binary scalar obtained from the activation function to represent the characteristics of a user's psychological state; A binary scalar for identifying trends in user psychological state changes, obtained from the activation function.
[0038] Optionally, when using a two-layer LSTM, the first layer is configured with 128 neurons to extract temporal features, and return_sequences=True indicates that the hidden state of each time step is output; the second layer is configured with 64 neurons to further reduce the dimensionality of the temporal features.
[0039] Optionally, the activation function can be sigmoid, softmax, or tanh, etc.
[0040] Optionally, the user's psychological state feature output is a scalar value representing the user's current psychological state. It can be converted into a value between 0 and 1 using the sigmoid activation function, representing the probability of a certain psychological state, such as the probability of a positive emotion.
[0041] The output of the user's psychological state change trend feature is also a scalar, representing the trend of the user's psychological state change. It can also be converted into a value between 0 and 1 using the sigmoid activation function. For example, if the output is close to 0, it means that the psychological state is stable or tends to be negative; close to 1 means that the psychological state is positive; and a value of 0.5 means that the psychological state has no obvious change.
[0042] The control unit is configured to, in response to a request from a target user and / or to detect that the target user has triggered a preset condition, read the real-time user state feature vector and input the user state feature vector into a generative dialogue model and / or a massage control algorithm to obtain an output for the target user's request.
[0043] Optional, preset conditions, including at least one of the following: Abnormal physiological parameters of the target user were detected; Anomalies were detected in the target user's behavioral data; Abnormal psychological state characteristics of the target user were detected; Abnormal trends in the psychological state of target users were detected.
[0044] Optionally, the above-mentioned anomaly may refer to the relevant parameters or features exceeding the system's default value, or it may refer to the relevant parameters or features exceeding the user-defined value in the system.
[0045] Optionally, when the preset condition is that a user falls, it also includes triggering at least one of the following: a local alarm, a remote notification, or an automatic emergency call.
[0046] Optionally, in response to a request from the target user and / or upon detecting that the target user has triggered a preset condition, a real-time user state feature vector is read, and the user state feature vector is input into a generative dialogue model and / or a massage control algorithm to obtain an output tailored to the target user's request, including: In response to the target user's dialogue request, read the real-time user state feature vector; Generate a state prompt based on the real-time user state feature vector; Input the target user's dialogue request and state prompt into the generative dialogue model to obtain the output for the target user's dialogue request.
[0047] Specifically, when the target user initiates a dialogue request and inputs statement A, a state prompt is generated based on the real-time user state feature vector, such as "The current user state is good, it's okay to joke around." This prompt is added to statement A and input into the generative dialogue model to obtain the output tailored to the target user's dialogue request. Using this method, the output of the generative dialogue model can more closely reflect the target user's psychological state.
[0048] Optionally, based on the real-time user state feature vector, the method for generating the state Prompt can be through a Prompt template. For example, the Prompt template could be "The current user's psychological state is __, please communicate with the user according to their psychological state," or "The current user's psychological state is __, please guide the user's psychological state to improve," or "The current user's psychological state is __, __ joking with the customer," etc., where the content of "__" is obtained based on the real-time user state feature vector. Optionally, the content of the placeholder "__" can be obtained based on the real-time user state feature vector. This can be done by directly filling in the user state feature vector into "__," or by filling in keywords mapped from the real-time user state feature vector into "__."
[0049] Optionally, a state Prompt is generated based on the real-time user state feature vector, including: The Prompt mapping corresponding to the target user is obtained based on the changes in the user's state during the historical dialogue between the target user and the generative dialogue model. A state prompt is generated based on the real-time user state feature vector and the Prompt mapping.
[0050] Optionally, methods for obtaining the Prompt mapping corresponding to the target user based on changes in user state during the historical dialogue between the target user and the generative dialogue model include: Based on the target user's historical dialogue with the generative dialogue model, obtain the historical Prompt used in the target user's historical dialogue with the generative dialogue model; Based on the target user's historical dialogue with the generative dialogue model, obtain historical user state change data corresponding to the historical dialogue; Based on the historical Prompt and the historical user status change data, analyze the impact of each keyword in the historical Prompt on the historical user status change; Based on the impact of each keyword in the historical Prompt on changes in the historical user's state, a Prompt mapping corresponding to the target user is obtained.
[0051] Optionally, methods for generating the state prompt based on the real-time user state feature vector and the Prompt mapping include: Based on the real-time user state feature vector and the expected user state feature vector, a state Prompt is generated using the Prompt mapping. The expected user state feature vector can be a default value or set by the user.
[0052] Optionally, methods for analyzing the impact of each keyword in the historical prompt on changes in historical user status include PCA analysis, linear regression, etc.
[0053] Optionally, in response to a request from the target user and / or upon detecting that the target user has triggered a preset condition, a real-time user state feature vector is read, and the user state feature vector is input into a generative dialogue model and / or a massage control algorithm to obtain an output tailored to the target user's request, including: In response to a massage request from a target user and / or upon detecting that the target user has triggered a preset condition, the system reads the real-time user state feature vector. The massage target is obtained based on the real-time user status feature vector; Based on the massage objectives, a massage control algorithm is used to control the massage device installed on the humanoid robot.
[0054] Example 3 A control method for an emotional companion humanoid robot based on multimodal interaction, comprising: Obtain physiological parameters and / or behavioral data of the target user; Based on physiological parameters and / or behavioral data, user state feature vectors are obtained through a user state recognition model. In response to a request from a target user and / or upon detecting that the target user has triggered a preset condition, the system reads the real-time user state feature vector and inputs the user state feature vector into a generative dialogue model and / or a massage control algorithm to obtain an output tailored to the target user's request.
[0055] Optionally, physiological parameters and / or behavioral data of the target user may be obtained through at least one of the following three methods: Physiological parameters and / or behavioral data of the target user are acquired by sensors installed on the humanoid robot device; Physiological parameters and / or behavioral data of the target user are acquired through sensors worn by the target user and communicating with the humanoid robot device; Physiological parameters and / or behavioral data of the target user are acquired through sensors worn by the target user and communicating with the Internet.
[0056] Optionally, the data acquisition unit can acquire the target user's physiological parameters and / or behavioral data from sensors on the humanoid robot device via a data bus, or acquire the target user's physiological parameters and / or behavioral data from other devices via near-field communication or network communication. Other devices can be dedicated extension devices, or wearable devices or smart devices such as smartwatches with physiological detection functions. Specifically, camera data of target users and their interactions with the environment can be obtained by installing cameras on humanoid robot devices; Blood oxygen saturation data and blood glucose data can be obtained by using an impedance spectral blood glucose monitoring sensor installed on a humanoid robot device; Body temperature data, respiratory rate data, etc. can be obtained by millimeter-wave vital sign sensors installed on humanoid robot devices; Optionally, heart rate data, blood pressure data, etc., can be obtained by communicating with the smartwatch worn by the target user.
[0057] Optionally, the user state recognition model is configured as follows: Temporal feature data is obtained by extracting local features of the target user's physiological parameters and / or behavioral data through a pre-model; The user state feature vector is obtained by analyzing the time series feature data through the LSTM model. The user state feature vector includes user psychological state features and user psychological state change trend features. The pre-model and LSTM model belong to the user state recognition model.
[0058] Optionally, the LSTM model is configured to include: At least two LSTM layers are used to extract time-dependent features from time-series data; Fully connected layer used to output the final user mental state feature vector; Activation functions used to represent probabilities or state strengths; The output of the LSTM model is configured to include: A binary scalar obtained from the activation function to represent the characteristics of a user's psychological state; A binary scalar for identifying trends in user psychological state changes, obtained from the activation function.
[0059] Optionally, when using a two-layer LSTM, the first layer is configured with 128 neurons to extract temporal features, and return_sequences=True indicates that the hidden state of each time step is output; the second layer is configured with 64 neurons to further reduce the dimensionality of the temporal features.
[0060] Optionally, the activation function can be sigmoid, softmax, or tanh, etc.
[0061] Optionally, the user's psychological state feature output is a scalar value representing the user's current psychological state. It can be converted into a value between 0 and 1 using the sigmoid activation function, representing the probability of a certain psychological state, such as the probability of a positive emotion.
[0062] The output of the user's psychological state change trend feature is also a scalar, representing the trend of the user's psychological state change. It can also be converted into a value between 0 and 1 using the sigmoid activation function. For example, if the output is close to 0, it means that the psychological state is stable or tends to be negative; close to 1 means that the psychological state is positive; and a value of 0.5 means that the psychological state has no obvious change.
[0063] Specifically, the framework of an LSTM model can be represented as: import tensorflow as tf from tensorflow.keras import layers, models # Define the model architecture def build_lstm_model(input_shape): model = models.Sequential() # LSTM layer: Used to process time-series data and capture time-dependent features model.add(layers.LSTM(128, activation='relu', input_shape=input_shape, return_sequences=True)) model.add(layers.LSTM(64, activation='relu', return_sequences=False)) # Output of User Psychological State Characteristics model.add(layers.Dense(32, activation='relu')) model.add(layers.Dense(1, activation='sigmoid', name="psychological_state")) # Output of User Psychological State Change Trend Characteristics model.add(layers.Dense(32, activation='relu')) model.add(layers.Dense(1, activation='sigmoid', name="psychological_trend")) return model # Assume the input shape is (time step, number of features). input_shape = (100, 10) # For example, 100 time steps, 10 features per time step model = build_lstm_model(input_shape) # Model Summary model.summary().
[0064] Optional, preset conditions, including at least one of the following: Abnormal physiological parameters of the target user were detected; Anomalies were detected in the target user's behavioral data; Abnormal psychological state characteristics of the target user were detected; Abnormal trends in the psychological state of target users were detected.
[0065] Optionally, the above-mentioned anomaly may refer to the relevant parameters or features exceeding the system's default value, or it may refer to the relevant parameters or features exceeding the user-defined value in the system.
[0066] Optionally, when the preset condition is that a user falls, it also includes triggering at least one of the following: a local alarm, a remote notification, or an automatic emergency call.
[0067] Optionally, in response to a request from the target user and / or upon detecting that the target user has triggered a preset condition, a real-time user state feature vector is read, and the user state feature vector is input into a generative dialogue model and / or a massage control algorithm to obtain an output tailored to the target user's request, including: In response to the target user's dialogue request, read the real-time user state feature vector; Generate a state prompt based on the real-time user state feature vector; Input the target user's dialogue request and state prompt into the generative dialogue model to obtain the output for the target user's dialogue request.
[0068] Specifically, when the target user initiates a dialogue request and inputs statement A, a state prompt is generated based on the real-time user state feature vector, such as "The current user state is good, it's okay to joke around." This prompt is added to statement A and input into the generative dialogue model to obtain the output tailored to the target user's dialogue request. Using this method, the output of the generative dialogue model can more closely reflect the target user's psychological state.
[0069] Optionally, based on the real-time user state feature vector, the method for generating the state Prompt can be through a Prompt template. For example, the Prompt template could be "The current user's psychological state is __, please communicate with the user according to their psychological state," or "The current user's psychological state is __, please guide the user's psychological state to improve," or "The current user's psychological state is __, __ joking with the customer," etc., where the content of "__" is obtained based on the real-time user state feature vector. Optionally, obtaining "__" based on the real-time user state feature vector can be done by directly filling in the user state feature vector into "__," or by filling in keywords mapped from the real-time user state feature vector into "__."
[0070] Optionally, a state Prompt is generated based on the real-time user state feature vector, including: The Prompt mapping corresponding to the target user is obtained based on the changes in the user's state during the historical dialogue between the target user and the generative dialogue model. A state prompt is generated based on the real-time user state feature vector and the Prompt mapping.
[0071] Optionally, methods for obtaining the Prompt mapping corresponding to the target user based on changes in user state during the historical dialogue between the target user and the generative dialogue model include: Based on the target user's historical dialogue with the generative dialogue model, obtain the historical Prompt used in the target user's historical dialogue with the generative dialogue model; Based on the target user's historical dialogue with the generative dialogue model, obtain historical user state change data corresponding to the historical dialogue; Based on the historical Prompt and the historical user status change data, analyze the impact of each keyword in the historical Prompt on the historical user status change; Based on the impact of each keyword in the historical Prompt on changes in the historical user's state, a Prompt mapping corresponding to the target user is obtained.
[0072] Optionally, methods for generating the state prompt based on the real-time user state feature vector and the Prompt mapping include: Based on the real-time user state feature vector and the expected user state feature vector, a state Prompt is generated using the Prompt mapping. The expected user state feature vector can be a default value or set by the user.
[0073] Optionally, methods for analyzing the impact of each keyword in the historical prompt on changes in historical user status include PCA analysis, linear regression, etc.
[0074] Optionally, in response to a request from the target user and / or upon detecting that the target user has triggered a preset condition, a real-time user state feature vector is read, and the user state feature vector is input into a generative dialogue model and / or a massage control algorithm to obtain an output tailored to the target user's request, including: In response to a massage request from a target user and / or upon detecting that the target user has triggered a preset condition, the system reads the real-time user state feature vector. The massage target is obtained based on the real-time user status feature vector; Based on the massage objectives, a massage control algorithm is used to control the massage device installed on the humanoid robot.
[0075] Example 4 This embodiment provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement any of the methods described above.
[0076] Specifically, such as Figure 2 As shown, Figure 2This is a schematic diagram of the structure of a computer device according to this application. The computer device may include: a processor 101, such as a central processing unit (CPU), a communication bus 102, a user interface 104, a network interface 103, and a memory 105. The communication bus 102 is used to enable communication between these components. The user interface 104 may include a display screen and an input unit such as a keyboard. Optionally, the user interface 104 may also include a standard wired interface or a wireless interface. The network interface 103 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 105 may be a storage device independent of the aforementioned processor 101. The memory 105 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as at least one disk storage device. The processor 101 may be a general-purpose processor, including a central processing unit, a network processor, etc., or it may be a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component.
[0077] Those skilled in the art will understand that the appendix Figure 2 The structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0078] like Figure 2 As shown, the memory 105, which serves as a storage medium, may include an operating system, a network communication module, a user interface module, and an application program for implementing a control method for an emotional companion humanoid robot based on multimodal interaction.
[0079] exist Figure 2 In the electronic device shown, the network interface 103 is mainly used for data communication with the network server; the user interface 104 is mainly used for data interaction with the user; the processor 101 and the memory 105 in this application can be set in the electronic device, and the electronic device can call the application program stored in the memory 105 through the processor 101 to implement a control method for an emotional companion humanoid robot based on multimodal interaction to implement the above method.
[0080] Example 5 This embodiment provides a computer-readable storage medium on which a computer program is stored, and a processor executes the computer program to implement any of the methods described above.
[0081] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a device including one or any combination of the above-mentioned memories. The computer may be a variety of computing devices, including smart terminals and servers.
[0082] In the above embodiments of this disclosure, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0083] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0084] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0085] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0086] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned non-volatile storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0087] The above are merely preferred embodiments of this disclosure. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this disclosure, and these improvements and modifications should also be considered within the scope of protection of this disclosure.
Claims
1. A humanoid robot system for emotional companionship based on multimodal interaction, characterized in that, The system is connected to a humanoid robot device, and includes a data acquisition unit, a data processing unit, and a control unit. The data acquisition unit is configured to acquire physiological parameters and / or behavioral data of the target user. The data processing unit is configured to obtain a user state feature vector based on the physiological parameters and / or behavioral data through a user state recognition model. The control unit is configured to, in response to a request from the target user and / or to detecting that the target user has triggered a preset condition, read a real-time user state feature vector and input the user state feature vector into a generative dialogue model and / or a massage control algorithm to obtain an output for the request from the target user.
2. The emotional companion robot system based on multimodal interaction according to claim 1, characterized in that, The data acquisition unit is configured to acquire physiological parameters and / or behavioral data of the target user through at least one of the following three methods: Physiological parameters and / or behavioral data of the target user are acquired by sensors installed on the humanoid robot device; Physiological parameters and / or behavioral data of the target user are acquired through sensors worn by the target user and communicating with the humanoid robot device; Physiological parameters and / or behavioral data of the target user are acquired through sensors worn by the target user and communicating with the Internet.
3. The emotional companion robot system based on multimodal interaction according to claim 1, characterized in that, The user state recognition model is configured as follows: Temporal feature data is obtained by extracting local features of the target user's physiological parameters and / or behavioral data through a pre-model; The time-series feature data is analyzed using an LSTM model to obtain a user state feature vector, which includes user psychological state features and user psychological state change trend features. The pre-model and the LSTM model belong to the user state recognition model.
4. The emotional companion robot system based on multimodal interaction according to claim 3, characterized in that, The LSTM model is configured to include: At least two LSTM layers are used to extract the time-dependent features of the time-series feature data; Fully connected layer used to output the final user mental state feature vector; Activation functions used to represent probabilities or state strengths; The output of the LSTM model is configured to include: A binary scalar obtained from the activation function to represent the characteristics of a user's psychological state; A binary scalar for identifying trends in user psychological state changes, obtained from the activation function.
5. The emotional companionship humanoid robot system based on multimodal interaction according to claim 1, characterized in that, The step of responding to the request of the target user and / or detecting that the target user has triggered a preset condition, reading the real-time user state feature vector, and inputting the user state feature vector into a generative dialogue model and / or massage control algorithm to obtain an output for the request of the target user includes: In response to the dialogue request from the target user, read the real-time user state feature vector; Based on the real-time user state feature vector, a state Prompt is generated; The target user's dialogue request and the state Prompt are input into the generative dialogue model to obtain an output for the target user's dialogue request.
6. The emotional companionship humanoid robot system based on multimodal interaction according to claim 5, characterized in that, The step of generating the state Prompt based on the real-time user state feature vector includes: Based on the changes in user state during the historical dialogue between the target user and the generative dialogue model, a Prompt mapping corresponding to the target user is obtained; A state Prompt is generated based on the real-time user state feature vector and the Prompt mapping.
7. The emotional companionship humanoid robot system based on multimodal interaction according to claim 1, characterized in that, The step of responding to the request of the target user and / or detecting that the target user has triggered a preset condition, reading the real-time user state feature vector, and inputting the user state feature vector into a generative dialogue model and / or massage control algorithm to obtain an output for the request of the target user includes: In response to the massage request from the target user and / or upon detecting that the target user has triggered a preset condition, the user's real-time state feature vector is read. The massage target is obtained based on the real-time user state feature vector; Based on the massage target, a massage control algorithm is used to control the massage device installed on the humanoid robot device.
8. The emotional companionship humanoid robot system based on multimodal interaction according to claim 1, characterized in that, The preset conditions include at least one of the following conditions: Abnormal physiological parameters of the target user were detected; Abnormal behavioral data of the target user was detected; Abnormal psychological state characteristics of the target user were detected. Abnormal trends in the psychological state of the target user were detected.
9. A control method for an emotional companion humanoid robot based on multimodal interaction, characterized in that, include: Obtain physiological parameters and / or behavioral data of the target user; Based on the physiological parameters and / or behavioral data, a user state feature vector is obtained through a user state recognition model; In response to the request of the target user and / or the detection that the target user has triggered a preset condition, the real-time user state feature vector is read and the user state feature vector is input into the generative dialogue model and / or massage control algorithm to obtain the output for the request of the target user.
10. A device, characterized in that, The device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to claim 9.