Accompanying robot control method and device, computer equipment and storage medium

By collecting multimodal emotional and environmental information and combining it with a pre-trained model for multi-source information fusion analysis, precise control commands are generated, solving the problem of inaccurate companion robot services and realizing personalized user interaction and environmental adaptation.

CN121374630APending Publication Date: 2026-01-23GUANGDONG KETYOO INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511880054.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing companion robots struggle to provide accurate services based on the user's actual state and environment, lacking personalization and precision.

Method used

By collecting environmental information, user multimodal emotional information, and user command information, and combining pre-trained emotion analysis models and intent recognition models, multi-source information fusion analysis is performed to generate context-related control commands to drive the companion robot to execute.

Benefits of technology

This enables companion robots to provide accurate and personalized services based on the user's actual state and environment, thereby improving human-computer interaction satisfaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121374630A_ABST
    Figure CN121374630A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an accompanying robot control method and device, computer equipment and a storage medium, and relates to the field of artificial intelligence, and the method comprises the steps: collecting environment information, multi-mode emotion information of a user, and user instruction information; identifying an emotional state of the user based on the multi-mode emotional information; identifying a user intention based on the user instruction information; and performing multi-source information fusion analysis based on the environment information, the emotional state and the user intention, generating a control instruction, and controlling an accompanying robot to execute the control instruction. According to the method, the limitation of a single information source is broken through, the problems that emotion recognition is easily interfered and intention understanding is rigid are solved, and the crossing from passive response to active understanding is realized, so that the accompanying robot can be controlled to provide accurate service according to the actual state and the actual environment condition of the user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a companion robot control method and device, computer equipment and a storage medium. BACKGROUND

[0002] A companion robot is an intelligent device that provides emotional support and daily companionship to users through interactive behaviors. With the development of social economy and the improvement of people's living standards, the application of companion robots in family, elderly care and education scenarios has gradually become popular, and the market demand continues to grow.

[0003] In the prior art, a companion robot usually adopts a fixed programmed interaction mode, such as a response mechanism based on simple speech recognition or pre-set actions. These functions can only achieve basic conversation or entertainment interaction, and the services provided lack personalization and precision. For example, when a user is in different emotional or complex environmental conditions, the existing robot cannot recognize these factors and make adaptive interaction, thereby affecting the service effect and user experience. This limitation makes it difficult for companion robots to fully meet the diverse needs of users in actual applications.

[0004] Therefore, there is a core technical problem in the prior art: it is difficult for a companion robot to provide accurate services according to the actual state of a user and the actual environmental conditions. SUMMARY

[0005] The embodiments of the present application provide a companion robot control method, device, computer equipment and storage medium, aiming to solve the problem that existing companion robots are difficult to provide accurate services according to the actual state of a user and the actual environmental conditions.

[0006] In a first aspect, the embodiments of the present application provide a companion robot control method, which includes:

[0007] Collecting environmental information, multi-modal emotional information of a user and user instruction information;

[0008] Identifying the emotional state of the user based on the multi-modal emotional information;

[0009] Identifying the user's intention based on the user instruction information;

[0010] Performing multi-source information fusion analysis based on the environmental information, the emotional state and the user's intention, generating a control instruction, and controlling the companion robot to execute the control instruction.

[0011] Optionally, the multi-modal emotional information includes user voice information, user facial expression information and user body movement information, and the identification of the emotional state of the user based on the multi-modal emotional information includes:

[0012] extracting a voice emotion feature from the user voice information, extracting an expression emotion feature from the user facial expression information, and extracting a motion emotion feature from the user body action information;

[0013] fusing the voice emotion feature, the expression emotion feature, and the motion emotion feature to obtain a fused emotion feature;

[0014] analyzing, by a pre-trained emotion analysis model, based on the fused emotion feature, and outputting the emotion state.

[0015] Optionally, the identifying a user intention based on the user instruction information comprises:

[0016] obtaining context information of the user instruction information;

[0017] analyzing, by a pre-trained intention recognition model, based on the user instruction information and the context information, and outputting the user intention.

[0018] Optionally, performing multi-source information fusion analysis based on the environment information, the emotion state, and the user intention to generate a control instruction comprises:

[0019] extracting an environment feature from the environment information, fusing the environment feature, the emotion state, and the user intention to obtain a multi-source information fusion feature;

[0020] analyzing, by a pre-trained instruction generation model, based on the multi-source information fusion feature, and outputting the control instruction.

[0021] Optionally, the method further comprises:

[0022] monitoring a preset scene mode;

[0023] if the preset scene mode is monitored, executing a preset control instruction set corresponding to the scene mode.

[0024] Optionally, the scene mode comprises a night mode and a leaving-home mode.

[0025] the preset control instruction set corresponding to the night mode comprises at least one of adjusting lighting to a preset intensity and playing a preset white noise audio;

[0026] the preset control instruction set corresponding to the leaving-home mode comprises at least one of controlling a curtain and lighting to be off and starting security alert monitoring.

[0027] Optionally, the method further comprises:

[0028] collecting state data of a user;

[0029] inference a potential demand of the user based on the state data through a pre-trained personalized service prediction model, and perform a personalized service operation corresponding to the potential demand, wherein the personalized service prediction model is trained based on historical data of the user;

[0030] And / or, the method further comprises:

[0031] receiving an environmental parameter sent by an environmental sensor;

[0032] determining a preset adjustment instruction of a target device based on the environmental parameter and a preset adjustment rule, and controlling the target device to execute the adjustment instruction;

[0033] And / or, the method further comprises:

[0034] monitoring a running state of a preset target device;

[0035] if the running state of the target device is a preset target state, issuing a reminder information corresponding to the target state.

[0036] In a second aspect, an embodiment of the present application further provides a companion robot control device, which comprises units for executing the above method.

[0037] In a third aspect, an embodiment of the present application further provides a computer device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the above method when executing the computer program.

[0038] In a fourth aspect, an embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program can implement the above method when being executed by a processor.

[0039] The embodiment of the present application provides a companion robot control method and device, computer equipment and a storage medium. The method comprises: collecting environmental information, multi-modal emotional information of a user and user instruction information; identifying the emotional state of the user based on the multi-modal emotional information; identifying the user intention based on the user instruction information; performing multi-source information fusion analysis based on the environmental information, the emotional state and the user intention, generating a control instruction, and controlling the companion robot to execute the control instruction. The present application synchronously collects environmental information, multi-modal emotional information and user instruction information, constructs a panoramic perception basis, then accurately identifies the emotional state of the user based on multi-modal fusion, understands the user intention in combination with the context, finally generates a context-related control instruction through multi-source information fusion analysis, and drives the companion robot to execute. The method breaks through the limitation of a single information source, overcomes the problems of emotional recognition being easily disturbed and intention understanding being rigid, realizes the leap from passive response to active understanding, and thus can control the companion robot to provide accurate services according to the actual state of the user and the actual environmental conditions. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0041] Figure 1 A flowchart of a companion robot control method provided by the embodiment of the present application is shown in the figure.

[0042] Figure 2 A schematic block diagram of a computer device provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0043] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0044] It should be understood that when used in the specification and the appended claims, the terms "comprise" and "include" indicate the presence of described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or sets thereof.

[0045] It should also be understood that the terms used herein are for the purpose of describing particular embodiments and are not intended to limit the application. As used in the specification and the appended claims, the singular forms "a," "an" and "the" are intended to include plural forms as well, unless the context clearly indicates otherwise.

[0046] It should further be understood that the term "and / or" as used herein refers to any combination of one or more of the associated listed items, and all possible combinations, and includes these combinations.

[0047] As used in the specification and the appended claims, the term "if' can be construed to mean "when" or "once," or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be construed to mean "once it is determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]," depending on the context.

[0048] Referring to Figure 1 The embodiment of the present application provides a companion robot control method, which comprises the following steps:

[0049] Referring to Figure 1 The embodiment of the present application provides a companion robot control method, which comprises the following steps:

[0050] S1, collecting environment information, multi-modal emotional information of a user and user instruction information.

[0051] In the implementation, in the information collection stage, the method synchronously collects environment information, multi-modal emotional information of a user and user instruction information, constructs a comprehensive and stereoscopic perception basis, breaks through the limitation of a traditional robot depending on a single information source (such as only voice), and provides rich and complementary data support for subsequent accurate analysis. For example, the environment information reflects the physical scene in which the user is located, the multi-modal emotional information reveals the inner state of the user, and the user instruction information directly expresses the explicit demand of the user, and synchronous acquisition of the three types of information enables the system to completely depict the user situation from three dimensions of environment, psychology and intention.

[0052] S2, identifying an emotional state of the user based on the multi-modal emotional information.

[0053] In the implementation, in the emotion recognition stage, the user emotional state is recognized based on the multi-modal emotional information, which significantly improves the accuracy and robustness of emotion judgment. Compared with the traditional method of relying only on facial expressions or voice tone, the comprehensive use of various emotional information (such as user voice information, user facial expression information, and user body movement information) for comprehensive analysis can effectively deal with the situation that part of the information is blocked or disturbed in the actual scene, and ensure that the real emotional state of the user can be stably and reliably judged in different environments, which is the basis for realizing emotional accompaniment.

[0054] In some preferred embodiments, the multi-modal emotional information includes user voice information, user facial expression information, and user body movement information. The step of "recognizing the emotional state of the user based on the multi-modal emotional information" specifically includes the following steps: extracting voice emotional features from the user voice information, extracting expression emotional features from the user facial expression information, and extracting action emotional features from the user body movement information; fusing the voice emotional features, the expression emotional features, and the action emotional features to obtain fused emotional features; and analyzing the fused emotional features based on a pre-trained emotional analysis model to output the emotional state.

[0055] In the implementation, the user voice information can be collected by a microphone, and the user facial expression information and the user body movement information can be collected by a camera, which is not specifically limited by the present application. By extracting corresponding emotional features from the user voice information, the user facial expression information, and the user body movement information, respectively, and fusing the multi-modal features, the limitations of single-modal emotion recognition in complex scenes are effectively overcome. For example, when the environmental noise is large, making it difficult to extract voice features, the system can still rely on facial expressions and body movement information to maintain the accuracy of emotion judgment; on the contrary, when the user is facing away from the camera, the voice features can still be used as an effective supplementary judgment basis. This multi-modal complementary mechanism greatly improves the robustness and reliability of emotion state recognition. Further, by comprehensively analyzing the fused emotional features through a pre-trained emotional analysis model, the system can capture more subtle and complex emotional changes, thereby laying a technical foundation for realizing truly personalized emotional care.

[0056] In an embodiment, the specific implementation is as follows:

[0057] First, multimodal emotional features are extracted. For user voice information, speech signal processing techniques are used to extract feature parameters representing emotions from the raw audio data, such as, but not limited to, Mel-frequency cepstral coefficients, fundamental frequency profile, energy distribution, and their temporal dynamic changes. These acoustic features collectively constitute the voice emotional features. Further, for user facial expression information, computer vision techniques are used, such as processing the acquired facial image sequences through convolutional neural networks, to extract feature vectors of facial key point motion units, texture changes in facial features, and overall expression morphology, forming facial expression emotional features. Further, for user body movement information, the coordinates of the user's body skeleton joints are reconstructed from the video stream, and then the joint movement trajectory, movement amplitude, movement speed, and body orientation changes are calculated to extract movement emotional features.

[0058] Next, multimodal emotional features are fused. The extracted speech emotional features, facial expression emotional features, and action emotional features are concatenated and aligned at the feature level to form a unified fused emotional feature vector. During the fusion process, an attention mechanism can be used to adaptively emphasize modal information with higher reliability in specific scenarios. For example, when there is high environmental noise, the weight of speech features can be automatically reduced, and more reliance can be placed on visual features.

[0059] Finally, the pre-trained sentiment analysis model analyzes the fused sentiment features and outputs the final sentiment state classification result. This sentiment analysis model can be, for example, a classifier built based on a deep neural network, such as a multilayer perceptron or a temporal neural network; this invention is not specifically limited to this. During the training phase, the sentiment analysis model undergoes supervised training using a large dataset of labeled multimodal sentiment data, learning a complex mapping relationship from fused sentiment features to discrete sentiment categories or continuous sentiment dimensions. During the inference phase, the real-time acquired fused sentiment features are input into the pre-trained sentiment analysis model, which then outputs a judgment of the user's current sentiment state, such as happy, sad, calm, or anxious.

[0060] S3, Identify the user's intent based on the user instruction information.

[0061] In specific implementation, user command information can be voice commands or text commands, etc., and this invention is not specifically limited. During the intent recognition stage, user intent is identified based on user command information, making human-computer interaction more natural and efficient. Through semantic parsing of user commands, the system can accurately understand the user's directly expressed operational needs, providing a clear direction for subsequent service execution. Furthermore, embodiments of this invention can combine contextual information for intent understanding, enabling the processing of complex commands that are referential or rely on historical dialogue.

[0062] In some preferred embodiments, the step of "identifying the user intent based on the user instruction information" specifically comprises the following steps: obtaining context information of the user instruction information; and analyzing the user instruction information and the context information based on a pre-trained intent recognition model to output the user intent.

[0063] In a specific implementation, the introduction of context information to assist in user intent recognition enables the system to understand user instructions with reference relationships or dependencies on dialogue history. This context-aware capability significantly improves the coherence and logic of the companion robot in multi-turn dialogue, making its interaction closer to natural communication between humans, and effectively solving the problem of rigid dialogue mode and inability to handle complex language scenarios of existing robots.

[0064] In an embodiment, the specific implementation is as follows:

[0065] The context information includes but is not limited to dialogue history records before the current dialogue turn, the interaction state of the user and the companion robot, and the running state of the related intelligent device. Specifically, the system maintains a dynamically updated dialogue context buffer for storing multi-turn dialogue content within a preset time window. Each turn of dialogue content is represented in a structured data, for example, including user statements, system replies, and corresponding timestamps. At the same time, the system continuously monitors the device state variables associated with the current interaction scenario, such as light switch state, media player content, etc., and integrates these state information as part of the environmental context.

[0066] Further, the pre-trained intent recognition model can adopt a neural network model based on the Transformer architecture, such as BERT or its variants, which is not specifically limited by the present application. The input of the intent recognition model can encode the current user instruction information and the related context information at the same time. Specifically, the current user instruction text, the historical dialogue sequence extracted from the dialogue context buffer, and the device state description text are spliced to form a complete input sequence. In the model training stage, the labeled training data is used for supervised pre-training, and the intent recognition model is trained to learn the mapping relationship from the instruction sequence with context to the predetermined intent label, such as "query weather", "control device", "on-demand media", etc., which is not specifically limited by the present application.

[0067] Further, in the reasoning stage, the current collected user instruction information and the obtained context information are combined into an input sequence in the above manner and input into the trained intent recognition model. The intent recognition model captures the association between the current instruction and the context through its multi-layer self-attention mechanism for deep semantic coding of the input sequence. Finally, the output layer of the intent recognition model produces a classification probability distribution of the corresponding user intent, and the system selects the intent category with the highest probability as the final recognized user intent output. In this way, the system can accurately understand the user's real intent in the referential sentence and continuous dialogue that depends on the context.

[0068] S4, based on the environmental information, the emotional state and the user intent, performing multi-source information fusion analysis to generate a control instruction, and controlling the companion robot to execute the control instruction.

[0069] In specific implementation, in the multi-source information fusion analysis and control instruction generation stage, the environmental information, the recognized emotional state and the user intent are fused and analyzed to realize the leap from "mechanical execution of instructions" to "intelligent understanding of the scene and active service". The system does not consider the user's instruction in isolation, but comprehensively considers it in the current environmental condition and the user's emotional background. For example, when the user issues the instruction "too bright", if the system simultaneously recognizes that the user is in a relaxed emotional state and the environmental light is indeed strong, it may generate the instruction "dim the light to a comfortable brightness"; if it recognizes that the user is in a state of impatience, it may simultaneously execute the combined instruction of dimming the light and playing soothing music. This decision mechanism based on multi-source information fusion makes the generated control instruction more context-related and personalized.

[0070] Further, by controlling the companion robot to execute the finally generated control instruction, a complete closed loop from perception, cognition to action is completed, so that the companion robot can provide a highly accurate, context-adaptive and active and natural service experience. The robot is no longer a simple tool, but an intelligent companion that can understand the user's complex state, adapt to environmental changes and make reasonable responses, significantly improving the satisfaction of human-computer interaction and the user's quality of life in the smart home environment.

[0071] In some preferred embodiments, the above step of "performing multi-source information fusion analysis based on the environmental information, the emotional state and the user intent to generate a control instruction" specifically includes the following steps: collecting environmental features from the environmental information; fusing the environmental features, the emotional state and the user intent to obtain multi-source information fusion features; and analyzing the multi-source information fusion features based on a pre-trained instruction generation model to output the control instruction.

[0072] In a specific implementation, the environment information includes at least one of video monitoring data, environment monitoring data, user setting information for the companion robot, user operation information for the companion robot, and alarm information. By including the video monitoring data, the environment monitoring data, the user setting information for the companion robot, the user operation information for the companion robot, and the alarm information in the category of the environment information, a comprehensive and three-dimensional environment perception system is constructed. The embodiment of the application can not only perceive physical environment parameters such as temperature and humidity, illumination, but also capture user behavior patterns and device state changes. At the same time, the access of the alarm information enables the system to have the ability to respond to emergencies in a timely manner. This all-round information integration effectively breaks the information silos among intelligent devices and provides a data foundation for realizing truly scenario-based intelligent services. It should be noted that the alarm information can be sent by a preset target device (for example, an environment monitoring device), and the application does not specifically limit it.

[0073] Further, by fusing the environment features, the emotional state, and the user intention to form a multi-source information fusion feature, and outputting a control instruction based on a pre-trained instruction generation model, a cross from “single instruction-action” mapping to “context-decision” generation is realized. For example, when the user intention is “reading”, the emotional state is “relaxation”, and the environment light is “dim”, the system may generate a composite control instruction of “turning on the reading light and playing light music”, instead of simply executing “turning on the light”. This decision mechanism based on multi-source information fusion enables the companion robot to provide more intelligent and personalized services, significantly improving the overall perception and response capability of the system.

[0074] In an embodiment, the specific implementation is as follows:

[0075] The collected environment information is subjected to feature extraction and standardization processing. Specifically, the environment information includes but is not limited to numerical data obtained through environment sensors, scene image data obtained through visual sensors, user setting information for the companion robot, user operation information for the companion robot, and alarm information. For numerical environment data, standardization preprocessing is performed and its time sequence features are extracted; for scene image data, a pre-trained convolutional neural network is used to extract its high-level semantic features; and for the setting information, the operation information, and the alarm information, a corresponding feature vector is encoded based on a preset encoding rule. These processed features jointly constitute an environment feature vector to represent the current environment state in a structured form.

[0076] Further, the above-mentioned environmental feature vector is fused with the encoded emotional state category and user intention category at the feature level. The specific fusion methods include but are not limited to feature splicing, weighted summation or attention mechanism-based fusion method, which is not specifically limited by the present application. In this process, a preset weight coefficient is assigned to different types of information to reflect its relative importance in the decision-making process, thereby generating a multi-source information fusion feature that can fully reflect the current user demand, emotional tendency and environmental conditions.

[0077] Further, the pre-trained instruction generation model is used to analyze the multi-source information fusion feature and output the final control instruction. The instruction generation model can adopt a deep neural network architecture, preferably based on an encoder-decoder framework, which is not specifically limited by the present application. The encoder is responsible for deep semantic encoding of the input multi-source information fusion feature, and the decoder generates the corresponding control instruction sequence according to the encoding result. The instruction generation model is supervised trained using a labeled data set containing multi-source information and corresponding control instructions in the training phase, learning the mapping relationship from complex context information to reasonable control instructions. In the inference phase, the real-time acquired multi-source information fusion feature is input into the trained instruction generation model, and the control instruction adapted to the current scene can be output, which can include specific operation types, target devices and corresponding parameters, which are not specifically limited by the present application.

[0078] In some preferred embodiments, the method further comprises the following steps: monitoring a preset scene mode; if the preset scene mode is monitored, executing a preset control instruction set corresponding to the scene mode.

[0079] In specific implementation, real-time environmental sensor data, user location information and system clock information are acquired. The environmental sensor data includes ambient light intensity and sound decibel value; the user location information is acquired through the geographic fence technology or indoor positioning system; and the system clock information provides the current time point and date information. The above-mentioned information is matched and calculated with the trigger condition of the predefined scene mode, which includes but is not limited to the ambient light intensity being lower than the threshold value and the system time being in the night period, or the user mobile device location exceeding the preset geographic fence range and continuously exceeding the timeout threshold. When any of the trigger conditions of the scene mode is identified, it is determined that the scene mode is monitored, and further, a preset control instruction set corresponding to the scene mode is executed.

[0080] In the embodiments of the present application, by monitoring the preset scene mode and automatically executing the corresponding preset control instruction set, the standardization and automation switching of the service mode of the companion robot are realized. For example, the system automatically switches to the night mode after sunset and performs operations such as reducing the brightness of the lighting, playing sleep-aiding music, and the like, without the need for the user to repeatedly issue instructions. This mode control not only reduces the operation burden of the user, but also guarantees the consistency and timeliness of the service experience. Especially for the elderly or the disabled, this function can provide unobtrusive and continuous environmental support, significantly improving the practicality and user dependence of the companion robot.

[0081] In some preferred embodiments, the scene mode includes a night mode and an away-from-home mode; the preset control instruction set corresponding to the night mode includes at least one of adjusting the lighting to a preset intensity and playing a preset white noise audio; and the preset control instruction set corresponding to the away-from-home mode includes at least one of controlling the window curtain and the light to be off and starting security alert monitoring.

[0082] In specific implementation, by specifically defining the control instruction set in the night mode and the away-from-home mode, the scene service has clear operation instructions and function guarantees. In the night mode, adjusting the lighting and playing white noise audio helps to create a comfortable sleep environment, which meets the needs of human physiological rhythms; and in the away-from-home mode, automatically turning off the window curtain and the light and starting security alert effectively improve the safety and energy utilization efficiency of the residence. This fine design for specific scenes makes the companion robot truly become a reliable environmental manager and safety guardian in the user's life.

[0083] In some preferred embodiments, the method further includes the steps of: collecting state data of the user; and inferring potential needs of the user based on the state data by using a pre-trained personalized service prediction model, and performing personalized service operations corresponding to the potential needs, wherein the personalized service prediction model is trained based on historical data of the user.

[0084] In specific implementation, by inferring potential needs based on the state data of the user and the historical training data and performing personalized service operations, the companion robot has forward-looking service capabilities. For example, the system can actively remind the user to pay attention to warmth when it detects rain, or even turn on the heating equipment in advance, by analyzing the historical data of the user and finding that the user has joint discomfort on rainy days. For example, when it detects that the user enters the balcony in the morning, the companion robot automatically slowly opens the window curtain, plays soft music, and broadcasts the weather and news summary of the day according to the historical habits of the user.

[0085] This transformation from "responding to commands" to "predicting needs" greatly improves the accuracy of services and user satisfaction. Especially for users who need long-term care, this function can form a continuously optimized personalized service loop, thereby establishing a deeper trust and dependence relationship.

[0086] In an embodiment, the specific implementation is as follows:

[0087] The state data includes but is not limited to the user's current behavior characteristics, real-time physiological parameters, and environmental interaction data, etc., which are not specifically limited by the present application. The behavior trajectory data of the user is continuously collected by multiple source heterogeneous sensors, real-time physiological parameters are collected, and the interaction data of the user with the accompanying robot and smart home equipment is recorded. The state data is processed by data cleaning and feature engineering, and is normalized and standardized to form a standardized user state feature vector.

[0088] Further, the personalized service prediction model adopts a sequence neural network architecture based on deep learning (for example, a long short-term memory network or a Transformer encoder structure, which is not specifically limited by the present application) to adapt to the time sequence characteristics of the user state data. The training process of the personalized service prediction model is based on the historical state data of the user and the corresponding service feedback records to construct a training sample set, and the model learns the mapping relationship from the user state features to the potential service needs through supervised learning.

[0089] Further, the user state feature vector collected and preprocessed in real time is input into the trained personalized service prediction model, and the personalized service prediction model outputs the probability distribution of each type of preset service need through forward calculation. The system filters out the high-probability potential needs according to the preset confidence threshold. Subsequently, according to the identified potential needs, the corresponding operation instruction sequence is matched from the pre-defined service operation library, and the related equipment is controlled to execute these personalized service operations. The whole system also establishes a continuous learning mechanism, and updates the parameters of the personalized service prediction model according to the explicit feedback of the user to the service and the implicit behavior response, to realize the continuous optimization of the service strategy.

[0090] In some preferred embodiments, the method further comprises the steps of: receiving an environmental parameter sent by an environmental sensor; determining a preset adjustment instruction of a target device based on the environmental parameter and a preset adjustment rule, and controlling the target device to execute the adjustment instruction.

[0091] In specific implementation, the smart linkage between the companion robot and the home environment is realized by receiving the environmental sensor parameters and automatically adjusting the target device based on preset rules. For example, when the environmental sensor detects that the indoor PM2.5 concentration exceeds the standard, the system automatically starts the air purifier; when the temperature sensor detects that the temperature is too high, the system links the air conditioner to cool down; the ambient light sensor can detect the ambient brightness, and the system adaptively adjusts the brightness of the robot lighting based on the ambient brightness. This rule-based automatic control not only improves the quality of the living environment in a timely manner, but also liberates the user from the tedious device operation, especially suitable for the elderly who are not familiar with the operation of scientific and technological devices, providing them with a consistent comfortable living experience.

[0092] In some preferred embodiments, the method further comprises the steps of: monitoring the running state of the preset target device; and issuing reminder information corresponding to the target state if the running state of the target device is the preset target state.

[0093] In specific implementation, the device management and reminder assistance functions of the companion robot are realized by monitoring the running state of the target device and actively issuing reminders when it reaches the preset target state. For example, when the system detects that the washing machine has completed the washing, it will remind the user in time to dry the clothes; when the filter life of the water purifier is about to expire, the system will push a replacement reminder in advance. This kind of function effectively solves the problem of low device use efficiency caused by the user's busy or forgetfulness, especially providing reliable home affair management support for forgetful elderly users.

[0094] In some preferred embodiments, the companion robot also has an emergency help function, specifically, when detecting user abnormalities, it can send alarm information to the preset contact person in time and wait for the arrival of the rescuer.

[0095] Further, the companion robot comprises a vehicle body base, a power drive system, an obstacle avoidance system and a robot arm. The vehicle body base is installed with a high-definition camera and a multi-microphone array, the multi-microphone array is composed of 8 microphone units, and 4G communication technology is used for data transmission.

[0096] The companion robot is configured with a communication module which supports multiple communication protocols such as WiFi and Bluetooth, realizing seamless connection with smart home devices. The communication module also integrates a smart home master control system, which can realize intelligent control of devices such as balcony clothes drying machine, lighting, washing machine and curtain.

[0097] The companion robot is configured with a sensor module, which includes an ambient light sensor, a temperature and humidity sensor, an air quality sensor, etc., for monitoring the surrounding environmental parameters and performing corresponding control tasks.

[0098] An embodiment of the present application provides a companion robot control method, comprising: collecting environment information, multi-modal emotional information of a user, and user instruction information; identifying an emotional state of the user based on the multi-modal emotional information; identifying a user intention based on the user instruction information; performing multi-source information fusion analysis based on the environment information, the emotional state, and the user intention, generating a control instruction, and controlling the companion robot to execute the control instruction. The present application synchronously collects environment information, multi-modal emotional information, and user instruction information, constructs a panoramic perception basis, then accurately identifies the emotional state of the user based on multi-modal fusion, understands the user intention in combination with a context, finally generates a context-related control instruction through multi-source information fusion analysis to drive the companion robot to execute. The method breaks through the limitation of a single information source, overcomes the problems of emotional recognition being easily disturbed and intention understanding being rigid, realizes a leap from passive response to active understanding, and thus can control the companion robot to provide accurate services according to the actual state of the user and the actual environment.

[0099] Corresponding to the above companion robot control method, the present application also provides a companion robot control device. The companion robot control device comprises units for executing the above-mentioned companion robot control method, and the companion robot control device can be configured in a desktop computer, a tablet computer, a laptop computer, or the like terminal. Specifically, the companion robot control device comprises:

[0100] a collection unit for collecting environment information, multi-modal emotional information of a user, and user instruction information;

[0101] a first identification unit for identifying an emotional state of the user based on the multi-modal emotional information;

[0102] a second identification unit for identifying a user intention based on the user instruction information;

[0103] a generation unit for performing multi-source information fusion analysis based on the environment information, the emotional state, and the user intention, generating a control instruction, and controlling the companion robot to execute the control instruction.

[0104] In some preferred embodiments, the multi-modal emotional information comprises user voice information, user facial expression information, and user body movement information, and the identifying an emotional state of the user based on the multi-modal emotional information comprises:

[0105] extracting voice emotional features from the user voice information, extracting expression emotional features from the user facial expression information, and extracting movement emotional features from the user body movement information;

[0106] fusing the voice emotional features, the expression emotional features, and the movement emotional features to obtain fused emotional features;

[0107] analyzing, by a pre-trained sentiment analysis model, based on the fused sentiment features, to output the sentiment state.

[0108] In some preferred embodiments, the identifying a user intention based on the user instruction information comprises:

[0109] obtaining context information of the user instruction information;

[0110] analyzing, by a pre-trained intention recognition model, based on the user instruction information and the context information, to output the user intention.

[0111] In some preferred embodiments, the multi-source information fusion analysis based on the environment information, the sentiment state and the user intention to generate a control instruction comprises:

[0112] collecting environment features from the environment information, fusing the environment features, the sentiment state and the user intention to obtain multi-source information fusion features;

[0113] analyzing, by a pre-trained instruction generation model, based on the multi-source information fusion features, to output the control instruction.

[0114] In some preferred embodiments, the method further comprises:

[0115] a mode monitoring unit configured to monitor a preset scene mode;

[0116] an execution unit configured to execute a preset control instruction set corresponding to the scene mode if the preset scene mode is monitored.

[0117] In some preferred embodiments, the scene mode comprises a night mode and a leaving-home mode.

[0118] The preset control instruction set corresponding to the night mode comprises at least one of adjusting lighting to a preset intensity and playing a preset white noise audio.

[0119] The preset control instruction set corresponding to the leaving-home mode comprises at least one of controlling a curtain and lighting to be off and starting security alert monitoring.

[0120] In some preferred embodiments, the method further comprises:

[0121] a state collecting unit configured to collect state data of a user;

[0122] An inference unit is configured to infer potential demands of the user based on the state data by using a pre-trained personalized service prediction model, and perform a personalized service operation corresponding to the potential demands, wherein the personalized service prediction model is trained based on historical data of the user.

[0123] And / or, further comprising:

[0124] A receiving unit is configured to receive environment parameters sent by an environment sensor.

[0125] A determining unit is configured to determine a preset adjustment instruction of a target device based on the environment parameters and a preset adjustment rule, and control the target device to execute the adjustment instruction.

[0126] And / or, further comprising:

[0127] A state monitoring unit is configured to monitor a running state of a preset target device.

[0128] A reminding unit is configured to issue reminding information corresponding to a preset target state if the running state of the target device is the preset target state.

[0129] It should be noted that the specific implementation process of the accompanying robot control device and each unit can be clearly understood by those skilled in the art, which can be referred to the corresponding description in the foregoing method embodiments. For the convenience and brevity of description, it will not be repeated here.

[0130] The accompanying robot control device can be realized in the form of a computer program, which can run on a computer device as shown in Figure 2 .

[0131] Please refer to Figure 2 , Figure 2 is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a terminal or a server, wherein the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, a wearable device, and an electronic device with a communication function. The server can be a stand-alone server or a server cluster composed of multiple servers.

[0132] The computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501, wherein the memory can include a non-volatile storage medium 503 and an internal memory 504.

[0133] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 can perform an accompanying robot control method.

[0134] The processor 502 is configured to provide computing and control capabilities to support the operation of the entire computer device 500.

[0135] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503, which, when executed by the processor 502, causes the processor 502 to perform a method for controlling a companion robot.

[0136] The network interface 505 is configured to communicate with other devices via a network. Those skilled in the art can understand that the above structure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device 500 to which the scheme of the present application is applied. The specific computer device 500 can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0137] The processor 502 is configured to run the computer program 5032 stored in the memory to implement the following steps:

[0138] Collecting environmental information, multi-modal emotional information of a user, and user instruction information;

[0139] Identifying the emotional state of the user based on the multi-modal emotional information;

[0140] Identifying the user's intention based on the user instruction information;

[0141] Performing multi-source information fusion analysis based on the environmental information, the emotional state, and the user's intention, generating a control instruction, and controlling the companion robot to execute the control instruction.

[0142] In some preferred embodiments, the multi-modal emotional information includes user voice information, user facial expression information, and user body movement information, and the identifying the emotional state of the user based on the multi-modal emotional information comprises:

[0143] Extracting voice emotional features from the user voice information, extracting expression emotional features from the user facial expression information, and extracting action emotional features from the user body movement information;

[0144] Fusing the voice emotional features, the expression emotional features, and the action emotional features to obtain fused emotional features;

[0145] Analyzing the fused emotional features based on a pre-trained emotional analysis model to output the emotional state.

[0146] In some preferred embodiments, the identifying the user's intention based on the user instruction information comprises:

[0147] acquire context information of the user instruction information;

[0148] analyze, by a pre-trained intent recognition model, based on the user instruction information and the context information, and output the user intent.

[0149] In some preferred embodiments, based on the environmental information, the emotional state, and the user intent, multi-source information fusion analysis is performed to generate a control instruction, including:

[0150] environmental features are collected from the environmental information, and the environmental features, the emotional state, and the user intent are fused to obtain multi-source information fusion features;

[0151] analyze, by a pre-trained instruction generation model, based on the multi-source information fusion features, and output the control instruction.

[0152] In some preferred embodiments, the method further includes:

[0153] monitor a preset scene mode;

[0154] If the preset scene mode is monitored, a preset control instruction set corresponding to the scene mode is executed.

[0155] In some preferred embodiments, the scene mode includes a night mode and a leaving-home mode;

[0156] The preset control instruction set corresponding to the night mode includes at least one of adjusting the lighting to a preset intensity and playing a preset white noise audio;

[0157] The preset control instruction set corresponding to the leaving-home mode includes at least one of controlling the curtains and lights to be turned off and starting security alert monitoring.

[0158] In some preferred embodiments, the method further includes:

[0159] collect state data of the user;

[0160] infer, by a pre-trained personalized service prediction model, potential needs of the user based on the state data, and perform a personalized service operation corresponding to the potential needs, wherein the personalized service prediction model is trained based on historical data of the user;

[0161] And / or, the method further includes:

[0162] receive environmental parameters sent by an environmental sensor;

[0163] Determine a preset adjustment instruction of a target device based on the environmental parameter and a preset adjustment rule, and control the target device to execute the adjustment instruction.

[0164] And / or, the method further comprises:

[0165] Monitoring a running state of a preset target device;

[0166] If the running state of the target device is a preset target state, issuing a reminder information corresponding to the target state.

[0167] It should be understood that, in the embodiments of the present application, the processor 502 can be a central processing unit (CPU), and the processor 502 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0168] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a storage medium, which is a computer readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the above-mentioned embodiments.

[0169] Therefore, the present application also provides a storage medium. The storage medium can be a computer readable storage medium. The storage medium stores a computer program. The computer program is executed by a processor to make the processor execute the following steps:

[0170] Collecting environmental information, multi-modal emotional information of a user and user instruction information;

[0171] Identifying an emotional state of the user based on the multi-modal emotional information;

[0172] Identifying a user intention based on the user instruction information;

[0173] Performing multi-source information fusion analysis based on the environmental information, the emotional state and the user intention, generating a control instruction, and controlling the companion robot to execute the control instruction.

[0174] In some preferred embodiments, the multi-modal emotional information comprises user voice information, user facial expression information, and user body action information, and the identifying the emotional state of the user based on the multi-modal emotional information comprises:

[0175] extracting voice emotional features from the user voice information, extracting expression emotional features from the user facial expression information, and extracting action emotional features from the user body action information;

[0176] fusing the voice emotional features, the expression emotional features, and the action emotional features to obtain fused emotional features;

[0177] analyzing, by a pre-trained emotional analysis model, based on the fused emotional features, and outputting the emotional state.

[0178] In some preferred embodiments, the identifying the user intention based on the user instruction information comprises:

[0179] obtaining context information of the user instruction information;

[0180] analyzing, by a pre-trained intention recognition model, based on the user instruction information and the context information, and outputting the user intention.

[0181] In some preferred embodiments, the multi-source information fusion analysis based on the environmental information, the emotional state, and the user intention to generate the control instruction comprises:

[0182] collecting environmental features from the environmental information, fusing the environmental features, the emotional state, and the user intention to obtain multi-source information fusion features;

[0183] analyzing, by a pre-trained instruction generation model, based on the multi-source information fusion features, and outputting the control instruction.

[0184] In some preferred embodiments, the method further comprises:

[0185] monitoring a preset scene mode;

[0186] if the preset scene mode is monitored, executing a preset control instruction set corresponding to the scene mode.

[0187] In some preferred embodiments, the scene mode comprises a night mode and a leaving-home mode;

[0188] the preset control instruction set corresponding to the night mode comprises at least one of adjusting lighting to a preset intensity and playing a preset white noise audio;

[0189] The preset control instruction set corresponding to the away mode comprises at least one of controlling a curtain and light to be turned off and starting security monitoring.

[0190] In some preferred embodiments, the method further comprises:

[0191] collecting state data of the user;

[0192] inference, by a pre-trained personalized service prediction model, potential demands of the user based on the state data, and performing a personalized service operation corresponding to the potential demands, wherein the personalized service prediction model is trained based on historical data of the user;

[0193] And / or, the method further comprises:

[0194] receiving an environmental parameter sent by an environmental sensor;

[0195] determining a preset adjustment instruction of a target device based on the environmental parameter and a preset adjustment rule, and controlling the target device to execute the adjustment instruction;

[0196] And / or, the method further comprises:

[0197] monitoring an operating state of a preset target device;

[0198] if the operating state of the target device is a preset target state, sending a reminder information corresponding to the target state.

[0199] The storage medium is an entity, non-transient storage medium, for example, can be a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a magnetic disk or an optical disk and various entity storage media that can store program codes. The computer readable storage medium can be non-volatile, or volatile.

[0200] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in the above description in general terms. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0201] In several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the described apparatus embodiments are merely schematic. For example, the division of the units is merely a logical function division. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In a possible implementation process, the steps of the described method can be performed in a different order, or can be omitted, or can be combined into another process, or can be implemented by using a corresponding hardware.

[0202] The steps in the method embodiments of the present application can be adjusted, combined and deleted in sequence according to actual needs. The units in the apparatus embodiments of the present application can be combined, divided and deleted according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.

[0203] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a terminal or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application.

[0204] In the above embodiments, the description of each embodiment is focused on, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0205] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, these modifications and variations also belong to the scope of the claims of the present application and their equivalent technologies, and the present application is intended to include these modifications and variations.

[0206] The above description is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A control method for a companion robot, characterized in that, include: Collect environmental information, user multimodal emotional information, and user command information; Identify the user's emotional state based on the aforementioned multimodal emotional information; Identify user intent based on the aforementioned user instruction information; Based on the environmental information, the emotional state, and the user's intent, multi-source information fusion analysis is performed to generate control commands, and the companion robot is controlled to execute the control commands.

2. The companion robot control method according to claim 1, characterized in that, The multimodal emotional information includes user voice information, user facial expression information, and user body movement information. The step of identifying the user's emotional state based on the multimodal emotional information includes: Extract voice emotion features from the user's voice information, extract facial expression emotion features from the user's facial expression information, and extract action emotion features from the user's body movement information; The voice emotion features, facial expression emotion features, and action emotion features are fused to obtain fused emotion features; The pre-trained sentiment analysis model analyzes the fused sentiment features and outputs the sentiment state.

3. The companion robot control method according to claim 1, characterized in that, The process of identifying user intent based on the user instruction information includes: Obtain the context information of the user instruction information; The pre-trained intent recognition model analyzes the user instruction information and the context information to output the user intent.

4. The companion robot control method according to claim 1, characterized in that, Based on the environmental information, the emotional state, and the user intent, multi-source information fusion analysis is performed to generate control commands, including: Environmental features are collected from the environmental information, and the environmental features, emotional state, and user intent are fused to obtain multi-source information fusion features; The pre-trained instruction generation model analyzes the multi-source information fusion features and outputs the control instructions.

5. The companion robot control method according to claim 1, characterized in that, The method further includes: Monitor preset scene modes; If a preset scene mode is detected, execute the preset set of control instructions corresponding to the scene mode.

6. The companion robot control method according to claim 5, characterized in that, The scene modes include night mode and away-from-home mode; The preset control command set corresponding to the night mode includes at least one of adjusting the lighting to a preset intensity and playing a preset white noise audio. The preset control command set corresponding to the "away from home" mode includes at least one of the following: controlling the curtains and lights to close, and activating security monitoring.

7. The companion robot control method according to claim 1, characterized in that, The method further includes: Collect user status data; The personalized service prediction model is trained based on the user's historical data to infer the user's potential needs from the state data and execute personalized service operations corresponding to the potential needs. And / or, the method further includes: Receive environmental parameters sent by environmental sensors; Based on the environmental parameters and preset adjustment rules, a preset adjustment command for the target device is determined, and the target device is controlled to execute the adjustment command. And / or the method further includes: Monitor the operating status of preset target equipment; If the target device is in a preset target state, a reminder message corresponding to the target state will be issued.

8. A control device for a companion robot, characterized in that, include: The data acquisition unit is used to collect environmental information, user multimodal emotional information, and user command information. The first identification unit is used to identify the user's emotional state based on the multimodal emotional information; The second identification unit is used to identify the user's intent based on the user instruction information; The generation unit is used to perform multi-source information fusion analysis based on the environmental information, the emotional state, and the user's intention, generate control commands, and control the companion robot to execute the control commands.

9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, can implement the method as described in any one of claims 1-7.

Citation Information

Cited By

  • Parkinson's disease emotion accompanying robot interaction method and system

    CN121870792A