Robot control method, device and equipment based on emotion enhancement and medium
By performing sentiment analysis on visual and natural language commands, generating sentiment embedding vectors, and using a large language model to generate comprehensive commands, the problem of emotional support for elderly care robots when providing physical assistance is solved. This achieves integrated emotional control of the robot when performing tasks, enhancing the psychological safety and user experience of the elderly.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-05
AI Technical Summary
Existing elderly care robots struggle to deeply integrate emotional interaction with motor control, resulting in an inability to provide effective emotional support while offering physical assistance, thus impacting the elderly's sense of psychological security and user experience.
By acquiring users' visual information and natural language commands, sentiment analysis is performed to generate sentiment embedding vectors. A large language model is then used to generate comprehensive commands that include task semantics and sentiment strategies. Action tokens, including action parameters and language response content, are generated to drive the robot to perform physical actions and simultaneously output voice content.
This achieves integrated control of the robot, providing both physical assistance and emotional support, enhancing the elderly's trust and humanized experience, and ensuring the consistency of actions and language and meeting the user's emotional needs.
Smart Images

Figure CN121973192A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a robot control method, apparatus, device, and medium based on emotion enhancement. Background Technology
[0002] With the continuous development of artificial intelligence and robotics, elderly care robots are gradually playing an important role in assistive services in fields such as finance, healthcare, insurance, and banking. Currently, the research and development of elderly care robots shows a trend of functional differentiation. One type of robot focuses on the execution of physical tasks, such as delivering medication, fetching water, or assisting with walking using robotic arms; its core technology lies in improving the accuracy and reliability of motion control. The other type of robot focuses on emotional companionship, such as biomimetic therapeutic robots or conversational robots with voice interaction capabilities, aiming to provide psychological comfort to the elderly through social interaction. Although these two types of robots have made some progress in their respective fields, they are usually independent in practical applications, lacking functional synergy and architectural unity.
[0003] In real-world elderly care scenarios, seniors not only need robots to assist with daily tasks but also crave emotional care and respect during the service process. For example, the gentleness of the robot's tone and the softness of its movements when delivering a water cup directly impact the senior's sense of security and service experience. However, existing technologies struggle to organically combine emotional interaction with motion control. While affective computing technology has made progress in areas such as facial expression recognition and voice emotion analysis, its functionality largely remains at the perception level, failing to achieve deep coupling with the robotic arm's motion generation process. Specifically, when an elderly person exhibits tension or fear, the robotic arm performing the task may still move rapidly along a preset trajectory, failing to alleviate the user's anxiety and potentially creating secondary risks due to abrupt movements. Conversely, when an elderly person expresses anxiety or loneliness verbally, while the robot can recognize the emotion, it cannot adjust its movement rhythm or provide verbal reassurance while performing physical operations.
[0004] The core needs of elderly care extend beyond efficiency and accuracy in task execution; they encompass multi-dimensional goals such as trust building, psychological safety, and emotional support. Existing robots that focus solely on accuracy while neglecting warmth often fail to gain genuine acceptance from the elderly. Simply relying on robotic arms to perform tasks can lead to a lack of human touch and a cold, impersonal experience; while companion robots with only conversational capabilities cannot meet the elderly's needs for practical care. Therefore, current technological architectures have significant limitations in addressing the unique challenges of elderly care, necessitating an integrated solution that deeply blends motion control and emotional interaction, enabling the elderly to receive practical assistance while also feeling understood, cared for, and respected. Summary of the Invention
[0005] This invention provides a robot control method, apparatus, device, and medium based on emotion enhancement, aiming to solve the problem of how to enable robots to provide effective emotional support while providing physical assistance.
[0006] In a first aspect, embodiments of the present invention provide a robot control method based on emotion enhancement, comprising: Acquire the user's visual information and natural language commands, perform sentiment analysis on the visual information and natural language commands, and generate sentiment embedding vectors; The visual information, the natural language instructions, and the emotion embedding vector are input into a preset large language model, which generates a comprehensive instruction that simultaneously includes task semantics and emotion strategy. An action token is generated based on the task semantics and the emotion strategy. The action token includes action parameters for controlling the robot's movement and language response content. The robot is driven to perform physical actions based on the motion parameters, and the language response is output simultaneously.
[0007] A further technical solution is that the step of performing sentiment analysis on the visual information and the natural language instructions to generate a sentiment embedding vector includes: Behavioral sentiment analysis is performed on the visual information to obtain the first sentiment feature; The natural language command is analyzed for intonation to generate a second sentiment feature; Semantic analysis is performed on the natural language instructions to generate a third sentiment feature; The emotion embedding vector is generated based on the first emotion feature, the second emotion feature, and the third emotion feature.
[0008] A further technical solution is that generating action tokens based on the task semantics and the emotion strategy includes: Determine the action trajectory in the action parameters based on the task semantics; The execution speed and acceleration in the action parameters are determined based on the emotional strategy. The language response content is generated based on the emotional strategy and the task semantics.
[0009] A further technical solution is that determining the execution speed and acceleration in the action parameters based on the emotion strategy includes: Obtain the preset speed threshold and acceleration threshold corresponding to the emotional strategy; The execution speed is limited to the speed threshold, and the acceleration is limited to the acceleration threshold.
[0010] A further technical solution is that generating the language response content based on the emotion strategy and the natural language instructions includes: Based on the aforementioned emotion strategy, the voice emotion type is determined; Based on the emotional type of the speech and the semantics of the task, the language response content is generated.
[0011] A further technical solution is that, after driving the robot to perform physical actions based on the motion parameters and simultaneously outputting the language response content, the method further includes: The system collects user feedback information in real time during the execution process and dynamically adjusts the robotic arm's movements or voice based on the user feedback information.
[0012] A further technical solution is that the dynamic adjustment of the robotic arm's movements or voice based on the user feedback information includes: Identify the user's emotional state based on the user feedback information; If the emotional state is a preset negative state, control the robotic arm to slow down its movement speed and / or output a soothing voice.
[0013] Secondly, embodiments of the present invention also provide an emotion-enhanced robot control device, which includes a unit for performing the above-described method.
[0014] Thirdly, embodiments of the present invention also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0015] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the above-described method.
[0016] This invention provides a robot control method, apparatus, device, and medium based on emotion enhancement. The method includes: acquiring a user's visual information and natural language commands; performing emotion analysis on the visual information and natural language commands to generate an emotion embedding vector; inputting the visual information, natural language commands, and emotion embedding vector into a preset large language model, which generates a comprehensive command simultaneously incorporating task semantics and emotional strategies; generating an action token based on the task semantics and emotional strategies, the action token including action parameters for controlling robot movement and language response content; driving the robot to perform physical actions based on the action parameters, and simultaneously outputting the language response content. This invention acquires the user's emotional state through multimodal perception and generates an emotion embedding vector, then uses a large language model to fuse visual information, natural language commands, and emotion embedding vectors to generate a comprehensive command, which is then transformed into an action token containing action parameters and language responses, ultimately driving the robot to perform physical actions and simultaneously outputting speech content. This method achieves integrated control of robot physical assistance and emotional support, allowing the elderly to feel understood and cared for while receiving practical help, effectively improving service trust and a more humanized experience. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating an emotion-enhanced robot control method provided in an embodiment of the present invention; Figure 2 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0021] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0022] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0023] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0024] Please see Figure 1 This invention provides a robot control method based on emotion enhancement, which includes the following steps: S1, acquire the user's visual information and natural language instructions, perform sentiment analysis on the visual information and natural language instructions, and generate a sentiment embedding vector.
[0025] In practice, visual information includes images captured by a camera, while natural language commands include speech signals captured by a microphone. By acquiring the user's visual information and natural language commands and performing sentiment analysis to generate sentiment embedding vectors, a precise foundation for sentiment perception is established for the entire system. This enables the robot to no longer passively receive operational commands, but to actively interpret the emotional state expressed by the user through visual information (e.g., facial expressions, body posture, etc., not specifically limited in this invention) and natural language commands (e.g., tone of voice, semantics, etc., not specifically limited in this invention), such as anxiety, loneliness, or unease. This transforms the originally abstract and easily ignored emotional signals into structured data that can be recognized and processed by the computing system, providing crucial emotional contextual information for subsequent intelligent decision-making.
[0026] In some preferred embodiments, the above step "performing sentiment analysis on the visual information and the natural language instruction to generate a sentiment embedding vector" specifically includes the following steps: performing behavioral sentiment analysis on the visual information to obtain a first sentiment feature; performing intonation analysis on the natural language instruction to generate a second sentiment feature; performing semantic analysis on the natural language instruction to generate a third sentiment feature; and generating the sentiment embedding vector based on the first sentiment feature, the second sentiment feature, and the third sentiment feature.
[0027] In practice, the single sentiment analysis process is broken down into three parallel analysis paths with different focuses: behavioral sentiment analysis of visual information, intonation analysis of natural language instructions, and semantic analysis. Specifically: Behavioral sentiment analysis of visual information yields primary sentiment features, enabling the system to capture emotional information conveyed by users through body posture and facial expressions rather than verbal expression. Examples include stiff postures due to physical discomfort or withdrawal movements due to fear. Specifically, behavioral sentiment analysis of visual information can be based on a pre-trained visual sentiment analysis model, which is not specifically limited in this invention.
[0028] By performing intonation analysis on natural language commands to generate secondary sentiment features, the system can identify the emotional nuances in a user's voice. For example, a trembling or hurried tone may reveal inner tension or urgency. Specifically, intonation analysis of natural language commands can be based on a pre-trained audio sentiment analysis model, which is not specifically limited in this invention.
[0029] Semantic analysis of natural language commands generates third-party sentiment features, allowing for an understanding of the user's emotions at the level of the language's inherent meaning, such as a user explicitly stating "I'm a little scared." Specifically, semantic analysis of natural language commands can be based on a pre-trained semantic analysis model, which is not specifically limited in this invention.
[0030] Finally, an emotion embedding vector is generated based on the first, second, and third emotion features, which essentially completes a feature-level fusion of multi-source information. This fusion can effectively overcome the limitations and uncertainties of single-modal perception. The fusion method can employ intention mechanisms, feature concatenation, etc., which are not specifically limited in this invention.
[0031] For example, if a user says "no problem" (positive semantic feature), but their tone trembles (negative intonation feature) and their body leans slightly backward (negative behavioral feature), the system can more accurately determine that the user's true emotion is tension rather than relaxation by combining these three factors, thus avoiding misjudgments that may occur when relying solely on semantic analysis. Therefore, the multi-dimensional, cross-modal emotional feature generation and fusion mechanism defined in this embodiment provides higher-quality and more comprehensive emotional input for the accurate decision-making of subsequent large language models, and is a solid foundation for ensuring that the entire system achieves personalized and humanized services.
[0032] S2, the visual information, the natural language instructions, and the emotion embedding vector are input into a preset large language model, and the large language model generates a comprehensive instruction that simultaneously includes task semantics and emotion strategy.
[0033] In practice, visual information, natural language commands, and emotion embedding vectors are input into a pre-defined large language model to generate comprehensive instructions that simultaneously incorporate task semantics and emotional strategies. Leveraging the powerful cross-modal information fusion and contextual reasoning capabilities of the large language model, specific task requirements are deeply semantically coupled with the captured user emotional state. This elevates the robot's decision-making logic from simply understanding "what to do" to comprehensively judging "how to do it under the current user's emotional state." For example, the basic task of "delivering a water cup" is transformed into a comprehensive instruction of "delivering the water cup slowly and steadily with reassuring words" upon recognizing the user's trembling hands and nervousness. This ensures that the robot's subsequent actions are driven by a composite goal that combines functional achievement with emotional care.
[0034] S3, Generate an action token based on the task semantics and the emotion strategy. The action token includes action parameters for controlling the robot's movement and language response content.
[0035] In practice, action tokens are generated based on the task semantics and emotional strategies contained in the comprehensive instructions. These tokens include the action parameters that control the robot's movement and the corresponding language response content, thus achieving a precise conversion and decoupling design from high-level decision-making to low-level execution instructions.
[0036] For example, abstract emotional strategies (such as "soothing") can be concretized into physical parameters that directly drive the robot's joint motors (such as reducing movement speed and acceleration) and carefully crafted speech text (such as gentle comforting statements). This structured instruction generation mechanism ensures that the robot's physical actions and its speech output are highly coordinated in intent and timing, avoiding the interactive confusion caused by the disconnect or even contradiction between actions and language.
[0037] In some preferred embodiments, the above step "generating action tokens based on the task semantics and the emotion strategy" specifically includes the following steps: determining the action trajectory in the action parameters based on the task semantics; determining the execution speed and acceleration in the action parameters based on the emotion strategy; and generating the language response content based on the emotion strategy and the task semantics.
[0038] In specific implementation, the action trajectory in the action parameters is determined based on the task semantics; the emotional strategy is concretized into two core control elements that can directly guide the robot's execution, namely the execution characteristics and language response content in the action parameters, thereby achieving precise mapping and decoupled control from emotional intent to specific execution parameters. Specifically, this invention explicitly decomposes the generation process of action tokens into three logically clear sub-steps: First, the motion trajectory in the action parameters is determined based on the task semantics. This ensures that the fundamental goal of physical task execution is achieved. For example, the trajectory endpoint of delivering a water cup must be within reach of the user's hand.
[0039] Secondly, determining the execution speed and acceleration in the motion parameters based on the emotional strategy is key to transforming the abstract emotional strategy into concrete physical motion control. For example, when the emotional strategy is "soothing," the system will automatically select a lower execution speed and acceleration, making the robotic arm's movement exhibit gentle and smooth characteristics, avoiding user tension caused by rapid, sudden movements; conversely, a higher speed is allowed when the emotional strategy is "emergency assistance."
[0040] Finally, based on emotional strategies and task semantics, the generated language response content ensures that the output speech remains consistent with the currently performed task and the intended emotion. This step-by-step generation and collaborative integration of "trajectory planning," "motion dynamics," and "language communication" enables the robot's behavioral output to possess both functional accuracy and emotional appropriateness in interaction. It transforms "emotional enhancement" from a vague concept into precisely controllable motion parameters (velocity, acceleration) and carefully designed language content.
[0041] For example, for the task of "delivering medication," while generating the action trajectory of "moving the medication in front of the user," the system determines whether to use a slow and gentle motion accompanied by a warm reminder such as "Here is your medication, please take it slowly," or a standard speed accompanied by a routine reminder such as "Please take your medication on time," depending on whether the emotional strategy is "encouragement" or "routine reminder." This structured action token generation mechanism significantly improves the predictability, controllability, and consistency of the system's output behavior and emotional expression.
[0042] In some preferred embodiments, the above step "determining the execution speed and acceleration in the action parameters based on the emotion strategy" specifically includes the following steps: obtaining a preset speed threshold and acceleration threshold corresponding to the emotion strategy; limiting the execution speed to within the speed threshold and the acceleration to within the acceleration threshold.
[0043] In this embodiment, a specific, preset threshold limitation mechanism is introduced for the execution speed and acceleration determined based on the emotion strategy. This greatly enhances the safety and robustness of the robot's action execution in emotionally sensitive scenarios and provides a hard guarantee to prevent physical risks or psychological fright caused by the robot's actions being too fast or too violent.
[0044] Specifically, specific emotional strategies (especially those corresponding to negative emotions such as anxiety or fear) are bound to explicit speed and acceleration limits. When the system identifies a user as being in such a state based on emotion analysis, regardless of the task's efficiency requirements, the execution speed and acceleration determined for the motion parameters must be limited to the corresponding preset thresholds. This is akin to installing a "soft limiter" triggered by emotional state on the robot's motion system. For example, preset rules can stipulate that when the emotional strategy is identified as "user anxiety," the maximum speed of the robotic arm's end effector must not exceed 50% of the normal value, and the maximum acceleration must not exceed 30% of the normal value. This restriction is mandatory and proactive, eliminating the possibility of generating dangerous motion parameters from the decision-making stage. It ensures that when facing emotionally unstable elderly users, the robot's primary behavioral principle is not "completing the task quickly," but rather "interacting safely and without coercion."
[0045] The threshold limitation mechanism defined in this embodiment essentially encodes the "safety first" nursing principle into the robot's action generation logic through a preset threshold, which is a key safety technology measure for building a reliable elderly care robot system.
[0046] In some preferred embodiments, the above step "generating the language response content based on the emotion strategy and the natural language instruction" specifically includes the following steps: determining the voice emotion type based on the emotion strategy, wherein the voice emotion type includes soothing, encouraging, or explanatory; and generating the language response content based on the voice emotion type and the task semantics.
[0047] In practice, a systematic and structured method for generating language content is provided, which ensures that the language responses output by the robot during the interaction process are not only relevant to the task, but also highly matched with the user's emotional state and the system's emotional strategy in terms of emotional tone, thereby significantly improving the appropriateness, affinity and psychological guidance effect of human-computer language interaction.
[0048] First, the emotional type of the speech is determined based on the emotional strategy, such as categorizing it as reassuring, encouraging, or explanatory. This step sets a clear emotional direction and communication goal for subsequent text generation, transforming abstract strategies into concrete speech behavior categories. For example, when the emotional strategy is to alleviate the user's tension, the speech emotional type is determined to be "reassuring"; when the user is frustrated due to task failure, the type is determined to be "encouraging." Second, based on the determined speech emotional type and task semantics, specific language response content is generated. This step is essentially a process of integrating emotional type with task information, ensuring that the generated language content carries a specific emotional tone without deviating from the current task context. For example, for the task of "helping to get up," if the speech emotional type is "reassuring," the generated content might be "Please don't be nervous, I will help you up slowly"; if the type is "encouraging," the generated content might be "You did very well, let's try again." This method overcomes the mechanical rigidity of simply playing pre-recorded audio or providing template-based responses, making each language response a personalized feedback to the specific current situation. This enables robots to choose appropriate speaking styles and content based on the emotional state of the other person, much like a human caregiver, thereby conducting more effective emotional communication and building user trust and reliance on the machine. Specifically, the language response content can be generated based on the emotional type of the speech and the semantics of the task through a pre-trained large language model; however, this invention does not specifically limit the specifics.
[0049] S4, based on the motion parameters, drive the robot to perform physical actions and simultaneously output the language response content.
[0050] In practice, the robot performs physical actions based on generated motion parameters, simultaneously outputting corresponding verbal responses. This is the key output link in transforming intelligent decision-making into a service that is perceptible to the end user, creating a human-like, multi-dimensional interactive experience. When the robot gently delivers items while speaking caring words in a soothing tone, this support provided simultaneously at both the physical and psychological levels greatly reduces the mechanical and impersonal feel of traditional service robots. It allows elderly users to receive tangible physical assistance while their psychological needs are simultaneously recognized and addressed, thereby building trust and reliance on the robot service.
[0051] Through the aforementioned interconnected technical steps, this invention successfully constructs an integrated robot control system capable of understanding emotions, developing thought strategies, and providing care. This transforms the robot in elderly care scenarios from a single-function tool into an intelligent companion that efficiently performs physical assistance tasks while providing timely and appropriate emotional support. This truly meets the core needs of the elderly for dignity, safety, and emotional belonging, enabling them to receive practical assistance while continuously feeling understood, cared for, and respected.
[0052] In some preferred embodiments, after the above step of "driving the robot to perform physical actions based on the motion parameters and synchronously outputting the language response content", the following steps are also included: collecting user feedback information in real time during the execution process, and dynamically adjusting the actions or voice of the robotic arm based on the user feedback information, wherein the user feedback information includes user posture feedback information and voice feedback information.
[0053] In practice, real-time collection of user feedback and dynamic adjustment transform the entire system from a static, one-way instruction-execution machine into a dynamic, environmentally adaptable, and continuously interactive intelligent agent. This endows the system with the ability to self-adjust and optimize based on real-time user responses when facing complex and unstructured real-world elderly care environments.
[0054] The system no longer relies solely on initial sensory information to make one-off decisions, but instead treats the execution process itself as a continuous window for perception and interaction. By collecting real-time user posture feedback (such as whether the body leans back to avoid an obstacle, or whether gestures indicate a stop) and voice feedback (such as the user saying "slow down" or "I'm scared"), the system can promptly obtain the actual effects of the initial actions and voice outputs, as well as the user's immediate psychological and physiological changes. Based on this dynamic feedback information, the system can adjust the robotic arm's movements or output voice online.
[0055] For example, even if the system initially sets a standard delivery speed based on a "neutral" emotion, if it detects a sudden withdrawal of the user's hand during execution (postural feedback), it can immediately trigger action adjustment, slowing down or even pausing the speed to avoid collisions or startling the user. Similarly, if the user issues a new voice command (voice feedback), the system can interrupt the current task flow to respond. This closed-loop mechanism greatly enhances the system's robustness, safety, and interactive flexibility. It acknowledges the potential for errors in initial emotion recognition and task understanding, and also recognizes that user emotions can dynamically change during task execution. By constructing such a continuous cycle of "perception-decision-execution-re-perception," the robot can better adapt to the uncertain service scenarios of the real world, providing a more sensitive, appropriate, and user-centric service experience, thereby continuously consolidating and enhancing user trust through dynamic interactions.
[0056] In some preferred embodiments, the above step "dynamically adjusting the movement or voice of the robotic arm based on the user feedback information" specifically includes the following steps: identifying the user's emotional state based on the user feedback information; if the emotional state is a preset negative state, controlling the robotic arm to slow down its movement speed and / or outputting a soothing voice, wherein the negative state includes a tense state or a resistant state.
[0057] In practice, the system identifies the user's emotional state based on user feedback information, and when a preset negative state is identified, it performs targeted operations such as slowing down the action speed and / or outputting soothing voice messages. This enables the system to proactively, promptly and appropriately mitigate risks and intervene in emotions when faced with negative user emotions, refining the granularity of safety assurance and emotional support from the task level to the real-time interaction moment.
[0058] Specifically, the system continuously inputs real-time user posture feedback information (images) and voice feedback information (voice signals), and analyzes them through a built-in emotion state recognition module (which can be a rule-based classifier, a lightweight machine learning model, or a large language model) to determine whether the user is currently in a preset negative state, such as tension or resistance. Once confirmed, the system no longer simply records the state but immediately triggers preset adjustment strategies. Slowing down the robotic arm's movement speed directly reduces the source of potential user anxiety or safety risks at the physical level, rebuilding the user's sense of security by making the movement more gentle and predictable. Simultaneously outputting soothing voice messages provides proactive intervention at the psychological level, explaining the current behavior and expressing concern through verbal communication, thereby easing the user's tension. The soothing voice messages can be output by a large language model based on user feedback information; this invention is not specifically limited to this.
[0059] For example, when a robotic arm is delivering an item, if the system detects through a camera that the user is frowning and leaning back (identifying a state of tension), it will immediately slow down the robotic arm's movement and simultaneously say through a speaker, "Please don't worry, my movements are slow and safe." This combination of "movement slowdown + voice reassurance" constitutes a three-dimensional, multimodal reassurance response. It is no longer a simple stop, but rather, while ensuring the task can continue (unless the user explicitly requests a stop), it proceeds in a gentler, more explanatory manner, thereby mitigating potential risks while maintaining the continuity of the interaction and the user's sense of control.
[0060] This invention proposes an emotion-enhanced robot control method, comprising: acquiring a user's visual information and natural language commands; performing emotion analysis on the visual information and natural language commands to generate an emotion embedding vector; inputting the visual information, natural language commands, and emotion embedding vector into a preset large language model, which generates a comprehensive command that simultaneously includes task semantics and emotion strategy; generating an action token based on the task semantics and emotion strategy, the action token including action parameters for controlling robot movement and language response content; driving the robot to perform physical actions based on the action parameters, and simultaneously outputting the language response content. This invention acquires the user's emotional state through multimodal perception and generates an emotion embedding vector, utilizes a large language model to fuse visual information, natural language commands, and emotion embedding vectors to generate a comprehensive command, and then transforms it into an action token containing action parameters and language response, ultimately driving the robot to perform physical actions and simultaneously outputting speech content. This method achieves integrated control of robot physical assistance and emotional support, enabling the elderly to feel understood and cared for while receiving practical help, effectively improving service trust and a humanized experience.
[0061] This invention can be applied to fields such as finance, healthcare, insurance, and banking. Specific application examples are as follows: Application Cases in the Financial Sector In bank branches or at the homes of elderly customers, when an elderly customer needs to handle complex financial product subscription transactions, this invention can be applied to intelligent service terminals or assistive robots. The system observes the customer through visual sensors, noticing that the customer repeatedly rubs their hands and furrows their brow while facing risk warnings on an electronic screen. Simultaneously, a voice sensor captures the customer muttering, "I don't quite understand these terms; I'm a little uneasy." The multimodal perception module analyzes this visual and linguistic information, generating an emotional embedding vector representing "confusion and anxiety." Subsequently, the visual information (the customer's confused expression), the natural language instruction ("submit financial product subscription"), and the emotional vector are input into a large language model. After comprehensive understanding, the large language model generates a comprehensive instruction: the core task is "assist in completing the financial product subscription process," and the emotional strategy is "slow down the pace and provide focused explanations to soothe emotions." Based on this, the motion generation module outputs corresponding motion tokens: its motion parameters control the robotic arm to guide and operate on the touchscreen in a slower, gentler manner, highlighting key terms and appropriately enlarging the font; its verbal response generates reassuring and explanatory statements such as, "Grandpa Wang, this is an explanation about expected returns. I'll show it to you slowly, please don't rush." By simultaneously executing these gentle physical guidance and clear verbal explanations, the robot effectively alleviates the cognitive stress and anxiety of elderly customers while assisting with financial transactions, enhancing their sense of security and trust in financial services.
[0062] Medical Application Cases In home care wards or community health centers, an elderly person with a chronic illness needs to take multiple medications daily. This invention can be integrated into a medical care robot. When the robot prepares to deliver pills and a water cup to the elderly person, its visual sensors detect that the elderly person is struggling to swallow, their face shows discomfort, and their body leans back slightly. Simultaneously, the elderly person's voice command, tinged with hesitation, says, "These pills are a bit big today..." The system's sentiment analysis module extracts the emotional features of "worry and resistance" from these signals, generating corresponding emotional embedding vectors. After receiving this vector and the task instruction of "delivering medication," the large language model generates a comprehensive instruction that not only includes the delivery task but also embeds emotional strategies of "extremely slow, reassuring, and providing alternatives." Action tokens are thus generated: the motion parameters are set to extremely low speed and acceleration, ensuring the robotic arm smoothly delivers the water cup and pills to the elderly person in a nearly non-threatening manner; the verbal response is, "Grandma Li, I'll hand it to you slowly. If you have difficulty swallowing, we can cut the pills into smaller pieces, would that be alright?" This kind of gentle care and alternative solutions, triggered by emotional perception and based on safe actions, makes the elderly feel respected and understood when receiving necessary medical assistance, thereby improving medication adherence and comfort.
[0063] Application Cases in the Insurance Field In an insurance claims or business consultation scenario, an elderly customer was struggling to submit claim materials due to unfamiliarity with the process. This is where the technology from our invention, applied to an insurance service robot, comes into play. The robot observes the customer frantically sorting through a pile of documents, sighing frequently, and hears the customer complaining through a microphone, "So many documents, I can't even tell which one I need." Based on this, the multimodal perception module generates an emotional embedding vector representing "frustration and helplessness." The large language model integrates the task ("assisting in sorting claim materials"), visual context, and emotional vectors, outputting a comprehensive instruction: execute the core task, employing a "step-by-step guidance and positive encouragement" emotional strategy. The subsequently generated action token instructs the robotic arm to help the customer sort and categorize documents with methodical and gentle movements, avoiding rapid flipping that could increase customer anxiety; simultaneously, the voice system outputs explanatory and encouraging content such as, "Mr. Zhang, it's okay, we'll take it one step at a time. See, the hospital bills are all here, this part is complete, you've done a great job." By combining physical assistance with material organization with patient guidance and positive feedback, the robot not only efficiently completes business support tasks, but also significantly reduces the helplessness and frustration of elderly customers when facing complex insurance matters, thus improving the service experience.
[0064] Application Cases in the Banking Sector In a bank's VIP room or a senior customer's home, when an elderly person with limited mobility needs to sign a large deposit certificate for confirmation, this invention can be deployed with a dedicated business assistance robot. The robot's vision sensors notice that the elderly person's hand is trembling slightly as they hold the pen, and their eyes reveal concern about the accuracy of the operation. Their natural language instruction is, "Do I need to sign here for confirmation?" The system's sentiment analysis module identifies the emotion of "tension and uncertainty" and generates an emotion embedding vector. After processing all the information, the large language model generates an instruction: the task semantic is "assist in locating the signature area and ensuring the process is completed," and the sentiment strategy is "extremely gentle, repeated confirmation, and enhanced confidence." An action token is then created: the robotic arm's motion parameters are strictly limited to extremely low speed and acceleration thresholds, using a stable and gentle force to precisely push the document to the most comfortable signing position for the elderly person, maintaining stable support; the language response generates reassuring and confirming words such as, "Yes, Aunt Wang, you can sign here after confirming everything is correct. The pen has been held steady for you, please rest assured, we have plenty of time." This kind of emotionally-based physical stability support and psychological comfort provided during critical financial operations greatly enhances elderly customers' sense of security and control when handling important financial matters, ensuring that business is completed smoothly in a controlled and secure environment.
[0065] Corresponding to the above-described emotion-enhanced robot control method, the present invention also provides an emotion-enhanced robot control device. This emotion-enhanced robot control device includes a unit for executing the aforementioned emotion-enhanced robot control method, and can be configured in a desktop computer, tablet computer, laptop computer, or other terminal. Specifically, the emotion-enhanced robot control device includes: The acquisition unit is used to acquire the user's visual information and natural language instructions, perform sentiment analysis on the visual information and natural language instructions, and generate a sentiment embedding vector. The input unit is used to input the visual information, the natural language instructions, and the emotion embedding vector into a preset large language model, and the large language model generates a comprehensive instruction that simultaneously contains task semantics and emotion strategy; A generation unit is configured to generate an action token based on the task semantics and the emotion strategy, wherein the action token includes action parameters for controlling robot movement and language response content; An execution unit is used to drive the robot to perform physical actions based on the motion parameters and simultaneously output the language response content.
[0066] In some preferred embodiments, the step of performing sentiment analysis on the visual information and the natural language instructions to generate a sentiment embedding vector includes: Behavioral sentiment analysis is performed on the visual information to obtain the first sentiment feature; The natural language command is analyzed for intonation to generate a second sentiment feature; Semantic analysis is performed on the natural language instructions to generate a third sentiment feature; The emotion embedding vector is generated based on the first emotion feature, the second emotion feature, and the third emotion feature.
[0067] In some preferred embodiments, generating action tokens based on the task semantics and the emotion strategy includes: Determine the action trajectory in the action parameters based on the task semantics; The execution speed and acceleration in the action parameters are determined based on the emotional strategy. The language response content is generated based on the emotional strategy and the task semantics.
[0068] In some preferred embodiments, determining the execution speed and acceleration in the action parameters based on the emotion strategy includes: Obtain the preset speed threshold and acceleration threshold corresponding to the emotional strategy; The execution speed is limited to the speed threshold, and the acceleration is limited to the acceleration threshold.
[0069] In some preferred embodiments, generating the language response content based on the emotion strategy and the natural language instructions includes: Based on the aforementioned emotion strategy, the voice emotion type is determined; Based on the emotional type of the speech and the semantics of the task, the language response content is generated.
[0070] In some preferred embodiments, it further includes: The feedback unit is used to collect user feedback information in real time during the execution process, and to dynamically adjust the actions or voice of the robotic arm based on the user feedback information.
[0071] In some preferred embodiments, the dynamic adjustment of the robotic arm's movements or voice based on the user feedback information includes: Identify the user's emotional state based on the user feedback information; If the emotional state is a preset negative state, control the robotic arm to slow down its movement speed and / or output a soothing voice.
[0072] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned emotion-enhanced robot control device and its various units can be referred to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity, these details will not be repeated here.
[0073] The aforementioned emotion-enhanced robot control device can be implemented as a computer program, which can, for example... Figure 2 It runs on the computer device shown.
[0074] Please see Figure 2 , Figure 2 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a terminal or a server. The terminal can be an electronic device with communication functions, such as a smartphone, tablet, laptop, desktop computer, personal digital assistant, or wearable device. The server can be a standalone server or a server cluster composed of multiple servers.
[0075] The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0076] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, it enables the processor 502 to execute an emotion-enhanced robot control method.
[0077] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0078] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute an emotion-enhanced robot control method.
[0079] The network interface 505 is used for network communication with other devices. Those skilled in the art will understand that the above structure is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. A specific computer device 500 may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.
[0080] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps: Acquire the user's visual information and natural language commands, perform sentiment analysis on the visual information and natural language commands, and generate sentiment embedding vectors; The visual information, the natural language instructions, and the emotion embedding vector are input into a preset large language model, which generates a comprehensive instruction that simultaneously includes task semantics and emotion strategy. An action token is generated based on the task semantics and the emotion strategy. The action token includes action parameters for controlling the robot's movement and language response content. The robot is driven to perform physical actions based on the motion parameters, and the language response is output simultaneously.
[0081] In some preferred embodiments, the step of performing sentiment analysis on the visual information and the natural language instructions to generate a sentiment embedding vector includes: Behavioral sentiment analysis is performed on the visual information to obtain the first sentiment feature; The natural language command is analyzed for intonation to generate a second sentiment feature; Semantic analysis is performed on the natural language instructions to generate a third sentiment feature; The emotion embedding vector is generated based on the first emotion feature, the second emotion feature, and the third emotion feature.
[0082] In some preferred embodiments, generating action tokens based on the task semantics and the emotion strategy includes: Determine the action trajectory in the action parameters based on the task semantics; The execution speed and acceleration in the action parameters are determined based on the emotional strategy. The language response content is generated based on the emotional strategy and the task semantics.
[0083] In some preferred embodiments, determining the execution speed and acceleration in the action parameters based on the emotion strategy includes: Obtain the preset speed threshold and acceleration threshold corresponding to the emotional strategy; The execution speed is limited to the speed threshold, and the acceleration is limited to the acceleration threshold.
[0084] In some preferred embodiments, generating the language response content based on the emotion strategy and the natural language instructions includes: Based on the aforementioned emotion strategy, the voice emotion type is determined; Based on the emotional type of the speech and the semantics of the task, the language response content is generated.
[0085] In some preferred embodiments, after driving the robot to perform physical actions based on the motion parameters and synchronously outputting the language response content, the method further includes: The system collects user feedback information in real time during the execution process and dynamically adjusts the robotic arm's movements or voice based on the user feedback information.
[0086] In some preferred embodiments, the dynamic adjustment of the robotic arm's movements or voice based on the user feedback information includes: Identify the user's emotional state based on the user feedback information; If the emotional state is a preset negative state, control the robotic arm to slow down its movement speed and / or output a soothing voice.
[0087] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0088] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program may be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0089] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program. When executed by a processor, the computer program causes the processor to perform the following steps: Acquire the user's visual information and natural language commands, perform sentiment analysis on the visual information and natural language commands, and generate sentiment embedding vectors; The visual information, the natural language instructions, and the emotion embedding vector are input into a preset large language model, which generates a comprehensive instruction that simultaneously includes task semantics and emotion strategy. An action token is generated based on the task semantics and the emotion strategy. The action token includes action parameters for controlling the robot's movement and language response content. The robot is driven to perform physical actions based on the motion parameters, and the language response is output simultaneously.
[0090] In some preferred embodiments, the step of performing sentiment analysis on the visual information and the natural language instructions to generate a sentiment embedding vector includes: Behavioral sentiment analysis is performed on the visual information to obtain the first sentiment feature; The natural language command is analyzed for intonation to generate a second sentiment feature; Semantic analysis is performed on the natural language instructions to generate a third sentiment feature; The emotion embedding vector is generated based on the first emotion feature, the second emotion feature, and the third emotion feature.
[0091] In some preferred embodiments, generating action tokens based on the task semantics and the emotion strategy includes: Determine the action trajectory in the action parameters based on the task semantics; The execution speed and acceleration in the action parameters are determined based on the emotional strategy. The language response content is generated based on the emotional strategy and the task semantics.
[0092] In some preferred embodiments, determining the execution speed and acceleration in the action parameters based on the emotion strategy includes: Obtain the preset speed threshold and acceleration threshold corresponding to the emotional strategy; The execution speed is limited to the speed threshold, and the acceleration is limited to the acceleration threshold.
[0093] In some preferred embodiments, generating the language response content based on the emotion strategy and the natural language instructions includes: Based on the aforementioned emotion strategy, the voice emotion type is determined; Based on the emotional type of the speech and the semantics of the task, the language response content is generated.
[0094] In some preferred embodiments, after driving the robot to perform physical actions based on the motion parameters and synchronously outputting the language response content, the method further includes: The system collects user feedback information in real time during the execution process and dynamically adjusts the robotic arm's movements or voice based on the user feedback information.
[0095] In some preferred embodiments, the dynamic adjustment of the robotic arm's movements or voice based on the user feedback information includes: Identify the user's emotional state based on the user feedback information; If the emotional state is a preset negative state, control the robotic arm to slow down its movement speed and / or output a soothing voice.
[0096] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk, or any other physical storage medium capable of storing program code. The computer-readable storage medium can be non-volatile or volatile.
[0097] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0098] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0099] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0100] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0101] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0102] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Since these modifications and variations fall within the scope of the claims and their equivalents, this invention also intends to include these modifications and variations.
[0103] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A robot control method based on emotion enhancement, characterized in that, include: Acquire the user's visual information and natural language commands, perform sentiment analysis on the visual information and natural language commands, and generate sentiment embedding vectors; The visual information, the natural language instructions, and the emotion embedding vector are input into a preset large language model, which generates a comprehensive instruction that simultaneously includes task semantics and emotion strategy. An action token is generated based on the task semantics and the emotion strategy. The action token includes action parameters for controlling robot movement and language response content. The robot is driven to perform physical actions based on the motion parameters, and the language response is output simultaneously.
2. The robot control method based on emotion enhancement according to claim 1, characterized in that, The step of performing sentiment analysis on the visual information and the natural language instructions to generate sentiment embedding vectors includes: Behavioral sentiment analysis is performed on the visual information to obtain the first sentiment feature; The natural language command is analyzed for intonation to generate a second sentiment feature; Semantic analysis is performed on the natural language instructions to generate a third sentiment feature; The emotion embedding vector is generated based on the first emotion feature, the second emotion feature, and the third emotion feature.
3. The robot control method based on emotion enhancement according to claim 1, characterized in that, The step of generating action tokens based on the task semantics and the emotion strategy includes: Determine the motion trajectory in the motion parameters based on the task semantics; The execution speed and acceleration in the action parameters are determined based on the emotional strategy. The language response content is generated based on the emotional strategy and the task semantics.
4. The robot control method based on emotion enhancement according to claim 3, characterized in that, The step of determining the execution speed and acceleration in the action parameters based on the emotion strategy includes: Obtain the preset speed threshold and acceleration threshold corresponding to the emotional strategy; The execution speed is limited to the speed threshold, and the acceleration is limited to the acceleration threshold.
5. The emotion-enhanced robot control method according to claim 3, characterized in that, The generation of the language response content based on the emotional strategy and the task semantics includes: Based on the aforementioned emotion strategy, the voice emotion type is determined; Based on the emotional type of the speech and the semantics of the task, the language response content is generated.
6. The robot control method based on emotion enhancement according to claim 1, characterized in that, After driving the robot to perform physical actions based on the motion parameters and synchronously outputting the language response content, the method further includes: The system collects user feedback information in real time during the execution process and dynamically adjusts the robotic arm's movements or voice based on the user feedback information.
7. The emotion-enhanced robot control method according to claim 6, characterized in that, The dynamic adjustment of the robotic arm's movements or voice based on the user feedback information includes: Identify the user's emotional state based on the user feedback information; If the emotional state is a preset negative state, control the robotic arm to slow down its movement speed and / or output a soothing voice.
8. A robot control device based on emotion enhancement, characterized in that, Includes a unit for performing the method as described in any one of claims 1-7.
9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, can implement the method as described in any one of claims 1-7.