AI humanoid robot motion regulation method and system based on digital twinning
By constructing an emotion-ethics joint twin model, ethical constraints are transformed into feasible domain boundaries of the emotional state space and hard constraints for motion planning, solving the problems of robot motion stuttering and boundary overstepping during interaction, and achieving natural, smooth, safe and compliant interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN YUNHENG INTELLIGENT CO LTD
- Filing Date
- 2026-04-03
- Publication Date
- 2026-06-19
AI Technical Summary
In existing technologies, the emotion understanding and expression module and the ethical constraint module are independent of each other, which may cause the robot to experience motion stuttering or overstepping of boundaries during interaction, making it unable to express emotions naturally and fluently while ensuring ethical norms.
An emotion-ethics joint twin model is constructed, which transforms ethical constraints into feasible domain boundaries of the emotional state space and hard constraints of kinematic planning. Information is collected through a multimodal perception module to generate whole-body joint motion trajectories that simultaneously satisfy kinematics, dynamics, spatial restricted areas, and taboo movements.
This enables robots to spontaneously adhere to ethical standards during motion control, avoiding motion lag and out-of-bounds behavior, ensuring natural and smooth interaction, and improving the adaptability and user experience of human-computer interaction.
Smart Images

Figure CN122239593A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of humanoid robot intelligent control and human-computer interaction technology, specifically to an AI humanoid robot motion control method and system based on digital twins. Background Technology
[0002] Digital twin technology, as a core link connecting the physical world and cyberspace, has been widely applied in the field of humanoid robot control in recent years. By constructing a virtual model that maps in real time to a physical entity, digital twins can monitor, simulate, and optimize the robot's motion state. As humanoid robots gradually move from industrial scenarios to social scenarios such as home services, medical rehabilitation, and public guidance, robots not only need to possess stable motion control capabilities but also need to exhibit emotional expressions that conform to social norms during interactions with humans. However, how to ensure the safety of robot movement while enabling it to understand human emotions and respond appropriately, and to consciously abide by laws, regulations, and social ethical standards, has become a pressing challenge in current technological development.
[0003] A search revealed that some existing research focuses on emotional understanding and expression during human-computer interaction. For example, patent publication number CN121148009A discloses a human-computer physical twin system and control method based on multimodal video analysis and adaptive mapping. This technology acquires human movements, expressions, and speech through a multimodal acquisition unit, extracts key points of the human skeleton and facial action unit coefficients through a semantic analysis module, and establishes a human-robot joint kinematic mapping through an adaptive mapping unit to capture and reproduce human movements. However, this technology focuses on kinematic mapping at the action level and does not address the robot's understanding of the emotional state of the interacting object or the regulation of its own emotional expression, nor does it consider the ethical constraints and social norms that the robot must adhere to during movement. On the other hand, some academic research fields have pointed out that current research suffers from two major separation problems: "emotional authenticity" and "ethical compliance." The emotion computing module and the ethical constraint module are often independent of each other, lacking deep integration at the design level, which may lead to situations where the robot's emotional expression conflicts with ethical requirements during interaction. However, this literature only analyzes the problem and does not propose specific technical solutions.
[0004] In summary, the technical shortcomings of existing technologies lie in the fact that the emotion understanding and expression module and the ethical constraint module operate independently, forming a sequential architecture of "understanding first, then filtering." Under this architecture, the emotion module generates the robot's emotional response based on the user's state, and the ethics module then reviews and restricts the content of the response. When the ethics module rejects the actions generated by the emotion module, the robot may experience movement stuttering or its expression being forcibly truncated, severely disrupting the natural fluency of human-computer interaction. More critically, existing ethical constraint technologies mostly remain at the content filtering level, failing to transform social norms such as social distance and cultural taboos into calculable and enforceable kinematic constraints. This means that even if the robot "knows" it shouldn't approach, it may still overstep boundaries due to the lack of consideration for these constraints in its motion planning. Therefore, how to deeply integrate emotion understanding and expression, ethical norms and constraints, and kinematic and dynamic feasibility to construct a unified digital twin control model, enabling physical humanoid robots to express emotions naturally and fluently in complex social environments while spontaneously adhering to ethical norms, has become a pressing technical problem for those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a method and system for motion control of AI humanoid robots based on digital twins. By constructing parameter coupling between the emotional and ethical dimensions at the bottom layer of the model, ethical constraints are transformed into feasible domain boundaries of the emotional state space and hard constraints of kinematic planning, thereby solving the conflict between emotional expression and ethical constraints and providing technical support for the safe and natural interaction of humanoid robots in social scenarios.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: On one hand, a method and system for motion control of an AI humanoid robot based on digital twins, the method comprising the following steps: Step 1: Construct an emotion-ethics joint twin model that is synchronized in real time with the physical humanoid robot. The emotion-ethics joint twin model includes a formal expression module for ethical constraints and an emotion computing and mapping module. Step two: Collect multimodal information of the interactive object and environmental information through the multimodal perception module deployed on the physical humanoid robot; Step 3: The emotion computing and mapping module generates the user's emotional state of the interactive object based on the multimodal information and the environmental information; Step four: The formal expression module of ethical constraints generates a set of ethical constraint parameters based on the environmental information and the preset cultural characteristic information. The set of ethical constraint parameters includes spatial forbidden zone parameters, taboo action parameters, and feasible domain parameters of emotional intensity. Step 5: The emotion calculation and mapping module generates the robot target emotion state within the range of the feasible domain parameters of emotion intensity based on the user's emotional state, the environmental information, the cultural feature information, and the ethical constraint parameter set. Step six: The emotion-ethics joint twin model generates a whole-body joint motion trajectory that simultaneously satisfies kinematic constraints, dynamic constraints, spatial forbidden zone parameters, and forbidden action parameters based on the robot target emotional state and the ethical constraint parameter set. Step 7: The physical humanoid robot executes the full-body joint movement trajectory.
[0007] By constructing an emotion-ethics joint twin model and generating full-body joint motion trajectories that simultaneously satisfy kinematic constraints, dynamic constraints, spatial forbidden zone parameters, and taboo action parameters, robots can spontaneously incorporate ethical norms as underlying constraints during motion control, achieving the synchronous generation of emotional expression and ethical compliance.
[0008] Furthermore, the formal expression module of ethical constraints includes an ethical knowledge graph, which stores the types of emotional expression and the range of emotional intensity allowed under different contextual conditions, indexed by scene type, cultural region, contact type and spatial distance. The spatial forbidden zone parameters are represented as a set of forbidden regions in three-dimensional space, the forbidden action parameters are represented as forbidden subspaces in joint space, the joint space is composed of all joint angles of the physical humanoid robot, and the emotional intensity feasible domain parameters are represented as arousal threshold range and valence threshold range.
[0009] By setting up an ethical knowledge graph indexed by scene type, cultural region, contact type, and spatial distance, and representing spatial forbidden zone parameters, taboo action parameters, and feasible domain parameters of emotional intensity as a three-dimensional forbidden region set, joint forbidden subspace, and arousal and valence threshold range, respectively, ethical constraints can be transformed from abstract norms into a structured, storable, and searchable form, providing a standardized constraint data foundation for the emotion-ethics joint twin model.
[0010] Furthermore, the set of prohibited areas in the three-dimensional space is determined based on the robot end-effector position coordinates, user position coordinates, a preset safe distance threshold for the scene type, and a contact type permission function; The forbidden subspace in the joint space is determined according to a preset set of forbidden joint angles; The arousal threshold range is determined by the lower and upper arousal thresholds based on the scene type and cultural characteristics information, and the valence threshold range is determined by the lower and upper valence thresholds based on the scene type and cultural characteristics information.
[0011] By determining the set of prohibited regions in three-dimensional space based on the robot's end-effector position coordinates and the user's position coordinates, determining the prohibited subspace in the joint space based on the preset set of taboo joint angles, and determining the range of arousal threshold and valence threshold based on scene type and cultural characteristic information, ethical requirements such as social distance, cultural taboos, and emotional intensity can be transformed into calculable and executable mathematical boundaries, thereby achieving precise quantification of ethical constraints at the motion planning level.
[0012] Furthermore, the emotion computing and mapping module includes an emotion understanding engine and an emotion-ethics joint mapper; The emotion understanding engine uses a multimodal fusion neural network to analyze the multimodal information and output the user's emotional state, which includes user arousal, user valence, and user dominance. The emotion-ethics joint mapper calculates the target emotion state of the robot based on the user's emotional state, the environmental information, the cultural characteristic information, and the set of ethical constraint parameters.
[0013] By setting up an emotion understanding engine and an emotion-ethics joint mapper to handle user emotion state recognition and robot target emotion state generation respectively, and by enabling the emotion understanding engine to use a multimodal fusion neural network to output user emotion states that include user arousal, user valence, and user dominance, the emotion-ethics joint mapper can ensure that the generated target emotion state always falls within the range defined by the ethical constraint parameter set, based on an accurate understanding of user emotions.
[0014] Furthermore, the emotion-ethics joint mapper determines a feasible region of emotion intensity based on the environmental information and the cultural characteristic information. By minimizing the deviation between the robot's target emotional state and the resonance function value, it selects the robot's target emotional state within the feasible region of emotion intensity. The resonance function is determined based on the psychological resonance theory.
[0015] By determining the feasible region of emotional intensity based on environmental and cultural information, and selecting the robot's target emotional state within the feasible region by minimizing the deviation between the robot's target emotional state and the resonance function value, the robot's emotional expression can be made as close as possible to the ideal psychological resonance state while meeting ethical constraints, thus achieving synergistic optimization of the naturalness of emotional expression and ethical norms.
[0016] Furthermore, the emotion-ethics joint twin model includes a motion-emotion co-planner, which contains an emotion-action primitive mapping library that stores whole-body motion trajectory primitives with emotion tags. The motion-emotion co-planner selects matching emotional action primitives from the emotion-action primitive mapping library based on the target emotional state of the robot, and adjusts the amplitude, velocity and duration parameters of the emotional action primitives to generate an initial whole-body joint motion trajectory.
[0017] By setting up an emotion-action primitive mapping library and adjusting the amplitude, velocity, and duration parameters of the primitives, abstract emotional states can be transformed into specific, executable whole-body movement trajectories, realizing the conversion from emotional semantics to motion control.
[0018] Furthermore, the motion-emotion co-planner optimizes the initial whole-body joint motion trajectory in a digital twin environment to generate the final whole-body joint motion trajectory. The trajectory optimization aims to minimize the deviation between the initial whole-body joint motion trajectory and the optimized trajectory, and uses spatial forbidden zone parameters, forbidden action parameters, dynamic feasibility constraints, self-obstacle avoidance constraints, and environmental obstacle avoidance constraints as constraints.
[0019] By optimizing the initial full-body joint motion trajectory in a digital twin environment with the goal of minimizing deviation, it can be ensured that the final executed trajectory maintains maximum consistency with the initial emotional expression intention while satisfying multiple constraints.
[0020] Furthermore, the method also includes: Step 8: The multimodal perception module collects feedback information from the interactive object after the robot executes the full-body joint motion trajectory. The emotion-ethics joint twin model includes an effect evaluation and parameter adaptor. The effect evaluation and parameter adaptor generates an acceptance score and a comfort score based on the feedback information, and updates the parameters of the emotion calculation and mapping module using a reinforcement learning algorithm based on the acceptance score and the comfort score.
[0021] By setting up an effect evaluation and parameter adaptor to update the parameters of the emotion computing and mapping module based on the feedback information from the interaction object, the robot can continuously optimize its emotion expression based on the user's actual reaction, thereby achieving a personalized human-computer interaction experience.
[0022] Furthermore, the multimodal information includes at least one of facial expression images, speech signals, posture and movement sequences, and physiological signals; The environmental information includes at least one of the following: scene type, spatial layout, distribution and status of other people in the environment; The preset cultural feature information includes at least one of geographical region information, language type information, and cultural custom tags.
[0023] By collecting multimodal information, including facial expression images, speech signals, posture and movement sequences, and physiological signals, as well as environmental information, including scene type and spatial layout, a more comprehensive data foundation can be provided for emotion understanding, thereby improving the accuracy of emotion state recognition.
[0024] On the other hand, the motion control system for AI humanoid robots based on digital twins is applicable to motion control methods for AI humanoid robots based on digital twins. The system comprises: A physical humanoid robot body, on which a multimodal perception module and a low-level motion execution module are deployed; An emotion-ethics joint twin that is synchronized in real time with the physical humanoid robot body, the emotion-ethics joint twin is deployed on an edge computing node or a cloud server, and includes an ethical constraint formal expression module, an emotion computing and mapping module, a motion-emotion co-planner, and an effect evaluation and parameter adaptor; A real-time synchronization communication module connects the physical humanoid robot body and the emotional-ethical joint twin for low-latency data synchronization. The multimodal perception module is used to collect multimodal information of interactive objects and environmental information; The formal expression module for ethical constraints is used to generate a set of ethical constraint parameters based on the environmental information and preset cultural characteristic information. The emotion computing and mapping module is used to generate the user's emotional state based on the multimodal information, and to generate the robot's target emotional state based on the user's emotional state, the environmental information, the cultural feature information, and the ethical constraint parameter set. The motion-emotion co-planner is used to generate whole-body joint motion trajectories based on the robot's target emotional state and the set of ethical constraint parameters. The underlying motion execution module is used to receive and execute the whole-body joint motion trajectory; The effect evaluation and parameter adaptor is used to generate acceptance and comfort scores based on the feedback information from the interaction object, and to update the parameters of the emotion calculation and mapping module.
[0025] By integrating the ethical constraint formalization expression module, the emotion computing and mapping module, the motion-emotion co-planner, and the effect evaluation and parameter adaptor into an emotion-ethics joint twin that is synchronized with the physical robot in real time, a complete closed loop of perception-understanding-planning-execution-optimization can be formed, realizing the emotional motion regulation under the constraints of social norms.
[0026] Compared with existing technologies, this AI humanoid robot motion control method and system based on digital twins has the following advantages: I. This invention constructs an emotion-ethics joint twin model that is synchronized in real time with a physical humanoid robot. It formalizes abstract ethical norms into feasible domain boundaries of the emotional state space and hard constraints of motion planning. It directly embeds ethical constraints into the underlying process of generating the robot's target emotional state and planning the trajectory of the whole-body joints. This replaces the serial architecture of the existing technology that generates emotions first and then filters ethics. It can fundamentally avoid the conflict between emotional expression and ethical constraints, eliminate the problems of robot motion stuttering and expression truncation during human-computer interaction, and prevent overstepping behavior caused by motion planning not incorporating ethical constraints. It achieves deep integration of emotional understanding, ethical compliance verification, and motion planning, effectively ensuring the natural fluency and safety compliance of humanoid robot interaction in social scenarios.
[0027] Second, the invention generates corresponding acceptance and comfort evaluations by setting up effect evaluation and parameter adaptors, combined with feedback information from interactive objects collected by the multimodal perception module, and uses reinforcement learning algorithms to complete the adaptive update of parameters of the emotion computing and mapping module. It can continuously optimize the robot's emotion expression strategy based on actual interaction effects. At the same time, relying on the built-in ethical knowledge graph and emotion-action primitive mapping library, it can flexibly adapt to the personalized needs of different scene types, cultural characteristics and interactive objects, greatly improve the adaptability of human-computer interaction and user experience, and effectively expand the application scope of humanoid robots in various social interaction scenarios.
[0028] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0030] Figure 1 This is a schematic diagram of the overall system architecture of the present invention; Figure 2 This is a schematic diagram illustrating the module interaction and data flow of the emotional-ethical joint twin model of the present invention; Figure 3 This is a schematic diagram of the trajectory generation and optimization process of the motion-emotion collaborative planning of the present invention; Figure 4 This is a diagram illustrating the method steps of the present invention. Detailed Implementation
[0031] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0032] Example This embodiment discloses a method and system for motion control of AI humanoid robots based on digital twins, which can be applied to social scenarios with high-frequency human-robot interaction, such as home services, medical rehabilitation, public guidance, and business reception. This embodiment constructs an emotion-ethics joint twin model that is synchronized in real time with the physical humanoid robot, transforming ethical constraints from abstract social norms into calculable parameter boundaries and hard constraints. This achieves a deep integration of emotion understanding, ethical compliance verification, and motion planning, avoiding the interaction stutters and boundary violations caused by the sequential architecture of emotion generation followed by ethical filtering in related technologies, thus ensuring the naturalness and safety of humanoid robot interactions in social scenarios.
[0033] like Figure 1 As shown, the control system in this embodiment includes three core components: a physical humanoid robot body, an emotion-ethics joint twin, and a real-time synchronous communication module.
[0034] The physical humanoid robot adopts a bipedal, full-body humanoid structure with no fewer than 20 degrees of freedom in its joints, including head joints, upper limb joints, lower limb joints, and trunk joints. Each joint is equipped with a high-precision servo motor and position encoder, enabling precise control of joint angles and real-time status feedback. The physical humanoid robot also features a multimodal perception module and a low-level motion execution module.
[0035] The multimodal perception module integrates multiple sensing units to collect multimodal information about the interacting object and environmental information. Specifically, the multimodal perception module includes a visual acquisition unit, a voice acquisition unit, a posture acquisition unit, and a physiological signal acquisition unit. The visual acquisition unit uses a combination of depth cameras and RGB cameras, both mounted at the robot's eyes, with a sampling frequency of at least 30 frames per second, to acquire facial expression images and posture sequences of the interacting object, while also acquiring information on the spatial layout and distribution of people in the environment. The voice acquisition unit uses a four-microphone array, mounted at the robot's ears, with a sampling frequency of at least 16kHz, to acquire the voice signals of the interacting object, while also achieving sound source localization and environmental sound information acquisition. The posture acquisition unit uses the robot's built-in inertial measurement unit, mounted at the core positions of the robot's torso and limbs, with a sampling frequency of at least 100Hz, to acquire the robot's own posture and motion state information. The physiological signal acquisition unit adopts a non-contact photoelectric sensing unit, which can acquire physiological signals such as heart rate and respiratory rate of the interactive object through the infrared channel of the visual acquisition unit, and can also acquire skin conductance signals of the interactive object through contact electrodes, providing supplementary data for emotional state recognition.
[0036] The underlying motion execution module includes a servo drive unit, a motion control unit, and a safety protection unit. The servo drive unit is connected one-to-one with the servo motors of each joint throughout the robot's body. It receives motion trajectory commands from all joints, outputs corresponding drive currents, and controls the servo motors to complete the corresponding actions. The motion control unit uses an embedded controller with a real-time performance of at least 1kHz to perform interpolation calculations of joint motion trajectories and forward and inverse kinematics operations, ensuring smooth and accurate motion execution. The safety protection unit has built-in torque thresholds, current thresholds, and position deviation thresholds. When the joint torque, drive current, or position deviation exceeds the corresponding threshold, it immediately triggers an emergency stop or compliant control mode to prevent damage to the interactive object or the robot itself.
[0037] The emotion-ethics joint twin is a digital twin virtual model that is synchronized in real time with the physical humanoid robot body. It is deployed on edge computing nodes or cloud servers. In this embodiment, it is preferentially deployed on edge computing nodes to reduce data transmission latency and ensure real-time interaction. The emotion-ethics joint twin includes four core functional modules: a formal expression module for ethical constraints, an emotion computing and mapping module, a motion-emotion co-planner, and an effect evaluation and parameter adaptive module.
[0038] The real-time synchronization communication module uses 5G LAN or wired Ethernet communication with a communication latency of no more than 10ms. It connects the physical humanoid robot body with the emotion-ethics joint twin to achieve bidirectional low-latency data synchronization. Specifically, the real-time synchronization communication module can upload multimodal information, environmental information, and joint motion state data collected by the physical humanoid robot body to the emotion-ethics joint twin in real time. At the same time, it can send the full-body joint motion trajectory commands generated by the emotion-ethics joint twin to the underlying motion execution module of the physical humanoid robot body in real time, ensuring that the virtual twin model and the physical entity are synchronized in real time.
[0039] like Figure 2 As shown in this embodiment, the four core modules of the emotion-ethics joint twin work collaboratively according to the preset data flow, realizing the entire process from multimodal information collection to target emotional state generation and then to compliant motion trajectory output.
[0040] The formal expression module for ethical constraints incorporates an ethical knowledge graph, which is used to transform abstract ethical norms, cultural customs, and social rules into structured, computable parameter data, ultimately generating a set of ethical constraint parameters.
[0041] Specifically, the ethics knowledge graph uses scenario type, cultural region, contact type, and spatial distance as indexes to store the permitted types and intensity ranges of emotional expression under different contextual conditions. The construction process of the ethics knowledge graph consists of three stages: knowledge acquisition, knowledge structuring, and knowledge verification. In the knowledge acquisition stage, social ethical norms, cultural customs, and interpersonal communication rules from different cultural regions and scenarios are collected to form a basic ethics knowledge base. In the knowledge structuring stage, the collected ethical norms are broken down into four core index dimensions: scenario type, cultural region, contact type, and spatial distance. Simultaneously, the permitted types of emotional expression, intensity ranges, safe distance thresholds, and prohibited actions are labeled for each index combination, forming structured graph entries. In the knowledge verification stage, professionals in sociology and ethics verify and correct the graph entries to ensure the accuracy and compliance of the graph content.
[0042] For example, scenario types can be categorized as private family scenarios, medical rehabilitation scenarios, public guided tour scenarios, and business reception scenarios. Cultural regions can be categorized as East Asian cultural regions, European and American cultural regions, and Southeast Asian cultural regions. Contact types can be categorized as non-contact, contact with non-sensitive areas, and contact with sensitive areas. Spatial distance can be categorized as intimate distance, personal distance, social distance, and public distance. Each combination of index dimensions corresponds to a unique graph entry, and each entry stores the constraint parameters under the corresponding conditions.
[0043] The specific process by which the formal expression module for ethical constraints generates a set of ethical constraint parameters is as follows: The module receives environmental information and preset cultural characteristic information, extracts scene type, spatial layout, and personnel distribution data from the environmental information, and extracts geographical region information, language type information, and cultural custom tags from the cultural characteristic information. The extracted data is then matched with the index dimensions of the ethical knowledge graph, and the graph entry with the highest matching degree is retrieved. Based on the retrieved entry data, the set of ethical constraint parameters is generated. The set of ethical constraint parameters includes three core parameters: spatial forbidden zone parameters, taboo action parameters, and emotional intensity feasible domain parameters.
[0044] Specifically, the spatial no-go zone parameters are represented as a set of prohibited areas in three-dimensional space. This set is determined based on the robot's end effector position coordinates, the user's position coordinates, a preset safety distance threshold for the scene type, and a contact type permission function. In this embodiment, the three-dimensional space uses a world coordinate system, with the origin fixed at the center of the ground at the robot's initial standing position. The X-axis points directly in front of the robot, the Y-axis points to the left of the robot, and the Z-axis is perpendicular to the ground and upwards. The prohibited area set is the set of areas in three-dimensional space that the robot's end effector is not allowed to enter. Specifically, it includes a spherical prohibited area centered on the interactive object's body with a radius equal to the safety distance threshold, as well as prohibited areas corresponding to preset fixed obstacles within the scene. The safety distance threshold is determined based on matched ethical knowledge graph entries, and the safety distance threshold varies for different scene types and cultural regions. The contact type permission function is used to determine the area permission status under different contact types. When the contact type is no contact, the prohibited area set covers the entire intimate distance range around the interactive object's body; when the contact type is non-sensitive part contact, only the area around the interactive object's sensitive parts is included in the prohibited area set.
[0045] Taboo action parameters are represented as forbidden subspaces within the joint space. The joint space comprises all the joint angles of the physical humanoid robot, with each joint's angle range corresponding to one dimension of the joint space. The forbidden subspace within the joint space is determined based on a pre-defined set of taboo joint angles. This set originates from the taboo action types marked in the corresponding entries of the ethics knowledge graph. Each taboo action type corresponds to a set of joint angle value ranges, and the space formed by these value ranges constitutes the forbidden subspace within the joint space. For example, an aggressive punching motion corresponds to a specific angle range for the shoulder and elbow joints of the upper limbs, and this range is included in the forbidden subspace; an excessive bowing motion that violates social etiquette corresponds to a specific angle range for the trunk and hip joints, and this range is also included in the forbidden subspace.
[0046] The feasible domain parameters for emotional intensity are represented by arousal threshold range and valence threshold range. The arousal threshold range is determined by the lower and upper limits of arousal based on scenario type and cultural characteristics. Similarly, the valence threshold range is determined by the lower and upper limits of valence based on scenario type and cultural characteristics. Specifically, arousal represents the level of emotional excitement, ranging from 0 to 1, with higher values indicating greater excitement and intensity. Valence represents the positive or negative tendency of an emotion, ranging from -1 to 1, with higher values indicating a more positive emotion. For example, in a medical rehabilitation scenario, the upper arousal threshold is set to 0.6 to prevent the robot's overly excited emotional expression from interfering with the patient; in a business reception scenario, the lower valence threshold is set to 0.2 to ensure the robot consistently maintains a positive emotional expression.
[0047] The emotion computing and mapping module includes an emotion understanding engine and an emotion-ethics joint mapper, which are used to identify the user's emotional state and generate the robot's target emotional state.
[0048] The emotion understanding engine employs a multimodal fusion neural network to analyze multimodal information and output the user's emotional state, which includes user arousal, valence, and dominance. Specifically, the multimodal fusion neural network uses a network architecture combining a multi-branch feature extraction structure and a cross-modal attention fusion structure. Each modality corresponds to an independent feature extraction branch: the branch for facial expression images uses a convolutional neural network structure to extract key facial expression features; the branch for speech signals uses a temporal convolutional network structure to extract pitch, speech rate, and volume-related features; the branch for posture and movement sequences uses a graph convolutional neural network structure to extract temporal motion features of key points in the human skeleton; and the branch for physiological signals uses a fully connected neural network structure to extract the variation features of physiological signals. The cross-modal attention fusion structure receives the single-modal features extracted from each branch and assigns corresponding weights to the features of different modalities through an attention mechanism, completing the fusion of multimodal features. Finally, it outputs the values of user arousal, valence, and dominance through a fully connected layer.
[0049] In this embodiment, user dominance is used to characterize the user's sense of control and dominance during the interaction process, with a value ranging from 0 to 1. The higher the value, the stronger the user's dominance. The emotion understanding engine is trained using a supervised training method. The training dataset uses a publicly available multimodal emotion recognition dataset and a custom-collected contextual emotion dataset. During training, the mean absolute error is used as the loss function, and the weight parameters of the network are updated through the backpropagation algorithm to ensure the accuracy of emotion state recognition.
[0050] The emotion-ethics co-mapping engine calculates the target emotional state of the robot based on the user's emotional state, environmental information, cultural characteristics, and a set of ethical constraint parameters. Specifically, the emotion-ethics co-mapping engine first determines a feasible region of emotional intensity based on environmental and cultural information. This feasible region is completely consistent with the feasible region of emotional intensity parameters in the ethical constraint parameter set, namely, the range of arousal threshold and valence threshold. Subsequently, the emotion-ethics co-mapping engine constructs a resonance function based on psychological resonance theory. By minimizing the deviation between the robot's target emotional state and the resonance function value, it selects the robot's target emotional state within the feasible region of emotional intensity.
[0051] The optimal formula for solving the robot target's emotional state in this embodiment is as follows: This formula is a constrained quadratic optimization problem. The goal is to find the robot's target emotional state with the smallest deviation from the psychological resonance function value within the feasible domain of emotional intensity, thereby achieving the optimal emotional resonance expression under ethical constraints.
[0052] The minimum value is the symbol used in mathematics to find the variable value that makes the subsequent expression minimum. The square operation is performed on the L2 norm, and the sum of squares of each element in the vector is calculated. This is used to characterize the Euclidean distance between two vectors, and here it is used to measure the degree of deviation between the robot's target emotional state and the resonance function value. The arousal value represents the target emotional state of the robot, and is the core variable to be optimized, with a value range of 0 to 1. denoted as the valence value of the robot's target emotional state, and denoted as the core variable to be optimized, with a value range of -1 to 1. The value represents the degree of dominance of the robot's target emotional state, and is the core variable to be optimized, with a value range of 0 to 1. The arousal value of the user's emotional state output by the emotion understanding engine is the input parameter, with a value range of 0 to 1. The valence value of the user's emotional state output by the emotion understanding engine is the input parameter, which ranges from -1 to 1. The dominant value of the user's emotional state output by the emotion understanding engine is the input parameter, which ranges from 0 to 1. is the scene type parameter in the environmental information, and is the input parameter, corresponding to the scene type index dimension in the ethics knowledge graph. is the preset cultural feature information parameter, and is the input parameter, corresponding to the cultural region index dimension in the ethics knowledge graph.
[0053] The resonance function, representing the arousal dimension, is determined based on psychological resonance theory and is used to calculate the arousal value that the robot should output under ideal conditions. Specifically, the value of the resonance function is determined according to the scenario type and cultural characteristics. In empathic scenarios, the value of the resonance function is positively correlated with the user's arousal value; in soothing scenarios, the value of the resonance function is negatively correlated with the user's arousal value, in order to soothe the user's emotions.
[0054] The resonance function, defined based on psychological resonance theory, is used to calculate the valence value that the robot should output under ideal conditions. Specifically, in positive emotional interaction scenarios, the value of the resonance function is positively correlated with the user's valence value, achieving emotional resonance; in negative emotional reassurance scenarios, the value of the resonance function is higher than the user's valence value, guiding the user's emotional recovery with positive emotions.
[0055] The resonance function, representing the dominance dimension, is determined based on psychological resonance theory and is used to calculate the dominance value that the robot should output under ideal conditions. Specifically, in service scenarios, the resonance function value is always lower than the user's dominance value to ensure the user's dominant position in the interaction process; in guidance scenarios, the resonance function value can be appropriately higher than the user's dominance value to achieve reasonable guidance for the user.
[0056] The symbol represents a constraint in mathematics, and the subsequent expression represents a hard constraint that must be satisfied during the optimization process. The lower threshold of arousal is a fixed parameter in the set of ethical constraint parameters, ranging from 0 to 1. The upper limit threshold of arousal is a fixed parameter in the set of ethical constraint parameters. Its value ranges from 0 to 1 and is always greater than or equal to the lower limit threshold of arousal. is the lower limit threshold of valence in the set of ethical constraint parameters, which is a fixed parameter with a value range of -1 to 1. The upper limit threshold of valence is a fixed parameter in the set of ethical constraint parameters, ranging from -1 to 1, and is always greater than or equal to the lower limit threshold of valence.
[0057] It is understandable that, through this optimization formula, the generation process of the robot's target emotional state directly uses ethical constraint parameters as hard boundary conditions, rather than filtering and verifying them after generation. This fundamentally avoids the conflict between emotional expression and ethical constraints, ensuring that the generated target emotional state always conforms to ethical norms and scenario requirements. In this embodiment, the optimization problem is solved using the gradient descent method. The solution process is completed in the digital twin environment of the emotion-ethics joint twin, and the solution time does not exceed 20ms, ensuring the real-time nature of the interaction.
[0058] The motion-emotion co-planner includes an emotion-action primitive mapping library, which is used to generate full-body joint motion trajectories that simultaneously satisfy multiple constraints based on the robot's target emotional state and ethical constraint parameter set.
[0059] Specifically, the emotion-motion primitive mapping library stores full-body motion trajectory primitives with emotion tags. Each motion trajectory primitive is a predefined, fixed-duration sequence of full-body joint movements. Each trajectory primitive corresponds to a unique emotion tag, which includes two dimensions: emotion type and emotion intensity. The emotion type corresponds to a combination of arousal, valence, and dominance, while the emotion intensity corresponds to the numerical value of arousal. For example, emotion tags can be categorized into types such as friendly greetings, gentle reassurance, enthusiastic welcome, and calm guidance, each corresponding to a different trajectory primitive. The construction process of the trajectory primitives involves capturing standardized social movements of professionals through motion capture, processing them through kinematic redirection, and adapting them to the joint structure of a physical humanoid robot, ultimately forming a standardized trajectory primitive library.
[0060] The specific process by which the motion-emotion co-planner generates the initial full-body joint motion trajectory is as follows: First, based on the arousal, valence, and dominance values of the robot's target emotional state, corresponding emotion tags are generated. Then, the emotion-action primitive with the highest matching degree with the emotion tag is selected from the emotion-action primitive mapping library. The matching degree calculation uses a cosine similarity algorithm to calculate the cosine similarity between the robot's target emotional state vector and the emotion tag vector corresponding to the trajectory primitive, and the trajectory primitive with the highest similarity is selected as the basic primitive. Subsequently, the motion-emotion co-planner adjusts the amplitude, velocity, and duration parameters of the emotion-action primitive to generate the initial full-body joint motion trajectory. Specifically, the adjustment ratio of the amplitude parameter is positively correlated with the arousal value of the robot's target emotional state; the higher the arousal value, the larger the movement amplitude. The adjustment ratio of the velocity parameter is also positively correlated with the arousal value of the robot's target emotional state; the higher the arousal value, the faster the movement speed. The adjustment ratio of the duration parameter is determined according to the rhythm of the interaction scene to ensure the synchronization between the action and the voice interaction.
[0061] The specific process by which the motion-emotion co-planner generates the final full-body joint motion trajectory is as follows: in a digital twin environment, the initial full-body joint motion trajectory is optimized to generate the final full-body joint motion trajectory. The trajectory optimization aims to minimize the deviation between the initial full-body joint motion trajectory and the optimized trajectory, and uses spatial forbidden zone parameters, forbidden action parameters, dynamic feasibility constraints, self-obstacle avoidance constraints, and environmental obstacle avoidance constraints as constraints.
[0062] The constraint optimization formula for trajectory optimization in this embodiment is as follows: This formula is a trajectory optimization problem with multiple constraints. The goal is to generate an optimized trajectory that deviates from the initial trajectory while satisfying all hard constraints. This ensures that the optimized trajectory not only conforms to the initial intention of emotional expression, but also fully meets ethical constraints and kinematic and dynamic requirements.
[0063] The minimum value is the symbol used in mathematics to find the variable value that makes the subsequent expression minimum. This is the definite integral operator, used to perform integration on an expression within the time interval from 0 to T. Here, it is used to calculate the cumulative deviation between the optimized trajectory and the initial trajectory over the entire motion cycle. Let t be the trajectory function of the robot's joint angles changing over time after optimization, and t be the core variable to be optimized, where t is the time variable, ranging from 0 to T, and T is the total duration of the entire motion process. The initial whole-body joint motion trajectory function is defined as , and the input parameters are generated after adjustment by the emotional action primitives. This is the square operation of the L2 norm, and this is used to calculate the sum of squares of each element in the vector. Here, it is used to measure the joint angle deviation between the optimized trajectory and the initial trajectory at the same moment. The symbol represents a constraint in mathematics, and the subsequent expression represents a hard constraint that must be satisfied during the optimization process. The position coordinate function of the robot's end effector in the three-dimensional world coordinate system is obtained by calculating the forward kinematics of the robot, with the joint angle trajectory as the input. . The set of three-dimensional prohibited regions corresponding to the spatial forbidden zone parameters in the set of ethical constraint parameters is a fixed parameter.
[0064] The constraint is that the elements on the left are not within the set on the right. This constraint means that throughout the entire motion, the position of the robot's end effector never enters a restricted space, thus satisfying the spatial distance requirement in the ethical constraints.
[0065] This is the prohibited subspace of joint space corresponding to the forbidden action parameters in the ethical constraint parameter set, and these are fixed parameters. The meaning of this constraint is that throughout the entire motion, the robot's full-body joint angle combinations must never enter the prohibited subspace of joint space, thus avoiding the execution of forbidden actions and satisfying the action specification requirements in the ethical constraints.
[0066] Let be the robot's inertia matrix, and be a known parameter of robot dynamics, derived from joint angles. The inertial characteristics of each joint of the robot were calculated. Joint angle trajectory The second derivative with respect to time, i.e., the joint angular acceleration function. Let be the Coriolis force and centrifugal force matrix of the robot, and be the known parameters of robot dynamics, derived from joint angles. With joint angular velocity The Coriolis force and centrifugal force characteristics of the robot during motion were calculated. Joint angle trajectory The first derivative with respect to time, i.e., the joint angular velocity function. Let be the robot's gravity vector, and be a known parameter of robot dynamics, derived from joint angles. The calculated torques characterize the force exerted by gravity on each joint of the robot. The minimum output torque threshold for each joint of the robot is a fixed parameter determined by the hardware parameters of the servo motor. The maximum output torque threshold for each joint of the robot is a fixed parameter determined by the hardware parameters of the servo motor, and it is always greater than or equal to the minimum output torque threshold. This constraint is a dynamic feasibility constraint for the robot, meaning that throughout the entire motion process, the driving torque required by each joint of the robot is always within the output capability range of the servo motor, ensuring the executability of the motion trajectory and avoiding situations such as motor step loss or overload.
[0067] The minimum distance function between the robot's limbs is given by the joint angles. Calculated. The preset minimum safe distance threshold for obstacle avoidance is a fixed parameter, set to 0.05 meters in this embodiment. This constraint is an obstacle avoidance constraint, meaning that throughout the entire movement, the distance between the robot's limbs is always greater than or equal to the minimum safe distance, preventing the robot from colliding with itself. The minimum distance function between the robot body and obstacles in the environment is given by the joint angles. It is calculated from the spatial layout data in the environmental information. The preset minimum safe distance threshold for obstacle avoidance is a fixed parameter, set to 0.1 meters in this embodiment. This constraint is an environmental obstacle avoidance constraint, meaning that throughout the entire movement process, the distance between the robot body and obstacles in the environment is always greater than or equal to the minimum safe distance, thus preventing the robot from colliding with objects or people in the environment.
[0068] It is understandable that through this trajectory optimization formula, ethical constraint parameters are directly embedded as hard constraints in the trajectory planning process, rather than being filtered and verified after trajectory generation. This fundamentally avoids ethical transgressions during motion execution, ensuring that the robot's motion not only conforms to the emotional expression intent but also fully meets ethical norms and safety requirements. In this embodiment, the trajectory optimization problem is solved using a sequential quadratic programming algorithm. The solution process is completed in the digital twin environment of the emotion-ethics joint twin, and kinematic simulation verification is performed simultaneously during the solution process to ensure that the optimized trajectory is fully executable.
[0069] like Figure 3 As shown, the process includes eight core steps: emotion tag generation, action primitive matching, primitive parameter adjustment, initial trajectory generation, constraint loading, trajectory optimization solution, simulation verification, and final trajectory output. It fully covers the entire process from target emotional state to executable motion trajectory.
[0070] The effect evaluation and parameter adaptor is used to adaptively update the parameters of the emotion calculation and mapping module based on the feedback information of the interaction object, so as to continuously optimize the interaction effect.
[0071] Specifically, after the humanoid robot executes its full-body joint motion trajectory, the multimodal perception module collects feedback information from the interacting object. This feedback information includes changes in facial expressions, voice feedback content, posture changes, and physiological signal changes. The effect evaluation and parameter adaptor receives this feedback information and generates an acceptance score and a comfort score. The acceptance score characterizes the degree to which the interacting object accepts the robot's emotional expressions and actions, ranging from 0 to 1, with higher values indicating higher acceptance. The comfort score characterizes the degree of comfort the interacting object experiences during the interaction, also ranging from 0 to 1, with higher values indicating higher comfort. The calculation of the acceptance and comfort scores employs a multimodal fusion evaluation model. The model input consists of the multimodal features of the feedback information, and the output consists of the two score values. The model is trained using a supervised training method, with the training dataset being a human-computer interaction feedback dataset labeled with acceptance and comfort scores.
[0072] The effect evaluation and parameter adaptor update the parameters of the emotion computing and mapping module using a reinforcement learning algorithm based on the acceptance and comfort scores. Specifically, the reinforcement learning algorithm employs a deep deterministic policy gradient algorithm, using the weighted sum of the acceptance and comfort scores as the reward function, and the resonance function parameters of the emotion-ethics joint mapper and the network weight parameters of the emotion understanding engine as the policy parameters to be optimized. Through continuous accumulation of human-computer interaction data, the parameters are iteratively updated, making the robot's emotional expression increasingly adaptable to the personalized needs of the interaction object, thereby improving the human-computer interaction experience.
[0073] The above describes the specific implementation methods of each component of the system. The following details the complete implementation steps of the AI humanoid robot motion control method based on digital twins in this embodiment, such as... Figure 4 As shown, this method is based on the above system, fully covering all technical features, and ensuring that technicians can implement it without pressure.
[0074] Step one involves constructing an emotion-ethics joint twin model that is synchronized in real-time with the physical humanoid robot. Specifically, a digital twin virtual environment is built in an edge computing node. Within this virtual environment, a virtual robot model is constructed that is completely identical to the physical humanoid robot in terms of structure, parameters, and motion characteristics. Simultaneously, a formal expression module for ethical constraints, an emotion computing and mapping module, a motion-emotion co-planner, and an effect evaluation and parameter adaptor are integrated into the virtual robot model to form the emotion-ethics joint twin model. After construction, a real-time synchronization communication module is used to perform real-time synchronization calibration between the physical humanoid robot and the emotion-ethics joint twin model. The calibration includes joint zero points, kinematic parameters, and dynamic parameters, ensuring that the virtual model and the physical entity are completely consistent in state.
[0075] Step two involves collecting multimodal information about the interacting object and environmental information through a multimodal perception module deployed on the physical humanoid robot. Specifically, the vision acquisition unit of the multimodal perception module acquires facial expression images and posture sequences of the interacting object, while simultaneously acquiring information about the spatial layout of the environment and the distribution and status of other people in the environment; the voice acquisition unit acquires the voice signal of the interacting object, while simultaneously acquiring sound information from the environment; and the physiological signal acquisition unit acquires physiological signals such as the heart rate and skin conductance of the interacting object. After the acquisition is completed, the acquired multimodal information and environmental information are uploaded in real time to the emotion-ethics joint twin model through a real-time synchronous communication module.
[0076] Step three: The sentiment computing and mapping module generates the user's emotional state for the interactive object based on multimodal and environmental information. Specifically, the sentiment understanding engine in the sentiment computing and mapping module receives multimodal information and preprocesses the information from different modalities. The preprocessing includes image denoising, speech denoising, and data normalization. Then, the preprocessed multimodal data is input into a multimodal fusion neural network. After feature extraction and cross-modal fusion, it outputs the user's emotional state in three dimensions: user arousal, user valence, and user dominance, thus completing the recognition of the user's emotional state.
[0077] Step four: The formal expression module for ethical constraints generates a set of ethical constraint parameters based on environmental information and preset cultural characteristic information. Specifically, the module receives environmental information and preset cultural characteristic information, extracts scene type, spatial layout, and personnel distribution data from the environmental information, and extracts geographical region information, language type information, and cultural custom tags from the cultural characteristic information. It then matches the extracted data with the index dimensions of the built-in ethical knowledge graph, retrieves the graph entry with the highest matching degree, and generates a set of ethical constraint parameters based on the retrieved entry data. The set of ethical constraint parameters includes three core parameters: spatial forbidden zone parameters, taboo action parameters, and emotional intensity feasible domain parameters.
[0078] Step five: The emotion computing and mapping module generates a target emotional state for the robot that falls within the feasible domain of emotional intensity parameters, based on the user's emotional state, environmental information, cultural characteristics, and ethical constraint parameters. Specifically, the emotion-ethics joint mapper in the emotion computing and mapping module receives the user's emotional state, environmental information, cultural characteristics, and ethical constraint parameters. It first determines the feasible domain of emotional intensity, then constructs a resonance function based on psychological resonance theory, and solves for the robot's target emotional state using a constrained quadratic optimization formula. This ensures that the arousal and valence values of the generated target emotional state always remain within the feasible domain of emotional intensity parameters.
[0079] Step Six: The emotion-ethics joint twin model generates a full-body joint motion trajectory that simultaneously satisfies kinematic constraints, dynamic constraints, spatial forbidden zone parameters, and taboo action parameters, based on the robot's target emotional state and ethical constraint parameter set. Specifically, the motion-emotion co-planner in the emotion-ethics joint twin model receives the robot's target emotional state and ethical constraint parameter set. First, it matches the corresponding emotional action primitives according to the robot's target emotional state. After adjusting the primitive parameters, it generates an initial full-body joint motion trajectory. Then, in the digital twin environment, with the goal of minimizing the deviation between the initial trajectory and the optimized trajectory, and with ethical constraint parameters, dynamic feasibility constraints, self-obstacle avoidance constraints, and environmental obstacle avoidance constraints as constraints, it completes the trajectory optimization solution. After simulation verification, the final full-body joint motion trajectory is generated.
[0080] Step seven: The physical humanoid robot executes the full-body joint motion trajectory. Specifically, the emotion-ethics joint twin model sends the generated full-body joint motion trajectory to the underlying motion execution module of the physical humanoid robot in real time through a real-time synchronous communication module. The motion control unit of the underlying motion execution module performs interpolation calculations on the trajectory commands, generating real-time control commands for each joint. The servo drive unit drives the servo motors of the corresponding joints according to the control commands, enabling the physical humanoid robot to complete the corresponding full-body movement and achieve human-computer interaction that conforms to emotional expression intentions and ethical norms.
[0081] Step eight involves collecting feedback information from the interactive object after the robot executes its full-body joint motion trajectory using the multimodal perception module, and then adaptively updating the module parameters. Specifically, after the robot completes its motion execution, the multimodal perception module continuously collects feedback information from the interactive object and uploads it to the effect evaluation and parameter adaptor of the emotion-ethics joint twin model. The effect evaluation and parameter adaptor generates an acceptance score and a comfort score based on the feedback information. Subsequently, a reinforcement learning algorithm is used, with the weighted sum of the two scores as the reward function, to update the parameters of the emotion calculation and mapping module, completing a full interaction loop.
[0082] In some optional implementations, the ethics knowledge graph can support online updates, regularly pushing the latest ethical norms and cultural customs data through a cloud server to update and expand the graph entries, enabling the robot to adapt to more scenarios and cultural regions.
[0083] In some optional implementations, the emotion-action primitive mapping library can support custom extensions. Users can add custom emotion-action primitives and label them with corresponding emotion tags according to the needs of actual application scenarios, making the robot's emotion expression richer and more personalized.
[0084] In some alternative implementations, the emotion-ethics joint twin model can simultaneously support the synchronous deployment of multiple robots. The twin model can be shared and synchronized through a cloud server, enabling multiple physical humanoid robots to simultaneously call the functions of the twin model and complete multi-robot collaborative human-robot interaction tasks.
[0085] In this embodiment, by formalizing ethical constraints into computable parameter boundaries and embedding them into the underlying processes of emotional state generation and motion trajectory planning, a deep integration of emotional understanding, ethical compliance, and motion planning is achieved. This avoids the interaction stuttering and boundary-crossing behaviors caused by the serial architecture in related technologies, ensuring the naturalness, fluency, and safety of humanoid robot interactions in social scenarios. Furthermore, this embodiment achieves adaptive optimization of robot emotional expression through a complete closed-loop optimization mechanism, adapting to the personalized needs of different interaction objects and possessing strong practicality and scalability.
[0086] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for motion control of an AI humanoid robot based on digital twins, characterized in that, The method includes the following steps: Step 1: Construct an emotion-ethics joint twin model that is synchronized in real time with the physical humanoid robot. The emotion-ethics joint twin model includes a formal expression module for ethical constraints and an emotion computing and mapping module. Step two: Collect multimodal information of the interactive object and environmental information through the multimodal perception module deployed on the physical humanoid robot; Step 3: The emotion computing and mapping module generates the user's emotional state of the interactive object based on the multimodal information and the environmental information; Step four: The formal expression module of ethical constraints generates a set of ethical constraint parameters based on the environmental information and the preset cultural characteristic information. The set of ethical constraint parameters includes spatial forbidden zone parameters, taboo action parameters, and feasible domain parameters of emotional intensity. Step 5: The emotion calculation and mapping module generates the robot target emotion state within the range of the feasible domain parameters of emotion intensity based on the user's emotional state, the environmental information, the cultural feature information, and the ethical constraint parameter set. Step six: The emotion-ethics joint twin model generates a whole-body joint motion trajectory that simultaneously satisfies kinematic constraints, dynamic constraints, spatial forbidden zone parameters, and forbidden action parameters based on the robot target emotional state and the ethical constraint parameter set. Step 7: The physical humanoid robot executes the full-body joint movement trajectory.
2. The motion control method for AI humanoid robots based on digital twins according to claim 1, characterized in that, The formal expression module of ethical constraints includes an ethical knowledge graph, which stores the types of emotional expression and the range of emotional intensity allowed under different contextual conditions, indexed by scene type, cultural region, contact type and spatial distance. The spatial forbidden zone parameters are represented as a set of forbidden regions in three-dimensional space, the forbidden action parameters are represented as forbidden subspaces in joint space, the joint space is composed of all joint angles of the physical humanoid robot, and the emotional intensity feasible domain parameters are represented as arousal threshold range and valence threshold range.
3. The motion control method for AI humanoid robots based on digital twins according to claim 2, characterized in that, The set of prohibited areas in the three-dimensional space is determined based on the robot end-effector position coordinates, user position coordinates, a preset safe distance threshold for the scene type, and a contact type permission function. The forbidden subspace in the joint space is determined according to a preset set of forbidden joint angles; The arousal threshold range is determined by the lower and upper arousal thresholds based on the scene type and cultural characteristics information, and the valence threshold range is determined by the lower and upper valence thresholds based on the scene type and cultural characteristics information.
4. The motion control method for AI humanoid robots based on digital twins according to claim 1, characterized in that, The emotion computing and mapping module includes an emotion understanding engine and an emotion-ethics joint mapper; The emotion understanding engine uses a multimodal fusion neural network to analyze the multimodal information and output the user's emotional state, which includes user arousal, user valence, and user dominance. The emotion-ethics joint mapper calculates the target emotion state of the robot based on the user's emotional state, the environmental information, the cultural characteristic information, and the set of ethical constraint parameters.
5. The motion control method for AI humanoid robots based on digital twins according to claim 4, characterized in that, The emotion-ethics joint mapper determines the feasible region of emotion intensity based on the environmental information and the cultural characteristic information. By minimizing the deviation between the robot's target emotional state and the resonance function value, it selects the robot's target emotional state within the feasible region of emotion intensity. The resonance function is determined based on the psychological resonance theory.
6. The motion control method for AI humanoid robots based on digital twins according to claim 1, characterized in that, The emotion-ethics joint twin model includes a motion-emotion co-planner, which contains an emotion-action primitive mapping library that stores whole-body motion trajectory primitives with emotion tags. The motion-emotion co-planner selects matching emotional action primitives from the emotion-action primitive mapping library based on the target emotional state of the robot, and adjusts the amplitude, velocity and duration parameters of the emotional action primitives to generate an initial whole-body joint motion trajectory.
7. The motion control method for AI humanoid robots based on digital twins according to claim 6, characterized in that, The motion-emotion co-planner optimizes the initial whole-body joint motion trajectory in the digital twin environment to generate the final whole-body joint motion trajectory. The trajectory optimization aims to minimize the deviation between the initial whole-body joint motion trajectory and the optimized trajectory, and uses spatial forbidden zone parameters, forbidden action parameters, dynamic feasibility constraints, self-obstacle avoidance constraints, and environmental obstacle avoidance constraints as constraints.
8. The motion control method for AI humanoid robots based on digital twins according to claim 1, characterized in that, The method further includes: Step 8: The multimodal perception module collects feedback information from the interactive object after the robot executes the full-body joint motion trajectory. The emotion-ethics joint twin model includes an effect evaluation and parameter adaptor. The effect evaluation and parameter adaptor generates an acceptance score and a comfort score based on the feedback information, and updates the parameters of the emotion calculation and mapping module using a reinforcement learning algorithm based on the acceptance score and the comfort score.
9. The motion control method for AI humanoid robots based on digital twins according to claim 1, characterized in that, The multimodal information includes at least one of facial expression images, speech signals, posture and movement sequences, and physiological signals; The environmental information includes at least one of the following: scene type, spatial layout, distribution and status of other people in the environment; The preset cultural feature information includes at least one of geographical region information, language type information, and cultural custom tags.
10. A motion control system for an AI humanoid robot based on digital twins, applicable to the motion control method for an AI humanoid robot based on digital twins as described in any one of claims 1 to 9, characterized in that, The system consists of: A physical humanoid robot body, on which a multimodal perception module and a low-level motion execution module are deployed; An emotion-ethics joint twin that is synchronized in real time with the physical humanoid robot body, the emotion-ethics joint twin is deployed on an edge computing node or a cloud server, and includes an ethical constraint formal expression module, an emotion computing and mapping module, a motion-emotion co-planner, and an effect evaluation and parameter adaptor; A real-time synchronization communication module connects the physical humanoid robot body and the emotional-ethical joint twin for low-latency data synchronization. The multimodal perception module is used to collect multimodal information of interactive objects and environmental information; The formal expression module for ethical constraints is used to generate a set of ethical constraint parameters based on the environmental information and preset cultural characteristic information. The emotion computing and mapping module is used to generate the user's emotional state based on the multimodal information, and to generate the robot's target emotional state based on the user's emotional state, the environmental information, the cultural feature information, and the ethical constraint parameter set. The motion-emotion co-planner is used to generate whole-body joint motion trajectories based on the robot's target emotional state and the set of ethical constraint parameters. The underlying motion execution module is used to receive and execute the whole-body joint motion trajectory; The effect evaluation and parameter adaptor is used to generate acceptance and comfort scores based on the feedback information from the interaction object, and to update the parameters of the emotion calculation and mapping module.
Citation Information
Patent Citations
Man-machine physical twin system based on multi-modal video analysis and adaptive mapping and control method
CN121148009A