A Robot Humanoid Behavior Planning and Control Method for Close Physical Interaction
Through motion capture and cluster analysis, a directed graph is constructed and a behavior planner is trained, which solves the problem of autonomous behavior planning and control of robots in close physical interaction scenarios, and realizes efficient and interpretable robot behavior.
Patent Information
- Application Number
- CN202411316401.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-09-20
AI Technical Summary
Robots lack the ability to handle emergencies in open dynamic environments and close collaboration scenarios, and their autonomy and efficiency need to be improved, especially in terms of feeling, understanding and reasoning their own embodied physical interactions with the world.
Through motion capture, the demonstration data of the close physical interaction of humans is obtained, and the human behavior pattern categories are obtained using prior knowledge and cluster analysis, a hierarchical directed graph is constructed and a behavior planner is trained to realize robot human-like behavior planning and control.
It realizes the autonomous behavior planning and control of robots in close physical interaction scenarios, improves the interpretability, anthropomorphism and adaptability of robot behavior, and meets the requirements of fast and dynamic human-machine collaboration.
Smart Images

Figure CN119077734B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotics technology, and in particular to a robot humanoid behavior planning and control method for close physical interaction. Background Art
[0002] With the development of robotics and artificial intelligence technology, machines have gradually evolved from pre-programmed automation systems to intelligent systems with perception, with the ability to adjust behavior based on environmental perception, and even the ability to predict human intentions. However, they still lack the ability to handle emergencies in open dynamic environments and close collaboration scenarios, and their autonomy and efficiency still need to be improved. Therefore, in the future, there is an urgent need for embodied intelligence that can make autonomous decisions and perform physical interaction tasks like humans, which is also one of the development directions foreseen at the beginning of the birth of artificial intelligence. One of the urgent problems to be solved in embodied intelligence is how robots can feel, understand and reason about their embodied physical interactions with the world. For example, in the field of home services, when service robots face close physical interaction tasks such as double-arm hugs and dexterous grasping, understanding their own behavior is of great significance to ensuring execution efficiency and service quality. However, due to the differences in cross-modal information and the difficulty of self-supervised learning, it is extremely challenging for robots to autonomously understand their own behavior and predict the perceived consequences caused.
[0003] In order to solve the problem of robot behavior understanding and perception consequence prediction, cognitive developmental robotics methods have gradually become the mainstream of embodied intelligence research. At present, the research on cognitive developmental robotics methods in the field of behavior understanding is still in its infancy. In addition, embodied intelligence requires robots to have the ability to develop autonomously and learn quickly in order to adapt to the dynamic uncertainty of the environment in real time. Therefore, how to reduce the method's dependence on large amounts of data, enhance the ability of autonomous development, and at the same time improve the task versatility of robot behavior understanding and perception consequence prediction has become a problem that needs to be solved in this field. Summary of the invention
[0004] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a robot humanoid behavior planning and control method for close physical interaction, so as to achieve close physical interaction between man and machine.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] One aspect of the present invention provides a method for planning and controlling a robot humanoid behavior for close physical interaction, comprising the following steps:
[0007] Step S1, obtaining demonstration data of close interaction between the robot and humans through motion capture;
[0008] Step S2: Based on the demonstration data, using prior knowledge and cluster analysis, obtain multiple human behavior pattern categories;
[0009] Step S3: Based on the human behavior pattern categories, by segmenting and calibrating the demonstration data, obtain multiple groups of motion primitive sequences including class behavior pattern labels and construct a hierarchical directed graph, and obtain a robot behavior planner through training;
[0010] Step S4: Construct a dynamically consistent mapping model between the target trajectory and the action space, and based on the mapping model and the planner, realize robot humanoid behavior planning and control.
[0011] As a preferred technical solution, in the step S2, the process of using prior knowledge and cluster analysis to obtain multiple human behavior pattern categories includes the following steps:
[0012] Step S201: Based on the demonstration data, through density-based cluster analysis, obtain a clustering result;
[0013] Step S202: Based on the clustering result and the previously obtained prior knowledge, obtain the key attributes in human close physical interaction behaviors, and construct human behavior pattern categories.
[0014] As a preferred technical solution, the step S202 includes:
[0015] Step S2021: Obtain prior knowledge based on the pre-acquired interdisciplinary literature;
[0016] Step S2022: Based on the prior knowledge and the clustering result, obtain the key attributes in human close physical interaction behaviors, where the key attributes include tightness, hug, and double-arm cooperation style;
[0017] Step S2023: Based on the key attributes, divide the hug action into multiple human behavior pattern categories according to whether the chest is in contact, the extension direction of the upper arm, and the force direction of the double arms.
[0018] As a preferred technical solution, in the step S3, the process of segmenting and calibrating the demonstration data based on the human behavior pattern categories includes:
[0019] Construct corresponding motion primitives for each human behavior pattern category, segment and calibrate the demonstration data, and obtain multiple groups of motion primitive sequences including class behavior pattern labels.
[0020] As a preferred technical solution, in the step S3, the hierarchical directed graph is constructed as:
[0021] G=(S,V,R,P,σ)
[0022] Among them, S represents the target hugging action, V is an AND node or an OR node in the directed graph, R represents the top-down production rule from the parent node α to its child node β, P represents the probability associated with each production rule, and σ is the behavior planning sequence defined by the grammar.
[0023] As a preferred technical solution, in the step S3, the process of obtaining the robot behavior planner through training includes:
[0024] Using the joint space of the two arms as the state space, taking the probability P between each node as the learning object, using the Q-learning rule in the way of time difference to learn the association between the motion primitive space and the state space, and restoring the grammar structure automatically through structural automatic distillation according to the posterior probability to obtain the robot behavior planner.
[0025] As a preferred technical solution, in the step S4, based on the mapping model and the planner, implementing the humanoid behavior planning and control of the robot includes:
[0026] For different joint states, using the robot behavior planner to generate the next motion primitive required to complete the target hugging action, and using the corresponding mapping model to execute the target hugging action.
[0027] As a preferred technical solution, the demonstration data includes the captured data of the hugging behavior actions of multiple participants of multiple age groups in multiple roles in multiple preset scenarios. Among them, the preset scenarios include social occasions, intimate relationships, emotional expressions, and motor functions, and the roles include the initiator and the recipient.
[0028] Another aspect of the present invention provides an electronic device, including: one or more processors and a memory, where the memory stores one or more programs, and the one or more programs include instructions for executing the aforementioned method for humanoid behavior planning and control of a robot for close physical interaction.
[0029] Another aspect of the present invention provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the aforementioned method for humanoid behavior planning and control of a robot for close physical interaction.
[0030] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0031] (1) Full-process coverage of the realization of robot humanoid behavior: The present invention is directed to close human-robot physical interaction, and based on human behavior demonstrations, realizes all processes of robot behavior learning, planning, and execution, enabling it to meet the requirements of fast and dynamic human-robot collaboration.
[0032] (2)Accurate modeling of behaviors: The present invention combines human demonstrations and prior knowledge to distinguish different behavior patterns. By adopting a clear and distinguishable interactive pattern classification method, the human close physical interaction behaviors are modeled as a sequence of motor primitives, which helps to simplify the planning and execution of robot behaviors.
[0033] (3)Effectively improves the interpretability of robot behaviors: The present invention learns a human motor primitive sequence planner based on a hierarchical directed graph method, and further combines a robot motor primitive control method derived from human behavior pattern analysis, effectively improving the interpretability, anthropomorphism and adaptability of robot behaviors, and effectively alleviating the contradiction between the planning interpretability and execution performance of robot interaction behaviors. Brief Description of the Drawings
[0034] Figure 1 It is a flowchart of a humanoid behavior planning and control method for a robot facing close physical interaction in the embodiment;
[0035] Figure 2 It is a schematic diagram of the idea of human behavior pattern analysis in the embodiment;
[0036] Figure 3 It is a schematic diagram of a directed graph planner and motor primitives provided in the embodiment;
[0037] Figure 4 It is a schematic diagram of the full-process framework of humanoid behavior planning and execution for a robot facing close physical interaction provided in the embodiment. Detailed Embodiment
[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] Embodiment 1
[0040] This embodiment provides a humanoid behavior planning and control method for a robot facing close physical interaction, with particular reference to human behavior demonstrations. The close physical interaction in this embodiment mainly refers to a binary hug, only focusing on the operations of the two arms. This method analyzes the human hug pattern categories based on human demonstration data and prior knowledge, effectively guiding the calibration of human behaviors, realizing the parsing of the robot hug behavior planning grammar, and guiding the setting and learning of Dynamic Movement Primitives (DMPs) to improve the anthropomorphism and interpretability of behavior planning and execution, and enabling the behavior to effectively adapt to complex and changing external interaction conditions.
[0041] As shown Figure 1 below, the method includes the following steps:
[0042] Step S1, based on the motion capture system, obtain the demonstration data of human close physical interaction behavior.
[0043] Specifically: For the collection of human hugging demonstration data, an accurate and effective motion capture environment is built indoors. The external measurement devices in the environment include optical sensors, including 20 infrared cameras with 60fps evenly arranged on the ceiling of a 6.5m×6.5m×2.5m space, located on the four sides of the square (five cameras on each side). For data accuracy, the room is kept isolated and quiet. In each experiment, two demonstrators wear motion capture suits with 53 markers and 17 inertial units and are required to freely hug each other. The inertial sensing and optical sensing data are combined and calculated to obtain real-time observation data of the angles and speeds of each joint.
[0044] To ensure the quality and comprehensiveness of human demonstration data, the demonstrator group consists of 6 people spanning the young, middle-aged, and elderly age groups. Each participant conducts a hugging demonstration experiment with other participants, forming 15 unique pairs. The experiment includes 4 preset scenarios (social occasions, intimate relationships, emotional expressions, and motor functions) and two participation roles (initiator and recipient), covering as comprehensively as possible all possible behaviors of humans in hugs. In this embodiment, each experimental condition is demonstrated 5 times, and a total of 15×2×4×5 = 600 groups of samples are obtained.
[0045] Step S2, based on the human behavior demonstration data, analyze the human behavior pattern by combining prior knowledge and clustering methods.
[0046] This step specifically includes the following sub-steps:
[0047] Step S201, based on the clustering analysis algorithm, preliminarily analyze the behavior patterns that appear in the human close physical interaction behavior demonstration data.
[0048] Specifically, use the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm to process the human hugging demonstration data, set the scan radius (eps) to 1, and the minimum number of points included (minPts) to 3. Organize the obtained clustering results, observe and analyze the samples within each category, explore the common characteristics between the samples, and focus on the emerging physical properties of the two arms.
[0049] Step S202, combine relevant literature in the fields of robotics, psychology, and behavioral science to refine the emerging key attributes in human close physical interaction behavior and obtain the categories of human behavior patterns. Specifically:
[0050] In this embodiment, as Figure 2 shown, the classification idea of human hugging behavior patterns is as follows: Based on the analysis of human demonstration data and the summary of interdisciplinary literature, three key attributes are refined and further composed of the category definition of human hugging behavior.
[0051] S2021, based on the research literature on human hugging behavior analysis in the fields of hugging robots, hugging psychology, and behavioral science, summarize the prior knowledge of existing research: all considered key attributes and the distinguished pattern categories;
[0052] S2022, combine the prior knowledge with the analysis results of the clustering result samples in step S201, and summarize and propose the emerging key attributes of human hugging behavior, including tightness, hugging, and the cooperation style of the two arms;
[0053] S2023, according to the above three emerging key attributes, design corresponding classification rules in sequence to guide and define the categories of human behavior patterns. Specifically: (1) Whether the chest is in contact; (2) The extension direction of the upper arm; (3) The force direction of the two arms. According to this rule, hugging is defined as 16 categories.
[0054] Step S3, based on the human behavior demonstration data and the pattern analysis results, obtain motion primitives and design a robot behavior planner based on a hierarchical directed graph.
[0055] Specifically, this step includes the following sub-steps:
[0056] Step S301, according to the human behavior pattern categories, calibrate and segment the human behavior demonstration data to obtain the corresponding motion primitives.
[0057] Specifically: For 16 human hugging pattern categories, set corresponding 16 motion primitives. According to the definitions of 16 human hugging pattern categories, segment and calibrate 600 groups of human demonstration samples to obtain 600 groups of motion primitive sequences with human behavior pattern category labels, and obtain a certain number of human teaching samples for each motion primitive.
[0058] Step S302, based on the calibrated demonstration data, use the hierarchical directed graph method to learn the motion primitive sequence planning.
[0059] Specifically, in this embodiment, the motion primitive sequence is processed into a hierarchical directed graph (Temporal And-Or Graph, T-AOG) data structure, and then according to the motion primitive sequence with action labels after calibration, learn the probability between the nodes of the directed graph and parse the behavior planning grammar. That is, the grammar can be described by the following sequence:
[0060] G=(S,V,R,P,σ) (1)
[0061] Where S represents a specific target action, i.e., hugging, V represents an AND-node or an OR-node in the directed graph, R represents a top-down production rule from the parent node α to its child node β, and P represents the probability associated with each production rule (i.e., the probability from the upper node to the lower node). The learning process is transformed into a process of learning the probability P between each node. σ is the behavior planning sequence defined by the grammar, i.e., the set of all valid sentences that the grammar can generate.
[0062] As Figure 3 shown, the outermost ends (terminal nodes) of the directed graph represent various motion primitives, indicating the process of changing the body from one state S t , through the transformation function (joint velocities) F, to another state S t+1 . Using the joint space of the two arms as the state space, the Q-learning rule in the way of time difference is used to learn the association between the motion primitive space and the state space, and the grammar structure is restored according to the posterior probability through Automatic Distillation of Structure (ADIOS). As shown in the following formula (2).
[0063] Q(S t , a i ) = (1 - α)·Q(S t , a i ) + α·[r(S t , a i ) + γ·max a′ Q(S t , a′)] (2)
[0064] Where a i represents the current action, a′ represents the possible next action, α represents the learning rate, r represents the reward function of the action, and γ represents the current value of the future reward (discount factor).
[0065] In summary, the behavior planner in the form of a directed graph is obtained through step S302.
[0066] Step S4, combine the behavior planner with the control method based on motion primitives to form a full-process framework for the planning and execution of the robot's humanoid behavior.
[0067] Specifically, this step includes the following sub-steps:
[0068] Step S401, for the motion primitive sequence obtained in step S301, learn a dynamically consistent mapping model between the target trajectory and the action space. Specifically, in this embodiment, for 16 hugging behavior pattern categories, learn a dynamically consistent mapping model between the joint angle trajectory and the control instruction.
[0069] Step S402: Combine the mapping model of each motion primitive with the motion primitive planner to implement the planning and execution of the robot's humanoid behavior. Specifically, this step includes the following sub-steps:
[0070] S4021: For different joint states, use the behavior planner to generate the next motion primitive required to complete the target hugging action;
[0071] S4022: Execute the mapping model of the motion primitive generated in S4021;
[0072] S4023: Repeat S4021 and S4022 until the target task is achieved, forming the entire process framework for the planning and execution of the robot's humanoid behavior.
[0073] Figure 4 This is the entire process framework for the planning and execution of the robot's humanoid behavior for close physical interaction provided in this embodiment. Combining with human demonstrations under the motion capture system, it can achieve the planning and execution of various tasks in the field of close physical interaction, with good task performance and broad application prospects.
[0074] Embodiment 2
[0075] This embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the method for planning and controlling the robot's humanoid behavior for close physical interaction as described in Embodiment 1.
[0076] Specifically, this electronic device may include a memory, a processor, and a program stored in the memory. When the processor executes the program, it implements the aforementioned method. The device processor includes a central processing unit (CPU), which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) or the computer program instructions loaded from the storage unit into the random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored.
[0077] The CPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus. Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc.
[0078] The communication unit allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks. The processing unit executes the various methods and processes described above, such as the foregoing steps. For example, in some embodiments, the foregoing steps may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more of the steps described above may be performed.
[0079] Alternatively, in other embodiments, the CPU may be configured to execute the foregoing method by any other suitable means (e.g., by means of firmware). The functions described above may be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and the like.
[0080] Embodiment 3
[0081] This embodiment provides a computer-readable storage medium including one or more programs for execution by one or more processors of an electronic device, the one or more programs including instructions for performing the robot anthropomorphic behavior planning and control method for close physical interaction as described in Embodiment 1.
[0082] Specifically, for the storage medium provided in this embodiment, a program is stored thereon, and when the program is executed, the foregoing method is implemented. The program code for implementing the method of the present invention may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes may be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0083] In the context of this embodiment, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing
[0084] As described above, only the specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims
Claims
1. A robot humanoid behavior planning and control method for close physical interaction, characterized in that: The steps include: Step S1, obtaining demonstration data of close interaction between the robot and humans through motion capture; Step S2, based on the demonstration data, using prior knowledge and cluster analysis, obtaining multiple human behavior pattern categories; Step S3, based on the human behavior pattern category, by segmenting and calibrating the demonstration data, multiple groups of motion primitive sequences including class behavior pattern labels are obtained and a hierarchical directed graph is constructed, and a robot behavior planner is obtained through training; Step S4, constructing a dynamically consistent mapping model between the target trajectory and the action space, and realizing the humanoid behavior planning and control of the robot based on the mapping model and the planner, In step S2, the process of obtaining multiple human behavior pattern categories by using prior knowledge and cluster analysis includes the following steps: Step S201, obtaining clustering results through density-based clustering analysis based on the demonstration data; Step S202, based on the clustering results and the prior knowledge obtained in advance, key attributes of human close physical interaction behaviors are obtained, and human behavior pattern categories are constructed. The step S202 includes: Step S2021, obtaining prior knowledge based on pre-acquired interdisciplinary literature; Step S2022, based on the prior knowledge and the clustering result, obtaining key attributes in human close physical interaction behavior, wherein the key attributes include tightness, hugging, and two-arm collaboration style; Step S2023, based on the key attributes, the hugging action is divided into multiple human behavior pattern categories according to whether the chest is in contact, the direction of the upper arm extension, and the direction of the force on the arms. In step S3, the hierarchical directed graph is constructed as follows: , in, S Indicates the target hug action, V is an AND node or an OR node in a directed graph, R Indicates that from the parent node α To its child nodes β The top-down production rules of P represents the probability associated with each production rule, σ is a sequence of action plans defined by the grammar, In step S3, the process of obtaining the robot behavior planner through training includes: Use the joint space of the two arms as the state space and the probability between each node P As the learning object, the Q learning rule in the form of time difference is used to learn the association between the motion primitive space and the state space, and the grammatical structure is restored through automatic structural distillation according to the posterior probability to obtain the robot behavior planner. The demonstration data includes motion capture data of hugging behaviors of multiple roles of participants of multiple age groups in multiple preset scenarios, wherein the preset scenarios include social occasions, intimate relationships, emotional expressions and motor functions, and the roles include initiators and recipients. Inertial sensing and optical sensing data are combined and calculated to obtain real-time observation data of the angle and speed of each joint.
2. A robot humanoid behavior planning and control method for close physical interaction according to claim 1, characterized in that: In step S3, the process of segmenting and calibrating the demonstration data based on the human behavior pattern category includes: A corresponding motion primitive is constructed for each human behavior pattern category, and the demonstration data is segmented and calibrated to obtain multiple groups of motion primitive sequences including class behavior pattern labels.
3. The method for humanoid robot behavior planning and control for close physical interaction according to claim 1, characterized in that: In the step S4, based on the mapping model and the planner, implementing the humanoid behavior planning and control of the robot includes: According to different joint states, the robot behavior planner is used to generate the next motion primitive required to complete the target hugging action, and the corresponding mapping model is used to execute the target hugging action.
4. An electronic device, characterized in that: include: One or more processors and a memory, wherein the memory stores one or more programs, wherein the one or more programs include instructions for executing the robot humanoid behavior planning and control method for close physical interaction as described in any one of claims 1-3.
5. A computer-readable storage medium, characterized in that: It includes one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the robot humanoid behavior planning and control method for close physical interaction as described in any one of claims 1-3.
Citation Information
Patent Citations
Quick imitation learning method, system and equipment for robot skill learning
CN113408621A
Unmanned system demonstration data set construction method based on prior knowledge and fuzzy reasoning
CN118295256A