An AI-based humanoid robot teaching method and system
By acquiring teaching images through cameras and updating the robot's joint motion state using motion recognition models, the problem of overly mechanical humanoid robot movements has been solved, resulting in more natural and fluid humanoid robot movements and improving the degree of anthropomorphism and adaptability.
Patent Information
- Application Number
- CN202511405066.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-29
AI Technical Summary
Due to the difference between simulation and the real world, humanoid robots in the current technology perform actions that are too mechanical and fail to achieve a good human-like effect.
By acquiring teaching images through a camera, using a motion recognition model to identify key points and joint movement states of the human body, generating simulation control commands, and updating the robot's joint movement states through intent tags, the naturalness and smoothness of the movements are improved by combining irrational features, velocity curves, and jitter simulation.
It improves the anthropomorphism of robot movements, enhances interactivity and collaboration with humans, adapts to the movement styles of different instructors, and makes the movement simulation more in line with actual needs.
Smart Images

Figure CN120862641B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robots, in particular to a humanoid robot teaching method and system based on AI. BACKGROUND
[0002] At present, due to the difference between simulation and the real world, the humanoid robot cannot actually be perfectly replicated, which leads to the situation that the robot is over-mechanized in the action execution process, so that the good humanization effect cannot be achieved. SUMMARY
[0003] In view of the defects, the embodiment of the present application discloses a humanoid robot teaching method based on AI, which can improve the humanization degree of the humanoid robot.
[0004] The first aspect of the embodiment of the present application discloses a humanoid robot teaching method based on AI, comprising:
[0005] obtaining the human teaching image of the teaching personnel in the current scene through the camera assembly;
[0006] inputting the continuous frame human teaching image into the pre-constructed action recognition model for recognition to determine the teaching action information of each movement stage and the human key point coordinates of each movement stage, determining the robot joints corresponding to each human key point, and calculating the joint motion state of each control joint of the corresponding robot;
[0007] determining the corresponding intention label according to the teaching action information of each movement stage, and updating the joint motion state of each control joint according to the intention label;
[0008] generating the simulation control instruction of each joint of the humanoid robot according to the joint motion state of each control joint after the action intention correction, and sending the simulation control instruction to the corresponding humanoid robot for action simulation.
[0009] As an optional implementation manner, in the first aspect of the embodiment of the present application, the determination of the corresponding intention label according to the teaching action information of each movement stage, and the updating of the joint motion state of each control joint according to the intention label, comprises:
[0010] determining the corresponding non-rational feature according to the teaching action information of each movement stage, wherein the non-rational feature is a feature model constructed based on a specific human action;
[0011] generating the corresponding disturbance residual according to the non-rational feature, and superimposing the generated disturbance residual to the joint motion state of each control joint for state updating. Through the above-mentioned manner, the naturalness of the robot action simulation is improved.
[0012] As an optional implementation, in the first aspect of the embodiment of the present application, the joint motion state includes speed curve information.
[0013] The intention label corresponding to each movement stage is determined according to the teaching action information of the movement stage, and the joint motion state of each control joint is updated according to the intention label.
[0014] The intention label corresponding to each movement stage is determined according to the teaching action information of the movement stage, and the speed curve template corresponding to the teaching action information is determined according to the intention label.
[0015] The speed curve information of each control joint is updated according to the speed curve template. In this way, the fluency and accuracy of the robot action simulation are improved.
[0016] As an optional implementation, in the first aspect of the embodiment of the present application, the teaching action information includes teaching trajectory information and acceleration curve information.
[0017] The intention label corresponding to each movement stage is determined according to the teaching action information of the movement stage, and the joint motion state of each control joint is updated according to the intention label.
[0018] The intention label corresponding to each movement stage is determined according to the teaching action information of the movement stage, and when the length of the teaching trajectory information of the corresponding movement stage exceeds a set value, a jitter superposition waveform is determined according to the intention label, the jitter superposition waveform is fused with the original teaching trajectory information, and the joint motion state of each control joint is updated according to the fused trajectory data.
[0019] The acceleration peak value of the corresponding action is determined according to the teaching action information of each movement stage, and the control point corresponding to the acceleration peak value is taken as a turning point. A pause with a set time interval is inserted at the teaching trajectory information at the turning point with a set probability. The joint motion state of each control joint is updated according to the updated data. Through the above-mentioned design of micro-jitter fusion and pause rhythm, the overall personification of the robot is improved, and the mechanical feeling is reduced.
[0020] As an optional implementation, in the first aspect of the embodiment of the present application, the human body key point coordinates of each movement stage are determined by inputting the continuous frame human body teaching image into a pre-constructed action recognition model for recognition, and the robot joints corresponding to each human body key point are determined.
[0021] The continuous frame human body teaching image is preprocessed to obtain a preprocessed image, and the preprocessing includes image size adjustment, normalization processing, and color space conversion.
[0022] The pre-processed image is analyzed by the human posture estimation model to obtain human key point coordinates and corresponding confidence levels of each movement stage, and low-confidence key point coordinates are filtered out;
[0023] According to the human posture estimation model, human key points are determined according to a set key point mapping rule to determine corresponding human joint information, and each human key point is determined to correspond to a robot joint. By accurately identifying human key point coordinates and reasonably mapping to robot joints, the robot can imitate more natural and smooth human movements.
[0024] As an optional implementation, in the first aspect of the embodiment of the present application, the input of the continuous frame human demonstration image into the pre-constructed action recognition model for recognition to determine the demonstration action information of each movement stage comprises:
[0025] The continuous frame human demonstration image is input into the pre-constructed action recognition model to determine the demonstration action information of each movement stage, and the action recognition model further comprises an attention module, which captures the correlation between joints to determine the corresponding power chain change information of the demonstration action information;
[0026] The corresponding intention label is determined according to the demonstration action information of each movement stage, and the joint motion state of each control joint is updated according to the intention label, comprising:
[0027] The corresponding intention label is determined according to the demonstration action information of each movement stage, and the standard power chain corresponding to the demonstration action information is determined according to the intention label;
[0028] The power chain change information is compared with the standard power chain, if the comparison is consistent, the joint motion state of each control joint is not updated, if the comparison is inconsistent, the joint motion state of each control joint is updated according to the comparison result. The action recognition model can identify the change process of the power chain, so as to more accurately determine the demonstration action information.
[0029] As an optional implementation, in the first aspect of the embodiment of the present application, after the simulation control instruction is sent to the corresponding humanoid robot for action simulation, it further comprises:
[0030] When the action simulation result meets the requirements, the simulation control instruction is saved. By saving the simulation control instruction meeting various individualized requirements, the robot can quickly adjust the action mode according to the specific requirements of the user.
[0031] The second aspect of the embodiment of the present application discloses a humanoid robot demonstration system based on AI, comprising:
[0032] An image acquisition module is configured to acquire the human demonstration image of the demonstrator in the current scene through the camera assembly;
[0033] An identification module is configured to input the human demonstration image of the continuous frames into a pre-constructed action recognition model for identification to determine the demonstration action information of each movement stage and the human key point coordinates of each movement stage, determine the robot joints corresponding to each human key point, and calculate the joint motion state of each control joint of the corresponding robot;
[0034] A state updating module is configured to determine the corresponding intention label according to the demonstration action information of each movement stage, and update the joint motion state of each control joint according to the intention label;
[0035] An instruction generation module is configured to generate the simulation control instruction of each joint of the humanoid robot according to the joint motion state of each control joint after the action preference correction, and send the simulation control instruction to the corresponding humanoid robot for action simulation.
[0036] The third aspect of the embodiment of the present application discloses an electronic device, comprising: a memory storing executable program code; a processor coupled with the memory; the processor invokes the executable program code stored in the memory, and is used for executing the AI-based humanoid robot demonstration method disclosed in the first aspect of the embodiment of the present application.
[0037] The fourth aspect of the embodiment of the present application discloses a computer readable storage medium storing a computer program, wherein the computer program causes a computer to execute the AI-based humanoid robot demonstration method disclosed in the first aspect of the embodiment of the present application.
[0038] Compared with the prior art, the embodiment of the present application has the following beneficial effects:
[0039] The AI-based humanoid robot demonstration method in the embodiment of the present application determines the corresponding intention label according to the demonstration action information of each movement stage, and updates the joint motion state of each control joint according to the intention label. This process integrates the understanding of the intention of the demonstration action, so that the robot can not only simply imitate the action, but also understand the purpose behind the action, thereby generating a joint motion state that is more in line with the actual needs, and improving the action rationality and adaptability of the robot. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described in the following embodiments are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0041] Figure 1 is a flowchart of the AI-based humanoid robot teaching method disclosed by the embodiments of the present application;
[0042] Figure 2 is a flowchart of the joint state updating disclosed by the embodiments of the present application;
[0043] Figure 3 is a flowchart of the human key point recognition disclosed by the embodiments of the present application;
[0044] Figure 4 is a structural diagram of an AI-based humanoid robot teaching system provided by the embodiments of the present application;
[0045] Figure 5 is a structural diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0047] It should be noted that the terms first, second, third, fourth, etc. in the specification and claims of the present application are used to distinguish different objects, not to describe a specific order. The terms of the embodiments of the present application include and have as well as any transformation thereof, which is intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices. EMBODIMENTS
[0048] Please refer to Figure 1 , Figure 1is a flowchart of the AI-based humanoid robot teaching method disclosed in the embodiments of the present application. Among them, the execution subject of the method described in the embodiments of the present application is an execution subject composed of software or / and hardware, which can receive relevant information through wired or / and wireless means and can send certain instructions. Of course, it can also have certain processing and storage functions. The execution subject can control multiple devices, such as remote physical servers or cloud servers and related software, or it can be a local host or server and related software that performs related operations on a device placed somewhere. In some scenarios, it can also control multiple storage devices, which can be placed in the same place or different places. As shown in Figure 1 The AI-based humanoid robot teaching method includes the following steps:
[0049] S101: Obtain the human teaching image of the teaching personnel in the current scene through the camera assembly;
[0050] S102: Input the continuous frame human teaching image into the pre-constructed action recognition model for recognition to determine the teaching action information of each movement stage and the human body key point coordinates of each movement stage, determine the robot joints corresponding to each human body key point, and calculate the joint motion state of each control joint of the corresponding robot;
[0051] S103: Determine the corresponding intention label according to the teaching action information of each movement stage, and update the joint motion state of each control joint according to the intention label;
[0052] S104: Generate simulation control instructions for each joint of the humanoid robot according to the joint motion state of each control joint after action intention correction, and send the simulation control instructions to the corresponding humanoid robot for action simulation.
[0053] In the embodiments of the present application, the human teaching image of the teaching personnel in the current scene is obtained through the camera assembly, which can directly capture the action details of the teaching personnel in the real environment, providing rich raw data for subsequent accurate recognition of teaching actions, and ensuring the integrity and authenticity of the teaching information.
[0054] Specifically, inputting the continuous frame human teaching image into the pre-constructed action recognition model can quickly and accurately identify the teaching action information of each movement stage and the human body key point coordinates. Furthermore, it can further determine the robot joints corresponding to each human body key point and calculate the joint motion state of each control joint of the corresponding robot, realizing accurate conversion from human body action to robot joint motion and improving the efficiency and accuracy of the teaching process.
[0055] According to the joint movement state of each control joint after the action intention is corrected, simulation control instructions of each joint of the humanoid robot are generated, and the instructions are sent to the corresponding humanoid robot for action simulation. In this way, the humanoid robot can imitate the action of the demonstrator in a more natural and realistic manner, enhancing the interactivity and collaboration between the robot and the human, and providing the possibility for the robot to be applied in more complex scenarios. Moreover, the degree of humanization of the robot action is higher.
[0056] For example, when the above-mentioned intention label is waving, if the conventional control mode is adopted, the hand will move back and forth between two specific points at a fixed frequency and speed, but in the actual waving process, the speed will change and the point will deviate. When the waving intention is recognized, the speed can be changed and the left and right end points can be randomly changed within a certain range to increase the randomness of the action. However, if it is a scene that requires fine movement, the precise point needs to be determined according to the mapping.
[0057] More preferably, as shown in Figure 2 According to the intention label, the joint movement state of each control joint is updated, including:
[0058] S1031: determining a corresponding non-rational feature according to the demonstration action information of each movement stage, the non-rational feature being a feature model constructed based on a specific human action;
[0059] S1032: generating a corresponding disturbance residual according to the non-rational feature, and superimposing the generated disturbance residual to the joint movement state of each control joint for state updating.
[0060] In the embodiment of the present application, by determining the non-rational features corresponding to the demonstration action information of each movement stage, these non-rational features are feature models constructed based on specific human actions, which can accurately depict the details of non-precise, personal habits or slight fluctuations in human actions. For example, there may be slight tremors or slight deviations in the angle when a human raises his hand. This non-rational feature model can accurately capture it, making the robot's understanding of human actions more in-depth, avoiding only recognizing idealized standard actions, and thus improving the accuracy of action simulation.
[0061] The non-rational characteristics are used to generate disturbance residuals, which are superimposed on the joint motion state of each control joint. The introduction of disturbance residuals is not arbitrary, but is based on the analysis of the non-rational characteristics of human motion, which simulates the natural fluctuations and uncertainties in human motion. This reasonable disturbance makes the update of the joint motion state of the robot closer to the real human motion, reduces the differences with human motion caused by too precise and mechanical motion, and improves the accuracy of motion simulation. In the specific implementation, the non-rational characteristics corresponding to different motions are different, so different non-rational characteristic disturbances are constructed according to different motions.
[0062] Human motion often has certain randomness and non-complete repeatability, that is, even if the same motion is performed, there will be slight differences each time. The introduction of non-rational characteristics and disturbance residuals makes the motion of the robot no longer a precise copy, but has certain randomness and changes, just like natural human motion. For example, when the robot imitates human walking, this way can make the size and frequency of the steps fluctuate naturally, making the motion look more natural and smooth, rather than mechanically repeating the same motion pattern as in traditional methods.
[0063] In addition to the above common non-rational characteristics, different characteristics can be constructed for different people in the specific implementation, because different people have their own habits when performing the same motion, and these differences are reflected in the non-rational characteristics. By recognizing and utilizing these non-rational characteristics to generate disturbance residuals to update the joint motion state, the robot can adapt to the motion styles of different demonstrators. The motion style of the robot is matched with the demonstrators, improving the adaptability of the robot when facing different demonstrators.
[0064] Different people may have similar joint angle change trends when performing the same motion (such as walking or waving), but the detailed characteristics (such as pause timing and acceleration peak) differ significantly, which is the core of style. It is difficult to capture this difference using only joint angle calculations, resulting in robots performing the same motion and lacking diversity in personification. Therefore, learning non-rational characteristics can improve the richness of personification.
[0065] In the specific implementation, motion scenarios and individual characteristics can also be used as conditional variables to input the probability model; for example: increasing the probability and amplitude of balance adjustment in a fatigue state; optimizing the conditional parameters through reinforcement learning to make the generated non-rational characteristics more consistent with the scenario logic.
[0066] More preferably, the joint motion state includes velocity curve information.
[0067] The intention label corresponding to each movement stage is determined according to the demonstration action information of the movement stage, and the joint motion state of each control joint is updated according to the intention label.
[0068] The intention label corresponding to each movement stage is determined according to the demonstration action information of the movement stage, and the joint motion state of each control joint is updated according to the intention label.
[0069] The speed curve information of each control joint is updated according to the speed curve template.
[0070] In real life, the speed of human action is not constant, but changes naturally according to the action intention and scene requirements. By determining the corresponding speed curve template according to the intention label and updating the speed curve information of the control joint accordingly, the robot can imitate the natural rhythm change in human action. For example, when imitating the action of reaching out to pick up an object, the robot can first accelerate slowly to reach out, and then appropriately decelerate when approaching the object, making the action more smooth and natural, avoiding the stiffness caused by uniform motion.
[0071] Smooth transition is required between demonstration actions of different movement stages, and reasonable adjustment of the speed curve is the key to achieving this goal. Updating the joint speed curve information with the speed curve template can ensure that the robot can naturally connect the speed when switching from one action to another, without abrupt pauses or speed mutations. For example, in the transition from standing to walking, the robot can gradually adjust the speed of the leg joints according to the speed curve template to achieve a smooth start, enhancing the overall coherence and smoothness of the action.
[0072] In practical applications, the corresponding relationship library of intention label and speed curve template can be enriched by continuously collecting and analyzing a large amount of demonstration action data. With the accumulation of data, the robot can learn more suitable speed curve patterns under different intentions and automatically adjust and optimize the speed curve information according to new demonstration actions. This intelligent learning ability enables the robot to adapt to more diverse action demonstrations and continuously improve the quality and effect of action simulation.
[0073] In addition to the above general simulation mode, in the specific implementation, it can be optimized and adjusted for individualized style, which can adapt to the style of different demonstrators. Different people may have different speeds and rhythms when performing the same action, which reflects the differences in personal style and habits. By determining the speed curve template according to the intention label and updating the joint speed curve information, the robot can adapt to the action style of different demonstrators. For example, for demonstrators with rapid actions and demonstrators with slow actions, the robot can select appropriate speed curve templates according to the differences in intention labels in their actions, so that the action style of the robot matches the demonstrators, and the adaptability of the robot to different demonstrators is improved.
[0074] More preferably, the demonstration action information includes demonstration trajectory information and acceleration curve information.
[0075] The intention label corresponding to each movement stage is determined according to the demonstration action information of each movement stage, and the joint motion state of each control joint is updated according to the intention label.
[0076] The intention label corresponding to each movement stage is determined according to the demonstration action information of each movement stage. When the length of the demonstration trajectory information of the corresponding movement stage exceeds a set value, the corresponding jitter superposition waveform is determined according to the intention label, the jitter superposition waveform is fused with the original demonstration trajectory information, and the joint motion state of each control joint is updated according to the fused trajectory data.
[0077] The acceleration peak value of the corresponding action is determined according to the demonstration action information of each movement stage, and the control point corresponding to the acceleration peak value is taken as a turning point. A pause with a set time interval is inserted at the turning point in the demonstration trajectory information with a set probability. The joint motion state of each control joint is updated according to the updated data.
[0078] In actual life, the actual action of human beings is not absolutely smooth, and there may be some small and unconscious jitter. When the length of the demonstration trajectory information exceeds a set value, the corresponding jitter superposition waveform is determined according to the intention label and fused with the original demonstration trajectory information, so that the robot can imitate the natural jitter in human action. For example, when imitating human operation of holding a tool, adding appropriate jitter superposition waveform can make the action of the robot look more like human real operation, avoid the stiffness and rigidity of mechanical action, and enhance the naturalness and reality of the action.
[0079] In addition to the slight jitter, humans often pause at certain key positions during the execution of an action according to task requirements or habits. By taking the control points corresponding to acceleration peaks as turning points and inserting pauses of a set time interval at the teaching trajectory information with a set probability, the robot can simulate this pause rhythm in human actions. For example, when imitating human writing, pauses are inserted at the turning points of the strokes, making the robot's writing action more in line with human writing habits and making the action more natural and smooth. Specifically, the time of change of motion direction is identified by calculating the rate of change of speed or acceleration peaks; a pause of 0.1-0.3 seconds is inserted at the turning point with a set probability (5%); and when implementing the specific embodiment, the probability can also be dynamically adjusted according to the type of action, such as increasing the pause probability when performing fine operations.
[0080] The teaching trajectory information and the acceleration curve information contain rich action details. Determining the intention label according to these information and further generating jitter superimposed waveforms and inserting pauses can more accurately capture and restore the subtle features of the teaching action. Different actions have different dynamic characteristics, including trajectory length, acceleration change, etc. The method can dynamically adjust the action update strategy according to the length of the teaching trajectory information and the acceleration peak and other parameters. When the trajectory is long, jitter is considered to be added; the pause position is determined according to the acceleration peak, so that the robot can adapt to various types of action teaching, whether the action is fast, slow, complex or simple, precise action simulation can be achieved.
[0081] The scheme of the embodiment of the application determines the jitter superimposed waveform and the pause strategy according to the intention label, which reflects the understanding of the robot for the intention of the teaching action. Different intention labels correspond to different action purposes, for example, the jitter and pause characteristics of the action under the grasping intention and the placing intention may be different. The robot can intelligently select the appropriate action update mode according to the intention label, so that the action is more in line with the actual requirements and exhibits a higher level of intelligence.
[0082] In actual application, a large amount of teaching action data can be analyzed by a machine learning algorithm to continuously optimize the parameters of the jitter superimposed waveform, the probability and time interval of the pause, etc. With the accumulation of data and the optimization of the algorithm, the robot can adaptively adjust the action parameters, improve the quality and effect of the action simulation, and realize intelligent adaptation to different scenes and tasks.
[0083] More preferably, as shown in Figure 3 The method comprises the steps of:
[0084] S1021: Preprocess the human demonstration image of continuous frames to obtain a preprocessed image, the preprocessing including image size adjustment, normalization processing and color space conversion;
[0085] S1022: Analyze the preprocessed image by a human pose estimation model to obtain human key point coordinates and corresponding confidence of each movement stage, and filter out low confidence key point coordinates;
[0086] S1023: Determine the human key point according to the human pose estimation model, determine the corresponding human joint information according to the set key point mapping rule, and determine the corresponding robot joint of each human key point.
[0087] The human pose estimation model in the embodiment of the application adopts MediaPipePose or OpenPose model, which detects 33 key skeleton points (such as head, shoulder, elbow, wrist, hip, knee, ankle, etc.) of a human body in real time, and outputs 2D pixel coordinates (x, y) and confidence (0-1) of each point. The advantages of MediaPipe are: lightweight design, supporting real-time inference on the edge side, and being suitable for dynamic motion capture.
[0088] The specific inference process is as follows: input the preprocessed image, extract the feature map through CNN; the graph convolution network (GCN) predicts the key point position, and optimizes the coordinates in combination with the human body topology (such as the connection relationship of shoulder-elbow-wrist); filter out low-precision points with confidence <0.6, and output high-reliability key point sequence.
[0089] In the embodiment of the application, the human demonstration image of continuous frames is preprocessed, including image size adjustment, normalization processing and color space conversion. Image size adjustment can make the image uniform to the specification suitable for the action recognition model processing, avoid the model recognition deviation caused by the image size difference; the normalization processing can eliminate the dimension influence in the image data, make the features of different images comparable, and help the model to extract the features more accurately; the color space conversion can convert the image to a color space more suitable for action recognition, highlight the color information related to human action, and reduce irrelevant color interference. Through these preprocessing operations, high-quality input images are provided for the subsequent human pose estimation model, thereby improving the accuracy of human key point coordinate recognition.
[0090] The human pose estimation model analyzes the preprocessed image to obtain the human key point coordinates and the corresponding confidence levels of each movement stage. The low-confidence key point coordinates are often caused by image blur, occlusion, or model recognition uncertainty, and have low accuracy. Filtering out these low-quality key point coordinates can avoid their adverse effects on subsequent robot joint mapping and motion imitation, ensuring that the human key point coordinates used have high reliability, thereby improving the overall motion recognition accuracy.
[0091] After determining the human key points according to the human pose estimation model, the corresponding human joint information is determined according to the set key point mapping rule, and the robot joints corresponding to each human key point are further determined. This mapping relationship based on human joint information fully considers the similarities and differences between human and robot body structures, and can more reasonably establish the correspondence between human motion and robot joint movement. For example, the arm joints of the human body and the arm joints of the robot have certain similarities in movement function and structure, and through this mapping rule, the motion of the human arm can be accurately mapped to the robot arm joints, making the robot motion imitation more in line with human movement habits and logic. A reasonable mapping relationship between human key points and robot joints is the key to achieving precise motion imitation. The mapping relationship determined through the above steps can ensure that the movement of each joint of the robot is highly consistent with the movement of the corresponding part of the human body when the robot imitates human motion. For example, when imitating the human action of bending over, the robot can accurately control the movement of the joints such as the waist and legs, achieving a natural and smooth bending action, and improving the precision and realism of motion imitation.
[0092] More preferably, the input of the continuous frame human demonstration image into the pre-constructed motion recognition model for recognition to determine the demonstration motion information of each movement stage comprises:
[0093] The continuous frame human demonstration image is input into the pre-constructed motion recognition model to determine the demonstration motion information of each movement stage. The motion recognition model further includes an attention module, which captures the correlation between joints to determine the corresponding power chain change information of the demonstration motion information;
[0094] The determination of the corresponding intention label according to the demonstration motion information of each movement stage, and the updating of the joint movement state of each control joint according to the intention label, comprises:
[0095] The determination of the corresponding intention label according to the demonstration motion information of each movement stage, and the determination of the standard power chain corresponding to the demonstration motion information according to the intention label;
[0096] The power chain change information is compared with a standard power chain, if the comparison is consistent, the joint movement state of each control joint is not updated, if the comparison is inconsistent, the joint movement state of each control joint is updated according to the comparison result.
[0097] The action recognition model can be constructed by combining a Transformer module and an LSTM module; the overall correlation between joints (such as the cooperative relationship of the arm, shoulder, and torso when waving) is captured through a self-attention mechanism; in addition to the above model, a bidirectional recurrent neural network can also be used to construct the action recognition model.
[0098] Specifically, human actions are not the superposition of isolated joint angles, but the result of the synergistic action of muscles, bones, and nerves, and there is a coupling relationship between joints, such as the natural fine-tuning of the torso to maintain balance when the arm swings. If only the single joint angle is mapped or calculated, it may result in stiff and disjointed actions; for example, when the robot arm is lifted, the torso is not synchronized and fine-tuned, which appears mechanical.
[0099] In the embodiment of the application, the attention module in the action recognition model can capture the correlation between joints. In human actions, each joint does not move independently, but cooperates to form an organic whole. For example, in a throwing action, the joints of the arm, shoulder, waist, and leg will move cooperatively in a certain order and intensity. The attention module can focus on the dynamic relationship between these joints, thereby more accurately identifying the demonstration action information of each movement phase. Compared with the traditional method of only focusing on the position of a single joint, this method can more comprehensively understand the nature of human actions, greatly improving the accuracy of action recognition.
[0100] By capturing the correlation between joints, the attention module can also determine the corresponding power chain change information of the corresponding demonstration action information. The power chain refers to the chain through which force is transmitted from one part of the body to another during movement. Understanding the changes in the power chain is crucial for accurately identifying the type, intensity, and intent of the action. For example, in a running action, the power chain is transmitted from the leg muscles through the hip joint, knee joint, and ankle joint to the ground, pushing the body forward. The action recognition model can identify the change process of this power chain, thereby more accurately determining the demonstration action information.
[0101] After determining the corresponding intention label according to the teaching action information of each movement stage, the standard power chain corresponding to the teaching action information can be further determined. The standard power chain is obtained based on a large amount of human motion data and motion physiology, and represents the power transmission mode of a certain action in an ideal state. For example, for a weight lifting action, the standard power chain specifies the correct power transmission path from the leg, through the waist and back to the arm, and finally lifting the weight. With the standard power chain as a reference, the action imitated by the robot can be more in line with the natural laws of human motion, and the rationality of action imitation can be enhanced.
[0102] The power chain change information is compared with the standard power chain, and the joint motion state of each control joint is updated according to the comparison result. If the comparison is consistent, it means that the current action conforms to the standard power chain, and there is no need to adjust the joint motion state; if the comparison is inconsistent, it means that the action has deviation, and the motion angle, speed and intensity of the joint need to be adjusted according to the comparison result, so that the action of the robot is closer to the standard action. This adjustment method based on power chain comparison can ensure that the action imitated by the robot is more accurate and reasonable, and avoid unnatural actions, so that the smoothness of the overall action is higher.
[0103] By comparing the power chain change information and the standard power chain, the joint that needs to be adjusted can be accurately located. For example, if it is found that the power chain is broken at the waist joint or the power transmission is not smooth, it can be determined that the motion state of the waist joint needs to be adjusted. This accurate positioning can avoid blind adjustment of all joints, and improve the efficiency and accuracy of joint control.
[0104] More preferably, after the simulation control instruction is sent to the corresponding humanoid robot for action simulation, it further comprises:
[0105] When the action simulation result meets the requirements, the simulation control instruction is saved.
[0106] Saving the simulation control instruction when the action simulation result meets the expected standard means that if the humanoid robot needs to perform the same or similar action again in the future, there is no need to generate the control instruction from scratch. For example, in an industrial production scene, the robot needs to repeatedly complete a series of standardized assembly actions. After saving the simulation control instruction that meets the requirements, the instruction can be directly called to make the robot quickly reproduce the successful action, greatly saving time and effort.
[0107] In addition to the above-mentioned manners, a personification generator can be further provided, which is specially designed for adding humanized features to perfect mechanical actions: personalized tremor model: learning the hand tremor characteristics of different individuals, which can be the small tremor of young and stable people, or the large tremor of old people. Contextual hesitation model: intelligently adding pauses according to the task scene, such as increasing small adjustment pauses during precise assembly, and maintaining smoothness in daily interaction.
[0108] In the machine simulation process, the general machine is to find the optimal route to simulate the action, but in the process of human movement, the path is often not optimal, so in the embodiment of the application, through the simulation of the humanoid path, the degree of humanization of the overall robot action is further improved.
[0109] The AI-based humanoid robot teaching method in the embodiment of the application determines the corresponding intention label according to the teaching action information of each movement stage, and updates the joint motion state of each control joint according to the intention label. This process integrates the understanding of the intention of the teaching action, so that the robot can not only simply imitate the action, but also understand the purpose behind the action, thereby generating a joint motion state that is more in line with the actual demand, and improving the rationality and adaptability of the robot action. Embodiments
[0110] Please refer to Figure 4 , Figure 4 is a structural schematic diagram of the AI-based humanoid robot teaching system disclosed by the embodiment of the application. As Figure 4 indicated, the AI-based humanoid robot teaching system can include:
[0111] The image acquisition module 21 is configured to acquire the human body teaching image of the teaching personnel in the current scene through the camera assembly.
[0112] The recognition module 22 is configured to input the continuous frame human body teaching image into the pre-constructed action recognition model for recognition to determine the teaching action information of each movement stage and the human body key point coordinates of each movement stage, determine the robot joint corresponding to each human body key point, and calculate the joint motion state of each control joint of the corresponding robot.
[0113] The state updating module 23 is configured to determine the corresponding intention label according to the teaching action information of each movement stage, and update the joint motion state of each control joint according to the intention label.
[0114] The instruction generation module 24 is configured to generate the simulation control instruction of each joint of the humanoid robot according to the joint motion state of each control joint after the action preference correction, and send the simulation control instruction to the corresponding humanoid robot for action simulation.
[0115] The AI-based humanoid robot demonstration method in the embodiment of the application determines the corresponding intention label according to the demonstration action information of each movement stage, and updates the joint motion state of each control joint according to the intention label. This process integrates the understanding of the intention of the demonstration action, so that the robot can not only simply imitate the action, but also understand the purpose behind the action, thereby generating a joint motion state that is more in line with the actual needs, and improving the action rationality and adaptability of the robot. Embodiments
[0116] Please refer to Figure 5 , Figure 5 is a structural schematic diagram of an electronic device disclosed by the embodiment of the application. The electronic device can be a computer, a server, etc. Of course, in certain cases, it can also be a smart device such as a mobile phone, a tablet computer, and a monitoring terminal, and an image acquisition device with processing function. As shown in the figure, Figure 5 The electronic device can include:
[0117] a memory 510 storing executable program codes;
[0118] a processor 520 coupled with the memory 510;
[0119] The processor 520 calls the executable program codes stored in the memory 510 to execute part or all of the steps of the AI-based humanoid robot demonstration method in the embodiment one.
[0120] The embodiment of the application discloses a computer readable storage medium storing a computer program, wherein the computer program causes a computer to execute part or all of the steps of the AI-based humanoid robot demonstration method in the embodiment one.
[0121] The embodiment of the application further discloses a computer program product, wherein when the computer program product runs on a computer, it causes the computer to execute part or all of the steps of the AI-based humanoid robot demonstration method in the embodiment one.
[0122] The embodiment of the application further discloses an application publishing platform, wherein the application publishing platform is used to publish a computer program product, wherein when the computer program product runs on a computer, it causes the computer to execute part or all of the steps of the AI-based humanoid robot demonstration method in the embodiment one.
[0123] In various embodiments of the application, it should be understood that the size of the serial number of the processes does not mean the inevitable sequence of execution, and the execution sequence of the processes should be determined according to their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the application.
[0124] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0125] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0126] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-accessible memory. Based on this understanding, the technical scheme of the present application or the part that contributes to the prior art or the whole or part of the technical scheme can be embodied in the form of a software product. The computer software product is stored in a memory, and includes some parts or all steps of the method for enabling a computer device (which can be a personal computer, a server or a network device, and specifically can be a processor in the computer device) to execute the method described in each embodiment of the present application.
[0127] In the embodiments provided by the present application, it should be understood that B corresponding to A means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean that B is determined only according to A, but B can also be determined according to A and / or other information.
[0128] Those skilled in the art can understand that part or all of the steps in the various methods of the embodiments can be completed by instructing the relevant hardware by a program, and the program can be stored in a computer readable storage medium, including a Read-Only Memory (ROM), a Random Access Memory (RAM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), a One-time Programmable Read-Only Memory (OTPROM), an Electrically-Erasable Programmable Read-Only Memory (EEPROM), a Compact Disc Read-Only Memory (CD-ROM) or other optical disk storage, a magnetic disk storage, a magnetic tape storage, or any other medium that can be used to carry or store data in a computer readable manner.
[0129] The AI-based humanoid robot teaching method and system, the electronic device and the storage medium disclosed in the embodiments of the present application are described in detail above, and the principles and implementation manners of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will be changed, and the above description of the present application should not be understood as a limitation.
Claims
1. An AI-based humanoid robot teaching method, characterized by, Comprising: acquiring a human demonstration image of a demonstrator in a current scene through a camera assembly; inputting the human demonstration image of the continuous frames into a pre-constructed action recognition model for recognition to determine the demonstration action information of each movement stage and the human key point coordinates of each movement stage, determine the robot joints corresponding to each human key point, and calculate the joint motion state of each control joint of the corresponding robot; the joint motion state includes velocity curve information; the inputting the human demonstration image of the continuous frames into a pre-constructed action recognition model for recognition to determine the human key point coordinates of each movement stage, determine the robot joints corresponding to each human key point, comprises: preprocessing the human demonstration image of the continuous frames to obtain a preprocessed image, the preprocessing includes image size adjustment, normalization processing and color space conversion; analyzing the preprocessed image through a human pose estimation model to obtain the human key point coordinates of each movement stage and the corresponding confidence, and filtering out the key point coordinates with low confidence; determining the human key point according to the human pose estimation model to determine its corresponding human joint information according to the set key point mapping rule, and determining the robot joints corresponding to each human key point; determining the corresponding intention label according to the demonstration action information of each movement stage, and updating the joint motion state of each control joint according to the intention label; the demonstration action information includes demonstration trajectory information and acceleration curve information; the determining the corresponding intention label according to the demonstration action information of each movement stage, and updating the joint motion state of each control joint according to the intention label, comprises: determining the corresponding non-rational feature according to the demonstration action information of each movement stage, the non-rational feature being a feature model constructed based on a specific human action; generating a corresponding disturbance residual according to the non-rational feature, and superimposing the generated disturbance residual to the joint motion state of each control joint for state updating; generating simulation control instructions of each joint of the humanoid robot according to the joint motion state of each control joint after action intention correction, and sending the simulation control instructions to the corresponding humanoid robot for action simulation.
2. The AI-based humanoid robot demonstration method of claim 1, wherein the determining the corresponding intention label according to the demonstration action information of each movement stage, and updating the joint motion state of each control joint according to the intention label, comprises: determining the corresponding intention label according to the demonstration action information of each movement stage, and determining the velocity curve template corresponding to the corresponding demonstration action information according to the intention label; updating the velocity curve information of each control joint according to the velocity curve template.
3. The AI-based humanoid robot demonstration method of claim 1, wherein the determining the corresponding intention label according to the demonstration action information of each movement stage, and updating the joint motion state of each control joint according to the intention label, comprises: According to the teaching action information of each movement stage, a corresponding intention label is determined, when the length of the teaching trajectory information of the corresponding movement stage exceeds a set value, then according to the intention label, a corresponding jitter superposition waveform is determined, the jitter superposition waveform is fused with the original teaching trajectory information, and the joint motion state of each control joint is updated according to the fused trajectory data; According to the teaching action information of each movement stage, the acceleration peak value of the corresponding action is determined, and the control point corresponding to the acceleration peak value is taken as a turning point, a pause with a set time interval is inserted at the teaching trajectory information at the turning point with a set probability; the joint motion state of each control joint is updated according to the updated data. 4.The AI-based humanoid robot teaching method of claim 1, wherein The continuous frame human body teaching image is input into the pre-constructed action recognition model to determine the teaching action information of each movement stage, and the action recognition model further includes an attention module, and the correlation between joints is captured through the attention module to determine the power chain change information corresponding to the corresponding teaching action information. The continuous frame human body teaching image is input into the pre-constructed action recognition model to determine the teaching action information of each movement stage, and the action recognition model further includes an attention module, and the correlation between joints is captured through the attention module to determine the power chain change information corresponding to the corresponding teaching action information. The continuous frame human body teaching image is input into the pre-constructed action recognition model to determine the teaching action information of each movement stage, and the action recognition model further includes an attention module, and the correlation between joints is captured through the attention module to determine the power chain change information corresponding to the corresponding teaching action information. The continuous frame human body teaching image is input into the pre-constructed action recognition model to determine the teaching action information of each movement stage, and the action recognition model further includes an attention module, and the correlation between joints is captured through the attention module to determine the power chain change information corresponding to the corresponding teaching action information. After the simulation control instruction is sent to the corresponding humanoid robot for action simulation, the simulation control instruction is further saved when the action simulation result meets the requirements. 5.The AI-based humanoid robot teaching method of claim 1, wherein After the simulation control instruction is sent to the corresponding humanoid robot for action simulation, the simulation control instruction is further saved when the action simulation result meets the requirements. It includes:
6. An AI-based humanoid robot teaching system, characterized by, The image acquisition module is used to acquire the human body teaching image of the teaching personnel in the current scene through the camera assembly; The recognition module is used to input the continuous frame human body teaching image into the pre-constructed action recognition model to determine the teaching action information of each movement stage and the human body key point coordinates of each movement stage, determine the robot joints corresponding to each human body key point, and calculate the joint motion state of each control joint of the corresponding robot; The joint motion state includes velocity curve information; the continuous frame human body teaching image is input into the pre-constructed action recognition model to determine the human body key point coordinates of each movement stage, determine the robot joints corresponding to each human body key point, and include: The continuous frame human body teaching image is preprocessed to obtain a preprocessed image, and the preprocessing includes image size adjustment, normalization processing and color space conversion; The preprocessed image is analyzed by a human posture estimation model to obtain human body key point coordinates and corresponding confidence of each movement stage, and low-confidence key point coordinates are filtered out; The preprocessed image is analyzed by a human posture estimation model to obtain human body key point coordinates and corresponding confidence of each movement stage, and low-confidence key point coordinates are filtered out; According to the human pose estimation model, human key points are determined, human joint information corresponding to the human key points is determined according to a set key point mapping rule, and robot joints corresponding to the human key points are determined. The state updating module is configured to determine an intention label according to the teaching motion information of each movement stage, and update the joint motion state of each control joint according to the intention label; the teaching motion information includes teaching trajectory information and acceleration curve information; the determination of the intention label according to the teaching motion information of each movement stage and the updating of the joint motion state of each control joint according to the intention label include: determining a non-rational feature corresponding to each movement stage according to the teaching motion information, the non-rational feature being a feature model constructed based on a specific human action; generating a disturbance residual according to the non-rational feature, and superimposing the generated disturbance residual on the joint motion state of each control joint to update the state; The instruction generation module is configured to generate simulation control instructions of each joint of the humanoid robot according to the joint motion state of each control joint after the action preference correction, and send the simulation control instructions to the corresponding humanoid robot for action simulation.
7. An electronic device, comprising: comprising: a memory storing executable program code; a processor coupled to the memory; the processor invokes the executable program code stored in the memory to execute the AI-based humanoid robot teaching method of any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program causes the computer to execute the AI-based humanoid robot teaching method of any one of claims 1 to 5. The computer readable storage medium stores a computer program, wherein the computer program causes the computer to execute the AI-based humanoid robot teaching method of any one of claims 1 to 5.
Citation Information
Patent Citations
Teaching robot control system and control method based on image guidance
CN115741714A
Robot control method, device and system
CN119772907A