Robot reorientation method, device, equipment and medium based on human foot-ground contact relationship
Patent Information
- Application Number
- CN202611096085.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-23
AI Technical Summary
然而,人体与机器人在骨长比例、关节自由度、足部结构、质量分布、关节活动范围和驱动方式方面存在明显差异
[0014]本申请中,通过预设人体运动重建算法对目标人体运动视频进行运动重建,以得到相应的人体运动数据,并基于所述人体运动数据进行人体模型前向计算,以得到所述目标人体运动视频中每个视频帧对应的人体三维坐标;所述人体三维坐标包括人体三维网格顶点坐标和人体三维骨骼关节坐标;基于所述人体三维坐标确定所述每个视频帧的人体足部代表点,并估计所述视频帧中的地面参考高度,以及根据所述人体足部代表点的运动特征和所述地面参考高度,为所述人体足部代表点生成相应的足地接触伪标签;所述人体足部代表点为人体的足部特征点位,包括脚尖、脚跟及脚底代表点中的至少一种;以所述足地接触伪标签为监督信号训练预设的时序接触预测模型,并通过训练后时序接触预测模型,预测所述每个视频帧的人体足地接触概率,以根据得到的预测结果确定所述目标人体运动视频的人体足地接触帧区间;将所述人体足地接触帧区间转换为目标机器人的支撑端的逐帧足地接触约束,并将所述逐帧足地接触约束引入所述目标机器人的重定向优化过程中,以得到所述目标机器人的目标动作序列。由上可见,本申请先利用预设人体运动重建算法解析目标人体运动视频,输出人体运动数据,再通过人体模型前向计算得到每一视频帧对应的人体三维网格顶点坐标与骨骼关节坐标,依托人体三维坐标提取脚尖、脚跟、脚底等人体足部代表点,结合所述人体足部代表点的运动特征与估算的地面参考高度生成足地接触伪标签,以所述足地接触伪标签作为监督信号训练时序接触预测模型,依靠模型输出各帧的足地接触概率,以确定人体足地接触帧区间,再把所述人体足地接触帧区间转化为机器人支撑端的逐帧足底接触约束,将约束融入机器人动作重定向优化计算,最终输出适配机器人执行的目标动作序列。这样一来,通过本申请的上述过程,通过人体运动重建算法搭配人体模型前向计算获取人体三维网格与骨骼坐标,能够仅凭视频素材还原人体三维运动信息,无需专业动作捕捉设备即可完成人体运动数字化还原,降低动作采集的硬件成本;从三维坐标提取脚尖、脚跟、脚底足部代表点,并结合点位运动特征与地面参考高度自动生成足地接触伪标签,无需人工逐帧标注脚部接触状态即可自主生成训练监督数据,解决足地接触标签获取成本高、标注效率低的问题;采用时序接触预测模型以伪标签完成训练并预测逐帧接触概率,同时划分连续接触帧区间,能够过滤单帧预测误判、平滑时序接触状态,输出稳定连贯的脚部支撑时段;将人体侧接触帧区间映射为机器人支撑端专属逐帧接触约束,并嵌入重定向优化流程同步求解,在动作映射阶段就同步约束机器人足底离地高度、水平滑移幅度与足底姿态,抑制机器人足部穿地、悬空、滑动、支撑逻辑错乱等缺陷,同时无需依赖机器人足底接触传感器即可生成有效接触约束,能够适配不同足部结构、无足底传感设备的各类足式机器人,保留原始人体视频的动作节奏与风格,优化后的机器人动作序列支撑逻辑贴合真实人体运动规律,可直接用于仿真验证、机器人模仿学习与实体机动作部署,进而优化机器人重定向方法以减少机器人动作中的足部滑动、足部穿地和支撑状态错误的问题。
Smart Images

Figure CN122598273B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot motion control technology, and in particular to robot repositioning methods, devices, equipment and media based on the human foot-ground contact relationship. Background Technology
[0002] With the development of humanoid robots, bipedal robots, and other legged robot technologies, generating robot-executable actions from human motion videos has become an important technical approach for robot skill acquisition. This type of technology typically includes steps such as human motion recovery, motion retargeting, simulation verification, and training through imitation learning or reinforcement learning. Existing technologies in the field of motion retargeting include geometric / optimization-based motion retargeting methods such as GMR (General Motion Retargeting), as well as methods for correcting foot slippage, ground penetration, and smoothing based on post-retargeting processing. These methods typically focus primarily on matching the position, rotation, or end-effector trajectory of key rigid bodies. However, there are significant differences between humans and robots in terms of bone length ratios, joint degrees of freedom, foot structure, mass distribution, joint range of motion, and actuation methods. If only geometric position and rotation matching are relied upon, robot retargeting results are prone to problems such as foot slippage, foot penetration, foot suspension, incorrect support states, root node height drift, and localized motion jitter.
[0003] In summary, optimizing robot repositioning methods to reduce foot slippage, foot treading, and support state errors during robot movements is a problem that urgently needs to be solved. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a robot repositioning method, apparatus, device, and medium based on the human foot-ground contact relationship, which can optimize the robot repositioning method to reduce problems such as foot slippage, foot penetration, and incorrect support state during robot movements. The specific solution is as follows: In a first aspect, this application provides a robot repositioning method based on human foot-ground contact relationship, including: The target human motion video is reconstructed using a preset human motion reconstruction algorithm to obtain corresponding human motion data. Based on the human motion data, a forward calculation of the human model is performed to obtain the three-dimensional coordinates of the human body corresponding to each video frame in the target human motion video. The three-dimensional coordinates of the human body include the coordinates of the vertex of the three-dimensional human body mesh and the coordinates of the three-dimensional human body skeletal joints. Based on the three-dimensional coordinates of the human body, the representative point of the human foot in each video frame is determined, and the ground reference height in the video frame is estimated. Based on the motion characteristics of the representative point of the human foot and the ground reference height, a corresponding foot-ground contact pseudo-label is generated for the representative point of the human foot. The representative point of the human foot is a characteristic point of the human foot, including at least one of the representative points of the toe, heel and sole. The preset temporal contact prediction model is trained using the foot-ground contact pseudo-label as a supervision signal, and the probability of human foot-ground contact in each video frame is predicted by the trained temporal contact prediction model, so as to determine the human foot-ground contact frame interval of the target human motion video based on the obtained prediction results. The human foot-to-ground contact frame interval is converted into a frame-by-frame foot-to-ground contact constraint of the support end of the target robot, and the frame-by-frame foot-to-ground contact constraint is introduced into the retargeting optimization process of the target robot to obtain the target action sequence of the target robot.
[0005] Optionally, the forward calculation of the human model based on the human motion data includes: If the posture parameter dimension of the human motion data is inconsistent with the target dimension required for the forward calculation of the human model, then the missing posture parameter dimension of the human motion data is filled in to obtain the filled human motion data. The human body model is forward-calculated based on the supplemented human motion data.
[0006] Optionally, determining the representative point of the human foot in each video frame based on the three-dimensional coordinates of the human body includes: From the vertex coordinates of the human body 3D mesh, the vertex coordinate set of the front end of the foot, the vertex coordinate set of the rear end of the foot, and the vertex coordinate set of the lowest region of the foot are determined; the vertex coordinate set of the front end of the foot is the set of vertex coordinates in the human foot mesh region where the axial coordinate data of the foot along the forward direction of the human body is greater than a first preset coordinate threshold; the vertex coordinate set of the rear end of the foot is the set of vertex coordinates in the human foot mesh region where the axial coordinate data of the foot along the forward direction of the human body is less than a second preset coordinate threshold, where the first preset coordinate threshold is greater than the second preset coordinate threshold; the vertex coordinate set of the lowest region of the foot is the set of vertex coordinates in the human foot mesh region where the vertical height axial coordinate data is less than a third preset coordinate threshold. The average coordinate positions of the set of coordinates of the front end vertex of the foot, the set of coordinates of the rear end vertex of the foot, and the set of coordinates of the lowest region vertex of the foot are respectively determined as the toe representative point, the heel representative point, and the sole representative point.
[0007] Optionally, estimating the ground reference height in the video frame includes: Statistical analysis of the temporal height distribution of the representative points of the human foot in the target human motion video; The ground reference height in the video frame is estimated based on the temporal height distribution.
[0008] Optionally, generating corresponding foot-ground contact pseudo-labels for the human foot representative points based on the motion characteristics of the human foot representative points and the ground reference height includes: If the human foot representative point continuously meets the preset foot-ground contact condition in several adjacent video frames, then it is determined that the human foot representative point is in the target foot-ground contact state in the several adjacent video frames. For the representative points of the human foot corresponding to the adjacent video frames, generate pseudo-labels of foot-ground contact indicating that the foot is in contact with the ground; and generate pseudo-labels of foot-ground contact indicating that the foot is not in contact with the ground for the representative points of the human foot corresponding to other video frames; the other video frames are the video frames other than the adjacent video frames in the target human motion video. The preset foot-ground contact conditions include the difference between the vertical height of the representative point of the human foot and the reference height of the ground being less than a preset height threshold, the velocity of the representative point of the human foot in the vertical direction being less than a preset vertical velocity threshold, and the velocity of the representative point of the human foot in the horizontal direction being less than a preset horizontal velocity threshold.
[0009] Optionally, the prediction results include the probability of toe contact and the probability of heel contact for each video frame; Accordingly, determining the human foot-ground contact frame interval of the target human motion video based on the obtained prediction results includes: The toe contact probability and the heel contact probability are smoothed frame by frame to obtain the corresponding smoothed toe contact probability and smoothed heel contact probability. The smooth toe contact probability and the smooth heel contact probability of the same foot in the same video frame are fused to obtain the corresponding foot support confidence. Based on the foot support confidence level, determine whether each video frame meets the preset support conditions to obtain the corresponding judgment result; Based on the judgment result, adjacent video frames that continuously meet the preset support conditions are merged into a human foot-ground contact frame interval.
[0010] Optionally, the step of converting the human foot-to-ground contact frame interval into frame-by-frame foot-to-ground contact constraints for the support end of the target robot, and introducing the frame-by-frame foot-to-ground contact constraints into the retargeting optimization process of the target robot to obtain the target action sequence of the target robot, includes: Based on the robot kinematics model of the target robot and the pre-established support end semantic mapping relationship, the human foot-ground contact frame interval is converted into the frame-by-frame foot-ground contact constraint of the target robot. The frame-by-frame foot contact constraint is introduced into the redirection optimization process of the target robot to obtain the initial action sequence of the target robot; The initial action sequence is subjected to temporal smoothing optimization to obtain the optimized target action sequence.
[0011] Secondly, this application provides a robot repositioning device based on the human foot-to-ground contact relationship, comprising: The forward calculation module is used to perform motion reconstruction on the target human motion video using a preset human motion reconstruction algorithm to obtain the corresponding human motion data, and to perform forward calculation on the human model based on the human motion data to obtain the human three-dimensional coordinates corresponding to each video frame in the target human motion video; the human three-dimensional coordinates include the human three-dimensional mesh vertex coordinates and the human three-dimensional skeletal joint coordinates. The tag generation module is used to determine the representative point of the human foot in each video frame based on the three-dimensional coordinates of the human body, estimate the ground reference height in the video frame, and generate corresponding foot-ground contact pseudo-tags for the representative point of the human foot according to the motion characteristics of the representative point of the human foot and the ground reference height; the representative point of the human foot is a characteristic point of the human foot, including at least one of the representative points of the toe, heel and sole. The probability prediction module is used to train a preset temporal contact prediction model with the foot-ground contact pseudo-label as a supervision signal, and to predict the probability of human foot-ground contact in each video frame through the trained temporal contact prediction model, so as to determine the human foot-ground contact frame interval of the target human motion video based on the obtained prediction results. The constraint introduction module is used to convert the human foot-to-ground contact frame interval into frame-by-frame foot-to-ground contact constraints of the support end of the target robot, and introduce the frame-by-frame foot-to-ground contact constraints into the retargeting optimization process of the target robot to obtain the target action sequence of the target robot.
[0012] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned robot repositioning method based on human foot-to-ground contact relationship.
[0013] Fourthly, this application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned robot repositioning method based on human foot-ground contact relationship.
[0014] In this application, a preset human motion reconstruction algorithm is used to reconstruct the motion of a target human motion video to obtain corresponding human motion data. Based on this human motion data, a forward calculation of the human model is performed to obtain the three-dimensional coordinates of the human body corresponding to each video frame in the target human motion video. The three-dimensional coordinates include the coordinates of the vertex of the three-dimensional human mesh and the coordinates of the three-dimensional human skeletal joints. Based on these three-dimensional coordinates, a representative point of the human foot in each video frame is determined, and the ground reference height in the video frame is estimated. Furthermore, based on the motion characteristics of the representative point of the human foot and the ground reference height, a corresponding foot-ground contact pseudo-point is generated for the representative point of the human foot. The labels are: the human foot representative points are the characteristic points of the human foot, including at least one of the toe, heel, and sole representative points; a preset temporal contact prediction model is trained using the foot-to-ground contact pseudo-labels as supervision signals, and the probability of human foot-to-ground contact in each video frame is predicted by the trained temporal contact prediction model, so as to determine the human foot-to-ground contact frame interval of the target human motion video according to the obtained prediction results; the human foot-to-ground contact frame interval is converted into frame-by-frame foot-to-ground contact constraints of the support end of the target robot, and the frame-by-frame foot-to-ground contact constraints are introduced into the retargeting optimization process of the target robot to obtain the target action sequence of the target robot. As can be seen from the above, this application first uses a preset human motion reconstruction algorithm to analyze the target human motion video and output human motion data. Then, it uses the human model to calculate the vertex coordinates and skeletal joint coordinates of the human three-dimensional mesh corresponding to each video frame. Based on the human three-dimensional coordinates, it extracts representative points of the human foot such as toes, heels, and soles. Combining the motion characteristics of the representative points of the human foot with the estimated ground reference height, it generates foot-to-ground contact pseudo-labels. The foot-to-ground contact pseudo-labels are used as supervision signals to train the temporal contact prediction model. The model outputs the foot-to-ground contact probability of each frame to determine the human foot-to-ground contact frame interval. Then, the human foot-to-ground contact frame interval is transformed into frame-by-frame sole contact constraints of the robot support end. The constraints are integrated into the robot motion retargeting optimization calculation, and finally, the target motion sequence adapted to the robot execution is output.In this way, through the process described above in this application, by using a human motion reconstruction algorithm combined with forward computation of a human model to obtain the human body's 3D mesh and skeletal coordinates, it is possible to reconstruct the human body's 3D motion information solely from video footage, completing the digital reconstruction of human motion without the need for professional motion capture equipment, thus reducing the hardware cost of motion acquisition. Representative points of the toes, heels, and soles of the feet are extracted from the 3D coordinates, and foot-ground contact pseudo-labels are automatically generated by combining the point motion characteristics with ground reference height. Training and supervision data can be generated autonomously without manual frame-by-frame annotation of foot contact states, solving the problems of high cost and low annotation efficiency in obtaining foot-ground contact labels. A temporal contact prediction model is used to complete training with pseudo-labels and predict frame-by-frame contact probabilities. Simultaneously, by dividing continuous contact frame intervals, it can filter out single-frame prediction misjudgments, smooth temporal contact states, and output stable and consistent foot data. During the support period, the human side contact frame interval is mapped to the robot's support end frame-by-frame contact constraints, and embedded in the redirection optimization process for synchronous solution. During the motion mapping stage, the robot's foot height off the ground, horizontal sliding amplitude, and foot posture are simultaneously constrained, suppressing defects such as robot feet penetrating the ground, dangling, sliding, and support logic errors. At the same time, effective contact constraints can be generated without relying on robot foot contact sensors, which can adapt to various legged robots with different foot structures and without foot sensors. It retains the motion rhythm and style of the original human video, and the optimized robot motion sequence support logic conforms to the real human motion law. It can be directly used for simulation verification, robot imitation learning, and physical robot motion deployment, thereby optimizing the robot redirection method to reduce problems such as foot sliding, foot penetrating the ground, and support state errors in robot motion. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0016] Figure 1 This is a flowchart of a robot repositioning method based on human foot-ground contact relationship disclosed in this application; Figure 2 This is a schematic diagram of the overall process of a robot repositioning method based on human foot-ground contact relationship disclosed in this application; Figure 3 This is a schematic diagram of a foot-ground contact prediction process disclosed in this application; Figure 4 A schematic diagram illustrating the process of introducing redirection optimization for frame-by-frame foot contact constraints disclosed in this application; Figure 5This is a schematic diagram of a robot repositioning device based on the human foot-ground contact relationship disclosed in this application; Figure 6 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Existing technologies in motion retargeting include geometric / optimization-based motion retargeting methods such as GMR, as well as methods for correcting foot slippage, ground penetration, and smoothing based on post-retargeting processing. These methods typically focus on matching the position, rotation, or end-effector trajectory of key rigid bodies. However, there are significant differences between humans and robots in terms of bone length ratios, joint degrees of freedom, foot structure, mass distribution, joint range of motion, and actuation methods. If only geometric position and rotation matching are relied upon, robot retargeting results are prone to problems such as foot slippage, foot penetration, foot suspension, incorrect support states, root node height drift, and local motion jitter.
[0019] To overcome the aforementioned technical problems, this application provides a robot repositioning method based on the human foot-ground contact relationship, which can optimize the robot repositioning method to reduce problems such as foot slippage, foot penetration, and incorrect support state during robot movements.
[0020] See Figure 1 As shown, this embodiment of the invention discloses a robot repositioning method based on human foot-ground contact relationship, including: Step S11: Perform motion reconstruction on the target human motion video using a preset human motion reconstruction algorithm to obtain corresponding human motion data, and perform forward calculation of the human model based on the human motion data to obtain the human three-dimensional coordinates corresponding to each video frame in the target human motion video; the human three-dimensional coordinates include the human three-dimensional mesh vertex coordinates and the human three-dimensional skeletal joint coordinates.
[0021] In this embodiment, a preset human motion reconstruction algorithm is used to process the target human motion video to generate human motion data. Then, based on the human motion data, a forward operation of the human model is performed to output the human three-dimensional coordinates corresponding to each frame in the video, including the human three-dimensional mesh vertex coordinates and three-dimensional skeletal joint coordinates.
[0022] It should be noted that, addressing the pain points of existing robot repositioning technologies, this application proposes a robot repositioning method based on human foot-to-ground contact relationships. This method can infer foot-to-ground contact relationships from human video movements without relying on robot foot contact sensors or discarding a large amount of original motion data, and transform this into robot side contact constraints. This is then explicitly introduced into the robot motion repositioning process, thereby reducing foot slippage, foot penetration, and support state errors in robot movements. Figure 2 The diagram shows the overall flow of a robot relocation method based on human foot-to-ground contact relationship provided in this application. The method first utilizes GVHMR (Gravity-View Human MotionRecovery, an AI video motion capture technology), HMR4D (4D Human Mesh Recovery, a four-dimensional video human mesh reconstruction), or other human motion recovery models to recover three-dimensional human motion data from human motion videos. Then, based on the human model parameters, forward computation is performed to obtain the vertices or joints of the human three-dimensional mesh, and representative points of the toe, heel, or sole are constructed from the human foot region. Subsequently, the human foot-to-ground contact state is predicted based on the height, velocity, and temporal motion characteristics of the representative foot points, generating contact confidence and contact intervals. Finally, the human-side contact interval is converted into robot-side contact constraints through a support-end semantic mapping relationship, and these contact constraints are added to GMR, inverse kinematics optimization, nonlinear optimization, geometric motion relocation, or other robot motion relocation optimization objectives to obtain a contact-aware robot motion sequence. This method is particularly suitable for scenarios where SMPL (Skinned Multi-Person Linear) motion parameters are recovered from human dance videos using a human motion video reconstruction model, and then robot motion is generated using GMR or other geometric / optimized motion retargeting frameworks. In such frameworks, the original retargeting process typically focuses primarily on the position and pose matching of key parts of the human and robot, while lacking explicit constraints on the height, slippage, and pose of the supporting foot within the contact zone. This application reduces slippage and ground-penetrating phenomena of the robot's supporting foot during the contact phase by introducing additional foot-ground contact constraints.
[0023] It is understood that the system architecture of the method in this application may include the following modules: human motion data acquisition module, human model forward calculation module, foot representative point construction module, foot-ground contact prediction module, contact interval generation module, support end semantic mapping module, contact constraint redirection module, temporal repair module, and motion output module. The system includes: a human motion data acquisition module for acquiring human motion data recovered from human motion videos; a human model forward calculation module for calculating human 3D mesh vertices or human 3D joints based on the human motion data; a foot representative point construction module for determining human foot representative points based on the human 3D mesh vertices or human 3D joints; a foot-to-ground contact prediction module for predicting human foot-to-ground contact relationships based on the motion characteristics of the human foot representative points; a contact interval generation module for generating contact confidence and contact intervals; a support end semantic mapping module for mapping human foot-to-ground contact relationships to time-varying contact constraints at the robot support end; a contact constraint redirection module for adding the time-varying contact constraints during robot motion redirection optimization; a temporal repair module for performing contact interval-level temporal repair on the contact-aware robot motion sequence; and an action output module for outputting the robot's target motion sequence.
[0024] Specifically, the process begins by acquiring human motion data: inputting a human motion video, and generating 3D human motion data using GVHMR, HMR4D, or other human motion reconstruction models. In one specific implementation, the human motion data includes SMPL human model parameters, at least one of body_pose, betas, global_orient, and transl. Here, body_pose represents the pose of each joint of the human body (human joint pose parameters), betas represents the human body shape parameters, global_orient represents the global orientation of the human root node, and transl represents the translation of the human root node. Then, forward computation is performed. Based on the human model parameters output by the human motion reconstruction model, forward computation of the human model is executed to obtain the vertices or joints of the human 3D mesh for each frame. It is understood that the human motion reconstruction model is not limited to GVHMR or HMR4D, and can be replaced with other monocular, multi-view, or multi-sensor human motion reconstruction models; the human model is not limited to SMPL, and can be replaced with SMPL-X or other parametric human models; the input action is not limited to dance, and can be extended to actions with foot-ground contact relationships such as walking, jumping, climbing stairs, and climbing platforms. It should be noted that the processing flow for forward calculation of the human body model based on the human motion data is as follows: If the posture parameter dimension of the human motion data is inconsistent with the target dimension required for forward calculation of the human body model, the missing posture parameter dimension of the human motion data is padded with posture to obtain padded human motion data; forward calculation of the human body model is performed based on the padded human motion data. That is, when the posture parameter dimension output by the upstream human motion recovery model is not completely consistent with the dimension required for forward calculation of the subsequent human body model, the missing joint posture dimension can be padded with a default posture. For example, for the missing hand-related posture dimension, zero value padded can be used. Since this invention mainly focuses on the lower limb and foot-ground contact relationship, the above-mentioned padded will not affect the main technical effect of foot-ground contact judgment.In this embodiment, a human motion reconstruction algorithm is used to parse the original video and generate standardized motion data for contact relationship modeling before robot motion retargeting. This allows for the reconstruction of complete human dynamic motion logic from a two-dimensional image, compensating for the lack of depth dimension information in a single two-dimensional image. This enables the recovery of foot-to-ground contact relationships from human motion videos and their application to robot motion retargeting, reducing reliance on high-precision motion capture equipment. Dimension consistency verification and posture completion correct for missing posture parameters and incomplete dimensions in the reconstructed motion data, eliminating computational errors and 3D coordinate distortion caused by mismatch between the original motion data and the human model input specifications. Furthermore, instead of treating human motion as a position and rotation trajectory to be matched, foot-to-ground contact semantics—which foot contacts the ground, when it contacts the ground, and whether the foot should remain stable during contact—are used as constraints for retargeting optimization. This ensures that the robot motion not only closely approximates human motion in overall posture but also is more rational in terms of supporting logic and foot-to-ground interaction.
[0025] Step S12: Determine the representative point of the human foot in each video frame based on the three-dimensional coordinates of the human body, estimate the ground reference height in the video frame, and generate corresponding foot-ground contact pseudo-labels for the representative point of the human foot according to the motion characteristics of the representative point of the human foot and the ground reference height; the representative point of the human foot is a characteristic point of the human foot, including at least one of the representative points of the toe, heel and sole.
[0026] In this embodiment, representative points of the human foot in each frame are constructed based on the three-dimensional coordinates of the human body in each frame. Simultaneously, the ground reference height for each frame is estimated. A corresponding foot-to-ground contact pseudo-label is generated by combining the motion characteristics of the representative points of the human foot with the ground reference height. The representative points of the human foot are characteristic points of the human foot, including at least one of the following: left toe representative point, left heel representative point, right toe representative point, right heel representative point, left sole representative point, or right sole representative point. These can be constructed from human joints, human mesh vertices, a local foot coordinate system, or a combination of the above data. The foot-to-ground contact pseudo-label can be generated from foot height, velocity, acceleration, ground distance, foot posture, or a combination of the above features.
[0027] It should be noted that the processing flow for determining the representative points of the human foot in each video frame based on the three-dimensional coordinates of the human body is as follows: The set of vertex coordinates for the front end of the foot, the set of vertex coordinates for the rear end of the foot, and the set of vertex coordinates for the lowest region of the foot are determined from the vertex coordinates of the three-dimensional human body mesh. The set of vertex coordinates for the front end of the foot is the set of vertex coordinates in the human foot mesh region where the axial coordinate data of the foot along the direction of human movement is greater than a first preset coordinate threshold. The set of vertex coordinates for the rear end of the foot is the set of vertex coordinates in the human foot mesh region where the axial coordinate data of the foot along the direction of human movement is less than a second preset coordinate threshold, where the first preset coordinate threshold is greater than the second preset coordinate threshold. The set of vertex coordinates for the lowest region of the foot is the set of vertex coordinates in the human foot mesh region where the vertical height axial coordinate data is less than a third preset coordinate threshold. The average coordinate positions of the set of vertex coordinates for the front end of the foot, the set of vertex coordinates for the rear end of the foot, and the set of vertex coordinates for the lowest region of the foot are respectively determined as the representative point of the toe, the representative point of the heel, and the representative point of the sole. That is, since standard human joint points typically only include skeletal points such as the hip, knee, and ankle, and may not directly include foot contact points such as the toes and heels, this invention preferably constructs representative foot points based on the vertices of the human foot mesh. In one embodiment, a set of vertices at the front end of the foot can be selected from the human foot mesh region, and their average position can be used as the representative toe point; a set of vertices at the rear end of the foot can be selected, and their average position can be used as the representative heel point; a set of vertices in the lowest region of the foot can be selected, and their average position can be used as the representative sole point. In this way, even if the upstream human motion recovery results do not directly output toe or heel joint points, this invention can still construct fine-grained foot reference points for contact judgment from the human model mesh. In one embodiment, the human motion recovery model output uses the Y-up coordinate system. In this case, the Y component in the three-dimensional coordinates represents the vertical height, the positive Y direction represents upward, the negative Y direction represents the direction of gravity, and the ground tangential plane is the XZ plane. Therefore, for any representative foot point, its vertical height is determined by its Y-coordinate, and its horizontal position is determined by its XZ coordinates. In other words, the displacement of the representative foot point in the XZ plane represents the tangential displacement of the ground. Subsequent foot-ground contact prediction, ground height estimation, and foot slip constraints can all be based on this coordinate definition.
[0028] It should be further noted that the processing flow for estimating the ground reference height in the video frame is as follows: The temporal height distribution of the representative points of the human foot in the target human motion video is statistically analyzed; the ground reference height in the video frame is estimated based on the temporal height distribution. That is, to determine whether the foot is in contact with the ground, it is necessary to estimate the ground reference in the human movement. In simple dance movements, the ground can usually be approximated as a fixed plane. The system can estimate the ground height based on the height distribution of the representative points of the foot in the time series. For example, candidate contact frames with lower foot height and lower foot speed can be selected, and the ground reference can be estimated based on the foot height of these candidate contact frames. In other embodiments, the ground reference can also be a time-varying height function, a piecewise plane, or a local support plane.
[0029] It should be noted that the processing flow for generating corresponding foot-to-ground contact pseudo-labels for the representative point of the human foot is as follows: If the representative point of the human foot continuously meets the preset foot-to-ground contact conditions in several adjacent video frames, then the representative point of the human foot is determined to be in the target foot-to-ground contact state in those several adjacent video frames; a foot-to-ground contact pseudo-label representing the presence of foot-to-ground contact is generated for the representative point of the human foot corresponding to the several adjacent video frames, and a foot-to-ground contact pseudo-label representing the absence of foot-to-ground contact is generated for the representative point of the human foot corresponding to other video frames; the other video frames are those other than the several adjacent video frames in the target human motion video; wherein, the preset foot-to-ground contact conditions include the difference between the vertical height of the representative point of the human foot and the ground reference height being less than a preset height threshold, the velocity of the representative point of the human foot in the vertical direction being less than a preset vertical velocity threshold, and the velocity of the representative point of the human foot in the horizontal direction being less than a preset horizontal velocity threshold. That is, in the absence of manually labeled foot-to-ground contact labels, this embodiment first automatically generates foot-to-ground contact pseudo-labels based on the motion characteristics of the representative point of the foot. For any foot representative point, it can be determined to be in contact state when the following conditions are met: the height of the foot representative point relative to the ground reference is less than a height threshold; the velocity of the foot representative point in the tangential plane of the ground is less than a horizontal velocity threshold; the velocity of the foot representative point in the vertical direction is less than a vertical velocity threshold; and this state remains continuous or stable within several neighboring frames. Therefore, frame-by-frame contact pseudo-labels for four channels—left toe, left heel, right toe, and right heel—can be automatically generated. To improve robustness, continuous contact scores can also be generated instead of directly generating binary labels. The continuous contact scores can be obtained by combining foot height, horizontal velocity, vertical velocity, and neighborhood temporal stability. In this way, this embodiment can obtain fine-grained contact reference points such as toes, heels, and soles by forward calculation of the human body model and construction of foot representative points, ensuring stable spatial position measurement of representative points and avoiding depth deviation caused by two-dimensional image positioning; it estimates the ground reference height for each frame, adapting to scenes of human undulation and slight changes in ground height in the video, and provides a frame-by-frame dynamic ground judgment benchmark; it generates foot-to-ground contact pseudo-labels based on the height, horizontal velocity, and vertical velocity of the foot representative points, which can distinguish between two approximate height scenarios: the foot quickly gliding across the ground at low altitude and the foot landing still, reducing misidentification caused by a single height threshold; it determines the contact state only after continuous compliance for multiple frames, filtering out false contact signals in a single frame caused by 3D reconstruction coordinate jitter and instantaneous measurement errors, and avoiding inter-frame jumps and fragmented noise in the labels.
[0030] Step S13: Train a preset temporal contact prediction model using the foot-to-ground contact pseudo-label as a supervision signal, and predict the probability of human foot-to-ground contact in each video frame using the trained temporal contact prediction model, so as to determine the human foot-to-ground contact frame interval of the target human motion video based on the obtained prediction results.
[0031] In this embodiment, the generated foot-to-ground contact pseudo-labels are used as supervisory signals to train the temporal contact prediction model. The trained temporal contact prediction model is then used to predict the probability of human foot-to-ground contact frame-by-frame. Based on the prediction results, the intervals of human foot-to-ground contact frames in the video where the foot makes contact with the ground are determined. Figure 3 The diagram illustrates a flowchart of a foot-ground contact prediction method provided in this application. The temporal contact prediction model can employ a multilayer perceptron, temporal convolutional network, recurrent neural network, graph neural network, Transformer network, or other temporal classification models. Its input is the three-dimensional motion features of the human body within a time window near the target frame, and it outputs four contact probabilities, corresponding to the contact probabilities of the left toe, left heel, right toe, and right heel, respectively. The three-dimensional motion features of the human body can include at least one of the following: the three-dimensional positions of the left and right toes and heels; the height of the left and right toes and heels relative to a ground reference; the horizontal and vertical velocities of the left and right toes and heels; the three-dimensional positions or rotations of the left and right ankle joints, knee joints, hip joints, and pelvis; root node height, root node velocity, and root node orientation; local foot posture or plantar normal; and confidence information output by the human motion recovery model.
[0032] It should be noted that the prediction results include the toe contact probability and heel contact probability of each video frame. Correspondingly, the processing flow for determining the human foot-to-ground contact frame interval of the target human motion video based on the obtained prediction results is as follows: The toe contact probability and heel contact probability are smoothed frame-by-frame to obtain corresponding smoothed toe contact probability and smoothed heel contact probability; the smoothed toe contact probability and smoothed heel contact probability of the same foot in the same video frame are fused to obtain the corresponding foot support confidence score; based on the foot support confidence score, it is determined whether each video frame meets the preset support conditions to obtain the corresponding judgment result; based on the judgment result, adjacent video frames that continuously meet the preset support conditions are merged into a human foot-to-ground contact frame interval. That is, the model outputs toe and heel contact probabilities frame-by-frame, and the frame-by-frame contact probabilities output by the temporal contact prediction model are smoothed temporally to obtain the smoothed contact confidence score. Subsequently, the toe contact probability and heel contact probability of the same foot are fused to obtain the whole foot support confidence score. Based on the confidence level of the entire foot support, the frame-by-frame support state can be further determined, including left foot support, right foot support, both feet support, both feet off the ground, or a transitional support state. Then, adjacent frames that continuously meet the support conditions are merged into a contact interval. To avoid short-term misjudgments, contact segments with excessively short durations can be deleted, or adjacent contact segments with short intervals can be merged. For the start and end points of the contact interval, a soft switching window can be set to gradually strengthen or weaken the contact constraints near the boundary. Through this step, the human foot-ground contact relationship is transformed from frame-by-frame prediction results into contact intervals with clear start and end times and confidence levels. In this way, this embodiment uses rule-based pseudo-labels to train a temporal contact prediction model. This model can obtain relatively stable contact prediction results even without manual contact annotation. Temporal smoothing of the contact probability, deletion of short segments, merging of adjacent segments, and soft switching of contact boundaries generate stable contact intervals, making the probability curve temporally continuous and smooth, improving the fault tolerance and accuracy of support state determination, and reducing the impact of single-frame misjudgments on robot motion redirection.
[0033] Step S14: Convert the human foot-ground contact frame interval into a frame-by-frame foot-ground contact constraint of the support end of the target robot, and introduce the frame-by-frame foot-ground contact constraint into the redirection optimization process of the target robot to obtain the target action sequence of the target robot.
[0034] In this embodiment, the human foot-to-ground contact frame interval is mapped to a frame-by-frame foot-to-ground contact constraint adapted to the support end of the target robot. This constraint is then introduced into the target robot motion redirection optimization operation to obtain the target motion sequence that the target robot can execute. The target robot can be a humanoid / bipedal / legged robot; the redirection framework is not limited to GMR, but can also incorporate inverse kinematics optimization, nonlinear optimization, geometric motion redirection, learned redirection, or other redirection methods that do not explicitly handle foot-to-ground contact relationships; the support end can be the left and right feet of a humanoid robot, or an end link, foot region, or support structure in other legged robots that can form support contact with the ground; the frame-by-frame foot-to-ground contact constraint can include foot height constraint, tangential slip constraint, foot posture constraint, support end stability constraint, or a combination of the above constraints, determined by the human foot-to-ground contact relationship, the robot kinematic model, and the support end semantic mapping relationship, without relying on robot foot contact sensors. In optional embodiments, the frame-by-frame foot-to-ground contact constraint also incorporates robot foot pressure, torque, or contact sensor information. In addition, this application can also be used in conjunction with contact rewards, foot slip penalties, or trajectory tracking rewards in subsequent reinforcement learning training.
[0035] Specifically, based on the target robot's kinematic model and the pre-established support end semantic mapping relationship, the human foot-to-ground contact frame interval is converted into frame-by-frame foot-to-ground contact constraints for the target robot; these frame-by-frame foot-to-ground contact constraints are introduced into the target robot's retargeting optimization process to obtain the target robot's initial motion sequence; the initial motion sequence is then subjected to temporal smoothing optimization to obtain the optimized target motion sequence. That is, as... Figure 4The diagram illustrates a frame-by-frame foot contact constraint redirection optimization process provided in this application. This application does not limit the specific robot model and does not rely on robot foot contact sensors. For different robot platforms, the robot support end set can be defined using robot kinematic models, URDF (Unified Robot Description Format), MJCF (MuJoCo XML Format), or manually configured profiles. In optional embodiments, if the robot platform has foot pressure, torque, or contact sensors, this sensor information can be fused with the human side contact prediction results as auxiliary information to further improve the reliability of the contact constraints. In one embodiment, the system establishes a semantic mapping relationship for the support ends. This mapping relationship includes at least the correspondence between the human left foot contact area and the robot's left support end, the correspondence between the human right foot contact area and the robot's right support end, the correspondence between the human toe / heel contact state and the robot's front / back / center point on the sole of the foot, as well as the contact reference point, local coordinate system, and support direction of the robot support end. When the robot's foot does not have a clear toe and heel structure, the contact state of the human toe and heel can be fused into the whole foot support confidence and mapped to the contact constraint of the robot's foot center point or a set of multiple points on the foot.
[0036] It is understood that the frame-by-frame foot-to-ground contact constraints can be incorporated into GMR, inverse kinematics optimization, nonlinear optimization, geometric motion redirection, learned motion redirection, or other robot motion redirection frameworks. Preferably, this application focuses on redirection frameworks where the original optimization objective is primarily key rigid body position matching, rotation matching, or end-effector trajectory matching, and where foot-to-ground contact relationships are not explicitly addressed. In one embodiment, the motion redirection framework is GMR. This invention incorporates soft contact constraints (frame-by-frame foot-to-ground contact constraints) into the refinement optimization stage of GMR, allowing the original key rigid body position and rotation matching objectives to participate in the optimization along with the foot-to-ground contact constraints. Thus, the redirection result not only satisfies the geometric matching between the human body and key robot parts but also satisfies the height, slip, and posture constraints of the robot's support end within the contact interval. In another embodiment, the motion redirection framework is a geometric motion redirection framework based on inverse kinematics or nonlinear optimization. This invention incorporates contact confidence and contact interval into its optimization objectives, causing the robot's support end to tend towards a stable support state within the contact interval. The soft contact constraints may include at least one of the following: plantar height consistency constraints, plantar tangential slip suppression constraints, plantar posture consistency constraints, and contact confidence modulation mechanisms. The frame-by-frame foot-ground contact constraints incorporated during robot motion retargeting optimization include at least one of the following: incorporating plantar height consistency constraints to reduce robot support end suspension or penetration of the ground; incorporating plantar tangential slip suppression constraints to reduce robot support end slippage within the contact area; incorporating plantar posture consistency constraints to ensure the robot support end posture remains consistent with the ground reference; and modulating the weights of the constraints based on contact confidence. This application does not rely on robot plantar contact sensors; the contact state and constraint strength of the robot support end are jointly determined by human-side contact prediction results, the robot kinematic model, and the semantic mapping relationship of the support end. In optional embodiments, robot plantar pressure, torque, or contact sensor information can also be fused as auxiliary constraints. In one embodiment, the contact-aware retargeting optimization objective can be expressed as: a weighted combination of the original retargeting objective and plantar height consistency constraints, plantar tangential slip suppression constraints, plantar posture consistency constraints, and regularization terms. By combining these objectives, the optimizer aims to keep the robot's movements close to human movements while ensuring that the robot's support end is close to the ground within the contact area, reducing slippage and maintaining a reasonable posture.
[0037] It should be noted that while frame-by-frame contact constraints can improve the foot position and posture in a single frame, problems such as slow drift within the contact interval, unsmooth root node compensation, and local joint jitter may still exist. Therefore, this embodiment can further perform contact interval-level temporal repair after redirection optimization. The temporal repair can include at least one of the following: a maintenance term, a support end stabilization term, a first-order smoothing term, a second-order smoothing term, a root node compensation smoothing term, and a contact boundary soft switching term. This step can further improve the continuity and contact stability of the robot's motion sequence. The final output is a contact-aware robot target motion sequence. This robot target motion sequence can be used for simulation tracking, reinforcement learning training, imitation learning training, motion library construction, or verification before actual deployment. Specifically, temporal repair of the contact-aware robot motion sequence includes at least one of the following: setting a preservation term to make the repaired motion sequence approximate the contact-aware robot motion sequence; setting a support end stabilization term within the contact interval to suppress support end drift; setting a first-order smoothing term to suppress excessive state changes between adjacent frames; setting a second-order smoothing term to suppress motion jitter or abrupt jumps; and setting a soft switching term at the contact boundary to ensure smooth changes in contact constraints near the start or end of contact. It is understood that temporal repair is not limited to offline optimization and can also be implemented using filtering, spline smoothing, model predictive control, or learning-based post-processing networks. In this embodiment, by mapping the human-side contact interval to the robot-side support end contact constraint through the robot's kinematic model and the support end semantic mapping relationship, and incorporating the contact constraint into GMR or other retargeting optimization objectives, the robot's foot slippage and ground penetration can be directly reduced during the retargeting process. Without relying on the robot's foot contact sensor, the human-side contact prediction result is used as the source of the robot-side contact constraint, making it suitable for robot platforms without foot contact sensors or where contact sensors are unavailable. The robot platform is not limited to a specific model and can be adapted to different robot foot structures through the support end semantic mapping relationship. Contact interval-level temporal repair is performed on the retargeted robot motion sequence. Compared to methods that filter out large amounts of bad motion data, this method can selectively repair unreliable parts of foot-to-ground contact while preserving the original motion style and rhythm, improving support stability and motion continuity. The robot motion sequence processed by this application is more suitable as a reference motion for reinforcement learning, imitation learning, simulation tracking, and pre-deployment verification.
[0038] As can be seen from the above, the embodiments of this application first use a preset human motion reconstruction algorithm to analyze the target human motion video and output human motion data. Then, the human model forward calculation is used to obtain the vertex coordinates and skeletal joint coordinates of the human three-dimensional mesh corresponding to each video frame. Based on the human three-dimensional coordinates, representative points of the human foot such as toes, heels, and soles are extracted. The motion characteristics of the human foot representative points and the estimated ground reference height are combined to generate foot-to-ground contact pseudo-labels. The foot-to-ground contact pseudo-labels are used as supervision signals to train the temporal contact prediction model. The foot-to-ground contact probability of each frame is output by the model to determine the human foot-to-ground contact frame interval. Then, the human foot-to-ground contact frame interval is transformed into frame-by-frame foot contact constraints of the robot support end. The constraints are integrated into the robot motion redirection optimization calculation, and finally, the target motion sequence adapted to the robot execution is output. In this way, through the above-described process of the embodiments of this application, by using a human motion reconstruction algorithm combined with forward computation of a human model to obtain the human body's three-dimensional mesh and skeletal coordinates, it is possible to reconstruct the human body's three-dimensional motion information based solely on video footage. This eliminates the need for specialized motion capture equipment, reducing the hardware cost of motion acquisition. Representative points of the toes, heels, and soles are extracted from the three-dimensional coordinates, and foot-ground contact pseudo-labels are automatically generated by combining the point motion characteristics with ground reference height. This eliminates the need for manual frame-by-frame annotation of foot contact states, autonomously generating training and supervision data and solving the problems of high cost and low annotation efficiency in obtaining foot-ground contact labels. Furthermore, a temporal contact prediction model is used to complete training with pseudo-labels and predict frame-by-frame contact probabilities. Simultaneously, by dividing continuous contact frame intervals, it can filter out single-frame prediction misjudgments, smooth temporal contact states, and output stable and consistent data. During the foot support phase, the human-side contact frame interval is mapped to a frame-by-frame contact constraint specific to the robot's support end, and this constraint is embedded in the redirection optimization process for synchronous solution. During the motion mapping stage, the robot's foot height off the ground, horizontal sliding amplitude, and foot posture are simultaneously constrained, suppressing defects such as robot feet slipping through the ground, hovering, sliding, and incorrect support logic. At the same time, effective contact constraints can be generated without relying on robot foot contact sensors, making it adaptable to various legged robots with different foot structures and without foot sensors. It retains the motion rhythm and style of the original human video, and the optimized robot motion sequence support logic conforms to the laws of real human motion. It can be directly used for simulation verification, robot imitation learning, and physical robot motion deployment, thereby optimizing the robot redirection method to reduce problems such as foot sliding, foot slipping through the ground, and incorrect support state in robot motion.
[0039] Accordingly, see Figure 5 As shown, this application embodiment also provides a robot repositioning device based on the human foot-ground contact relationship, including: The forward calculation module 11 is used to perform motion reconstruction on the target human motion video using a preset human motion reconstruction algorithm to obtain corresponding human motion data, and to perform forward calculation on the human model based on the human motion data to obtain the human three-dimensional coordinates corresponding to each video frame in the target human motion video; the human three-dimensional coordinates include the human three-dimensional mesh vertex coordinates and the human three-dimensional skeletal joint coordinates. The tag generation module 12 is used to determine the representative point of the human foot in each video frame based on the three-dimensional coordinates of the human body, estimate the ground reference height in the video frame, and generate corresponding foot-ground contact pseudo-tags for the representative point of the human foot according to the motion characteristics of the representative point of the human foot and the ground reference height; the representative point of the human foot is a feature point of the human foot, including at least one of the toe, heel and sole representative points; The probability prediction module 13 is used to train a preset temporal contact prediction model with the foot-ground contact pseudo-label as a supervision signal, and to predict the probability of human foot-ground contact in each video frame through the trained temporal contact prediction model, so as to determine the human foot-ground contact frame interval of the target human motion video based on the obtained prediction results. The constraint introduction module 14 is used to convert the human foot-to-ground contact frame interval into frame-by-frame foot-to-ground contact constraints of the support end of the target robot, and introduce the frame-by-frame foot-to-ground contact constraints into the retargeting optimization process of the target robot to obtain the target action sequence of the target robot.
[0040] In some specific embodiments, the forward computation module 11 may specifically include: The posture completion unit is used to complete the missing posture parameter dimensions of the human motion data if the posture parameter dimension of the human motion data is inconsistent with the target dimension required for the forward calculation of the human model, so as to obtain the completed human motion data. The forward calculation unit is used to perform forward calculations on the human body model based on the supplemented human motion data.
[0041] In some specific embodiments, the label generation module 12 may specifically include: The set determination unit is used to determine the set of vertex coordinates of the front end of the foot, the set of vertex coordinates of the rear end of the foot, and the set of vertex coordinates of the lowest region of the foot from the vertex coordinates of the human body three-dimensional mesh. The set of vertex coordinates of the front end of the foot is the set of vertex coordinates in the human body foot mesh region where the axial coordinate data of the foot along the direction of human movement is greater than a first preset coordinate threshold. The set of vertex coordinates of the rear end of the foot is the set of vertex coordinates in the human body foot mesh region where the axial coordinate data of the foot along the direction of human movement is less than a second preset coordinate threshold, where the first preset coordinate threshold is greater than the second preset coordinate threshold. The set of vertex coordinates of the lowest region of the foot is the set of vertex coordinates in the human body foot mesh region where the axial coordinate data of the vertical height is less than a third preset coordinate threshold. The representative point determination unit is used to determine the average coordinate positions of the set of coordinates of the front end vertex of the foot, the set of coordinates of the rear end vertex of the foot, and the set of coordinates of the lowest region vertex of the foot as the representative point of the toe, the representative point of the heel, and the representative point of the sole of the foot, respectively.
[0042] In some specific embodiments, the label generation module 12 may specifically include: The distribution statistics unit is used to statistically analyze the temporal height distribution of the representative points of the human foot in the target human motion video; A height estimation unit is used to estimate the ground reference height in the video frame based on the temporal height distribution.
[0043] In some specific embodiments, the label generation module 12 may specifically include: The state determination unit is used to determine that the human foot representative point is in the target foot-ground contact state in the adjacent video frames if the human foot representative point continuously meets the preset foot-ground contact condition in the adjacent video frames. The pseudo-label generation unit is used to generate foot-ground contact pseudo-labels representing the presence of foot-ground contact for the human foot representative points corresponding to the adjacent several video frames, and to generate foot-ground contact pseudo-labels representing the absence of foot-ground contact for the human foot representative points corresponding to other video frames; the other video frames are other video frames in the target human motion video besides the adjacent several video frames. The preset foot-ground contact conditions include the difference between the vertical height of the representative point of the human foot and the reference height of the ground being less than a preset height threshold, the velocity of the representative point of the human foot in the vertical direction being less than a preset vertical velocity threshold, and the velocity of the representative point of the human foot in the horizontal direction being less than a preset horizontal velocity threshold.
[0044] In some specific implementations, the prediction results include the probability of toe contact and the probability of heel contact for each video frame; Accordingly, the probability prediction module 13 may specifically include: A temporal smoothing unit is used to perform frame-by-frame temporal smoothing on the toe contact probability and the heel contact probability respectively, so as to obtain the corresponding smoothed toe contact probability and smoothed heel contact probability. The probability fusion unit is used to fuse the smooth toe contact probability and the smooth heel contact probability of the same foot in the same video frame to obtain the corresponding foot support confidence. The condition judgment unit is used to determine whether each video frame meets the preset support conditions based on the foot support confidence level, so as to obtain the corresponding judgment result; The frame merging unit is used to merge adjacent video frames that continuously meet the preset support conditions into a human foot-ground contact frame interval based on the judgment result.
[0045] In some specific embodiments, the constraint introduction module 14 may specifically include: The interval conversion unit is used to convert the human foot-ground contact frame interval into the frame-by-frame foot-ground contact constraint of the target robot according to the robot kinematic model of the target robot and the pre-established support end semantic mapping relationship. A constraint introduction unit is used to introduce the frame-by-frame foot contact constraint into the retargeting optimization process of the target robot to obtain the initial action sequence of the target robot. The smoothing optimization unit is used to perform temporal smoothing optimization on the initial action sequence to obtain the optimized target action sequence.
[0046] Furthermore, embodiments of this application also disclose an electronic device, Figure 6 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the robot repositioning method based on human foot-to-ground contact relationship disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0047] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0048] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0049] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the robot repositioning method based on human foot-to-ground contact relationship executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0050] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned robot repositioning method based on human foot-to-ground contact relationship. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0051] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0052] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0053] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0054] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0055] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A robot repositioning method based on human foot-ground contact relationship, characterized in that, include: The target human motion video is reconstructed using a preset human motion reconstruction algorithm to obtain corresponding human motion data. Based on the human motion data, a forward calculation of the human model is performed to obtain the three-dimensional coordinates of the human body corresponding to each video frame in the target human motion video. The three-dimensional coordinates of the human body include the coordinates of the vertex of the three-dimensional human body mesh and the coordinates of the three-dimensional human body skeletal joints. Based on the three-dimensional coordinates of the human body, the representative point of the human foot in each video frame is determined, and the ground reference height in the video frame is estimated. Based on the motion characteristics of the representative point of the human foot and the ground reference height, a corresponding foot-ground contact pseudo-label is generated for the representative point of the human foot. The representative point of the human foot is a characteristic point of the human foot, including at least one of the representative points of the toe, heel and sole. A preset temporal contact prediction model is trained using the foot-ground contact pseudo-label as a supervision signal, and the probability of human foot-ground contact in each video frame is predicted by the trained temporal contact prediction model, so as to determine the human foot-ground contact frame interval of the target human motion video based on the obtained prediction results. The human foot-ground contact frame interval is converted into a frame-by-frame foot-ground contact constraint of the support end of the target robot, and the frame-by-frame foot-ground contact constraint is introduced into the redirection optimization process of the target robot to obtain the target action sequence of the target robot. The step of generating corresponding foot-ground contact pseudo-labels for the human foot representative points based on the motion characteristics of the human foot representative points and the ground reference height includes: If the human foot representative point continuously meets the preset foot-ground contact condition in several adjacent video frames, then it is determined that the human foot representative point is in the target foot-ground contact state in the several adjacent video frames. For the representative points of the human foot corresponding to the adjacent video frames, generate pseudo-labels of foot-ground contact indicating that the foot is in contact with the ground; and generate pseudo-labels of foot-ground contact indicating that the foot is not in contact with the ground for the representative points of the human foot corresponding to other video frames; the other video frames are the video frames other than the adjacent video frames in the target human motion video. The preset foot-ground contact conditions include the difference between the vertical height of the representative point of the human foot and the reference height of the ground being less than a preset height threshold, the velocity of the representative point of the human foot in the vertical direction being less than a preset vertical velocity threshold, and the velocity of the representative point of the human foot in the horizontal direction being less than a preset horizontal velocity threshold.
2. The robot repositioning method based on human foot-ground contact relationship according to claim 1, characterized in that, The forward calculation of the human model based on the human motion data includes: If the posture parameter dimension of the human motion data is inconsistent with the target dimension required for the forward calculation of the human model, then the missing posture parameter dimension of the human motion data is filled in to obtain the filled human motion data. The human body model is forward-calculated based on the supplemented human motion data.
3. The robot repositioning method based on human foot-ground contact relationship according to claim 1, characterized in that, The step of determining the representative point of the human foot in each video frame based on the three-dimensional coordinates of the human body includes: From the vertex coordinates of the human body 3D mesh, the vertex coordinate set of the front end of the foot, the vertex coordinate set of the rear end of the foot, and the vertex coordinate set of the lowest region of the foot are determined; the vertex coordinate set of the front end of the foot is the set of vertex coordinates in the human foot mesh region where the axial coordinate data of the foot along the forward direction of the human body is greater than a first preset coordinate threshold; the vertex coordinate set of the rear end of the foot is the set of vertex coordinates in the human foot mesh region where the axial coordinate data of the foot along the forward direction of the human body is less than a second preset coordinate threshold, where the first preset coordinate threshold is greater than the second preset coordinate threshold; the vertex coordinate set of the lowest region of the foot is the set of vertex coordinates in the human foot mesh region where the vertical height axial coordinate data is less than a third preset coordinate threshold. The average coordinate positions of the set of coordinates of the front end vertex of the foot, the set of coordinates of the rear end vertex of the foot, and the set of coordinates of the lowest region vertex of the foot are respectively determined as the toe point, the heel point, and the sole point.
4. The robot repositioning method based on human foot-ground contact relationship according to claim 1, characterized in that, The estimation of the ground reference height in the video frame includes: Statistical analysis of the temporal height distribution of the representative points of the human foot in the target human motion video; The ground reference height in the video frame is estimated based on the temporal height distribution.
5. The robot repositioning method based on human foot-ground contact relationship according to claim 1, characterized in that, The prediction results include the probability of toe contact and the probability of heel contact for each video frame; Accordingly, determining the human foot-ground contact frame interval of the target human motion video based on the obtained prediction results includes: The toe contact probability and the heel contact probability are smoothed frame by frame to obtain the corresponding smoothed toe contact probability and smoothed heel contact probability. The smooth toe contact probability and the smooth heel contact probability of the same foot in the same video frame are fused to obtain the corresponding foot support confidence. Based on the foot support confidence level, determine whether each video frame meets the preset support conditions to obtain the corresponding judgment result; Based on the judgment result, adjacent video frames that continuously meet the preset support conditions are merged into a human foot-ground contact frame interval.
6. The robot repositioning method based on human foot-ground contact relationship according to any one of claims 1 to 5, characterized in that, The process of converting the human foot-to-ground contact frame interval into frame-by-frame foot-to-ground contact constraints for the support end of the target robot, and incorporating these frame-by-frame foot-to-ground contact constraints into the retargeting optimization process of the target robot to obtain the target robot's target action sequence, includes: Based on the robot kinematics model of the target robot and the pre-established support end semantic mapping relationship, the human foot-ground contact frame interval is converted into the frame-by-frame foot-ground contact constraint of the target robot. The frame-by-frame foot contact constraint is introduced into the redirection optimization process of the target robot to obtain the initial action sequence of the target robot; The initial action sequence is subjected to temporal smoothing optimization to obtain the optimized target action sequence.
7. A robot repositioning device based on the human foot-ground contact relationship, characterized in that, include: The forward calculation module is used to perform motion reconstruction on the target human motion video using a preset human motion reconstruction algorithm to obtain the corresponding human motion data, and to perform forward calculation on the human model based on the human motion data to obtain the human three-dimensional coordinates corresponding to each video frame in the target human motion video; the human three-dimensional coordinates include the human three-dimensional mesh vertex coordinates and the human three-dimensional skeletal joint coordinates. The tag generation module is used to determine the representative point of the human foot in each video frame based on the three-dimensional coordinates of the human body, estimate the ground reference height in the video frame, and generate corresponding foot-ground contact pseudo-tags for the representative point of the human foot according to the motion characteristics of the representative point of the human foot and the ground reference height; the representative point of the human foot is a characteristic point of the human foot, including at least one of the representative points of the toe, heel and sole. The probability prediction module is used to train a preset temporal contact prediction model with the foot-ground contact pseudo-label as a supervision signal, and to predict the probability of human foot-ground contact in each video frame through the trained temporal contact prediction model, so as to determine the human foot-ground contact frame interval of the target human motion video based on the obtained prediction results. The constraint introduction module is used to convert the human foot-ground contact frame interval into frame-by-frame foot-ground contact constraints of the support end of the target robot, and introduce the frame-by-frame foot-ground contact constraints into the retargeting optimization process of the target robot to obtain the target action sequence of the target robot. The tag generation module includes: The state determination unit is used to determine that the human foot representative point is in the target foot-ground contact state in the adjacent video frames if the human foot representative point continuously meets the preset foot-ground contact condition in the adjacent video frames. The pseudo-label generation unit is used to generate foot-ground contact pseudo-labels representing the presence of foot-ground contact for the human foot representative points corresponding to the adjacent several video frames, and to generate foot-ground contact pseudo-labels representing the absence of foot-ground contact for the human foot representative points corresponding to other video frames; the other video frames are other video frames in the target human motion video besides the adjacent several video frames. The preset foot-ground contact conditions include the difference between the vertical height of the representative point of the human foot and the reference height of the ground being less than a preset height threshold, the velocity of the representative point of the human foot in the vertical direction being less than a preset vertical velocity threshold, and the velocity of the representative point of the human foot in the horizontal direction being less than a preset horizontal velocity threshold.
8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the robot repositioning method based on human foot-ground contact relationship as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when the computer programs are executed by a processor, they implement the robot repositioning method based on human foot-to-ground contact relationship as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Humanoid robot anthropomorphic walking method fusing human motion redirection and model prediction control
CN121696985A
Motion model refinement based on contact analysis and optimization
US20210335028A1