Method and system for generating intelligent motion of body and electronic equipment
By optimizing and mapping the sequence of key points of the human skeleton, a natural and coherent sequence of virtual human motion images is generated and mapped onto the robot, solving the problem of disjointed movements in existing technologies and realizing the natural motion output of the embodied intelligent agent.
Patent Information
- Application Number
- CN202511379663.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-01-13
AI Technical Summary
In existing embodied intelligent motion generation methods, the movements often violate kinematic constraints, resulting in disjointed and unnatural generated movements. Furthermore, due to discrepancies between the human body in the video and the virtual intelligent agent or robot system, the movements are disjointed or jump abruptly.
By acquiring human motion video data, target motion information and initial human skeleton key point sequence are obtained based on motion capture data. The skeleton key point sequence is optimized using human structure priors, source skeleton motion data is obtained, and it is mapped and aligned to the virtual human skeleton structure to generate optical flow map and mask map. Finally, a natural and coherent motion image sequence is generated on the virtual human image and mapped onto the robot.
It enables the embodied intelligent agent to output natural and coherent actions, reduces the cost of the action redirection process, and avoids action distortion.
Smart Images

Figure CN121328610A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and graphics technology, and in particular to an embodied intelligent motion generation method, system, electronic device, and computer-readable storage medium thereof. Background Technology
[0002] The Innate Method, a motion generation method for bio-intelligent systems, aims to generate motion by mapping human motion patterns from reference videos onto virtual intelligent agents or robots through motion redirection.
[0003] However, existing methods for generating embodied intelligent motion often result in movements that violate kinematic constraints, leading to disjointed and unnatural motion. Furthermore, due to discrepancies between the human body in the video and the virtual intelligent robot, the generated motion may be disjointed or exhibit sudden jumps.
[0004] Therefore, how to achieve coherent and natural motion generation has become a technical problem that urgently needs to be solved by existing technologies. Summary of the Invention
[0005] This invention provides a method, system, electronic device, and computer-readable storage medium for generating embodied intelligent motion, thereby enabling coherent and natural motion generation.
[0006] According to a first aspect of the present invention, the present invention provides a method for generating embodied intelligent motion, the method comprising: Acquire video data of human movement; Motion capture data is obtained based on the video data; Based on the motion capture data, acquire target motion information and initial human skeleton key point sequence; Based on prior knowledge of human structure, the sequence of key points of the human skeleton is optimized to obtain source skeleton motion data; Based on the target motion information, a virtual human image is obtained through a large visual language model; Obtaining the skeleton structure of a virtual human based on virtual human images; The source skeleton motion data is mapped and aligned to the virtual human skeleton structure to obtain the redirected skeleton data; Optical flow map and mask map are obtained based on the virtual human image and the redirected skeleton data; Based on the optical flow map and the mask map, an initial sequence of virtual human motion images is generated on the basis of the virtual human image; A virtual human motion image sequence is obtained based on the initial virtual human motion image sequence; The motion image sequence of the virtual human is mapped onto the robot to obtain the action output of the biological intelligent agent.
[0007] Optionally, the key point information of the human skeleton is optimized based on prior knowledge of human structure to obtain source skeleton motion data, including: The human skeleton constraint conditions are obtained based on the prior knowledge of the human body structure. Based on the aforementioned human skeleton reservation conditions, the key point information of the human skeleton is corrected to obtain the source skeleton motion data.
[0008] Optionally, the method for mapping and aligning the source skeleton motion data to the virtual human skeleton structure to obtain the redirected skeleton data includes: Based on the source skeleton motion data and the virtual human skeleton structure, obtain the affine transformation function of key points; The driving actions represented by the source skeleton motion data are aligned to the virtual human skeleton structure based on the key point affine transformation function to obtain the redirected skeleton data.
[0009] Optionally, the method for obtaining the optical flow map and mask map based on the virtual human image and the redirected skeleton data includes: The virtual human image is input into an image encoder to obtain a virtual human feature sequence; The virtual human feature sequence and the redirected skeleton data are input into the motion generator to obtain the optical flow map and the mask map.
[0010] Optionally, a method for generating a sequence of virtual human motion images based on the optical flow map and the mask map, on the basis of the virtual human image, includes: The optical flow map and the mask map are input into an image synthesizer to obtain a synthesized motion map; The synthesized motion graph is combined with the virtual human image to obtain the virtual human motion image sequence.
[0011] Optionally, obtaining a virtual human motion image sequence based on the initial virtual human motion image sequence includes: Obtain the original image sequence based on the video data of the human movement; Calculate the reconstruction loss between the initial virtual human motion image sequence and the original image sequence; Based on the reconstruction loss and the initial virtual human motion image sequence, a virtual human motion image sequence is obtained.
[0012] Optionally, the method for obtaining an initial sequence of key points of the human skeleton based on the motion capture data includes: The motion capture data is input into a keypoint detector to extract the initial human skeleton keypoint sequence.
[0013] According to a second aspect of the present invention, the present invention provides an embodied intelligent motion generation system for implementing the above-described embodied intelligent motion generation method, the system comprising: The video data acquisition module is used to acquire video data of human movement. A motion capture data acquisition module is used to acquire motion capture data based on the video data; The motion capture data parsing module is used to obtain target motion information and an initial sequence of key points of the human skeleton based on the motion capture data. The skeleton motion data optimization and processing module is used to optimize the sequence of key points of the human skeleton based on prior knowledge of human structure to obtain source skeleton motion data. The virtual human image acquisition module is used to acquire a virtual human image based on the target motion information and through a large visual language model. The virtual human skeleton structure acquisition module is used to acquire the virtual human skeleton structure based on virtual human images. A redirection module is used to map and align the source skeleton motion data to the virtual human skeleton structure to obtain redirected skeleton data. An optical flow map and mask map generation module is used to obtain an optical flow map and a mask map based on the virtual human image and the redirected skeleton data; The initial virtual human motion image sequence acquisition module generates an initial virtual human motion image sequence based on the optical flow map and the mask map, on the basis of the virtual human image. A virtual human motion image sequence acquisition module is used to acquire a virtual human motion image sequence based on the initial virtual human motion image sequence; The motion mapping module is used to map the sequence of motion images of the virtual human onto the robot to obtain the motion output of the biological intelligent agent.
[0014] According to a third aspect of the present invention, an electronic device is provided for implementing the above-described embodied intelligent motion generation method, the system comprising: According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the above-described embodied intelligent motion generation method.
[0015] Compared with the prior art, the technical solution of the present invention has the following beneficial effects: In the embodied intelligent motion generation method, system, electronic device, and readable storage medium of this invention, the source skeleton motion data is obtained by optimizing the sequence of key points of the human skeleton based on prior knowledge of human structure. Therefore, the obtained source skeleton motion data is more consistent with conventional kinematics. Since this invention maps and aligns the source skeleton motion data to the virtual human skeleton structure to obtain redirected skeleton data, the redirected skeleton data obtained through motion redirection can fit the body shape of the virtual human. Based on the virtual human image and the redirected skeleton data, optical flow maps and mask maps are obtained. Based on the optical flow maps and mask maps, an initial virtual human motion image sequence is generated on the virtual human image. A virtual human motion image sequence is obtained based on the initial virtual human motion image sequence. The virtual human motion image sequence is mapped to the robot to obtain the action output of the embodied intelligent agent. Therefore, the action output of the embodied intelligent agent is more natural and coherent. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the bio-intelligent motion generation method in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of the bio-intelligent motion generation system in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the electronic device in an embodiment of the present invention. Detailed Implementation
[0018] As mentioned in the background, how to achieve coherent and natural motion generation has become a technical problem that urgently needs to be solved by existing technologies.
[0019] In view of this, the present invention provides an embodied intelligent motion generation method, which involves: acquiring video data of human motion; acquiring motion capture data based on the video data; acquiring target motion information and an initial sequence of key points of the human skeleton based on the motion capture data; optimizing the sequence of key points of the human skeleton based on prior knowledge of human structure to obtain source skeleton motion data; acquiring a virtual human image based on the target motion information using a large visual language model; acquiring a virtual human skeleton structure based on the virtual human image; mapping and aligning the source skeleton motion data to the virtual human skeleton structure to obtain redirected skeleton data; obtaining an optical flow map and a mask map based on the virtual human image and the redirected skeleton data; generating an initial sequence of virtual human motion images based on the optical flow map and the mask map; acquiring a virtual human motion image sequence based on the initial sequence of virtual human motion images; and mapping the virtual human motion image sequence to a robot to obtain the motion output of the embodied intelligent agent. This achieves coherent and natural motion generation. The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0020] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0021] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0022] To make the above-mentioned objectives, features and beneficial effects of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Please refer to Figure 1This invention provides a method for generating embodied intelligent motion, the method comprising: S1: Acquire video data of human movement; S2: Obtain motion capture data based on the video data; S3: Based on the motion capture data, obtain the target motion information and the initial human skeleton key point sequence; S4: Based on prior knowledge of human structure, optimize the sequence of key points of the human skeleton to obtain source skeleton motion data; S5: Based on the target motion information, obtain a virtual human image through a large visual language model; S6: Obtain the virtual human skeleton structure based on virtual human images; S7: Map and align the source skeleton motion data to the virtual human skeleton structure to obtain the redirected skeleton data; S8: Obtain an optical flow map and a mask map based on the virtual human image and the redirected skeleton data; S9: Based on the optical flow map and the mask map, generate an initial virtual human motion image sequence based on the virtual human image; S10: Obtain a virtual human motion image sequence based on the initial virtual human motion image sequence; S11: Map the virtual human motion image sequence to the robot to obtain the action output of the biological intelligent agent.
[0024] As an example, video data of human motion includes video data of a human body performing any one or a combination of two or more of the following movements: walking, running, jumping, standing, dribbling, and launching.
[0025] As an example, motion capture data refers to the human motion trajectory and posture information collected by motion capture devices (such as optical tracking systems, inertial sensors, or depth cameras).
[0026] As an example, step S2, the method for obtaining an initial sequence of key points of the human skeleton based on the motion capture data, includes: The motion capture data is input into a keypoint detector to extract the initial human skeleton keypoint sequence.
[0027] In this embodiment, because the key point sequence of the human skeleton is optimized based on prior knowledge of human structure to obtain source skeleton motion data, the obtained source skeleton motion data is more consistent with conventional kinematics. Since this invention maps and aligns the source skeleton motion data to the virtual human skeleton structure to obtain redirected skeleton data, the redirected skeleton data obtained through action redirection can fit the virtual human's body shape. Based on the virtual human image and the redirected skeleton data, optical flow maps and mask maps are obtained. Based on the optical flow maps and mask maps, an initial virtual human motion image sequence is generated on the virtual human image. A virtual human motion image sequence is obtained based on the initial virtual human motion image sequence. The virtual human motion image sequence is mapped to the robot to obtain the action output of the bionic agent. Therefore, the action output of the bionic agent is more natural and coherent.
[0028] As a specific implementation method, in step S4, the key point information of the human skeleton is optimized based on the prior knowledge of human structure to obtain source skeleton motion data.
[0029] S41: Obtain human skeleton constraints based on the prior knowledge of the human body structure; S42: Based on the human skeleton reservation conditions, correct the initial human skeleton key point sequence to obtain source skeleton motion data.
[0030] The mechanism for optimizing and constraining the initial human skeleton key point sequence has the core function of ensuring that the generated movements conform to the natural movement laws of the human body and avoiding abnormal postures or physical inconsistencies.
[0031] Specifically, step S42, the method for correcting the initial human skeleton key point sequence based on the human skeleton reservation conditions to obtain source skeleton motion data, includes: S421: Initialize the initial human skeleton key point sequence to obtain the initialized human skeleton key point sequence; S422: Calculate the skeletal length of a reference person in human motion video data; S423: Construct a loss function based on the bone length of a reference human in human motion video data; S424: Iteratively update the initial human skeleton keypoint sequence based on the loss function to obtain the source skeleton motion data.
[0032] The method for initializing the initial human skeleton key point sequence can be the low-pass filtering initial value method.
[0033] A method for calculating the skeletal length of a reference person in human motion video data can be used to calculate the skeletal length of the reference person in the human motion video data by calculating the average frame rate in the human motion video data.
[0034] The method for iteratively updating the initialized human skeleton keypoint sequence based on the loss function to obtain the source skeleton motion data includes: using gradient descent, L-BFGS, or quadratic programming to iteratively update the initialized human skeleton keypoint sequence.
[0035] As an example, the aforementioned human skeleton constraints can refer to predefined models or rules regarding human skeletal structure, joint range of motion, and movement constraints, derived from prior human anatomy. The prior human anatomy refers to known laws of human anatomy, kinematics, and biomechanics. For example, human skeleton constraints might include: the spine being composed of 24 vertebrae; the limbs being connected to the trunk via joints; and joint types (such as ball-and-socket joints and hinge joints), ensuring the biological rationality of the skeleton model. Limiting the maximum range of motion of each joint prevents the generation of movements that violate physiological limits. For instance, the maximum bending angle of a human elbow joint is 150 degrees. If the initial human skeleton keypoint sequence shows an elbow bending angle of 180 degrees, the system will correct the elbow bending angle to 150 degrees.
[0036] As an example, step S7, the method for mapping and aligning the source skeleton motion data to the virtual human skeleton structure to obtain the redirected skeleton data includes: S71: Obtain the affine transformation function of key points based on the source skeleton motion data and the virtual human skeleton structure; S72: Align the driving action represented by the source skeleton motion data to the virtual human skeleton structure based on the key point affine transformation function to obtain the redirected skeleton data.
[0037] This embodiment aligns the source skeleton motion data and the virtual human skeleton structure through a redirection mechanism, eliminating the need to recapture motion data for each virtual human image, thus reducing costs and avoiding motion distortion.
[0038] As a specific embodiment, step S8, the method for obtaining the optical flow map and mask map based on the virtual human image and the redirected skeleton data, includes: S81: Input the virtual human image into the image encoder to obtain the virtual human feature sequence; S82: Input the virtual human feature sequence and the redirected skeleton data into the motion generator to obtain the optical flow map and the mask map.
[0039] As a specific implementation method, step S9, the method for generating a sequence of virtual human motion images based on the virtual human image, according to the optical flow map and the mask map, includes: S91: Input the optical flow map and the mask map into the image synthesizer to obtain a synthesized motion map.
[0040] S92: Combine the synthesized motion graph with the virtual human image to obtain the virtual human motion image sequence.
[0041] As an example, the optical flow graph is a graph that represents the motion information of an object in an image, and the optical flow graph describes the motion vector of each pixel from the first frame image to the second frame image.
[0042] As an example, the mask image is a map used to identify specific regions in an image. A mask image is typically a binary map or label map of the same size as the original image, where each pixel value indicates whether the pixel belongs to a certain category or region of interest.
[0043] Specifically, step S10, obtaining the virtual human motion image sequence based on the initial virtual human motion image sequence, includes: S101: Obtain the original image sequence based on the video data of the human body movement; S102: Calculate the reconstruction loss between the initial virtual human motion image sequence and the original image sequence; S103: Based on the reconstruction loss and the initial virtual human motion image sequence, obtain the virtual human motion image sequence.
[0044] Please refer to Figure 2 The present invention provides an embodied intelligent motion generation system for realizing the above-mentioned embodied intelligent motion generation, the system comprising: The video data acquisition module 100 is used to acquire video data of human motion. Motion capture data acquisition module 200 is used to acquire motion capture data based on the video data; The motion capture data parsing module 300 is used to obtain target motion information and an initial sequence of key points of the human skeleton based on the motion capture data. The skeleton motion data optimization and processing module 400 is used to optimize the sequence of key points of the human skeleton based on prior knowledge of human structure to obtain source skeleton motion data. The virtual human image acquisition module 500 is used to acquire a virtual human image based on the target motion information and through a large visual language model; Virtual human skeleton structure acquisition module 600 is used to acquire virtual human skeleton structure based on virtual human image; A redirection module is used to map and align the source skeleton motion data to the virtual human skeleton structure to obtain redirected skeleton data. The optical flow map and mask map generation module 700 is used to obtain an optical flow map and a mask map based on the virtual human image and the redirected skeleton data; The initial virtual human motion image sequence acquisition module 800 generates an initial virtual human motion image sequence based on the optical flow map and the mask map, on the basis of the virtual human image. The virtual human motion image sequence acquisition module 900 is used to acquire a virtual human motion image sequence based on the initial virtual human motion image sequence. The motion mapping module 1100 is used to map the virtual human motion image sequence to the robot to obtain the motion output of the biological intelligent agent.
[0045] According to a third aspect of the invention, please refer to Figure 3 The present invention also provides an electronic device, which includes a memory 2100, a processor 2200 and a computer program 2110 stored in the memory 2100, for running the above-described embodied intelligent motion generation method.
[0046] According to a fourth aspect of the present invention, embodiments of the present invention also provide a computer-readable storage medium for storing a computer program that, when executed by a processor, implements the above-described embodied intelligent motion generation method.
[0047] The computer storage medium may include various media that can store computer programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0048] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.
Claims
1. A method for generating embodied intelligent motion, characterized in that, The method includes: Acquire video data of human movement; Motion capture data is obtained based on the video data; Based on the motion capture data, acquire target motion information and initial human skeleton key point sequence; Based on prior knowledge of human structure, the sequence of key points of the human skeleton is optimized to obtain source skeleton motion data; Based on the target motion information, a virtual human image is obtained through a large visual language model; Obtaining the skeleton structure of a virtual human based on virtual human images; The source skeleton motion data is mapped and aligned to the virtual human skeleton structure to obtain the redirected skeleton data; Optical flow map and mask map are obtained based on the virtual human image and the redirected skeleton data; Based on the optical flow map and the mask map, an initial sequence of virtual human motion images is generated on the basis of the virtual human image; A virtual human motion image sequence is obtained based on the initial virtual human motion image sequence; The motion image sequence of the virtual human is mapped onto the robot to obtain the action output of the biological intelligent agent.
2. The embodied intelligent motion generation method as described in claim 1, characterized in that, Based on prior knowledge of human anatomy, the initial sequence of key points of the human skeleton is optimized to obtain source skeleton motion data, including: The human skeleton constraint conditions are obtained based on the prior knowledge of the human body structure. The initial human skeleton key point sequence is corrected based on the aforementioned human skeleton reservation conditions to obtain source skeleton motion data.
3. The embodied intelligent motion generation method as described in claim 1, characterized in that, The method for mapping and aligning the source skeleton motion data to the virtual human skeleton structure to obtain the redirected skeleton data includes: Based on the source skeleton motion data and the virtual human skeleton structure, obtain the affine transformation function of key points; The driving actions represented by the source skeleton motion data are aligned to the virtual human skeleton structure based on the key point affine transformation function to obtain the redirected skeleton data.
4. The embodied intelligent motion generation method as described in claim 1, characterized in that, The method for obtaining optical flow maps and mask maps based on the virtual human image and the redirected skeleton data includes: The virtual human image is input into an image encoder to obtain a virtual human feature sequence; The virtual human feature sequence and the redirected skeleton data are input into the motion generator to obtain the optical flow map and the mask map.
5. The embodied intelligent motion generation method as described in claim 1, characterized in that, A method for generating a sequence of motion images of a virtual human based on the optical flow map and the mask map includes: The optical flow map and the mask map are input into an image synthesizer to obtain a synthesized motion map; The synthesized motion graph is combined with the virtual human image to obtain the virtual human motion image sequence.
6. The embodied intelligent motion generation method as described in claim 1, characterized in that, Obtaining a virtual human motion image sequence based on the initial virtual human motion image sequence includes: Obtain the original image sequence based on the video data of the human movement; Calculate the reconstruction loss between the initial virtual human motion image sequence and the original image sequence; Based on the reconstruction loss and the initial virtual human motion image sequence, a virtual human motion image sequence is obtained.
7. The embodied intelligent motion generation method as described in claim 1, characterized in that, The method for obtaining the initial human skeleton key point sequence based on the motion capture data includes: The motion capture data is input into a keypoint detector to extract the initial human skeleton keypoint sequence.
8. A body-integrated intelligent motion generation system, characterized in that, The system includes: The video data acquisition module is used to acquire video data of human movement. A motion capture data acquisition module is used to acquire motion capture data based on the video data; The motion capture data parsing module is used to obtain target motion information and an initial sequence of key points of the human skeleton based on the motion capture data. The skeleton motion data optimization and processing module is used to optimize the sequence of key points of the human skeleton based on prior knowledge of human structure to obtain source skeleton motion data. The virtual human image acquisition module is used to acquire a virtual human image based on the target motion information and through a large visual language model. The virtual human skeleton structure acquisition module is used to acquire the virtual human skeleton structure based on virtual human images. A redirection module is used to map and align the source skeleton motion data to the virtual human skeleton structure to obtain redirected skeleton data. An optical flow map and mask map generation module is used to obtain an optical flow map and a mask map based on the virtual human image and the redirected skeleton data; The initial virtual human motion image sequence acquisition module generates an initial virtual human motion image sequence based on the optical flow map and the mask map, on the basis of the virtual human image. A virtual human motion image sequence acquisition module is used to acquire a virtual human motion image sequence based on the initial virtual human motion image sequence; The motion mapping module is used to map the sequence of motion images of the virtual human onto the robot to obtain the motion output of the biological intelligent agent.
9. An electronic device, the electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, When the processor executes the computer program, it implements the embodied intelligent motion generation method as described in any one of claims 1-7.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the embodied intelligent motion generation method as described in any one of claims 1-8.
Citation Information
Cited By
Operation demonstration video processing method, device and system, medium and program product
CN122223476A