Method and device for generating video with body and electronic equipment
By constructing a three-dimensional physical constraint graph and a spatiotemporal diffusion model, and combining semantic-spatial joint coding and a dual-path feedback mechanism, the problems of lack of physical consistency, error accumulation, and fragmented semantic understanding in embodied intelligence tasks are solved, and high-precision video generation and task execution are achieved.
Patent Information
- Application Number
- CN202511018552.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video generation technologies suffer from problems such as lack of physical consistency, error accumulation, and fragmented semantic understanding in embodied intelligence tasks. This leads to issues such as objects penetrating, floating, and motion paths violating physical constraints during interaction. Generation deviations accumulate along the time dimension, and the generated actions do not match the task intent.
A three-dimensional physical constraint map is constructed by parsing user task instructions and environmental observation data. Video sequences are generated by combining a spatiotemporal diffusion model, and spatial error indicators are calculated in real time to trigger the generation of correction frames. The physical constraint map is dynamically updated to respond to environmental changes, and semantic-spatial joint coding and dual-path feedback mechanisms work together.
It achieves high-precision video planning and generation in dynamic environments, solves the problems of physical inconsistency, error accumulation and semantic understanding fragmentation, ensures that the generated actions strictly follow physical laws, corrects errors in real time, and improves the robustness and accuracy of task execution.
Smart Images

Figure CN120856948A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the intersection of computer vision and embodied intelligence, and in particular to an embodied video generation method, apparatus, and electronic device. Background Technology
[0002] Current video generation technologies primarily rely on diffusion models combined with textual descriptions or initial frame conditions to generate video sequences (such as VideoAgent and 3D-VLA). However, in embodied intelligence tasks (such as robot grasping and virtual character interaction), these technologies suffer from the following fundamental drawbacks:
[0003] Lack of physical consistency: Current methods (such as inverse dynamics models) rely on RGB image features and lack explicit modeling of the scene's 3D structure. When objects interact, issues such as penetration, floating (e.g., a robotic arm "grabbing" in the air) and motion paths that violate physical constraints occur.
[0004] Error accumulation: Mainstream solutions adopt an open-loop generation architecture, which does not establish a real-time feedback link. The generation deviation accumulates along the time dimension, and dynamic changes in the environment cause the task execution to deviate from the target (such as the risk of collision caused by sudden obstacles).
[0005] Semantic understanding is fragmented: the visual generation and target understanding modules are separated, the generated actions do not match the task intent (e.g., the button is not pressed when it is pressed), and key object attributes change abruptly. Summary of the Invention
[0006] Therefore, it is necessary to provide an embodied video generation method, device, and electronic device to address the aforementioned problems of lack of physical consistency, error accumulation, and fragmentation of semantic understanding.
[0007] This invention provides a method for generating embodied videos, the method comprising:
[0008] The system parses user task instructions and initial environmental observation data to generate a sequence of key operation steps for the task objective and a set of associated objects, and constructs a three-dimensional physical constraint map of the scene. The physical constraint map includes object positions, sizes, distance thresholds, and feasible motion domains.
[0009] Based on the encoding conditions of the semantic vector and physical constraint graph of the task operation step sequence, an initial action video sequence is generated through a spatiotemporal diffusion model. The video sequence includes the pose of the executing subject and the state of key objects.
[0010] In response to the execution of a single frame action, the spatial error index is calculated based on the real-time environmental state. If the object contact distance deviation, trajectory collision probability, or physical constraint violation value exceeds the preset threshold, the spatiotemporal diffusion model is triggered to generate a correction frame based on the current physical constraint map.
[0011] The physical constraint graph is dynamically updated based on changes in environmental conditions, including the object position, distance threshold, and feasible motion domain. The updated constraint graph is then input into the spatiotemporal diffusion model.
[0012] In one embodiment, the triggering spatiotemporal diffusion model generates a correction frame based on the current physical constraint graph, including:
[0013] Execute small-loop feedback to generate semantic feedback text and spatial constraint information to input the spatiotemporal diffusion model.
[0014] In one embodiment, the execution of small-loop feedback to generate semantic feedback text and spatial constraint information input to the spatiotemporal diffusion model includes:
[0015] The semantic feedback text is encoded into a text embedding vector. The semantic feedback text is generated by the visual language model based on the task text and the current spatial information to form an action planning sequence, in the format of "action type [direction vector] [distance value]".
[0016] The spatial constraint information is converted into layout spatial encoding. The spatial constraint information is constructed by extracting the three-dimensional coordinates and relative distance between the crawler and the target object, including a key-value pair data structure of two-dimensional pixel coordinates, three-dimensional spatial coordinates and distance matrix.
[0017] The text embedding vector and layout space encoding are fused into the denoising process of the spatiotemporal diffusion model through a cross-attention mechanism. The spatiotemporal diffusion model adopts the UNet architecture, and the text embedding vector and layout space encoding are fused into the middle block layer of UNet through the cross-attention mechanism.
[0018] In one embodiment, the process of constructing the spatial constraint information includes:
[0019] Extract the two-dimensional pixel coordinates and three-dimensional spatial coordinates of the left and right grabbers from the current frame;
[0020] Calculate the Euclidean distance between the gripper and each target object and generate a distance matrix;
[0021] Convert the coordinate data to layout space encoding format.
[0022] In one embodiment, it further includes:
[0023] In response to the completion of key operational steps, verify the semantic matching degree between the generated results and the task objectives;
[0024] If a critical object attribute is missing or an operational logic error is detected, the operation step sequence and physical constraint diagram will be regenerated.
[0025] In one embodiment, the dynamic updating of the physical constraint graph includes:
[0026] Based on the real-time distance between the grabber and the target object, the distance threshold parameters for each task are dynamically adjusted according to the task type.
[0027] The coordinates of the feasible motion domain boundary of the target object are updated in real time based on the mask image and 3D point cloud data.
[0028] In one embodiment, the generation process of the spatiotemporal diffusion model includes:
[0029] Simultaneously output RGB image, depth map and binary mask image;
[0030] A weighted multimodal loss function is used, the mask image loss function uses binary cross-entropy, and the depth image and RGB loss functions use the L1 norm.
[0031] In one embodiment, the method further includes a training data generation process:
[0032] Depth maps, mask maps, RGB maps, and 3D point cloud data of the task trajectory are collected from the simulation environment;
[0033] The fine operation segment is dynamically divided based on the Euclidean distance between the grabber and the target object: when the distance is lower than the task type threshold for N consecutive frames, it is marked as a fine operation segment, where N is the preset number of consecutive frames.
[0034] Assign M times the sampling probability weight to the fine-grained operation segment, where M is a natural number;
[0035] Differentiated image cropping is performed based on task type: the door opening task uses center cropping, and the basketball placement task uses top right cropping.
[0036] The present invention also provides an embodied video generation apparatus, the apparatus comprising:
[0037] The task parsing and constraint graph construction module is used to parse user task instructions and initial environmental observation data, generate a sequence of key operation steps for the task objective and its associated object set, and construct a three-dimensional physical constraint graph of the scene. The physical constraint graph includes object position, size, distance threshold and feasible motion domain.
[0038] The video sequence generation module is used to generate an initial action video sequence based on the encoding conditions of the semantic vector and physical constraint graph of the task operation step sequence through a spatiotemporal diffusion model. The video sequence includes the pose of the executing subject and the state of key objects.
[0039] The feedback correction module is used to respond to the execution of a single frame action and calculate the spatial error index based on the real-time environmental state. If the object contact distance deviation, trajectory collision probability, or physical constraint violation value exceeds the preset threshold, the spatiotemporal diffusion model is triggered to generate a correction frame based on the current physical constraint map.
[0040] The dynamic constraint graph update module is used to dynamically update the object positions, distance thresholds, and motion feasible regions in the physical constraint graph according to changes in environmental state, and input the updated constraint graph into the conditional encoding process of the spatiotemporal diffusion model.
[0041] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the embodied video generation method as described above.
[0042] The aforementioned embodied video generation method, device, and electronic equipment systematically address the physical inconsistencies, error accumulation, and semantic fragmentation issues in the background technology through the synergistic effect of semantic-spatial joint encoding and dual-path feedback mechanisms. First, by explicitly defining three-dimensional spatial relationships based on physical constraint graphs, each frame of action generated by the spatiotemporal diffusion model strictly adheres to physical laws. Combined with a single-frame-level spatial feedback mechanism, quantitative indicators such as contact distance deviation and collision probability are calculated in real time, triggering action correction and completely eliminating object penetration and floating phenomena in traditional solutions. Second, when the spatial error indicators of a single-frame action exceed the standard, a correction frame is immediately triggered, completely resolving the deviation amplification problem of the open-loop architecture. Simultaneously, by dynamically updating the physical constraint graph in real time to respond to environmental changes, the action sequence continuously satisfies spatial constraints in dynamic scenes. Furthermore, a step-level semantic feedback mechanism verifies the consistency of task logic at key operation nodes, proactively preventing the chain reaction of errors caused by action deviations. Finally, the joint encoding of task semantics and physical constraints directly binds abstract operational intentions to spatial parameters, and semantic vector encoding maps abstract task instructions to a sequence of operation steps. By dynamically updating the constraint graph in real time to bind spatial parameters, the fragmentation between semantic understanding and physical execution is resolved. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0044] Figure 1 This is a flowchart of an embodiment of a video generation method;
[0045] Figure 2 This is a schematic diagram of an embodiment of a video generation device;
[0046] Figure 3 This is an internal structural diagram of a computer device according to one embodiment. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] Current video generation technologies primarily rely on diffusion models combined with textual descriptions or initial frame conditions to generate video sequences (such as VideoAgent and 3D-VLA). However, in embodied intelligence tasks (such as robot grasping and virtual character interaction), these technologies suffer from the following fundamental drawbacks:
[0049] Lack of physical consistency: Current methods (such as the inverse dynamics model of ICLR 2025) rely solely on RGB image features and lack explicit modeling of the 3D structure of the scene. This results in objects penetrating, floating, and misaligning during interactions (e.g., a robotic arm grabbing an object without calculating the precise distance between the gripper and the target, leading to "empty grabbing"). Action paths violate physical constraints (e.g., a virtual character opening a door without touching the doorknob, or a basketball throwing trajectory ignoring the effects of gravity).
[0050] Error accumulation: The mainstream solution adopts an open-loop generation architecture ("initial input → one-way prediction"), without establishing a real-time feedback link between the generated result and the environmental state. The generation deviation accumulates along the time dimension (e.g., autonomous vehicles generate collision trajectories because they do not perceive changes in pedestrian positions in real time). Dynamic changes in the environment (e.g., moving obstacles) cause the task execution to deviate from the target (e.g., a robotic arm continues to execute the wrong path due to sudden obstacle occlusion).
[0051] Semantic understanding is fragmented: the generation process is disconnected from the task semantics, the visual generation is separated from the target understanding module, the generated action does not match the task intent (e.g., in the "press button" task, the robotic arm only hovers above the button without pressing it down), and key object attributes change abruptly or disappear (e.g., in the assembly task, the screw color / position is inconsistent).
[0052] This invention achieves high-precision video planning and generation in dynamic environments by integrating a dual-feedback mechanism that combines semantic task understanding with spatial physical constraints, along with a collaborative control strategy for large and small loops. It is primarily applied to embodied intelligence scenarios such as robot operation, autonomous driving decision-making, and virtual reality interaction, addressing the problems of physical inconsistency, error accumulation, and fragmented semantic understanding caused by insufficient environmental understanding and closed-loop feedback in existing video generation methods.
[0053] The following is combined with Figures 1-3 The present invention describes a method, apparatus, and electronic device for generating embodied videos.
[0054] like Figure 1 As shown, in one embodiment, an embodied video generation method includes the following steps:
[0055] Step S110: parse the user task instructions and initial environmental observation data, generate the key operation step sequence of the task objective and its associated object set, and construct a three-dimensional physical constraint map of the scene. The physical constraint map includes object position, size, distance threshold and motion feasible domain.
[0056] Task instruction parsing: The Visual Language Model (VLM) is used to parse user task instructions (such as "open the door"), and combined with initial environmental observation data (RGB image, depth map, mask image, 3D point cloud), the key operation step sequence (such as "move → rotate → push") and the set of associated objects (such as door handle, door body).
[0057] Spatial data extraction: Point cloud data is processed using Open3D and semantic masks are generated using Grounding DINO to extract the two-dimensional pixel coordinates, three-dimensional spatial coordinates, and Euclidean distance between the grabber and the target object.
[0058] Physical constraint graph construction: Construct a structured constraint graph containing the following elements: grippers: 2D / 3D coordinates of the left and right grippers; targets: 2D / 3D coordinates of the target objects; distances: distance matrix between the grippers and the target objects; dynamic parameters: initialize distance thresholds (e.g., 0.1 meters for door opening) and feasible motion domains (based on real-time point cloud boundary updates) according to task type.
[0059] Step S120: Based on the semantic vector of the task operation step sequence and the encoding conditions of the physical constraint graph, an initial action video sequence is generated through a spatiotemporal diffusion model. The video sequence includes the pose of the executing subject and the state of key objects.
[0060] The spatiotemporal diffusion model simultaneously outputs RGB images, depth maps, and binary mask images during the generation process, providing geometric and semantic information about the scene. A weighted multimodal loss function is used during training, with a total loss L. total Defined as:
[0061] L total =w rgb ·L rgb +w mask ·L mask +w depth ·L depth ,
[0062] Among them, w rgb =1.0,w mask =1.0,w depth =0.6 (weight configuration, emphasizing RGB and mask precision); L rgb =L1_Loss(RGB) pred RGB gt (RGB image loss, L1 norm); L depth =L1_Loss(Depth pred Depth gt (Depth map loss, L1 norm); L mask =BCEWithLogitsLoss(Mask) pred Mask gt (Mask image loss, binary cross-entropy loss function, because it is a binary segmentation task). RGB loss weight is 1.0, mask loss weight is 1.0, and depth map loss weight is 0.6; the mask image loss function uses binary cross-entropy, while the depth map and RGB loss functions use the L1 norm. The loss weight allocation (RGB: 1.0, mask: 1.0, depth map: 0.6) and the dedicated loss functions (binary cross-entropy for the mask, L1 norm for depth / RGB) balance multi-task learning and solve the problems of generated object color change and disappearance.
[0063] In step S130, in response to the execution of a single frame action, the spatial error index is calculated based on the real-time environmental state. If the object contact distance deviation, trajectory collision probability, or physical constraint violation value exceeds the preset threshold, the spatiotemporal diffusion model is triggered to generate a correction frame based on the current physical constraint map.
[0064] During the execution of the video sequence, closed-loop feedback is performed in real time. Specifically, in response to a spatial error index exceeding a threshold during a single-frame action, a correction frame is generated. The spatial error index includes at least one of object contact distance deviation, trajectory collision probability, and physical constraint violation value. Upon completion of key operation steps, the semantic matching degree between the generated result and the task objective is verified. If missing key object attributes or operational logic errors are detected, the task operation step sequence and physical constraint diagram are regenerated.
[0065] The physical constraint violation value is a quantitative synthesis of the following three types of spatial error indices:
[0066] Object contact distance deviation: The absolute value of the distance between the real-time grasper and the target object exceeding the task type threshold (calculation formula: |d actual -d threshold |), where the threshold is dynamically set according to the task (e.g., 0.1 meters for the door opening task);
[0067] Trajectory collision probability: The percentage of overlap between the execution path and obstacles / boundaries calculated based on a 3D coordinate encoder (calculation formula: );
[0068] Motion feasible domain violation: The Euclidean distance by which the pose coordinates of the executing agent exceed the boundary of the physical constraint graph (calculation formula: ).
[0069] When any of the above indicators exceeds a preset threshold (such as distance deviation > 0.05 meters, collision probability > 10%, feasible region violation > 0.01 meters), a correction frame is triggered.
[0070] Execute the following closed-loop operation in real time until the task is completed:
[0071] The small-loop feedback mechanism performs physical consistency or task deviation detection on the predicted video frame sequence generated by the spatiotemporal diffusion model (e.g., through visual language model parsing). If the detection fails (e.g., detecting floating objects, penetration, or missing key actions), semantic feedback text and spatial constraint information are generated, triggering the spatiotemporal diffusion model to generate a correction frame based on the current physical constraint map (i.e., performing prediction-level correction before the actual action is executed). During the generation phase, the small loop detects physical inconsistencies (e.g., object penetration) or task semantic deviations (e.g., grasping failure), triggering the diffusion model to make immediate corrections through semantic feedback text and spatial constraint information, thus avoiding error accumulation.
[0072] The large loop feedback mechanism monitors the environmental state after actual execution. When a displacement change is detected to be less than a displacement threshold (e.g., 0.001) for a preset number of consecutive times (e.g., 15 times), a spatial reconstruction command is generated, triggering the regeneration of the task operation step sequence (i.e., replanning at the task level when encountering obstacles during actual execution). The large loop monitors displacement changes after the robotic arm executes (determining "stuck" when displacement is <0.001 for 15 consecutive times), triggering task step reconstruction to resolve target deviation issues in dynamic environments.
[0073] By dividing feedback into small loops (real-time understanding of predicted video frames) and large loops (monitoring of the post-execution environment), a hierarchical error correction mechanism is achieved, and the dual loops work together to improve physical consistency and task robustness.
[0074] The trigger condition for small-loop feedback is: parsing and predicting video frames using a visual language model, outputting a triplet containing generation quality error codes, observation difference descriptions, and regeneration decisions. Specifically, parsing and predicting the video frame sequence using a visual language model (VLM) outputs a structured judgment containing the following information:
[0075] Generate quality error codes: identify the type of problem generated, such as TextureOrColorChange, ObjectDisappearance, InconsistentAppearance, etc. (see the key failure modes analyzed in document 3).
[0076] Observational discrepancy description: Specifically describes the difference between the predicted frame and the expected or physical constraints, such as NoEffectiveContactEstablished (no effective contact established, i.e., "empty grab"), CollisionRiskDetected (collision risk detected), TargetObjectMissing (target object missing), etc.
[0077] Regeneration Decision: Based on the above error codes and descriptions, determine whether a correction frame (Reject) needs to be generated or the current prediction (Accept) should be accepted. If it is Reject, the correction frame generation process is triggered.
[0078] The small loop triggers corrections by parsing triples using VLM, such as generating feedback text when detecting "empty grab" to avoid invalid actions. The trigger condition for feedback from the large loop is: when the displacement of the robotic arm end effector changes by less than 0.001 in 15 consecutive observations, it is determined to be in a stuck state. The large loop determines the "stuck" state based on the displacement threshold (0.001) and the number of consecutive frames (15) to prevent misjudgment (too low a threshold leads to frequent reconstruction) or delayed response (too high a threshold misses the opportunity for error correction), improves dynamic adaptation efficiency, and ensures the accuracy of error correction by quantifying the feedback trigger condition.
[0079] The generation of the correction frame for small loop feedback includes:
[0080] The semantic feedback text is encoded as a text embedding vector. This text is generated by the visual language model based on the task text and current spatial information (such as key coordinates / distances extracted from spatial constraints). Essentially, it is an action planning sequence in the format "action type [direction vector][distance value]", for example, "move[-1,0,0][0.19]", "move[0,0,-1][0.21]", "push[0,0,0][0.00]". This text explicitly instructs the spatiotemporal diffusion model how to adjust actions to correct detected problems (e.g., moving a certain distance in which direction to contact the target). The semantic feedback text generates action sequence instructions (e.g., "move[-1,0,0][0.19]"), specifying the direction and distance of movement, guiding the diffusion model to generate frames that conform to the task semantics.
[0081] Spatial constraint information is converted into layout spatial encoding. This spatial constraint information is constructed by extracting the 3D coordinates and relative distance between the grasper and the target object, including a key-value pair data structure of 2D pixel coordinates, 3D spatial coordinates, and distance matrices. The spatial constraint information stores coordinates and distances (e.g., grasper 3D coordinates, target object Euclidean distance) as key-value pairs, providing real-time data for the physical constraint graph and solving the problem of missing spatial perception. A cross-attention mechanism is used to fuse text embedding vectors and layout spatial encoding into the denoising process of the spatiotemporal diffusion model. The spatiotemporal diffusion model adopts the UNet architecture, and the text embedding vectors and layout spatial encoding are fused to the middle block layer of UNet through the cross-attention mechanism. This method is also applicable to diffusion models with the DiT (Diffusion Transformer) architecture. By fusing dual-path feedback through cross-attention, the diffusion model understands and utilizes environmental feedback, dynamically adjusting the generation strategy (e.g., correcting the grasping direction), forming a closed-loop "generation-understanding-correction" process to reduce physical inconsistencies.
[0082] The process of constructing spatial constraint information includes: extracting the two-dimensional pixel coordinates and three-dimensional spatial coordinates of the left and right grabbers from the current frame, and updating the boundary of the feasible motion domain by combining the point cloud; calculating the Euclidean distance between the grabber and each target object and generating a distance matrix; converting the coordinate data into layout spatial encoding format, generating an Euclidean distance matrix and converting it into layout encoding, and guiding the diffusion model to execute according to the updated plan through cross attention, avoiding collisions or misalignments, significantly reducing object floating and penetration problems, and ensuring physical consistency.
[0083] The spatial constraint information is structured as key-value pairs, including: `grippers` field: stores the 2D pixel coordinates and 3D spatial coordinates of the left and right grippers; `targets` field: stores the 2D pixel coordinates and 3D spatial coordinates of each target object; `distances` field: stores the Euclidean distance matrix between the grippers and the target objects, in the format {gripper_id, target_id, distance}. The key-value pair format explicitly stores the gripper coordinates (grippers), target object coordinates (targets), and distance matrix (distances), providing machine-readable layout encoding input, solving the problem of diffusion models being unable to parse spatial information, and ensuring effective injection of physical constraints.
[0084] The training data for the spatiotemporal diffusion model is generated through the following process:
[0085] Data acquisition: Collect depth maps, mask maps, RGB maps, and 3D point cloud data of the successful mission trajectory from the simulation environment (such as Meta-World), and record information such as mission name and time step.
[0086] Dynamic segmentation of fine operation phases: The fine operation phase is dynamically identified based on the Euclidean distance between the grabber and the target object: When the distance of N consecutive frames (N is set according to the task type, see the distance threshold table) is lower than the distance threshold of the task, this time period is marked as a fine operation phase.
[0087] Frame samples marked as fine-grained operation segments are assigned a 5x sampling probability weight, making the model pay more attention to these critical, error-prone interaction stages during training.
[0088] Differentiated image cropping: Execute differentiated image cropping strategies based on the regions of interest for different tasks.
[0089] For example, the door opening task uses center clipping.
[0090] For example, the basketball placement task uses the top right crop.
[0091] The cropping frame parameters are dynamically set based on the task type and viewpoint, for example:
[0092] Third-person perspective of the basketball shooting task: The clipping box coordinates are (100, 10, 228, 138);
[0093] First-person perspective for basketball shooting missions: Center-based cropping;
[0094] Door opening task: Use center clipping;
[0095] The remaining task and perspective combinations use center clipping by default.
[0096] (The specific cropping method is determined based on the position of the most relevant area of the task within the image frame).
[0097] Specific values can also be provided for the specific cutting method:
[0098]
[0099] All other combinations of tasks and perspectives are center-cropped.
[0100] Sample construction: Every 3 frames is used as an initial frame, and sampling is performed according to the fine sampling time period. The final cropped and processed images (RGB, Depth, Mask) and corresponding task text, spatial information, etc. constitute the training samples.
[0101] Based on the dynamic division of fine-grained operation segments according to the distance between the grasper and the target object (e.g., distance < 0.1 meters for 3 consecutive frames), a sampling probability of 5 times is assigned, focusing on key action frames. Task-customized cropping strategies (cropping the center when the door opens, cropping the upper right corner when the basketball is placed) preserve task-related areas, enhancing the accuracy of the generated video, and differentiated data sampling strategies improve training efficiency.
[0102] Step S140: Dynamically update the object position, distance threshold, and feasible motion domain in the physical constraint diagram according to changes in environmental state, and input the updated constraint diagram into the spatiotemporal diffusion model.
[0103] Dynamically update the physical constraint graph, including:
[0104] Based on the real-time distance between the grabber and the target object, the key distance threshold parameters for each task in the physical constraint diagram are dynamically adjusted according to the task type. These thresholds are used for: determining whether a fine-tuning state has been entered (distance is below the threshold for N consecutive frames); serving as a benchmark for spatial error indicators (such as contact distance deviation); and affecting the boundary calculation of the feasible motion domain. Specifically, the distance threshold for the door opening task is 0.1 meters, the basketball placement task is 0.15 meters, and the assembly task is 0.1 meters. Examples of dynamic distance threshold settings for typical tasks are shown in the table below.
[0105]
[0106] The feasible motion domain boundary coordinates of the target object are updated in real time based on the mask image and 3D point cloud data. Task thresholds (e.g., 0.1 meters for door opening, 0.1 meters for assembly) are updated according to real-time distance, and the feasible motion domain is updated based on the mask and point cloud. Parameters are customized for different tasks (e.g., a basketball placement threshold of 0.15 meters) to avoid action failures caused by one-size-fits-all constraints, improving the success rate in dynamic scenes. Physical constraint parameters are dynamically adjusted to enhance environmental adaptability.
[0107] The embodied video generation method in this embodiment systematically solves the problems of physical inconsistency, error accumulation, and semantic fragmentation in the background technology through the synergistic effect of semantic-spatial joint coding generation and a two-level feedback mechanism. First, based on the physical constraint graph, the three-dimensional spatial relationships (such as the object contact distance threshold and the feasible motion domain) are explicitly defined, ensuring that each frame of action generated by the spatiotemporal diffusion model strictly follows physical laws. Combined with a single-frame-level spatial feedback mechanism, quantitative indicators such as contact distance deviation and collision probability are calculated in real time and action correction is triggered, completely eliminating the object penetration and floating phenomena in traditional solutions (the physical violation rate was reduced from 21.3% to 1.2% in robotic arm grasping tests). Second, when the spatial error index of a single frame action exceeds the standard, the generation of a correction frame is immediately triggered, completely solving the deviation amplification problem of the open-loop architecture. Simultaneously, by dynamically updating the physical constraint graph in real time to respond to environmental changes (such as adjusting the feasible domain of moving obstacles), the action sequence is ensured to continuously meet spatial constraints in dynamic scenarios. At the same time, the step-level semantic feedback mechanism verifies the consistency of task logic at key operation nodes (such as detecting the downward displacement when "pressing the button"), actively preventing the chain propagation of errors caused by action deviation. Finally, the joint encoding of task semantics and physical constraints directly binds the abstract operation intention (such as "opening the door") to spatial parameters (door handle coordinates + rotation angle), and verifies and maintains key object attributes (such as screw position / color in assembly tasks) through semantic matching degree verification, so that the generated actions are accurately aligned with the task target (the virtual door opening task completion rate increases from 68% to 99%), realizing a closed-loop control of the entire link from task understanding to physical execution.
[0108] The embodied video generation apparatus provided by the present invention will be described below. The embodied video generation apparatus described below can be referred to in correspondence with the embodied video generation method described above.
[0109] like Figure 2 As shown, in one embodiment, an embodied video generation device includes a task parsing and constraint graph construction module 210, a video sequence generation module 220, a feedback correction module 230, and a dynamic constraint graph update module 240.
[0110] The task parsing and constraint graph construction module 210 is used to parse user task instructions and initial environmental observation data, generate a sequence of key operation steps of the task objective and its associated object set, and construct a three-dimensional physical constraint graph of the scene. The physical constraint graph includes object position, size, distance threshold and feasible motion domain.
[0111] The video sequence generation module 220 is used to generate an initial action video sequence based on the encoding conditions of the semantic vector and physical constraint graph of the task operation step sequence through a spatiotemporal diffusion model. The video sequence includes the pose of the executing subject and the state of key objects.
[0112] The feedback correction module 230 is used to respond to the execution of a single frame action and calculate the spatial error index based on the real-time environmental state. If the object contact distance deviation, trajectory collision probability, or physical constraint violation value exceeds the preset threshold, the spatiotemporal diffusion model is triggered to generate a correction frame based on the current physical constraint map.
[0113] The dynamic constraint graph update module 240 is used to dynamically update the object position, distance threshold and motion feasible region in the physical constraint graph according to the changes in environmental state, and input the updated constraint graph into the condition encoding process of the spatiotemporal diffusion model.
[0114] Figure 3 This example illustrates a schematic diagram of the physical structure of an electronic device, which can be a smart terminal. Its internal structure diagram can be as follows: Figure 3 As shown. The electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an embodied video generation method, which includes:
[0115] The system parses user task instructions and initial environmental observation data to generate a sequence of key operation steps for the task objective and a set of associated objects, and constructs a three-dimensional physical constraint map of the scene. The physical constraint map includes object positions, sizes, distance thresholds, and feasible motion domains.
[0116] Based on the encoding conditions of the semantic vector and physical constraint graph of the task operation step sequence, an initial action video sequence is generated through a spatiotemporal diffusion model. The video sequence includes the pose of the executing subject and the state of key objects.
[0117] During the execution of the video sequence, closed-loop feedback is performed in real time. Specifically, in response to the execution of a single frame action, spatial error indicators are calculated based on the real-time environmental state. If the object contact distance deviation, trajectory collision probability, or physical constraint violation value exceeds a preset threshold, the spatiotemporal diffusion model is triggered to generate a correction frame based on the current physical constraint map. In response to the completion of key operation steps, the semantic matching degree between the generated result and the task objective is verified. If the missing key object attributes or operation logic error is detected, the task operation step sequence and physical constraint map are regenerated.
[0118] The physical constraint graph is dynamically updated based on changes in environmental conditions, including the object position, distance threshold, and feasible motion domain. The updated constraint graph is then input into the conditional encoding process of the spatiotemporal diffusion model.
[0119] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the electronic device to which the present invention is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0120] On the other hand, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements an embodied video generation method, the method comprising:
[0121] The system parses user task instructions and initial environmental observation data to generate a sequence of key operation steps for the task objective and a set of associated objects, and constructs a three-dimensional physical constraint map of the scene. The physical constraint map includes object positions, sizes, distance thresholds, and feasible motion domains.
[0122] Based on the encoding conditions of the semantic vector and physical constraint graph of the task operation step sequence, an initial action video sequence is generated through a spatiotemporal diffusion model. The video sequence includes the pose of the executing subject and the state of key objects.
[0123] During the execution of the video sequence, closed-loop feedback is performed in real time. Specifically, in response to the execution of a single frame action, spatial error indicators are calculated based on the real-time environmental state. If the object contact distance deviation, trajectory collision probability, or physical constraint violation value exceeds a preset threshold, the spatiotemporal diffusion model is triggered to generate a correction frame based on the current physical constraint map. In response to the completion of key operation steps, the semantic matching degree between the generated result and the task objective is verified. If the missing key object attributes or operation logic error is detected, the task operation step sequence and physical constraint map are regenerated.
[0124] The physical constraint graph is dynamically updated based on changes in environmental conditions, including the object position, distance threshold, and feasible motion domain. The updated constraint graph is then input into the conditional encoding process of the spatiotemporal diffusion model.
[0125] In another aspect, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, it implements an embodied video generation method, the method comprising:
[0126] The system parses user task instructions and initial environmental observation data to generate a sequence of key operation steps for the task objective and a set of associated objects, and constructs a three-dimensional physical constraint map of the scene. The physical constraint map includes object positions, sizes, distance thresholds, and feasible motion domains.
[0127] Based on the encoding conditions of the semantic vector and physical constraint graph of the task operation step sequence, an initial action video sequence is generated through a spatiotemporal diffusion model. The video sequence includes the pose of the executing subject and the state of key objects.
[0128] During the execution of the video sequence, closed-loop feedback is performed in real time. Specifically, in response to the execution of a single frame action, spatial error indicators are calculated based on the real-time environmental state. If the object contact distance deviation, trajectory collision probability, or physical constraint violation value exceeds a preset threshold, the spatiotemporal diffusion model is triggered to generate a correction frame based on the current physical constraint map. In response to the completion of key operation steps, the semantic matching degree between the generated result and the task objective is verified. If the missing key object attributes or operation logic error is detected, the task operation step sequence and physical constraint map are regenerated.
[0129] The physical constraint graph is dynamically updated based on changes in environmental conditions, including the object position, distance threshold, and feasible motion domain. The updated constraint graph is then input into the conditional encoding process of the spatiotemporal diffusion model.
[0130] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory.
[0131] By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0132] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0133] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
Claims
1. A method for generating embodied videos, characterized in that, The method includes: The system parses user task instructions and initial environmental observation data to generate a sequence of key operation steps for the task objective and a set of associated objects, and constructs a three-dimensional physical constraint map of the scene. The physical constraint map includes object positions, sizes, distance thresholds, and feasible motion domains. Based on the encoding conditions of the semantic vector and physical constraint graph of the task operation step sequence, an initial action video sequence is generated through a spatiotemporal diffusion model. The video sequence includes the pose of the executing subject and the state of key objects. In response to the execution of a single frame action, the spatial error index is calculated based on the real-time environmental state. If the object contact distance deviation, trajectory collision probability, or physical constraint violation value exceeds the preset threshold, the spatiotemporal diffusion model is triggered to generate a correction frame based on the current physical constraint map. The physical constraint graph is dynamically updated based on changes in environmental conditions, including the object position, distance threshold, and feasible motion domain. The updated constraint graph is then input into the spatiotemporal diffusion model.
2. The embodied video generation method according to claim 1, characterized in that, The triggering spatiotemporal diffusion model generates a correction frame based on the current physical constraint graph, including: Execute small-loop feedback to generate semantic feedback text and spatial constraint information to input the spatiotemporal diffusion model.
3. The embodied video generation method according to claim 2, characterized in that, The execution of small-loop feedback, generating semantic feedback text and spatial constraint information input to the spatiotemporal diffusion model, includes: The semantic feedback text is encoded into a text embedding vector. The semantic feedback text is generated by the visual language model based on the task text and the current spatial information to form an action planning sequence, in the format of "action type [direction vector] [distance value]". The spatial constraint information is converted into layout spatial encoding. The spatial constraint information is constructed by extracting the three-dimensional coordinates and relative distance between the crawler and the target object, including a key-value pair data structure of two-dimensional pixel coordinates, three-dimensional spatial coordinates and distance matrix. The text embedding vector and layout space encoding are fused into the denoising process of the spatiotemporal diffusion model through a cross-attention mechanism. The spatiotemporal diffusion model adopts the UNet architecture, and the text embedding vector and layout space encoding are fused into the middle block layer of UNet through the cross-attention mechanism.
4. The embodied video generation method according to claim 3, characterized in that, The process of constructing the spatial constraint information includes: Extract the two-dimensional pixel coordinates and three-dimensional spatial coordinates of the left and right grabbers from the current frame; Calculate the Euclidean distance between the gripper and each target object and generate a distance matrix; Convert the coordinate data to layout space encoding format.
5. The embodied video generation method according to claim 1, characterized in that, Also includes: In response to the completion of key operational steps, verify the semantic matching degree between the generated results and the task objectives; If a critical object attribute is missing or an operational logic error is detected, the operation step sequence and physical constraint diagram will be regenerated.
6. The embodied video generation method according to claim 1, characterized in that, The dynamic updating of the physical constraint graph includes: Based on the real-time distance between the grabber and the target object, the distance threshold parameters for each task are dynamically adjusted according to the task type. The coordinates of the feasible motion domain boundary of the target object are updated in real time based on the mask image and 3D point cloud data.
7. The embodied video generation method according to claim 1, characterized in that, The generation process of the spatiotemporal diffusion model includes: Simultaneously output RGB image, depth map and binary mask image; A weighted multimodal loss function is used, the mask image loss function uses binary cross-entropy, and the depth image and RGB loss functions use the L1 norm.
8. The embodied video generation method according to any one of claims 1 to 7, characterized in that, The method also includes a training data generation process: Depth maps, mask maps, RGB maps, and 3D point cloud data of the task trajectory are collected from the simulation environment; The fine operation segment is dynamically divided based on the Euclidean distance between the grabber and the target object: when the distance is lower than the task type threshold for N consecutive frames, it is marked as a fine operation segment, where N is the preset number of consecutive frames. Assign M times the sampling probability weight to the fine-grained operation segment, where M is a natural number; Differentiated image cropping is performed based on task type: the door opening task uses center cropping, and the basketball placement task uses top right cropping.
9. A device for generating embodied videos, characterized in that, The device includes: The task parsing and constraint graph construction module is used to parse user task instructions and initial environmental observation data, generate a sequence of key operation steps for the task objective and its associated object set, and construct a three-dimensional physical constraint graph of the scene. The physical constraint graph includes object position, size, distance threshold and feasible motion domain. The video sequence generation module is used to generate an initial action video sequence based on the encoding conditions of the semantic vector and physical constraint graph of the task operation step sequence through a spatiotemporal diffusion model. The video sequence includes the pose of the executing subject and the state of key objects. The feedback correction module is used to respond to the execution of a single frame action and calculate the spatial error index based on the real-time environmental state. If the object contact distance deviation, trajectory collision probability, or physical constraint violation value exceeds the preset threshold, the spatiotemporal diffusion model is triggered to generate a correction frame based on the current physical constraint map. The dynamic constraint graph update module is used to dynamically update the object positions, distance thresholds, and motion feasible regions in the physical constraint graph according to changes in environmental state, and input the updated constraint graph into the condition encoding process of the spatiotemporal diffusion model.
10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the embodied video generation method according to any one of claims 1 to 8.
Citation Information
Cited By
Robot VLA cooperative control method based on space-time enhancement and trajectory smoothing
CN122334339A
A robot VLA cooperative control method based on space-time enhancement and trajectory smoothing
CN122334339B