World motion model enhancement method and device fusing structured prior and continuous dynamics
By combining a dual-perspective collaborative and structured prior world action model with continuous dynamics and spectral analysis, the problem of unstable action learning in existing technologies is solved, and robustness and closed-loop stability of action generation in complex scenarios are achieved.
Patent Information
- Application Number
- CN202610689861.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-25
AI Technical Summary
Existing world models struggle to stably learn the correspondence between actions and state changes in human video action learning. They lack a unified expression of scene semantics, object relationships, and execution constraints, resulting in insufficient robustness of action generation and poor long-term action stability and closed-loop execution performance.
By employing a dual-perspective collaborative spatiotemporal aligned input, combining structured priors and continuous dynamics, we characterize the evolution of latent states from a continuous-time perspective through spectral analysis and neural control differential equations (NCDEs). We introduce knowledge graph constraints and causal constraints to construct a world action model that integrates structured priors and continuous dynamics.
It improves the feasibility and stability of action learning, enhances the ability to recognize periodic and micro-dynamic actions, improves the rationality and executability of action generation, and ensures robustness and closed-loop stability in complex scenarios.
Smart Images

Figure CN122635533A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of world modeling, action learning, structured prior constraint reasoning, continuous dynamics modeling, and embodied intelligence, and more specifically, to a method and apparatus for enhancing world action models by integrating structured priors and continuous dynamics. Background Technology
[0002] In recent years, world model methods, which learn the patterns of observation changes over time, have attracted widespread attention in tasks such as robot control, action generation, and agent decision-making. Compared with methods that directly regress actions from observations, world models can learn the state evolution process to a certain extent, thus providing a basis for future changes in action generation. Therefore, they are considered an important technical approach to improve continuous decision-making capabilities.
[0003] However, existing world models still have significant shortcomings when learning actions from human videos. First, human videos typically lack fine-grained action labels that strictly correspond to the target actor. Models are more likely to learn appearance changes or temporal co-occurrence relationships, but struggle to stably learn the corresponding rules of "what action causes what state change," resulting in insufficient conditionality and weak executability of output actions. Second, existing methods mostly rely on statistical correlations in continuous representations for modeling, lacking explicit utilization of task semantics, object relationships, pre- and post-action conditions, causal dependencies, and contact and motion constraints. Consequently, they are prone to misinterpreting statistically common but actually unreasonable changes as executable actions. Third, existing video-action joint modeling methods often continuously use the model's own predictions as subsequent inputs during inference. As the number of scrolling steps increases, errors tend to accumulate, affecting long-term action stability and closed-loop execution performance.
[0004] Furthermore, action generation in complex operational scenarios not only depends on local visual changes but is also influenced by scene semantics, object interaction relationships, and execution constraints. World models that rely solely on data-driven training typically lack the ability to uniformly express and utilize these constraints, resulting in insufficient robustness under conditions of scene changes, occlusion interference, task switching, or weak annotation. They also struggle to provide clear and reliable constraints for the action generation process. Meanwhile, most existing video-action joint modeling methods still approximate state transitions using fixed discrete time steps, making it difficult to accurately characterize the continuous changes in latent states in scenarios with varying observation intervals, uneven action execution rhythms, or strong intra-block continuous evolution. Summary of the Invention
[0005] Based on existing analysis, the following technical problems urgently need to be solved: how to effectively learn latent actions under weakly labeled human video conditions; how to integrate semantic, causal, and physical constraints into the training and inference process of the world model; how to use frequency domain information to enhance the ability to identify rhythmic and micro-dynamic actions; how to characterize the smooth evolution of latent states with control paths from a continuous-time perspective through action-conditional NCDE (Neural Controlled Differential Equation); and how to achieve stable closed-loop action generation under real observation feedback.
[0006] To address the aforementioned technical challenges, this invention proposes an enhancement method for world action models that integrates structured priors and continuous dynamics. This method uses first-person perspective video as the primary input for action learning and synchronous third-person perspective video as supplementary scene input, constructing structured prior constraints and continuous spatiotemporal latent variables. The former is responsible for generating rule constraints, causal constraints, and counterfactual constraints, while the latter is responsible for learning continuous world evolution under action conditions. Furthermore, within the continuous spatiotemporal latent variable branch, it further distinguishes between the first-person perspective action mainstream, the third-person perspective state auxiliary flow, the frequency domain auxiliary action representation branch, and the continuous-time latent dynamics branch based on action condition NCDE. This allows the world action model to maintain local action sensitivity while gaining a larger scene receptive field and enhancing its modeling ability for periodic, oscillatory, or rhythmic actions as well as continuous state evolution under irregular time intervals.
[0007] The technical solution adopted in this invention is as follows: The first aspect provides a method for enhancing world action models by integrating structured priors and continuous dynamics, including: Construct a spatiotemporally aligned input for dual-view collaboration, wherein the spatiotemporally aligned input for dual-view collaboration includes first-view video and third-view video; Using first-person perspective videos as the main input for action learning and third-person perspective videos as supplementary scene input, a structured prior basic constraint representation is constructed. Under weak labeling conditions, spectral analysis is performed on the local temporal features of the action, and the results are jointly encoded with the spatiotemporal features into a potential action. The system integrates dual-view video semantics and task conditions to generate a continuous semantic conditional representation corresponding to the current time block, and maps it to a unified conditional representation of the world action model. The underlying action representation, unified conditional representation, and structured prior basic constraints are input into the world action model, and a continuous-time latent dynamics branch based on neural control differential equations is introduced in the basic pre-training stage. Based on the results of basic pre-training, the dynamic priors output by the continuous-time latent dynamics branch are invoked, and combined with causal constraints, counterfactual constraints, and execution constraints, joint video and action reasoning training is performed under the guidance of unified conditional representation to obtain an augmented world action model that integrates structured priors and continuous dynamics.
[0008] In one implementation, constructing a spatiotemporally aligned input for dual-view collaboration includes: Collect human first-person and third-person perspective video sequences; Frame sampling, segmentation, timestamp alignment, and unified encoding are performed on the acquired first-view and third-view motion video sequences. The encoded dual-view spatiotemporal feature sequences are used as spatiotemporal alignment inputs for dual-view collaboration. The first-view video is used to provide local motion changes and fine operation processes, while the third-view video is used to provide overall spatial relationships and the state of the scene outside occlusion.
[0009] In one implementation, first-view video is used as the primary input for action learning, and third-view video is used as supplementary scene input to construct structured prior constraints, including: Perform object detection, instance tracking and event recognition on dual-view videos to obtain entity nodes, event nodes and their relational edges, and construct a scene-event graph corresponding to the current video segment; Retrieve knowledge graphs based on video topics, object sets, and event sets to form knowledge subgraphs related to the current video; By transforming the knowledge subgraph into a rule mask, continuous bias, and candidate causal path scores, the knowledge subgraph is converted into discrete structural constraints, which serve as the basic constraint representation of the structured prior.
[0010] In one implementation, under weak labeling conditions, spectral analysis is performed on the local temporal features of an action, and the results are jointly encoded with spatiotemporal features into a potential action, including: A latent action extractor is constructed using a dual-frame spatiotemporal Transformer architecture. Spectral analysis is performed on the temporal features of action-related local regions to jointly learn a time-frequency consistent low-dimensional continuous latent action representation from the video.
[0011] In one implementation, based on the structured prior constraint representation, a continuous semantic condition representation corresponding to the current time block is extracted and mapped to a unified condition representation of the world action model, including; Extracting short-term local semantic summaries, long-term semantic guidance, and task language prompts from dual-view videos; The extracted short-term local semantic summary, long-term semantic guidance, task language prompts, and dual-view coding features are jointly mapped into a unified conditional representation of the world action model, forming a block-level conditional input aligned with the current trajectory block.
[0012] In one implementation, the underlying action representation, unified conditional representation, and structured prior basic constraints are jointly input into the world action model. During the basic pre-training phase, a continuous-time latent dynamics branch based on neural control differential equations is introduced, including: First-person perspective features, third-person perspective features, frequency-domain enhanced latent action representations, and unified conditional representations are input into the world action model. A continuous-time latent dynamics branch based on neural control differential equations is introduced. The main branch of the continuous-time latent dynamics branch uses an autoregressive conditional diffusion Transformer with causal modeling capabilities to uniformly model the dependencies between historical and current state blocks. Flow matching is used as the main loss, which enables the model to learn the continuous transmission direction from noisy state to real state and the fine-grained correspondence between actions and consequences.
[0013] In one implementation, based on the results of basic pre-training, the dynamic priors output by the continuous-time latent dynamics branch are invoked, and combined with causal constraints, counterfactual constraints, and execution constraints, joint video and action reasoning training is performed under the guidance of a unified conditional representation to obtain an augmented world action model that integrates structured priors and continuous dynamics, including: Using the preceding trajectory block as a clean context, the current trajectory block representation, along with structured prior constraints, frequency domain action priors, continuous-time latent dynamics priors given by the continuous-time latent dynamics branch, and counterfactual conditions, are input into the joint inference backbone of the world action model. Based on the historical context, joint inference is performed on the current block. Under the constraints of temporal causality, the model comprehensively utilizes object relationships, rule constraints, and counterfactual change information to recover the state evolution and action results of the current block, obtaining a future action representation that is consistent with the task conditions and guided by structured priors.
[0014] In one implementation, the method further includes asynchronous closed-loop execution based on an enhanced world action model that integrates structured priors and continuous dynamics, specifically: Before deployment, the augmented world action model is subjected to few-step distillation and rapid reparameterization to form two switchable inference modes: a few-step joint generation mode and a rapid action mode. The augmented world action model is deployed as an online inference module in the closed-loop system, forming an asynchronous parallel working mechanism with the execution module.
[0015] Based on the same inventive concept, a second aspect of the present invention provides a world action model enhancement device that integrates structured priors and continuous dynamics, comprising: A dual-view collaborative spatiotemporal alignment input construction module is used to construct a dual-view collaborative spatiotemporal alignment input, wherein the dual-view collaborative spatiotemporal alignment input includes a first-view video and a third-view video; The basic constraint construction module for structured priors is used to construct basic constraints for structured priors by using first-view video as the main input for action learning and third-view video as a supplementary input for the scene. The time-frequency enhanced latent action coding module is used to perform spectral analysis on the local temporal features of actions under weak labeling conditions, and jointly encode them with spatiotemporal features into latent actions; The joint semantic condition representation construction module is used to integrate dual-view video semantics and task conditions to generate a continuous semantic condition representation corresponding to the current time block and map it to a unified condition representation of the world action model; The dynamics modeling module is used to input the underlying action representation, unified conditional representation, and structured prior basic constraints into the world action model, and introduce a continuous-time latent dynamics branch based on neural control differential equations during the basic pre-training stage. The joint video-action reasoning module is used to call the dynamic priors output by the continuous-time latent dynamics branch based on the results of basic pre-training, and combine causal constraints, counterfactual constraints and execution constraints to perform joint video and action reasoning training under the guidance of unified conditional representation, so as to obtain an augmented world action model that integrates structured priors and continuous dynamics.
[0016] Based on the same inventive concept, a third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it is used to implement the world action model enhancement method that integrates structured priors and continuous dynamics as described in the first aspect.
[0017] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows: (1) This invention can extract latent action representations from human videos with weak annotations or no fine-grained action labels, and combine them with a world model to learn the correspondence between actions and state evolution, thereby improving the feasibility of using human videos for action learning and alleviating the problem of insufficient action learning in existing methods. (2) In the latent action extraction stage, the present invention introduces spectrum analysis to extract frequency domain features such as main frequency, spectral energy and phase change of the temporal changes of the action-related local area, thereby enhancing the ability to identify periodic, oscillatory and micro-dynamic actions and improving the stability of the latent action representation.
[0018] (3) This invention introduces continuous-time latent dynamics modeling based on NCDE into the action condition world model, enabling the model to treat dual-view video states, frequency-domain enhanced potential actions and task conditions as control paths, thereby more accurately describing the smooth changes of latent states under irregular observation intervals, different action rhythms and continuous evolution within blocks.
[0019] (4) This invention introduces knowledge graph constraints, semantic constraints, causal constraints and physical constraints into the action generation process, so that action learning no longer depends only on visual statistical correlation, but can be constrained by task semantics, object relations and execution conditions at the same time, thereby improving the rationality, executability and stability of action generation.
[0020] (5) By combining video and motion modeling, motion recovery can be kept consistent with future state changes, and can maintain good robustness and consistency under complex scenes, occlusion interference and task changes.
[0021] (6) By suppressing the accumulation of errors in the rolling inference of the model through the closed-loop execution method of writing back real observations, the system can continuously correct the subsequent action generation results based on the feedback of the real environment, thereby improving the closed-loop stability under online deployment conditions. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating the world action model enhancement method that integrates structured priors and continuous dynamics in an embodiment of the present invention. Figure 2 This is a block diagram of the world action model enhancement device that integrates structured priors and continuous dynamics in an embodiment of the present invention. Detailed Implementation
[0024] Example 1 This embodiment provides a method for enhancing world action models by integrating structured priors and continuous dynamics. Please refer to [link to relevant documentation]. Figure 1 ,include: S1: Construct a spatiotemporally aligned input for dual-view collaboration, wherein the spatiotemporally aligned input for dual-view collaboration includes a first-view video and a third-view video.
[0025] Specifically, S1 is the construction of spatiotemporally aligned input for dual-view collaboration, which can be achieved in the following way: S1.1: Acquire human first-person perspective video and third-person perspective motion video sequences; S1.2: Perform frame sampling, segmentation, timestamp alignment, and unified encoding on the acquired first-view video and third-view motion video sequences. Use the encoded video segments as spatiotemporal alignment input for dual-view collaboration. The first-view video is used to provide local motion changes and fine operation processes, while the third-view video is used to provide overall spatial relationships and occluded scene states.
[0026] This embodiment preferably uses both first-person and third-person perspective videos simultaneously. The first-person perspective video is denoted as... Third-person perspective video is recorded as The two are aligned using a unified time base to form a synchronized segment. First-person perspective is used for the main task of action learning, while third-person perspective is used for supplementing the overall scene. Indicates the first frame, This represents the total number of video frames. The first-person perspective is better suited for providing short-term changes in fine movements such as local variations near the hand or execution end, contact initiation positions, and grasping and insertion. The third-person perspective is better suited for providing the overall posture of the human or robot, the global relative position between the target and obstacles, the motion path, and the scene state outside of occlusion. To avoid diluting the primary action perspective with the third-person perspective, this implementation does not directly replace the first-person perspective, but instead uses it as an auxiliary state flow integrated into the world action model.
[0027] For the input first-person perspective video clip and third-person perspective video clips (ego represents first-person perspective, exo represents third-person perspective), respectively encode first-person perspective frames. With third-person perspective frames Both share the same video encoder. The shared parameters yield first-view and third-view features, i.e. and Parameter sharing refers to mapping two perspective inputs to a unified continuous spatiotemporal latent representation space through the same set of network weights, so that subsequent alignment, splicing, weighted fusion, or cross-attention calculations can be performed. To preserve perspective differences, first-view and third-view identifiers are injected at the input or during the encoding process, or independent positional encodings are used, so that the model can still distinguish different perspective sources within the shared representation space.
[0028] Then the gating vector is constructed. :
[0029] Let be the gate vector of the t-th frame. () represents the Sigmoid activation function, MLP() is a multilayer perceptron structure, pool() is a pooling operator, and it forms the fused features of frame t. :
[0030] in, Always retain it as the primary component. This involves mapping third-person perspective features to the same dimension as the first-person perspective, and supplementing the overall state only within the range allowed by the gating coefficients. Therefore, the world action model primarily uses the first-person perspective for local action modeling, only introducing third-person perspective information when assessing occlusion enhancement, overall pose changes, or environmental layout. Thus, the introduction of the third-person perspective does not disrupt the system's basic closed-loop capability, but rather enhances scene coverage and state integrity. If the third-person perspective is temporarily unavailable in the runtime environment, then... The system automatically reverts to a first-person perspective-only driving mode.
[0031] S2: Using first-person perspective video as the primary input for action learning and third-person perspective video as supplementary scene input, a structured prior representation of basic constraints is constructed. It should be noted that structured priors include knowledge graph constraints, counterfactual constraints, causal constraints, etc., with knowledge graph constraints being a fundamental constraint representation of structured priors.
[0032] Specifically, S2 is a knowledge graph-driven discrete constraint construction, which can be implemented in the following way: S2.1: Perform object detection, instance tracking, and event recognition on the dual-view video to obtain entity nodes, event nodes, and their relational edges, and construct a scene-event graph corresponding to the current video segment; S2.2: Retrieve the knowledge graph according to the video topic, object set, and event set to form a knowledge subgraph related to the current video; S2.3: By transforming the knowledge subgraph into a discrete structural constraint through rule masks, continuous biases, and candidate causal path scores, it serves as a structured prior constraint representation.
[0033] In the specific implementation process, considering that the third-person perspective has a stronger expressive ability in terms of object spatial relationships, overall posture, and environmental layout, while the first-person perspective has a higher sensitivity in terms of local contact, operational details, and action changes, features extracted frame by frame were used to determine the appropriate perspective. and Building object-relational features Event action characteristics The object relationship feature is defined as follows:
[0034] Primarily using third-person perspective features to highlight object entities and overall layout information; event action features are defined as:
[0035] Primarily using first-person perspective features, it is used to highlight the changes in motion and interaction cues during local operations such as grasping, placing, and contact. and These represent the projection transformations that map the features from another viewpoint onto the representation space of the current branch. and These represent the gating coefficients for object relationship branches and event action branches, respectively, used to control the injection intensity of information from another perspective.
[0036] In object branches, based on Object detection and instance segmentation are performed on each frame, and cross-frame instance tracking is used to associate observations of the same object at different times as object trajectories. Furthermore, by combining visual feature pooling, category embedding, and geometric trajectory statistics within the object region, an object node set is constructed. In the event branch, based on Event recognition is performed on the sliding frame segment to detect operations such as grabbing, placing, inserting, twisting, tilting, pressing, opening and closing, and contact. An event node set is constructed by combining the trajectory of the participating object, the event type, and the start and end time periods. Further define the node set of the unified scene-event graph. And construct a candidate edge set by using three types of candidate nodes: object-object, event-object, and event-event. .
[0037] For any pair of candidate nodes Construction relationship features:
[0038] in For node features, Represents relative position, contact geometry, or spatial statistics of the objects involved. Indicates time interval, Indicates a contact or interaction marker. Input a relation classifier, output a relation score :
[0039] in , These are the weight matrix and bias vector of the relation classifier. And the maximum relation probability exceeds a threshold. Create typed edges to obtain the scene-event graph. :
[0040] To align the topic information in the video with the knowledge graph, topic vectors can be calculated for each video segment. Then, based on topic vectors and knowledge graph concept embedding The search is based on similarity. Based on weight... Before choosing Each knowledge node, followed by restrictions on the set of relation types, maximum hop count L, and width reserved per hop. Under certain conditions, restricted multi-hop expansion is performed to obtain a knowledge subgraph related to the current video content. This knowledge subgraph includes object attributes, action preconditions, action postconditions, structural constraints, etc. Subsequently, it is converted into rule constraints, including hard constraint masks. Continuous bias With candidate causal path set .in This indicates whether information exchange connections are allowed or prohibited. This indicates that additional biases are applied to node pairs that satisfy the rules. When injecting constraints into a joint network, the following attention form can be used:
[0041] Here, Γ is a sufficiently large positive number used to suppress connections that conflict with the rules. The candidate causal path set P is only used to constrain and rank training or candidate actions; the knowledge graph part does not directly output action values.
[0042] S3: Under weak labeling conditions, perform spectral analysis on the local temporal features of the action and encode them together with the spatiotemporal features into a potential action.
[0043] Specifically, S3 is a temporal-frequency augmentation latent action learning method for weakly labeled videos, which can be implemented in the following way: A latent action extractor is constructed using a dual-frame spatiotemporal Transformer architecture. Spectral analysis is performed on the temporal features of action-related local regions to jointly learn a time-frequency consistent low-dimensional continuous latent action representation from the video.
[0044] Specifically, to address the issue that human demonstration videos often lack fine-grained motion labels that strictly correspond to the robot's motion space, a dual-frame spatiotemporal Transformer architecture is employed to construct a latent motion extractor. Specifically, latent motion encoding is performed primarily on two consecutive first-person perspective frames. The latent motion from the previous frame is then used to conditionally reconstruct the subsequent frame, obtaining a unified proxy motion. (Latent Actions); The third-person perspective provides auxiliary information when there is significant occlusion, significant changes in overall posture, or large changes in the relative position of the object, enabling human videos without fine action labels to be used as training samples for action conditions.
[0045] In practice, human videos typically lack fine-grained motion labels consistent with the target robot's motion space. Therefore, this implementation first establishes a latent motion extractor. The latent motion extractor employs a two-frame spatiotemporal Transformer architecture based on the dual-view video encoding results, sharing features from two consecutive first-view frames output by the video encoder. and As the primary input, third-person perspective features are introduced as auxiliary context when needed. A spatiotemporal Transformer encoder, consisting of alternating spatial and temporal attention layers, extracts global motion features. Spatial multi-head self-attention models the spatial relationships within a single frame, while temporal multi-head self-attention models the dynamic changes across frames, ultimately yielding high-level motion representations of the first and third perspectives at frame t. and .
[0046] Building upon this, this embodiment further performs spectral analysis on the short-term temporal features of the action-related local regions. Preferably, a temporal segment of the action region is first constructed based on the hand region, the execution end neighborhood, the contact region, or the target object neighborhood, and local temporal features of the corresponding region are extracted within a continuous L-frame range. After regional pooling, cross-frame splicing, and normalization, a regional time-series signal is formed. Subsequently, a short-time Fourier transform, wavelet transform, or other frequency domain mapping is performed on the regional time-series signal along the time dimension to obtain frequency domain descriptions such as dominant frequency, spectral energy, bandwidth, and phase changes. For action segments with frequent local contact, significant oscillations, or strong rhythms, the short-time Fourier transform is preferred to characterize the frequency distribution within a fixed window.
[0047] in It is the time spectrum of the r-th action region. For the r-th region in the th... Local time-series signal at time 10:00. Indicates the center of the time window. Represents angular frequency. The window function is used; for action segments with large duration variations, uneven velocity variations, or significant local abrupt changes, continuous wavelet transform is preferred to obtain multi-scale time-frequency representation.
[0048] By combining the frequency domain coding results of the local regions related to the current action, a frequency domain action prior is formed. As intermediate prior conditions, these are used to enhance the latent action representation. Subsequently, the frequency-domain action priors are projected, aligned, and gatedly fused with the time-domain spatiotemporal features to form a joint time-frequency action representation. This representation, along with the knowledge constraint encoding, serves as auxiliary condition information for the latent action extraction stage. In the latent action extraction stage, first-person perspective action features are used as the main input, with the aforementioned auxiliary condition information as additional conditions. After outputting the posterior distribution parameters of the latent actions through the action head, low-dimensional continuous latent action vectors are obtained through reparameterization sampling. During the decoding stage, the previous frame... The encoding features are used as the query, with As a condition, motion information is injected and the next frame is reconstructed using cross-attention. Therefore, the action itself is first manifested as local changes near the execution end, and the first-person perspective can more stably reflect such changes; for operations with periodicity, oscillation or rhythm, such as twisting, stirring, reciprocating wiping, tapping, shaking, etc., frequency domain features can further enhance the ability to recognize subtle differences in action and repetitive action patterns; the third-person perspective is used to supplement the context when there is local occlusion or large changes in overall posture.
[0049] Based on the aforementioned latent action extraction and condition reconstruction process, the latent action extractor is jointly trained, and its training objective is defined as follows:
[0050]
[0051] In the formula, This represents the total loss from predicting potential actions, with the first term being the reconstruction term. It generates model parameters. It is a posteriori inference of network parameters. It is the expectation of the posterior distribution. The first term is the conditional generation distribution, used to improve the decoder's ability to predict the next frame given the previous frame and potential actions; the second term is the KL divergence constraint term. Indicates the weight of the KL term. It is the KL divergence. It is the posterior distribution of the latent actions. The first term is the prior distribution of latent actions, used to control the degree of compression of the latent action distribution to maintain a compact and continuous representation; the third term is the dual-view consistency term. The first term represents the loss and its weight for dual-view consistency, used to ensure that the potential actions extracted primarily from the first viewpoint are consistent with the overall context provided by the third viewpoint; the fourth term is the graph constraint consistency term. The first term represents the loss and its weight for graph constraint consistency, used to ensure that potential actions are consistent with the event semantics and relational states in the current scene-event graph; the fifth term is the action region enhancement term. These are the action region enhancement loss and its weights, used to enhance the model's attention to local action changes, contact region changes, and target object motion; the sixth term can be a frequency domain consistency term. These are the frequency domain consistency loss and its weight, used to constrain the consistency between the frequency domain representation of the potential action and the actual temporal changes in the action region in terms of dominant frequency, spectral energy distribution, or phase change, so as to enhance the stable representation of periodic and micro-dynamic actions.
[0052] S4: Integrate dual-view video semantics and task conditions to generate a continuous semantic conditional representation corresponding to the current time block and map it to a unified conditional representation of the world action model.
[0053] Specifically, S4 is the joint conditional representation construction of task language and video semantics. Based on the discrete knowledge constraints that have been formed, it further extracts the continuous semantic conditional representation corresponding to the current time block.
[0054] S4 includes: S4.1: Extract short-term local semantic summaries, long-term semantic guidance, and task language prompts from dual-view videos; S4.2: Extract short-term local semantic summaries, long-term semantic guidance, task language prompts, and dual-view coding features and jointly map them into a unified conditional representation to form a block-level conditional input aligned with the current trajectory block.
[0055] In the specific implementation process, after obtaining the latent action representation, in order to ensure that the latent action is no longer just an isolated control signal participating in subsequent modeling, but is always under the constraints of a clear task objective, object relationship and multi-view scene context, a joint representation of task language and video semantics is further constructed, which is used as the conditional input for the subsequent pre-training of the action conditional world model.
[0056] First, a two-level semantic representation is constructed for the dual-view video. The first level is a short-time local semantic summary. From the scene-event diagram The system obtains object, event, contact state, and local scene information, and forms object markers, event markers, contact markers, and scene markers. These markers are then pooled and fused to obtain... Short-term local semantic summaries are used to characterize the objects, events, and contact relationships involved in the current local operation, providing an explicit description of the proximate variables of the action intervention.
[0057] The second level is long-term semantic guidance. It is extracted from a frozen visual language model and used to supplement task stages, cross-temporal dependencies, and global scene constraints, providing semantic references for the distant causal background in the state evolution chain. In implementation, it integrates first-person video clips, third-person video clips, and task language cues. A common input-frozen visual language model is used to extract joint semantics in a unified cross-modal coding space.
[0058] The Visual Language Model (VLM) first segments and embeds image blocks from two video frames, then superimposes spatial and temporal location codes to form corresponding visual tag sequences. Next, the text cues are encoded by a tokenizer to obtain a text tag sequence. Finally, the two tag sequences are concatenated into a unified input sequence. The data is fed into the Transformer backbone of the VLM. After multiple layers of cross-modal interaction, the intermediate hidden states not only retain the image content itself, but also gradually incorporate high-level information such as task instructions, object attributes, spatial relationships, action stages, and long-term context. To simultaneously preserve shallow spatial details and deep semantic reasoning information, pyramid features are preferably extracted from multiple intermediate layers of the VLM. Each layer extracts corresponding features from the hidden states of visual and textual tags and performs pooling, then merges them into joint semantic features, and aligns them to the unified representation dimension required for subsequent conditional modeling through a projection layer. Then, the features from each layer are aggregated to form a long-term semantic guide aligned with the current time block. .thus, It is no longer a simple visual encoding result, but a semantic summary that includes object categories, action intentions, interaction relationships, task constraints, and long-term contextual information obtained from the video.
[0059] Short-time local semantic summarization and long-term semantic guidance Building upon this foundation, event-level, object-level, and relation-level semantic representations extracted from dual-view videos are further incorporated to construct a complete task semantic representation. .in, This is used to uniformly represent the task objectives, object relationships, interaction states, and semantic constraints corresponding to the current time block. It further enhances the complete task semantic representation. Together with the dual-view encoded features, the block-level conditional inputs of the subsequent action-conditional world model are constructed. The conditional representation of the i-th block is defined as follows: :
[0060] in This indicates that the conditional projection network can serve as a unified conditional input for the pre-training of subsequent action-conditional world models.
[0061] S5: Input the underlying action representation, unified condition representation, and structured prior basic constraints into the world action model, and introduce a continuous-time latent dynamics branch based on neural control differential equations during the basic pre-training stage.
[0062] Specifically, S5 is a continuous-time latent dynamics modeling of action conditions under structured prior constraints. It inputs the latent action representation, unified condition representation, and structured prior constraints into the world action model, and introduces a continuous-time latent dynamics branch based on neural control differential equations in the basic pre-training stage to learn the dynamic laws of the evolution of world state driven by latent actions.
[0063] S5 includes: First-person perspective features, third-person perspective features, frequency-domain enhanced latent action representations, and unified condition representations are uniformly input into the action-conditional world model. A continuous-time latent dynamics module based on action-conditional NCDE is introduced. The backbone of the continuous-time latent dynamics module uses an autoregressive conditional diffusion Transformer with causal modeling capabilities to uniformly model the dependencies between historical and current state blocks. Flow matching is used as the main loss, which enables the model to learn the continuous transmission direction from noisy state to real state and the fine-grained correspondence between actions and consequences.
[0064] In practice, after obtaining the frequency-domain enhanced latent actions and joint semantics, a basic pre-training of the action-conditional world model is performed. The purpose of this stage is not to directly output the target's action, but to learn the fundamental principles of "how latent actions drive changes in the world state," that is, to enable the model to possess stable state transition modeling capabilities given the current state, action, and task context. Its basic objective can be expressed as:
[0065] in, This represents the hidden state at time t. This indicates that potential actions are treated as explicit interventions in the current hidden state. Furthermore, this implementation introduces continuous-time latent dynamics modeling based on action-conditional non-deterministic evolution (NCDE), enhancing the modeling capability of continuous state evolution processes within a model block.
[0066] Let the original frame-level time be t', the trajectory block index be k, and the time step within the block be i'. Specifically, a trajectory is divided into M trajectory blocks, and a clean context is constructed based on the preceding trajectory blocks. The clean fusion of video latent representation and action latent representation is as follows: and The noise distribution of the video basis and motion basis is and For the current k-th trajectory block, the continuous video latent representation is divided into several time blocks according to the temporal compression ratio, and actions are also divided into corresponding action blocks according to the same time window. To reduce the action space span, relative action representation is preferred, and the actions corresponding to each latent representation frame are recoded based on the starting point of that latent representation frame, so that the model prioritizes local action increments rather than long-range absolute pose changes. At the same time, to reduce the interference of future actions on the current prediction, it is also preferred to inject actions in blocks according to the video temporal compression ratio, injecting only the action blocks within their time window into the corresponding latent representation frame. For the third-view branch, it is not allowed to generate actions independently, but only participates in state fusion and attention supplementation, thereby ensuring that the main chain of action learning is not dispersed.
[0067] In addition to discrete-time block-level modeling, a continuous-time latent dynamics branch is introduced. Let the continuous-time hidden state corresponding to this trajectory block be... Furthermore, the first-person video block features, third-person scene features, and frequency-domain enhanced potential actions are combined. And the joint semantics together construct a piecewise smooth control path ,in It can be obtained from natural cubic splines, Hermite cubic splines, or other piecewise differentiable interpolation methods. The k-th trajectory block in continuous time... Implicit state satisfy:
[0068] in, It is the start time of the kth trajectory block. Let be the hidden state at the start of the k-th trajectory block. () is the NCDE vector field function, and s is the integration variable. This is the control path for the k-th trajectory block. The NCDE solver is used to solve the path from a given initial state. Numerical integration of the above equations under the given conditions yields the terminal continuous-time state. This is used to characterize the continuous evolution process within the current trajectory block from the initial observation to the final observation. Thus, the model no longer relies solely on approximate state transitions at fixed discrete time steps, but can continuously advance the state changes within the block based on actual time intervals and action rhythms. When the observation time interval changes, the action speed is uneven, or there are obvious continuous transitions within the action block, the NCDE branch can provide smoother and more stable hidden state evolution priors. Preferably, for long-duration trajectory blocks or cases where the control path exhibits significant local high-frequency changes, a Log-NCDE form can be further adopted to compress and represent the higher-order cumulative changes of the control path in local sub-intervals, thereby preserving richer path sequence information and higher-order interaction information to enhance the expressive ability for long-term complex paths and higher-order change information.
[0069] Take the latent representation of the noisy fused video respectively With action potential representation , and conditional representation The time location identifier and modality type identifier are uniformly mapped to the same representation space, along with the preceding clean context. The autoregressive conditional diffusion Transformer backbone is input together with the input. Simultaneously, the continuous-time state trajectory summary and terminal state embedding obtained from the NCDE branch integration are mapped as continuous-time dynamics priors, which participate in subsequent recovery along with the current block input. This backbone performs causal modeling of the dependencies between "historical state block—current action block—current state block" at the time block level, ensuring that the noise recovery process of the current block can both reference the evolution trend of historical world states and be guided by the task objective, object relationships, and rule constraints. For each Transformer block, inter-layer modulation parameters generated by the long-term semantic inference branch through a small mapping network are introduced to apply sample-related semantic modulation, as follows:
[0070]
[0071] () represents the characteristic linear modulation operator, and u represents the feature to be modulated. It is the first Layer scaling vector, It is the first Layer bias vector, () is the first A layered modulation parameter generation network is used. The backbone achieves modeling through temporal causal self-attention and cross-modal conditional fusion. Within each layer, causal self-attention is first performed, followed by cross-attention based on historical context. Cross-modal attention establishes the coupling relationship between action blocks, state blocks, and conditional constraints. Finally, the reconstruction vector field of the current block is predicted using video reconstruction head and action reconstruction head. and It is used to characterize the direction in which the current video state and action state converge toward the distribution of the real target.
[0072] In terms of training objectives, flow matching is used as the main loss in the basic pre-training stage to learn the continuous transmission direction from the base distribution to the real data distribution.
[0073]
[0074] in It is the main loss of flow matching. and These represent the velocity field prediction network outputs for the video branch and the action branch, respectively. This is used to balance the contributions of action branches and video branches to the total loss. While using only the main loss described above allows the model to learn the direction of regression from noisy states to the target state, it is insufficient to adequately constrain the fine-grained correspondence between actions and consequences. Therefore, three auxiliary constraints are further introduced: firstly, temporal consistency loss. First, the "inter-frame difference of predicted velocity" must be consistent with the "inter-frame difference of actual velocity"; second, action alignment loss. Aligning the i-th frame block and the i-th action block in the intermediate feature space strengthens the correspondence between actions and consequences; thirdly, the condition space consistency term. This maintains a compact model of the current task context, object relationships, and constraints. The final overall pre-training objective definition is as follows:
[0075] Through this stage of training, the model acquires basic spatiotemporal priors on interactive physics, continuous motion of objects, contact sequence, and action consequences. The learned dynamic priors will serve as the initialization basis for subsequent joint world action model training.
[0076] S6: Based on the results of basic pre-training, the dynamic priors output by the continuous-time latent dynamics branch are invoked, and combined with causal constraints, counterfactual constraints, and execution constraints, joint video and action reasoning training is performed under the guidance of unified conditional representation to obtain an augmented world action model that integrates structured priors and continuous dynamics.
[0077] Specifically, S6 is a joint video-action reasoning system guided by structured priors and counterfactual conditions. Based on the basic pre-training results, joint video-action denoising training is performed on a per-trajectory-block basis.
[0078] S6 includes: Using the preceding trajectory block as a clean context, the current trajectory block representation, along with structured prior constraints, frequency domain action priors, continuous-time potential dynamics priors given by the NCDE branch, and counterfactual conditions, are input into the joint inference backbone of the world action model. Based on the historical context, joint inference is performed on the current block. Under the constraints of time causality, the model comprehensively utilizes object relationships, rule constraints, and counterfactual change information to recover the state evolution and action results of the current block, obtaining a future action representation that is consistent with the task conditions and guided by structured priors.
[0079] In the specific implementation process, based on the action conditional dynamics prior obtained in step S5, this stage further conducts joint video-action reasoning for the final action generation task. Given historical observations, semantic conditions, knowledge graph constraints, and counterfactual conditions, the world action model performs conditional recovery and generation of future H-step states and actions. For local regions related to the current action, frequency domain action priors can also be invoked. The continuous-time latent dynamics prior obtained from the NCDE branch enables the model to utilize not only semantic and causal information, but also rhythmic and periodic motion pattern information and intra-block continuous-time evolution information during the joint inference stage. Let... For historical observation, It is the future observation sequence from time l to l+H. It is the sequence of future actions from time l to l+H. This is a complete semantic representation extracted from the internal structure of the dual-view video. External language is an optional condition. This represents the current state of the entity. The structured graph prior for the current moment includes rule constraints, object relationships, and causal priors provided by the knowledge graph branches. For counterfactual intervention conditions, the corresponding condition generation distribution can be expressed as:
[0080] First, the scene-event graph extracted from the current dual-view video is fused with the knowledge subgraph to form a structured graph prior representation of the current moment. Perform graph encoding on it to obtain a graph representation. Furthermore, based on the consistency score between the action and the high-scoring causal path, the degree of violation of rules, constraints, and contact order, and the degree of matching with the main frequency stable interval, spectral energy concentration interval, and phase continuity, and combined with the intra-block state trajectory summary or terminal continuous-time state given by the NCDE continuous-time latent dynamics branch, candidate actions are soft-ranked, thereby forming the action feasible region prior at the current moment. Before generating an action, the graph and the ontology state jointly provide a soft prior that "the current action should roughly fall into which feasible domain". This allows the space for the current action to be narrowed in the early stages of action recovery, rather than performing posterior filtering only at the end.
[0081] In addition, counterfactual reasoning is incorporated for the structured intermediate representation, which is then fed back into the video-action network inference. Specifically, this is applied to the prior representation of the current structured graph. Applying explicit counterfactual intervention conditions The structured graph sequence after intervention was obtained. in, This can be represented as changes in object attributes, object removal, relationship changes, target changes, or changes in the action's feasible domain. Therefore, the fact trajectory graph representation encoding and the counterfactual reasoning graph representation encoding are fed into the joint video-action reasoning, enabling the model to simultaneously learn factual and counterfactual trajectory recovery, thereby obtaining the counterfactual consistency capability of "how the future state and future action of the counterfactual trajectory should change synchronously when the object, relationship, or target conditions change."
[0082] Considering that video and action branches have different noise tolerances, especially under low-step inference conditions, action output typically needs to converge to an executable state more quickly, while video representation can still retain a certain degree of uncertainty. If both modalities still share the same noisy time step, action prediction will be constrained by insufficiently cleaned visual conditions, leading to decreased action stability under low-step inference. Therefore, this implementation employs decoupled noise scheduling for video and action. Action time steps remain uniformly distributed, while a high-noise bias continues to be applied to the video time step, but this bias strength is further correlated with the actual number of inference steps. When the system uses fewer denoising steps... At this time, the video time step will be pushed into a higher noise region to allow the model to adapt in advance to the operating condition that "the vision is still not completely clear but the action needs to be reliably given"; when the inference steps As the bias increases, it automatically weakens, thus gradually returning to the normal joint denoising training distribution.
[0083] Correspondingly, let the video time step and motion time step of the k-th block be respectively:
[0084] in
[0085] Furthermore, at this decoupling time step, the noisy states of the video and the action are respectively constructed as follows:
[0086]
[0087] , It is the noisy state of the k-th video and motion. , It's a clean video, showing proper motion. , It is the video and motion branch time step. , Video and motion noise base state It is a power-law coefficient. For the video branch, linear interpolation is performed between the clean video latent representation and the video basis distribution noise according to the video time step; for the action branch, in order to make it converge to the clean action end faster under the same time step, a power-law type time reparameterization is introduced to improve the speed at which the action representation converges to the clean action end. Thus, in one-step or less-step denoising scenarios, the intermediate states in the inference process will be more concentrated in the region where "the video information still contains noise but the action is close to being recoverable", thereby improving the robustness of action generation under low-latency deployment.
[0088] In the k-th trajectory block, the noisy video and noisy action latent representations in the current block are uniformly mapped to the same representation space along with the semantic conditional representation, ontology state, knowledge graph encoding results, frequency domain action prior encoding results, and counterfactual intervention encoding results. Temporal location identifiers and modality type identifiers are added to form the unified input sequence for the current block, which is then fed into the autoregressive conditional diffusion Transformer. In the k-th trajectory block... The layer performs attention computation with temporal causality, knowledge graph rule bias, and action selective masking. The current action block uses selective attention masking, accessing only temporally close and causally related preceding action representations, rather than indiscriminately absorbing all old action information. The current noisy video block and noisy action block first undergo self-attention interaction under strict causal constraints, then query clean history blocks, knowledge graph memory, and counterfactual intervention graphs. This allows the reconstruction to reference not only the evolutionary trend of the historical world state but also the current task rules, object relationships, and causal path information, thereby determining the appropriate shifts in future states and actions under counterfactual conditions. After stacking multiple Transformer layers, the future video and future actions of the current block are recovered, ultimately outputting three types of coupled results: the video reconstruction field of the current block. , movement recovery field and video-action coordination consistency field . Characterizes the direction of recovery from the current video state to the true future video state. This represents the recovery direction predicted by the joint network on the action branch, used to characterize the direction in which the current noisy action representation converges towards the clean target action representation. This is used to constrain the consistency between the action recovery direction and the future visual evolution direction, as well as the frequency domain rhythm constraints. The action recovery field is obtained. Then, based on this, the current noisy action representation can be restored to a clean action. express: And through the motion decoder You can get the future action block of the current block. : .
[0089] In one implementation, the method further includes S7: asynchronous closed-loop execution based on an enhanced world action model that integrates structured priors and continuous dynamics, specifically: Before deployment, the augmented world action model is subjected to few-step distillation and rapid reparameterization to form two switchable inference modes: a few-step joint generation mode and a rapid action mode. The augmented world action model is deployed as an online inference module in the closed-loop system, forming an asynchronous parallel working mechanism with the execution module.
[0090] Specifically, S7 is based on asynchronous closed-loop execution and real-time deployment of real observation write-back. It performs low-step distillation and rapid reparameterization on the augmented world action model, and performs asynchronous closed-loop execution and deployment based on real observation write-back. The new real observations and ontology state write-back cache suppress error accumulation.
[0091] To adapt to real-time closed-loop deployment, after completing the joint training of S6, the joint model undergoes further processing including few-step distillation and rapid reparameterization. This implementation sets two switchable operating modes for the same world action model during the inference and deployment phase: a "few-step joint generation mode" and a "rapid action mode." At any given time, the system preferentially selects one of these modes as the current primary execution mode. The few-step joint generation mode is used when the task has high requirements for future state consistency, complex operation planning, or constraints on action consequences; the rapid action mode is used when the task has higher requirements for online response latency.
[0092] In the few-step joint generation mode, an autoregressive diffusion teacher model is obtained through teacher-forced training, and PF-ODE trajectories are sampled given a true history prefix. Let the joint state of the k-th trajectory block be denoted as... .in, This represents the potential representation block of the current video. This represents the potential representation block of the current action. Then, it uses intermediate noise states in the trajectory. and its corresponding clean targets As a supervised sample, for causal students Perform ODE initialization training:
[0093] in, It is the ODE initialization training loss. For expectation operator, This represents the causal context formed by the preceding true trajectory blocks. For the conditional representation of the world model, This represents the current ontology state. This training allows the few-step student model to learn the basic ability to recover the current clean joint block from the current noisy joint block while maintaining the causal autoregressive structure. Subsequently, a few-step distribution matching distillation is performed to ensure that the student-generated distribution continues to approximate the target distribution. Let the student-generated joint state object be written as... The corresponding joint distribution matching distillation objective representation is:
[0094] This represents the gradient of the joint distribution matching distillation objective with respect to the student model parameters θ. The score function is the true distribution. By using the difference between the true distribution score and the student distribution score, the student model is subjected to distribution correction, so that the student model can not only maintain future visual consistency during short-step reasoning, but also maintain the executability of action blocks and the correspondence between them and visual consequences.
[0095] In fast-action mode, the student model can be reparameterized into a fast inference structure consisting of a world encoder and a motion expert decoder. During the training phase, the video co-training objective is retained to encourage the world encoder to learn physical and temporal representations related to actions. However, during the online inference phase, instead of explicitly performing multi-step denoising generation of future videos, the current world representation is obtained through a single forward encoding using current observations, ontology state, cached context, and relevant prior information. The motion expert head then directly outputs the next action block. Therefore, the main value of video modeling lies in shaping the world representation during the training phase, without requiring explicit generation of future videos in each control cycle during online execution.
[0096] After completing the aforementioned steps-based deployment preparations, the system enters the online closed-loop execution phase. In each closed-loop cycle, the system starts from the joint noise state under the current context conditions, performs several steps of joint denoising, and obtains the clean action of the current trajectory block and the corresponding potential representation of the future video. In addition, the system can also call the NCDE continuous-time dynamics branch based on the actual time interval between two adjacent real observations to continuously advance the hidden state within the block, obtaining a continuous-time dynamics prior that matches the current time span; for long-term complex control paths, the Log-NCDE form can be further used to enhance the expression of higher-order changes and long-term dependencies. The potential representation of the future video is only used as an intermediate future constraint for action prediction to help the model maintain consistency between actions and future state changes, and is not directly continued as the state of the next round of real environment.
[0097] After obtaining the current action block, the system preferably smooths the action sequence before sending it to the execution layer. To suppress the risk of historical erroneous actions propagating to subsequent action blocks during autoregressive action generation, selective action attention masking is introduced during the action block-level generation process to mitigate error accumulation.
[0098] After action blocks are issued, the execution and inference modules preferably operate in an asynchronous parallel manner. That is, the controller continuously executes the most recently generated action block, while the inference module concurrently computes the next action block in the background using the latest available observations. In this way, real-time distillation primarily reduces the effective denoising steps required for a single generation round, while the runtime asynchronous parallel mechanism reduces idle time spent waiting for the current action to execute. Simultaneously, the system can further reduce redundant computations in each closed-loop round and the risk of visual residual noise interference in action branches by combining cached state reuse, causal attention forward sharing, and video and action decoupling noise scheduling, ensuring that the model runs continuously and stably according to the action block rhythm in a real closed-loop system.
[0099] After the current action block finishes executing, the system reads the new real first-person perspective observation. and the state of the body and will Encoded as Write back to cache; if a new third-person view is observed If available, then synchronously encode it as... The fusion state is then updated. Simultaneously, the predicted video latent representations used only for intermediate planning in the previous round are discarded and no longer directly used as the world state for the next round. Therefore, subsequent inference by the system is always anchored to the actual perceived results, rather than the pseudo-environment states generated by the model itself.
[0100] The world action model of this invention does not output a simple future video sequence, but rather a next action or short-term action block that can be directly invoked by the execution layer. Upon receiving an action or action block, the agent's execution layer can execute it through a joint space controller, a Cartesian trajectory tracker, an impedance controller, or other underlying robot control modules. During execution, if the external environment changes, the perception module will collect new real-world visual observations and the agent's state, and write them back to the world action model's cache context in the next inference cycle.
[0101] Example 2 Based on the same inventive concept, this embodiment discloses a world action model enhancement device that integrates structured priors and continuous dynamics. Please refer to [link to relevant documentation]. Figure 2 ,include: The dual-view collaborative spatiotemporal alignment input construction module 101 is used to construct the dual-view collaborative spatiotemporal alignment input, wherein the dual-view collaborative spatiotemporal alignment input includes a first-view video and a third-view video. The basic constraint construction module 102 of the structured prior is used to construct structured prior constraints by using the first-view video as the main input for action learning and the third-view video as the supplementary input for the scene. The time-frequency enhanced latent action coding module 103 is used to perform spectral analysis on the local temporal features of an action under weak labeling conditions, and jointly encode them with spatiotemporal features into a latent action; The joint conditional representation construction module 104 is used to fuse dual-view video semantics and task conditions to generate a continuous semantic conditional representation corresponding to the current time block, and map it to a unified conditional representation of the world action model. The dynamics modeling module 105 is used to input the underlying action representation, unified condition representation and structured prior basic constraints into the world action model, and introduce a continuous-time underlying dynamics branch based on neural control differential equations in the basic pre-training stage. The joint video-action reasoning module 106 is used to call the dynamic priors output by the continuous-time latent dynamics branch based on the results of basic pre-training, and combine causal constraints, counterfactual constraints and execution constraints to perform joint video and action reasoning training under the guidance of unified conditional representation, so as to obtain an augmented world action model that integrates structured priors and continuous dynamics.
[0102] In one embodiment, the apparatus further includes an asynchronous closed-loop execution module for: performing asynchronous closed-loop execution based on an enhanced world action model that integrates structured priors and continuous dynamics, specifically: Before deployment, the augmented world action model is subjected to few-step distillation and rapid reparameterization to form two switchable inference modes: a few-step joint generation mode and a rapid action mode. The augmented world action model is deployed as an online inference module in the closed-loop system, forming an asynchronous parallel working mechanism with the execution module.
[0103] Since the device in Embodiment 2 of this invention is the same device used in the world action model enhancement method that integrates structured priors and continuous dynamics in Embodiment 1, those skilled in the art can understand the specific structure and variations of this device based on the method described in Embodiment 1 of this invention, and therefore will not be repeated here. All devices used in the method of Embodiment 1 of this invention fall within the scope of protection of this invention.
[0104] Example 3 Based on the same inventive concept, the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in Embodiment 1.
[0105] Since the computer device described in Embodiment 3 of this invention is the same computer device used to implement the world action model enhancement method integrating structured priors and continuous dynamics in Embodiment 1 of this invention, those skilled in the art can understand the specific structure and variations of this computer device based on the method described in Embodiment 1 of this invention, and therefore will not be repeated here. All computer devices used in the method of Embodiment 1 of this invention fall within the scope of protection of this invention.
[0106] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0107] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0108] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various modifications and variations to the embodiments of the invention without departing from the spirit and scope of the invention. Thus, if these modifications and variations of the embodiments of the invention fall within the scope of the claims of the invention and their equivalents, the invention also intends to include these modifications and variations.
Claims
1. A method for enhancing world action models by integrating structured priors and continuous dynamics, characterized in that, include: Construct a spatiotemporally aligned input for dual-view collaboration, wherein the spatiotemporally aligned input for dual-view collaboration includes first-view video and third-view video; Using first-person perspective videos as the main input for action learning and third-person perspective videos as supplementary scene input, a structured prior basic constraint representation is constructed. Under weak labeling conditions, spectral analysis is performed on the local temporal features of the action, and the results are jointly encoded with the spatiotemporal features into a potential action. The system integrates dual-view video semantics and task conditions to generate a continuous semantic conditional representation corresponding to the current time block, and maps it to a unified conditional representation of the world action model. The underlying action representation, unified conditional representation, and structured prior basic constraints are input into the world action model, and a continuous-time latent dynamics branch based on neural control differential equations is introduced in the basic pre-training stage. Based on the results of basic pre-training, the dynamic priors output by the continuous-time latent dynamics branch are invoked, and combined with causal constraints, counterfactual constraints, and execution constraints, joint video and action reasoning training is performed under the guidance of unified conditional representation to obtain an augmented world action model that integrates structured priors and continuous dynamics.
2. The method for enhancing the world action model by fusing structured priors and continuous dynamics as described in claim 1, characterized in that, Constructing a spatiotemporally aligned input for dual-view collaboration includes: Collect human first-person and third-person perspective video sequences; Frame sampling, segmentation, timestamp alignment, and unified encoding are performed on the acquired first-view and third-view motion video sequences. The encoded dual-view spatiotemporal feature sequences are used as spatiotemporal alignment inputs for dual-view collaboration. The first-view video is used to provide local motion changes and fine operation processes, while the third-view video is used to provide overall spatial relationships and the state of the scene outside occlusion.
3. The method for enhancing the world action model by integrating structured priors and continuous dynamics as described in claim 1, characterized in that, Using first-person perspective video as the primary input for action learning and third-person perspective video as supplementary scene input, a structured prior constraint is constructed, including: Perform object detection, instance tracking and event recognition on dual-view videos to obtain entity nodes, event nodes and their relational edges, and construct a scene-event graph corresponding to the current video segment; Retrieve knowledge graphs based on video topics, object sets, and event sets to form knowledge subgraphs related to the current video; By transforming the knowledge subgraph into a rule mask, continuous bias, and candidate causal path scores, the knowledge subgraph is converted into discrete structural constraints, which serve as the basic constraint representation of the structured prior.
4. The method for enhancing the world action model by fusing structured priors and continuous dynamics as described in claim 1, characterized in that, Under weak labeling conditions, spectral analysis is performed on the local temporal features of actions, and these features are jointly encoded with spatiotemporal features to form potential actions, including: A latent action extractor is constructed using a dual-frame spatiotemporal Transformer architecture. Spectral analysis is performed on the temporal features of action-related local regions to jointly learn a time-frequency consistent low-dimensional continuous latent action representation from the video.
5. The method for enhancing the world action model by fusing structured priors and continuous dynamics as described in claim 1, characterized in that, Based on the structured prior constraint representation, the continuous semantic condition representation corresponding to the current time block is extracted and mapped to a unified condition representation of the world action model, including; Extracting short-term local semantic summaries, long-term semantic guidance, and task language prompts from dual-view videos; The extracted short-term local semantic summary, long-term semantic guidance, task language prompts, and dual-view coding features are jointly mapped into a unified conditional representation of the world action model, forming a block-level conditional input aligned with the current trajectory block.
6. The method for enhancing the world action model by fusing structured priors and continuous dynamics as described in claim 1, characterized in that, The underlying action representation, unified conditional representation, and structured prior constraints are jointly input into the world action model. During the basic pre-training phase, a continuous-time latent dynamics branch based on neural control differential equations is introduced, including: First-person perspective features, third-person perspective features, frequency-domain enhanced latent action representations, and unified conditional representations are input into the world action model. A continuous-time latent dynamics branch based on neural control differential equations is introduced. The main branch of the continuous-time latent dynamics branch uses an autoregressive conditional diffusion Transformer with causal modeling capabilities to uniformly model the dependencies between historical and current state blocks. Flow matching is used as the main loss, which enables the model to learn the continuous transmission direction from noisy state to real state and the fine-grained correspondence between actions and consequences.
7. The method for enhancing the world action model by fusing structured priors and continuous dynamics as described in claim 6, characterized in that, Based on the results of basic pre-training, the dynamic priors output by the continuous-time latent dynamics branch are invoked, and combined with causal constraints, counterfactual constraints, and execution constraints, joint video and action reasoning training is performed under the guidance of a unified conditional representation to obtain an augmented world action model that integrates structured priors and continuous dynamics, including: Using the preceding trajectory block as a clean context, the current trajectory block representation, along with structured prior constraints, frequency domain action priors, continuous-time latent dynamics priors given by the continuous-time latent dynamics branch, and counterfactual conditions, are input into the joint inference backbone of the world action model. Based on the historical context, joint inference is performed on the current block. Under the constraints of temporal causality, the model comprehensively utilizes object relationships, rule constraints, and counterfactual change information to recover the state evolution and action results of the current block, obtaining a future action representation that is consistent with the task conditions and guided by structured priors.
8. The method for enhancing the world action model by fusing structured priors and continuous dynamics as described in claim 1, characterized in that, The method also includes asynchronous closed-loop execution based on an enhanced world action model that integrates structured priors and continuous dynamics, specifically: Before deployment, the augmented world action model is subjected to few-step distillation and rapid reparameterization to form two switchable inference modes: a few-step joint generation mode and a rapid action mode. The augmented world action model is deployed as an online inference module in the closed-loop system, forming an asynchronous parallel working mechanism with the execution module.
9. A world action model enhancement device integrating structured priors and continuous dynamics, characterized in that, include: A dual-view collaborative spatiotemporal alignment input construction module is used to construct a dual-view collaborative spatiotemporal alignment input, wherein the dual-view collaborative spatiotemporal alignment input includes a first-view video and a third-view video. The basic constraint construction module for structured priors is used to construct basic constraints for structured priors by using first-view video as the main input for action learning and third-view video as a supplementary input for the scene. The time-frequency enhanced latent action coding module is used to perform spectral analysis on the local temporal features of actions under weak labeling conditions, and jointly encode them with spatiotemporal features into latent actions; The joint semantic condition representation construction module is used to integrate dual-view video semantics and task conditions to generate a continuous semantic condition representation corresponding to the current time block, which serves as a unified condition representation. The dynamics modeling module is used for basic pre-training of the world action model for action conditionation. It introduces a continuous-time latent dynamics branch based on neural control differential equations to learn the dynamic laws of latent actions driving the evolution of world states. The joint video-action reasoning module is used to call the dynamic priors output by the continuous-time latent dynamics branch based on the results of basic pre-training, and combine causal constraints, counterfactual constraints and execution constraints to perform joint video and action reasoning training under the guidance of unified conditional representation, so as to obtain an augmented world action model that integrates structured priors and continuous dynamics.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the world action model enhancement method that integrates structured priors and continuous dynamics as described in any one of claims 1 to 8.