World model based control method and system

CN122652997APending Publication Date: 2026-08-28GUOHAO CARBON (BEIJING) ENERGY TECH RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610840981.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

然而,这种方法存在以下缺陷:首先,像素级预测需要处理数千万级别的视觉特征块,计算量和内存占用呈二次方增长,严重限制了实时应用能力;其次,基于相关性的预测方法缺乏对物理因果关系的理解,在分布偏移的测试数据上表现不佳;此外,传统模型在长时序规划中面临严重的累积误差问题,如基线方法在五步任务序列中成功率从第一步的 93.3% 骤降至第五步的 51.1%

Benefits of technology

[0027]The control method and system based on the world model provided in this invention generate current causal constraints based on a perception dataset including tactile data and a perception causal relationship structure. The current perception dataset is mapped to a causal prediction latent space to obtain current causal latent features. Future causal latent states are inferred by combining historical time-series features and causal matrices. Feature weight parameters are calculated using a task-oriented dynamic attention mechanism to filter effective latent features. Multi-source data fusion and multi-worldline future state inference are performed to obtain the causal latent features of the target future state. Control commands are generated to complete the corresponding operations. This method is based on three-dimensional heterogeneous data input of visual appearance, tactile physics, and language rules to achieve comprehensive environmental information collection from explicit visual observation, intrinsic physical properties, and high-level rule constraints. It solves the defect of pure visual models that cannot perceive the material, force, and physical constraints of objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122652997A_ABST
    Figure CN122652997A_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of control method and system based on world model, it is related to automatic control technical field, the method generates current causal constraint based on the perception data set including tactile data and the perception causal relationship structure, current perception data set is mapped to causal prediction latent space, obtain current causal latent feature, combine historical time series characteristics and causal matrix and deduce future causal latent state, utilize task-oriented dynamic attention mechanism to calculate feature weight parameter, filter effective latent feature, carry out multi-source data fusion and multi-world line future state deduction, obtain target future state causal latent feature, generate control instruction and complete corresponding operation, the method is based on the three-dimensional heterogeneous data input of visual appearance, tactile physics and language rule, realize all-around environmental information collection of explicit visual observation, intrinsic physical property, high-level rule constraint, solve the defect that pure vision model cannot perceive object material, stress, physical constraint.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of control technology based on world models, and specifically to a control method and system based on world models. Background Technology

[0002] Traditional world models primarily rely on pixel-level prediction, directly forecasting future video frames using deep learning models. However, this approach suffers from several drawbacks: First, pixel-level prediction requires processing tens of millions of visual feature blocks, leading to a quadratic increase in computational and memory usage, severely limiting real-time application capabilities. Second, correlation-based prediction methods lack an understanding of physical causality and perform poorly on test data with distribution shifts. Furthermore, traditional models face severe cumulative error problems in long-term planning; for example, the success rate of baseline methods in a five-step task sequence drops sharply from 93.3% in the first step to 51.1% in the fifth step. Summary of the Invention

[0003] To address the aforementioned problems, embodiments of the present invention provide a control method and system based on a world model.

[0004] In a first aspect, embodiments of the present invention provide a control method based on a world model, comprising:

[0005] The current sensory dataset is acquired in real time, and the current sensory dataset includes current visual data, current tactile data, and current language data;

[0006] Based on the current sensing dataset, current causal constraints are generated using a sensing causal relationship structure, which includes the causal relationship between the previously sensed data and the subsequently sensed data.

[0007] The current perception dataset is mapped to the causal prediction latent space to obtain the current causal latent features. The future causal latent state is inferred by combining the historical time series features and the causal matrix in the perception causal relationship structure.

[0008] Based on real-time task instructions, feature weight parameters are calculated using a task-oriented dynamic attention mechanism, and effective latent features are selected.

[0009] Multi-source data fusion and multi-worldline future state deduction are performed on the current causal constraints, the current causal latent features, the future causal latent states, the feature weight parameters, and the effective latent features to obtain the causal latent features of the target future state;

[0010] Control instructions are generated based on the causal latent features of the target future state, and the corresponding operations are performed using the control instructions.

[0011] Furthermore, before generating the current causal constraints using the perceptual causal relationship structure based on the current perceptual dataset, the method further includes: constructing the perceptual causal relationship using a differentiable directed acyclic graph.

[0012] Furthermore, before mapping the current perceptual dataset to the causal prediction latent space, the method further includes: constructing the causal prediction latent space using the perceptual causal relationship structure, wherein the causal prediction latent space includes static causal latent features and dynamic causal latent features, wherein the static causal latent features are used to describe static structural causal features, and the dynamic causal latent features are used to describe dynamic motion causal features.

[0013] Furthermore, before calculating the feature weight parameters using the task-oriented dynamic attention mechanism according to the real-time task instructions, the method further includes: constructing the task-oriented dynamic attention mechanism based on the language data, the perceptual causal relationship structure, and the causal prediction latent space.

[0014] Furthermore, constructing the causal prediction latent space using the perceptual causal relationship structure includes: acquiring a historical perceptual dataset, which includes historical visual data, historical tactile data, and historical language data; performing two-layer, two-branch latent feature encoding based on the historical perceptual data in the historical perceptual dataset; causal-constrained attention mapping temporal modeling; and constructing a temporal prediction and loss function.

[0015] Furthermore, the two-layer, two-branch latent feature encoding based on the historical sensing data in the historical sensing dataset includes: through... Construct a 32-dimensional static causal latent feature, in which, It is a static causal latent feature. These are standardized historical visual data, historical tactile data, and historical language data; through Construct a 32-dimensional dynamic causal latent feature, in which, As a dynamic causal latent feature, These are multimodal temporal historical visual data, historical tactile data, and historical language data; through... A 64-dimensional total causal latent feature is constructed, and the historical perception data in the historical perception dataset is subjected to dimensionality reduction processing so that the feature data volume of the historical perception data reaches a preset threshold. This represents the overall static causal latent characteristic.

[0016] Furthermore, constructing the perceived causal relationship using a differentiable directed acyclic graph includes: constructing multimodal causal variables based on the historical perception dataset to achieve the structuring of causal nodes in the physical scene; and generating a dynamic causal matrix based on the historical perception dataset.

[0017] Furthermore, the dynamic causal matrix is ​​constructed using the following formula: A = Sinkhorn((log(α) + ε) / τ), where α∈RN×N, representing the learnable parameter matrix, which represents the potential causal strength between variables; ε~Gumbel (0,1) is a noise term used to introduce randomness to optimize the discrete structure search; τ represents the annealing temperature parameter used to control the smoothness of the matrix, and when τ→0, the matrix converges to a hard double random matrix; Sinkhorn (・) represents the Sinkhorn operator used to convert the input matrix into a double random matrix through iterative row and column normalization.

[0018] Furthermore, constructing the perceived causal relationship using a differentiable directed acyclic graph also includes:

[0019] Construct a differentiable directed acyclic constraint loss function; construct a total loss function based on the directed acyclic constraint loss function.

[0020] Secondly, embodiments of the present invention provide a control system based on a world model, comprising:

[0021] The data acquisition module is used to acquire the current sensory dataset in real time, which includes current visual data, current tactile data and current language data;

[0022] The causal constraint acquisition module is used to generate current causal constraints based on the current perception dataset and using a perception causal relationship structure, wherein the perception causal relationship structure includes the causal relationship between the previously perceived data and the subsequently perceived data.

[0023] The causal latent feature acquisition module is used to map the current perception dataset to the causal prediction latent space to obtain the current causal latent features, and combine the historical time series features and the causal matrix in the perception causal relationship structure to infer the future causal latent state.

[0024] The latent feature filtering module is used to calculate feature weight parameters and filter effective latent features based on real-time task instructions and a task-oriented dynamic attention mechanism.

[0025] The multi-source data fusion module is used to perform multi-source data fusion and multi-worldline future state inference on the current causal constraints, the current causal latent features, the future causal latent states, the feature weight parameters, and the effective latent features to obtain the causal latent features of the target future state.

[0026] The control execution module is used to generate control instructions based on the causal latent features of the target's future state, and to use the control instructions to complete the corresponding operations.

[0027] The control method and system based on the world model provided in this invention generate current causal constraints based on a perception dataset including tactile data and a perception causal relationship structure. The current perception dataset is mapped to a causal prediction latent space to obtain current causal latent features. Future causal latent states are inferred by combining historical time-series features and causal matrices. Feature weight parameters are calculated using a task-oriented dynamic attention mechanism to filter effective latent features. Multi-source data fusion and multi-worldline future state inference are performed to obtain the causal latent features of the target future state. Control commands are generated to complete the corresponding operations. This method is based on three-dimensional heterogeneous data input of visual appearance, tactile physics, and language rules to achieve comprehensive environmental information collection from explicit visual observation, intrinsic physical properties, and high-level rule constraints. It solves the defect of pure visual models that cannot perceive the material, force, and physical constraints of objects. Attached Figure Description

[0028] The accompanying drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the embodiments and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:

[0029] Figure 1 A flowchart illustrating a control method segment based on a world model provided in an embodiment of the present invention;

[0030] Figure 2 A flowchart illustrating another control method based on a world model provided in an embodiment of the present invention;

[0031] Figure 3 This is a functional structure diagram of a world model-based control system provided in an embodiment of the present invention. Detailed Implementation

[0032] To enable those skilled in the art to better understand the technical solutions of the embodiments of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the embodiments of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0033] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.

[0034] As used herein, the term “and / or” includes any and all combinations of one or more of the associated enumerated entries. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the embodiments disclosed herein. As used herein, the singular forms “a” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Terms such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0035] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and embodiments of this disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.

[0036] The world-model-based control method of this disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be an in-vehicle device, user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in memory. Alternatively, the method can be executed by a server.

[0037] The present invention will be further described below with reference to the accompanying drawings. It should be understood that the following embodiments are only for illustrating the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Various modifications or substitutions made to the present invention by those skilled in the art without departing from the spirit and substance of the present invention should fall within the scope of protection of the present invention.

[0038] Existing multimodal world models have the following shortcomings: First, most methods only focus on the visual modality and ignore the key role of tactile feedback in understanding physical interactions; second, causal discovery algorithms usually rely on discrete optimization, resulting in non-differentiability and low computational efficiency; third, they lack effective task-oriented mechanisms and cannot dynamically adjust prediction strategies according to specific tasks.

[0039] Based on the current state of technology, the inventors have innovatively proposed a control scheme based on a world model. This scheme uses three types of multimodal heterogeneous physical observation data—visual, tactile, and linguistic—to cover scene apparent spatial information, intrinsic mechanical and physical information, and high-level task rule information, forming a complete environmental physical representation system. This provides comprehensive data support for causal modeling and state inference, and can optionally supplement with temporal interactive action data to improve the causal link of action intervention in environmental changes.

[0040] To achieve the above objectives, embodiments of the present invention provide a control method based on a world model, such as... Figure 1 As shown, the method includes:

[0041] S11. Real-time acquisition of the current perception dataset, which includes current visual data, current tactile data, and current language data.

[0042] S12. Based on the current perception dataset, generate the current causal constraints using the perception causal relationship structure. The perception causal relationship structure includes the causal relationship between the previously perceived data and the subsequently perceived data.

[0043] S13. Map the current perception dataset to the causal prediction latent space to obtain the current causal latent features. Combine the historical time series features and the causal matrix in the perception causal relationship structure to infer the future causal latent state.

[0044] S14. Based on the real-time task instructions, calculate the feature weight parameters using the task-oriented dynamic attention mechanism, and screen effective latent features.

[0045] S15. Perform multi-source data fusion and multi-worldline future state deduction on the current causal constraints, current causal latent features, future causal latent states, feature weight parameters, and effective latent features to obtain the causal latent features of the target future state.

[0046] S16. Generate control instructions based on the causal latent features of the target's future state, and use the control instructions to complete the corresponding operations.

[0047] The world-model-based control method provided in this embodiment generates current causal constraints based on a perception dataset including tactile data and a perception causal relationship structure. It maps the current perception dataset to a causal prediction latent space to obtain current causal latent features. Combining historical time-series features and a causal matrix, it infers future causal latent states. It uses a task-oriented dynamic attention mechanism to calculate feature weight parameters, filters effective latent features, performs multi-source data fusion and multi-worldline future state inference, obtains the causal latent features of the target future state, and generates control commands to complete the corresponding operations. This method is based on three-dimensional heterogeneous data input of visual appearance, tactile physics, and language rules to achieve comprehensive environmental information collection from explicit visual observation, intrinsic physical properties, and high-level rule constraints, solving the deficiency of pure visual models in being unable to perceive object material, force, and physical constraints.

[0048] Among related technologies, multimodal world models mainly suffer from the following shortcomings: First, they only focus on the visual modality and ignore the key role of tactile feedback in understanding physical interactions; second, causal discovery algorithms usually rely on discrete optimization, resulting in non-differentiability and low computational efficiency; third, they lack an effective task-oriented mechanism and cannot dynamically adjust prediction strategies according to specific tasks.

[0049] As an improvement to the above embodiments, this invention provides another control method based on a world model. Through multimodal causal graph learning, causal latent space prediction, and a task-oriented dynamic attention mechanism, it achieves future state deduction by fusing multi-source data from visual, tactile, and linguistic sources. This method relates to the fields of artificial intelligence and robotics, and can be widely applied to intelligent robots, autonomous driving, virtual reality, etc. Figure 2 As shown, the method includes:

[0050] S21. Differentiable directed acyclic graphs are used to construct perceptual causal relationships.

[0051] Specifically, in some optional embodiments, S21 can be implemented through the following process (not shown in the figure):

[0052] S211. Construct multimodal causal variables based on historical perception datasets to achieve the structuring of causal nodes in physical scenes.

[0053] It should be noted that the perception data in the embodiments of this invention includes both historical perception data and current perception data, both of which include multimodal heterogeneous physical observation data of three types: visual, tactile, and linguistic. This covers scene apparent spatial information, intrinsic mechanical and physical information, and high-level task rule information, constituting a complete environmental physical representation system, providing comprehensive data support for causal modeling and state inference.

[0054] Visual data includes temporal RGB video frames and depth images, which are unstructured temporal spatial appearance data representing the outlines, positions, textures, and spatial layouts of objects in a scene. Tactile data includes 3D contact point clouds, triaxial force, and triaxial torque sensing data, which are semi-structured mechanical observation data representing the contact forces, deformations, and material physical properties of objects. Language data includes natural language task instructions and physical rule description text, which are discrete symbolic semantic data representing high-level task objectives, physical interaction constraints, and prior operational rules.

[0055] In this embodiment, to address the issue of inconsistent sampling frequencies for visual, tactile, and language data, global temporal synchronization is achieved using inverse time-distance weighted interpolation.

[0056] , It is the sampling time of the reference visual frame, serving as a globally unified timing clock. These are the original sampling moments for tactile and language data. It is unaligned raw multimodal sensor data. It is the set of original sampling points near the reference time. It is a time-distance inverse weighted coefficient to achieve smooth temporal interpolation alignment.

[0057] To eliminate dimensional differences, the scales of two-dimensional pixel coordinates and three-dimensional world coordinates are unified in the following way:

[0058] , It is the original spatial coordinate data. These are the extreme values ​​of coordinates for a single batch. It is a uniform scale coordinate normalized to [0,1].

[0059] This embodiment addresses the issues of temporal misalignment and spatial dimension inconsistency in heterogeneous modalities by designing a spatiotemporal dual-dimensional adaptive alignment mechanism to achieve high-precision synchronization of data with different frequencies, dimensions, and forms. To address the problem of multimodal small-sample training, a differentiated modality enhancement strategy adapted to the characteristics of vision, touch, and language is designed to avoid modal information distortion caused by a single enhancement method. This eliminates spatiotemporal misalignment and dimension conflicts in multimodal data, avoiding information biases in subsequent causal learning and feature fusion; simultaneously, it effectively suppresses model overfitting, improves the model's adaptability and stability in complex and unknown physical scenarios, and provides high-quality, aligned, and unified input data for backend causal modeling.

[0060] S212. Generate a dynamic causal matrix based on historical perception dataset.

[0061] Constructing a unified set of causal variables It covers visual spatial variables, tactile mechanical variables, and language rule variables, and realizes the structured definition of causal nodes in physical scenes.

[0062] Gumbel-Sinkhorn dynamic causality matrix generation:

[0063] ,in, It is a variable right One-way causal weights These are uniformly distributed random sampled values ​​between 0 and 1. It uses Gumbel random noise to achieve continuous relaxation of discrete structures. It is the temperature coefficient, which decays from 1.0 to 0.1 during training, realizing the convergence of the causal structure from continuous approximation to discrete.

[0064] S213. Construct a differentiable directed acyclic constrained loss function.

[0065] DAG (Directed Acyclic Graph) acyclic constraint loss function:

[0066] ,in, It is the Hadamard product of matrices. It is a matrix exponentiation operation. It is a matrix trace operation. It is the total number of causal variables. It is an acyclic constraint loss, and when the loss approaches 0, the causal graph is a legal physical DAG structure.

[0067] S214. Construct the total loss function based on the directed acyclic constraint loss function.

[0068] The total loss function is: .

[0069] It should be understood that the Gumbel-Sinkhorn approximation method can also be used to define the dynamic causal matrix A. Specifically, the dynamic causal matrix can be constructed using the following formula:

[0070] A = Sinkhorn((log(α) + ε) / τ), where α∈RN×N, represents the learnable parameter matrix, representing the potential causal strength between variables; ε~Gumbel (0,1) is the noise term, which is a noise term from the Gumbel distribution, used to introduce randomness to optimize the discrete structure search; τ represents the annealing temperature parameter, used to control the smoothness of the matrix. When τ approaches 0, the matrix converges to a hard doubly random matrix; Sinkhorn (·) represents the Sinkhorn operator, used to convert the input matrix into a doubly random matrix through iterative row and column normalization.

[0071] To ensure the causal graph learned remains acyclic during training, a NOTEARS-based acyclicity constraint is introduced: tr(e^(W·W)) - d = 0, where W is the weighted adjacency matrix and d is the number of variables. This continuous acyclicity representation is incorporated as a penalty term into the loss function, guiding the optimization away from cyclic structures.

[0072] In terms of multimodal data processing, this embodiment of the invention employs a dual-pipeline architecture, SPOTS (Simultaneous Prediction of Optical and Tactile Sensations). This architecture explicitly models cross-modal interactions while preserving modality-specific inductive biases. Unlike traditional single-pipeline fusion methods, SPOTS utilizes two independent frame prediction models: one for the visual modality and the other for the tactile modality, achieving information interaction through cross-modal connections.

[0073] It should be understood that for language modalities, a pre-trained large language model can be used as a semantic feature extractor to convert natural language instructions into high-dimensional semantic vectors. Through spatial location mapping and token-level contrastive learning, RGB images, 3D point clouds, and tactile signals are directly aligned within the large language model to construct a unified feature space.

[0074] This invention abandons the traditional approach of fixed causal structures and manually defined causal rules, proposing a cross-modal dynamic causal discovery mechanism based on differentiable DAGs. This mechanism updates the causal topology in real time according to scene changes, introduces the Gumbel-Sinkhorn continuous relaxation strategy to overcome the technical bottlenecks of traditional discrete causal graphs, such as the inability to perform gradient backpropagation and end-to-end training. It also incorporates NO TEARS acyclic hard constraints to mathematically ensure that the mined causal structure conforms to physical logic, eliminating causal loop paradoxes. This achieves automatic mining, dynamic updating, interpretability, and trainability of physical causal structures, ensuring that all subsequent feature interactions, temporal predictions, and state deductions follow real physical causal relationships. This solves the problems of black-box prediction, lack of physical logic, and causal inconsistencies in traditional deep learning models, significantly improving the rationality and interpretability of world model deduction results.

[0075] S22. Construct a causal prediction latent space using the perceived causal relationship structure. The causal prediction latent space includes static causal latent features and dynamic causal latent features. Static causal latent features are used to describe static structural causal features, and dynamic causal latent features are used to describe dynamic motion causal features.

[0076] This embodiment compresses and encodes all past observations and actions into fixed-size memory weights, rather than storing every historical image. It employs a causal variational autoencoder architecture, where the encoding process conditionally depends on the latent states of previous frames, ensuring the temporal coherence of the encoded video sequence.

[0077] The specific implementation adopts a hierarchical latent space structure, which decomposes the video into two parts: static structure and dynamic motion. The static elements of the encoded scene can be represented by z_s, including background, object shape and color, etc., corresponding to the first frame of the video. z_m represents the dynamic elements, i.e. motion patterns, representing the changes from one frame to the next. Through this decomposition, the model can learn a more efficient representation.

[0078] During the prediction phase, the model predicts chain-like sequences of motion in the latent space, rather than pixel-level prediction of complete frames. This embodiment of the invention reduces the amount of feature data to 1.02% of traditional models while maintaining prediction accuracy.

[0079] The latent space dynamics model takes the following form: z_{t+1} = f(z_t, a_t), where z_t is the current latent state, a_t is the action, and f (・) can be a deterministic, stochastic, or counterfactual transformation function. The model employs a causal attention Transformer architecture to accurately model the temporal causal characteristics of real-world space, stripping away visual textures and other superficial interferences within the latent space to precisely capture the essence of physical dynamics.

[0080] Specifically, in some optional embodiments, S22 can be implemented through the following process (not shown in the figure):

[0081] S221. Obtain the historical perception dataset, which includes historical visual data, historical tactile data, and historical language data.

[0082] S222. Perform two-layer, two-branch latent feature encoding based on historical sensing data in the historical sensing dataset.

[0083] Specifically, S222 may include (not shown in the figure):

[0084] pass Construct a 32-dimensional static causal latent feature, in which, It is a static causal latent feature. These are standardized historical visual data, historical tactile data, and historical language data, respectively.

[0085] pass Construct a 32-dimensional dynamic causal latent feature, in which, As a dynamic causal latent feature, These are multimodal temporal historical visual data, historical tactile data, and historical language data, respectively.

[0086] pass A 64-dimensional total causal latent feature was constructed, and the historical perception data in the historical perception dataset was dimensionality reduced so that the feature data volume of the historical perception data reached a preset threshold. This represents the overall static causal latent characteristic.

[0087] S223, Causal Constraint Attention Mapping Temporal Modeling.

[0088] Specifically, the following methods can be used to implement causal constraint attention mapping temporal modeling.

[0089] , It is a query, key, and value matrix of latent feature mapping. It is the key feature dimension, used for attention score normalization. It is a normalized dynamic causal matrix, and the constraint attention only interacts between causal nodes.

[0090] S224. Construct time series prediction and loss function.

[0091] Specifically, the time series prediction and loss function can be constructed in the following way:

[0092] ;

[0093] .

[0094] This embodiment employs a static structure combined with a dynamic motion hierarchical latent space decoupling mechanism. It decomposes the scene's physical attributes into inherent invariant attributes and temporally changing attributes, aligning with the operational laws of the real physical world. It utilizes a Transformer attention mechanism with causal matrix constraints, abandoning traditional global indiscriminate attention and allowing only interaction of features with physical causal relationships. This achieves extreme feature lightweight compression, reducing the traditional 6272-dimensional features to 64 dimensions, achieving the core technical indicator that the feature data volume is only 1.02% of the traditional model. Specifically, feature decoupling modeling makes the model's understanding of scene structure and motion changes more closely aligned with the physical essence, significantly improving temporal prediction accuracy. Causal attention completely eliminates interference from invalid features, enhancing the physical logic of model inference. Extreme dimensionality reduction significantly reduces the number of model parameters, computational load, and storage overhead, increasing model inference speed and greatly lowering the deployment threshold, balancing high accuracy and lightweight design.

[0095] S23. Construct a task-oriented dynamic attention mechanism based on language data, perceptual causal relationship structure, and causal prediction latent space.

[0096] The task-oriented dynamic attention mechanism is a key feature of this invention for achieving efficient reasoning. Based on routing network predictions, this mechanism enables the model to selectively focus on the most informational visual streams. Specifically, the adaptive computation mechanism is an innovative attention computation mechanism that connects the instantaneous application of perceptual computation with its impact on decision-making outcomes. This mechanism can dynamically allocate perceptual computation resources and adjust attention weights in real time according to task requirements. Task relevance is achieved by defining a task relevance function to dynamically weight different modalities and features. The degree of attention weight given to map points depends on the robot's current state, enabling the encoder to dynamically adjust its focus on the environment based on position, velocity, and posture. Multi-granularity memory network integration integrates neural processes with multi-granularity memory networks, achieving an inference accuracy of 83.4% using only 5% labeled data in missing modality benchmark tests. It can dynamically select appropriate memory granularity and inference strategies based on task type and environmental changes.

[0097] Specifically, this step includes:

[0098] Task routing weights are generated using the following method:

[0099] ;

[0100] .

[0101] Adaptive computing power allocation is performed in the following ways:

[0102] .

[0103] Multi-granularity memory updates can be performed in the following ways:

[0104] Short-term sequential memory: ;

[0105] Long-term causal memory: ,in, It is a task-adaptive feature weight matrix. It is a dynamic filtering threshold. It is the memory-gated fusion coefficient. It is a pool of short-term temporal memory and long-term causal memory.

[0106] This embodiment employs a dynamic weight control mechanism for task routing. Based on language task instructions, it adaptively determines the task value of each feature, dynamically filters effective physical features, and uses an adaptive perceptual computing resource allocation strategy to precisely compute high-value features and sparsely compute low-value features, achieving on-demand scheduling of computing power. It constructs a dual-granularity memory system combining temporal short-term memory and causal long-term memory, solidifying the physical causal priors of the scene and avoiding redundant computations. This effectively addresses the shortcomings of traditional fixed attention mechanisms in adapting to multi-task scenarios. The model's generalization ability for different physical interaction tasks such as grasping, pushing, and placing is significantly improved. Adaptive computing power allocation reduces invalid computational overhead by 30%~50%, enhancing the model's real-time inference capability. The multi-granularity memory mechanism accelerates the model's inference convergence speed and strengthens the model's ability to remember and reuse dynamic scenes and long-term physical rules.

[0107] S24. Acquire the current perception dataset in real time. The current perception dataset includes current visual data, current tactile data, and current language data.

[0108] S25. Based on the current perception dataset, generate the current causal constraints using the perception causal relationship structure. The perception causal relationship structure includes the causal relationship between the previously perceived data and the subsequently perceived data.

[0109] S26. Map the current perception dataset to the causal prediction latent space to obtain the current causal latent features. Combine the historical time series features and the causal matrix in the perception causal relationship structure to deduce the future causal latent state.

[0110] S27. Based on real-time task instructions, calculate feature weight parameters using a task-oriented dynamic attention mechanism and filter effective latent features.

[0111] S28. Perform multi-source data fusion and multi-worldline future state deduction on the current causal constraints, current causal latent features, future causal latent states, feature weight parameters, and effective latent features to obtain the causal latent features of the target future state.

[0112] In this embodiment of the invention, multi-source data fusion innovatively employs an encoder-free multimodal alignment mechanism. Through spatial location mapping and token-level comparative learning, it directly aligns RGB images, 3D point clouds, and tactile signals within a large language model, constructing a unified feature space. This avoids the information loss caused by traditional encoders and improves the efficiency of cross-modal information transmission.

[0113] In terms of future state projection, this invention achieves a complete closed loop from world understanding and prediction to world intervention. Specifically, it includes: directly encoding physical laws such as gravity, friction, and rigid body constraints into a neural network through physical causal prior encoding, providing physical constraints for prediction; constructing a multi-worldline search mechanism, starting from the underlying physical prior constraints and task objectives, generating key worldlines causally related to the core objective in parallel in the latent space, with each worldline labeled with a causal relationship; and through goal-oriented prediction optimization, specifically, based on a goal evaluation mechanism, the model understands and predicts based on the task objective, automatically cutting off computational branches that violate common sense and ignoring irrelevant details, thereby significantly reducing computational power consumption.

[0114] Specifically, the implementation process of this step may include:

[0115] In the process of encoderless token-level multimodal alignment, the following method is used, employing unified projection of modal tokens:

[0116] .

[0117] The cross-modal contrast alignment loss can be calculated as follows:

[0118] .

[0119] The following methods can be used to select the optimal trajectory across multiple world lines:

[0120] ;

[0121] .

[0122] This embodiment employs an encoder-free multimodal alignment mechanism, eliminating the feature distortion and information loss problems caused by the global fusion of traditional Transformer encoders. It achieves accurate alignment of heterogeneous modalities through token-level contrastive learning. A multi-worldline search and extrapolation strategy under physical causal prior constraints generates multiple future scene trajectories in parallel. The optimal solution is selected from three dimensions: causal compliance, task matching, and physical rationality, achieving deep fusion of spatial location mapping and semantic rules. This ensures that multimodal fusion strictly conforms to real physical space and task constraints. It effectively solves the problems of misalignment and feature distortion in the fusion of visual, tactile, and linguistic heterogeneous modalities. The cross-modal fusion accuracy is significantly better than traditional encoder fusion schemes. Multi-worldline extrapolation avoids the randomness and irrationality of single trajectory prediction, greatly improving the stability, accuracy, and physical compliance of future physical state predictions. It achieves a unified triple constraint of physical causality, task requirements, and kinematic rules, ensuring that the model extrapolation results fully conform to the physical operating logic of the real world.

[0123] The optimal latent space future state features are restored to the original multimodal physical state using the following decoder:

[0124] It outputs future time-series visual images, tactile mechanical states, object motion trajectories, and scene interaction states, completing the closed loop of physical world model deduction.

[0125] This embodiment employs a modality-specific decoding and mapping mechanism, using differentiated decoders for two different types of data: visual images and tactile sensing data. This enables precise reverse mapping from low-dimensional causal latent features to real physical states. It ensures that the abstract physical features of the latent space can be accurately reconstructed into observable and verifiable real-world scene states, achieving a complete closed loop from causal modeling and latent space deduction to the output of real physical states, thus guaranteeing the usability and authenticity of the model's output results.

[0126] S29. Generate control instructions based on the causal latent features of the target's future state, and use the control instructions to complete the corresponding operations.

[0127] It should be noted that, in this embodiment, the following overall total loss calculation formula can be used:

[0128] .

[0129] The world-model-based control method provided in this embodiment creates a dynamic causal graph learning mechanism. Based on a differentiable DAG and a dynamic causal discovery module with counterfactual intervention, it achieves real-time updates of the causal strength matrix, with a causal identification F1 score of 83.9%. It innovatively simulates quantum entanglement operations in a classical environment using quantum state fusion technology, achieving efficient cross-modal feature entanglement through phase gate control, with a quantum entanglement degree (QED) of 73.0%, significantly improving environmental robustness. All past observations and actions are compressed and encoded into fixed-size memory weights, reducing the feature data volume to only 1.02% of traditional models and increasing planning speed by more than 8 times. By dynamically adjusting the prediction strategy based on task relevance, it significantly reduces computational complexity while maintaining prediction accuracy. In other words, this method, based on three-dimensional heterogeneous data input of visual appearance, tactile physics, and linguistic rules, achieves comprehensive environmental information collection from explicit visual observations, intrinsic physical properties, and high-level rule constraints, overcoming the limitations of pure visual models in perceiving object material, force, and physical constraints. By introducing a causal reasoning mechanism and multimodal fusion technology, this method addresses the technical bottlenecks of traditional pixel-level prediction methods in terms of computational efficiency, generalization ability, and long-term planning. This method can extract causal relationships from visual, tactile, and linguistic data, construct a low-dimensional causal feature space, and achieve efficient future state deduction through a task-oriented dynamic attention mechanism.

[0130] It should be understood that the control method based on the world model provided by the present invention is not limited to the implementation process described in the foregoing embodiments. The specific process steps can be designed by those skilled in the art based on the technical concept of the embodiments of the present invention and according to engineering needs.

[0131] It is understood that the various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, these embodiments will not be described in detail here. Those skilled in the art will understand that the specific execution order of each step in the above methods of specific implementation should be determined by its function and possible internal logic.

[0132] To facilitate the implementation of the aforementioned world-model-based control method, embodiments of the present invention provide a world-model-based control system, such as... Figure 3 As shown, the system includes:

[0133] Data acquisition module 31 is used to acquire the current sensory dataset in real time. The current sensory dataset includes current visual data, current tactile data, and current language data.

[0134] The causal constraint acquisition module 32 is used to generate current causal constraints based on the current sensing dataset and using the sensing causal relationship structure. The sensing causal relationship structure includes the causal relationship between the previously sensing data and the subsequently sensing data.

[0135] The causal latent feature acquisition module 33 is used to map the current perception dataset to the causal prediction latent space to obtain the current causal latent features, and to infer the future causal latent state by combining the historical time series features and the causal matrix in the perception causal relationship structure.

[0136] The latent feature screening module 34 is used to calculate feature weight parameters and screen effective latent features based on real-time task instructions and a task-oriented dynamic attention mechanism.

[0137] The multi-source data fusion module 35 is used to perform multi-source data fusion and multi-worldline future state deduction on the current causal constraints, current causal latent features, future causal latent states, feature weight parameters and effective latent features to obtain the causal latent features of the target future state.

[0138] The control execution module 36 is used to generate control instructions based on the causal latent features of the target's future state, and to use the control instructions to complete the corresponding operations.

[0139] The world-model-based control system provided in this invention unifies autonomous driving data through intermediate representation markers, enabling the structured and solidified constraint verification inputs and outputs generated during the online phase. This effectively avoids inconsistencies in verification caused by data drift, improving the reliability and stability of autonomous driving decision constraint verification. By introducing a causal reasoning mechanism and multimodal fusion technology, it addresses the technical bottlenecks of traditional pixel-level prediction methods in terms of computational efficiency, generalization ability, and long-term planning. This method can extract causal relationships from visual, tactile, and linguistic data, construct a low-dimensional causal feature space, and achieve efficient future state deduction through a task-oriented dynamic attention mechanism.

[0140] This invention provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the control method steps based on a world model in any of the foregoing embodiments. The computer-readable storage medium may be volatile or non-volatile.

[0141] This disclosure provides an electronic device comprising: at least one processor; at least one memory; and one or more I / O interfaces connected between the processor and the memory; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the aforementioned world-model-based control method.

[0142] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the system, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).

[0143] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0144] Computer program instructions used to perform the operations of embodiments of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of embodiments of this disclosure.

[0145] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK).

[0146] Various aspects of embodiments of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0147] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0148] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, wherein the module, segment, or portion of an instruction contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0150] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the embodiments of this disclosure as set forth in the appended claims.

Claims

1. A control method based on a world model, characterized in that, include: The current sensory dataset is acquired in real time, and the current sensory dataset includes current visual data, current tactile data, and current language data; Based on the current sensing dataset, current causal constraints are generated using a sensing causal relationship structure, which includes the causal relationship between the previously sensed data and the subsequently sensed data. The current perception dataset is mapped to the causal prediction latent space to obtain the current causal latent features. The future causal latent state is inferred by combining the historical time series features and the causal matrix in the perception causal relationship structure. Based on real-time task instructions, feature weight parameters are calculated using a task-oriented dynamic attention mechanism, and effective latent features are selected. Multi-source data fusion and multi-worldline future state deduction are performed on the current causal constraints, the current causal latent features, the future causal latent states, the feature weight parameters, and the effective latent features to obtain the causal latent features of the target future state; Control instructions are generated based on the causal latent features of the target future state, and the corresponding operations are performed using the control instructions.

2. The method according to claim 1, characterized in that, Before generating the current causal constraints based on the current perceptual dataset and the perceptual causal relationship structure, the method further includes: The perceived causal relationship is constructed using a differentiable directed acyclic graph.

3. The method according to claim 2, characterized in that, Before mapping the current perceptual dataset to the causal prediction latent space, the method further includes: The causal prediction latent space is constructed using the perceived causal relationship structure. The causal prediction latent space includes static causal latent features and dynamic causal latent features. The static causal latent features are used to describe static structural causal features, and the dynamic causal latent features are used to describe dynamic motion causal features.

4. The method according to claim 3, characterized in that, Before calculating feature weight parameters using a task-oriented dynamic attention mechanism based on real-time task instructions, the method further includes: The task-oriented dynamic attention mechanism is constructed based on the language data, the perceptual causal relationship structure, and the causal prediction latent space.

5. The method according to claim 3, characterized in that, Constructing the causal prediction latent space using the aforementioned perceptual causal relationship structure includes: Obtain a historical perception dataset, which includes historical visual data, historical tactile data, and historical language data; Two-layer, two-branch latent feature encoding is performed based on the historical sensing data in the aforementioned historical sensing dataset; Causal-constrained attention mapping temporal modeling; Construct time series prediction and loss functions.

6. The method according to claim 5, characterized in that, The two-layer, two-branch latent feature encoding based on the historical sensing data in the aforementioned historical sensing dataset includes: pass Construct a 32-dimensional static causal latent feature, in which, It is a static causal latent feature. These are standardized historical visual data, historical tactile data, and historical language data, respectively. pass Construct a 32-dimensional dynamic causal latent feature, in which, As a dynamic causal latent feature, These are multimodal temporal historical visual data, historical tactile data, and historical language data, respectively. pass A 64-dimensional total causal latent feature is constructed, and the historical perception data in the historical perception dataset is subjected to dimensionality reduction processing so that the feature data volume of the historical perception data reaches a preset threshold. This represents the overall static causal latent characteristic.

7. The method according to claim 5, characterized in that, Constructing the perceived causal relationship using a differentiable directed acyclic graph includes: Multimodal causal variables are constructed based on the historical perception dataset to realize the structuring of causal nodes in the physical scene; A dynamic causal matrix is ​​generated based on the historical perception dataset.

8. The method according to claim 7, characterized in that, The dynamic causal matrix is ​​constructed using the following formula: A = Sinkhorn((log(α) + ε) / τ), where α∈RN×N, represents the learnable parameter matrix, representing the potential causal strength between variables; ε~Gumbel (0,1) is a noise term used to introduce randomness to optimize the discrete structure search; τ represents the annealing temperature parameter used to control the smoothness of the matrix. When τ approaches 0, the matrix converges to a hard double random matrix; Sinkhorn (·) represents the Sinkhorn operator, used to convert the input matrix into a double random matrix through iterative row and column normalization.

9. The method according to claim 7, characterized in that, Constructing the perceived causal relationship using a differentiable directed acyclic graph also includes: Construct a differentiable directed acyclic constrained loss function; The total loss function is constructed based on the directed acyclic constraint loss function.

10. A control system based on a world model, characterized in that, include: The data acquisition module is used to acquire the current sensory dataset in real time, which includes current visual data, current tactile data and current language data; The causal constraint acquisition module is used to generate current causal constraints based on the current perception dataset and using a perception causal relationship structure, wherein the perception causal relationship structure includes the causal relationship between the previously perceived data and the subsequently perceived data. The causal latent feature acquisition module is used to map the current perception dataset to the causal prediction latent space to obtain the current causal latent features, and combine the historical time series features and the causal matrix in the perception causal relationship structure to infer the future causal latent state. The latent feature filtering module is used to calculate feature weight parameters and filter effective latent features based on real-time task instructions and a task-oriented dynamic attention mechanism. The multi-source data fusion module is used to perform multi-source data fusion and multi-worldline future state inference on the current causal constraints, the current causal latent features, the future causal latent states, the feature weight parameters, and the effective latent features to obtain the causal latent features of the target future state. The control execution module is used to generate control instructions based on the causal latent features of the target's future state, and to use the control instructions to complete the corresponding operations.