A spatial world multi-modal basic model system with meta-evolution and capability emergence and a training and reasoning method
Patent Information
- Application Number
- CN202610646244.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-12
- Publication Date
- 2026-09-25
AI Technical Summary
[0002]当前通用大模型主要聚焦文本语义与二维图像理解,缺乏对真实三维物理世界的建模能力,无法实现空间坐标推理、三维场景重建、轨道推演与物理规则拟合
相较于现有技术,本发明具备如下优势:
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence large-scale models, multimodal fusion intelligence, three-dimensional Gaussian neural field reconstruction, space physical modeling, autonomous evolutionary artificial intelligence, large-scale distributed training, and high-concurrency inference service technology, specifically to a multimodal basic model system for a spatial world with meta-evolution and emergent capabilities, and a training and inference method. Background Technology
[0002] Current general-purpose large-scale models mainly focus on text semantics and 2D image understanding, lacking the ability to model the real 3D physical world. They cannot achieve spatial coordinate reasoning, 3D scene reconstruction, trajectory deduction, and physical rule fitting. Existing intelligent models do not have an endogenous autonomous evolution mechanism. The model's capabilities rely on external training iterations, and it cannot achieve self-reflection, self-correction, and self-verification. This makes them prone to problems such as semantic illusion, feature distortion, and logical contradictions. Meanwhile, existing space intelligence models suffer from weak modal fusion capabilities, with text, images, point clouds, and spatial coordinate information being fragmented and unable to construct a unified world representation. Traditional training frameworks struggle to support training models with trillions of parameters, exhibiting issues such as memory overflow, gradient instability, breakpoint loss, and multi-GPU synchronization anomalies. Existing inference systems are functionally limited, unable to simultaneously support multiple tasks including language generation, embodied control, aerospace orbit simulation, and 3D rendering, resulting in insufficient engineering and commercial deployment capabilities. In summary, existing technologies suffer from numerous drawbacks, including insufficient spatial modeling capabilities, lack of autonomous evolution capabilities, susceptibility to AI illusions, inconsistent modal fusion, poor training stability, limited inference functionality, and difficulties in commercial deployment. Summary of the Invention
[0003] 1. Purpose of the invention The purpose of this invention is to overcome the shortcomings of the prior art and provide a multimodal basic model system and training and inference method for the spatial world with meta-evolution and capability emergence. It constructs a unified basic model for the spatial world, realizes deep multimodal fusion, accurate modeling of the physical world, autonomous evolution and iteration of the model, and automatic emergence and solidification of capabilities, while achieving industrial-grade stable training and commercial-grade inference deployment. 2. Overall Technical Solution The complete system of this invention includes multimodal input encoding, three-dimensional neural field fusion, spatiotemporal trajectory extrapolation, physical constraint correction, spatial memory evolution, meta-evolutionary autonomous iteration, capability emergence solidification, spatial Transformer modeling, multi-task output, distributed training, weight merging, and a closed-loop commercial inference service. The core innovation of this invention lies in constructing a completely endogenous meta-evolutionary closed-loop system and a capability emergence solidification mechanism, which allows the model to self-optimize, self-correct, and self-upgrade without manual iteration; at the same time, it constructs a unified spatial world representation to achieve deep alignment between text semantics and the three-dimensional physical world. 3. Beneficial effects Compared with existing technologies, the present invention has the following advantages: (1) It has autonomous meta-evolutionary capabilities. The model can automatically regularize, self-evaluate, self-correct, and self-iterate, completely suppressing AI illusions; (2) It has the ability to automatically detect and solidify emerging intelligent capabilities, and can continuously generate and retain new intelligent capabilities; (3) The model is intelligently upgraded with zero structural cost by adopting a logical layer expert evolution mechanism; (4) Achieve unified modeling of four-dimensional modalities including text, images, 3D point clouds, and spatial coordinates; (5) Integrating physical rules with spatiotemporal trajectory deduction, the output results closely match the real physical world; (6) Supports industrial-grade stable training with trillions of parameters, breakpoint resume training, and lossless weight export; (7) It can output language, action, track and rendering results in one stop, and has strong commercial application capabilities. Detailed Implementation The implementation process of this invention is divided into seven parts: model architecture implementation, meta-evolution mechanism implementation, capability emergence solidification implementation, expert routing evolution implementation, distributed training implementation, weight derivation implementation, and commercial inference deployment implementation. The model's overall hidden dimension is set to 12,288, and the backbone network has 80 layers. It employs a spatially aware Transformer structure, adaptable to modeling ultra-long sequences. A multimodal encoder uniformly receives four types of input features, achieving dimensional alignment and semantic fusion. A 3D Gaussian neural field encoder performs differentiable field encoding on point cloud features, realizing implicit reconstruction of the 3D scene. The spatiotemporal trajectory engine learns spatial temporal variation patterns, outputting continuous and stable spatial inference features. The meta-evolutionary core is involved in the entire training and inference process, performing regularization cleaning, quality assessment, adversarial verification, error penalty, and post-mortem correction on each round of features to continuously optimize the model's representation distribution. The system monitors changes in feature distribution in real time and automatically solidifies new capabilities after triggering emergent conditions. The expert routing matrix is dynamically updated based on historical call data to achieve optimal resource allocation. The training phase employs distributed training across multiple machines and GPUs, enabling gradient accumulation, gradient pruning, and mixed-precision training, and automatically saving training breakpoints. After training, the distributed shard weights are merged into a single-file universal weight, adaptable for single-machine inference deployment. Finally, a high-concurrency interface service is deployed to achieve intelligent commercial output in a multimodal space. Attached Figure Description Figure 1 is a flowchart of the overall architecture of a multimodal basic model of a spatial world with meta-evolution and capability emergence according to the present invention. Figure 2 is a flowchart of the autonomous closed-loop iterative process within the meta-evolutionary core of this invention. Figure 3 is a flowchart of the emergent capability detection and automatic solidification process of the model of the present invention. Figure 4 is a flowchart of the logically seamless expert routing adaptive evolution of the present invention. Figure 5 is a flowchart of the entire process of distributed training, weight derivation, and commercial inference service of this invention. Specific Work Process The complete closed-loop workflow of this invention is as follows: The system receives four types of input: text tokens, spatial coordinates, image blocks, and 3D point clouds. These are then uniformly represented by a multimodal encoder. The system integrates 3D Gaussian neural field spatial features to deduce the spatiotemporal trajectory variation patterns. Long-term spatial memory knowledge is injected. This knowledge is then fed into the meta-evolutionary core to complete autonomous cleaning, evaluation, error correction, iteration, and evolution. Emergent capabilities are detected and solidified. Global modeling is completed through a multi-layered spatial Transformer. Output deviations are corrected through physical rule constraints. The system outputs language generation, embodied actions, space orbits, and 3D rendering results in parallel. During the training phase, a distributed framework is used to complete large-scale parameter optimization and breakpoint saving. During the inference phase, standardized weights are loaded, providing stable commercial multimodal intelligent services to external users.
Claims
1. A multimodal basic model system for a spatial world possessing meta-evolution and emergent capabilities, characterized in that, include: Multimodal spatial coding module, 3D Gaussian neural field feature coding module, spatiotemporal trajectory tensor inference module, physical rule constraint engine module, spatial long-term memory evolution module, meta-evolution core control module, spatial perception Transformer backbone module, multi-task output head module, distributed training and commercial inference service module; The multimodal spatial coding module is used to receive four types of inputs: text tokens, spatial three-dimensional coordinates, image block features, and three-dimensional point cloud features, and to complete the unified dimension coding and semantic alignment of multimodal information. The three-dimensional Gaussian neural field feature encoding module is used to perform dimensional transformation and temporal alignment on the three-dimensional point cloud Gaussian features, and to achieve deep fusion of the three-dimensional spatial features with the semantic features of the main text. The spatiotemporal trajectory tensor inference module is used to learn the spatial temporal change law, output spatiotemporal trajectory features and generate trajectory inference auxiliary constraint loss; The physical rule constraint engine module is used to introduce prior physical rules in real space, perform physical correction on the model feature output, and generate physical constraint loss. The spatial long-term memory evolution module is used to store, cluster, filter, and iteratively update long-term spatial feature memories, thereby realizing spatial knowledge accumulation and autonomous evolution. The meta-evolution core control module incorporates a feature regularization unit, a self-evaluation unit, an adversarial referee unit, a self-deception punishment unit, a self-review unit, a self-correction unit, a memory evolution unit, an ability emergence detection unit, an emergence solidification unit, and a logic expert routing evolution unit, enabling autonomous iterative optimization of the model across the entire chain. The spatial perception Transformer backbone module is used for iterative modeling of multi-layer spatial perception features, integrating semantic, spatial, physical, and trajectory multi-dimensional information; The multi-task output head module is used to output language prediction results, humanoid body action parameters, aerospace orbit control parameters, three-dimensional neural rendering images, and physical offset correction parameters in parallel. The distributed training and commercial inference service module is used to implement multi-machine and multi-card distributed training, gradient accumulation optimization, breakpoint saving, shard breakpoint weight merging, single-machine standardized weight export, and high-concurrency multimodal inference service deployment.
2. The system according to claim 1, characterized in that, The self-evaluation unit of the meta-evolutionary core control module is used to perform multi-dimensional quantitative scoring on the hidden features output by the model. The adversarial referee unit is used to generate an independent and objective third-party feature quality score. The self-deception penalty loss is constructed by calculating the difference between the self-scoring score and the referee score to suppress the model's self-deception and false feature generation behavior.
3. The system according to claim 1, characterized in that, The capability emergence detection unit is used to quantify the degree of change in the model feature distribution in real time. When the change in feature distribution exceeds a preset threshold, it is determined to be capability emergence, triggering the emergence solidification unit. The emergence solidification unit completes adaptive fine-tuning of feature distribution through a learnable feature projection matrix, locks and solidifies the emerging capability features, and ensures stable retention of emerging capabilities.
4. The system according to claim 1, characterized in that, The logical expert routing evolution unit does not add, delete, or replace physical expert modules. Instead, it dynamically updates the logical routing association weight matrix between experts by statistically analyzing the historical call frequency of each expert, thereby achieving autonomous evolution of expert intelligent routing with zero hardware overhead and zero structural changes.
5. A training and inference method for a multimodal basic model of a spatial world with meta-evolution and emergent capabilities, characterized in that, Includes the following steps: Step 1: Construct a multimodal training dataset by collecting text sequences, spatial 3D coordinates, image block features, and 3D point cloud feature samples, and complete data cleaning and standardization. Step 2: Input the multimodal features into the multimodal spatial coding module to complete the unified semantic coding; Step 3: Three-dimensional spatial feature fusion is completed through three-dimensional Gaussian neural field encoding, and temporal inference and trajectory loss calculation are completed through spatiotemporal trajectory engine; Step four: Introduce spatial long-term memory features. Complete the infusion of historical knowledge and feature enhancement; Step 5: The data is fed into the meta-evolution core module, where feature regularization, self-evaluation, adversarial verification, self-deception penalty, debriefing and error correction, memory iteration, and expert route evolution are completed in sequence. Step 6: Real-time detection of the model's emergent capabilities and automatic solidification of the emergent capabilities; Step 7: Complete global feature iterative modeling through a multi-layer spatial perception Transformer, and superimpose physical rule constraint correction; Step 8: Output language, action, trajectory, rendering, and physical offset results in parallel through the multi-task output head, calculate the multi-target joint loss to complete the reverse update; Step 9: Use a distributed training framework to complete multi-machine, multi-GPU iterative training, gradient pruning, gradient accumulation, and breakpoint saving. Step 10: After training is complete, merge the distributed sharded breakpoints into single-file standard weights; Step 11: Load standardized weights, deploy commercial inference services, and provide text generation and multimodal spatial intelligent inference capabilities to external users.