A multimodal manifold autonomous intelligent robot control system

CN122559966APending Publication Date: 2026-08-14黄承斌
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

克服现有技术多模态割裂、无物理约束、模型不可进化、集群协同弱、控制精度低的缺陷,提供一种多模态流形自治智能机器人控制系统,实现文本、视频、点云、音频四模态统一流形感知、物理世界动力学建模、分层任务决策、机器人关节精准控制、模型自主进化升级、故障自愈、分布式智能体集群协同全闭环工业级自治智能

Benefits of technology

四模态统一流形编码融合,实现全维度环境精准认知,彻底解决模态特征割裂问题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This invention discloses a multimodal manifold autonomous intelligent robot control system, belonging to the field of embodied intelligence technology. The system integrates modules for multimodal coding, cross-modal fusion, high-dimensional manifold perception, heterogeneous expert reasoning, physical world modeling, hierarchical task planning, and joint motion decoding. It can access multiple types of sensory signals, including text, video, point clouds, and audio. It also features a fault self-healing, autonomous evolution, version iteration, self-developed architecture, and distributed cluster collaboration system, along with a distributed training and management mechanism. This system integrates the advantages of mainstream embodied models such as VLA, JEPA, and WAM, possessing capabilities in environmental cognition, physical deduction, autonomous decision-making, cluster operation, and self-optimization, effectively improving the control accuracy and scene adaptability of industrial robots.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of large-scale artificial intelligence, multimodal perception fusion, physical world modeling, hierarchical task planning, robot joint servo control, distributed heterogeneous expert hybrid network, and autonomous evolution and iteration technology, and in particular to an autonomous intelligent robot closed-loop control system based on high-dimensional manifold topology coding. Background Technology

[0002] Existing industrial intelligent robot control systems suffer from drawbacks such as a single perception mode, inability to uniformly represent cross-modal characteristics, lack of physical and dynamic constraints, fixed and non-evolvable decision-making links, inefficient model computing power scheduling, lack of autonomous fault repair capabilities, and inability of multiple agents to collaborate in a cluster. Traditional independent encoding methods for vision, audio, point cloud, and text are fragmented and cannot build a unified environmental cognition; conventional neural networks lack the ability to reason about the topology of high-dimensional manifold spaces, resulting in large deviations between physical simulations and actual robot movements; fixed network structures cannot adaptively iterate and upgrade according to tasks, and large-scale model training is prone to gradient divergence, numerical anomalies, and waste of computing resources; without hierarchical task planning and joint trajectory decoding closed loops, the robot's execution accuracy and autonomous adaptability are low. Summary of the Invention

[0003] Purpose of the invention Overcoming the shortcomings of existing technologies such as multimodal fragmentation, lack of physical constraints, non-evolvable models, weak cluster collaboration, and low control precision, this paper provides a multimodal manifold autonomous intelligent robot control system. This system enables unified manifold perception across four modalities (text, video, point cloud, and audio), dynamic modeling of the physical world, hierarchical task decision-making, precise control of robot joints, autonomous evolution and upgrading of the model, fault self-healing, and a fully closed-loop industrial-grade autonomous intelligence system with distributed intelligent agent cluster collaboration. Technical solution A multimodal manifold autonomous intelligent robot control system includes a multimodal manifold encoding module, a cross-modal topology fusion module, a high-dimensional manifold perception module, a heterogeneous MoE expert parallel reasoning module, a physical world modeling module, a hierarchical task planning module, a robot joint motion decoding module, an autonomous evolution iteration engine, a fault self-healing and repair module, an automatic version upgrade module, an architecture self-development module, a distributed intelligent agent bee colony cluster module, and a distributed training and management module. Multimodal manifold coding module Each component is configured with a text token embedding encoder, a video manifold visual encoder, a 3D point cloud topology encoder, and an audio field encoder. These components independently perform high-dimensional feature mapping on the input text commands, time-series video frames, spatial 3D point clouds, and environmental audio waveforms, and output unified-dimensional single-modal topological features. Cross-modal topology fusion module The four single-modal features are concatenated along the temporal dimension, and a global topological attention mechanism is used to complete the deep association and fusion of the four modalities, eliminating the modal semantic barrier and generating a globally unified environment-aware feature. High-dimensional manifold sensing module By performing high-dimensional manifold space transformation on the fusion features, we can extract deep representations of environmental geometric topology, semantic associations, and spatiotemporal evolution, and construct the cognitive base features of the abstract world. Heterogeneous MoE Expert Parallel Inference Module Multiple heterogeneous expert networks are set up, and a small number of effective experts are dynamically activated through routing algorithms to complete large-scale and efficient inference. At the same time, routing balance loss is constrained to avoid expert idleness and failure. Physical World Modeling Module The environment structure is analyzed based on the scene graph network, and the world state features with physical constraints are output by the dynamic prediction network. Force value truncation is used to prevent abnormal deviation of the physical state. Hierarchical task planning module Complete the high-level global task planning and low-level action instruction mapping, and output two-level decision results: planning semantic features and execution action instruction features. Robot joint motion decoding module Using historical action sequences as queries and task planning features as memory encoding, trajectory features are generated through a Transformer decoder, mapped to robot joint control variables, and speed safety constraints are applied. Autonomous Evolution Iteration Engine Based on the model's comprehensive loss and fitness score, the manifold feature space is driven to iteratively evolve, thereby enabling the model to autonomously improve its cognitive ability. Fault self-healing repair module The forward inference process detects numerical anomalies, gradient anomalies, and module malfunctions in real time, automatically triggering repair processes to ensure uninterrupted system operation. Automatic version upgrade module The model version is automatically determined based on the evolutionary fitness threshold, and high-quality inference weights and decision-making logic are solidified. Architectural self-developed modules It autonomously searches for the optimal network layer structure, expert allocation strategy, and attention topology to continuously optimize the performance of the model's inherent architecture. Distributed intelligent agent bee colony cluster module It dynamically generates distributed intelligent agent units, constructs a global topology network, completes multi-agent collaborative decision-making through cluster consensus voting, and adaptively adjusts the cluster size and cooperation strategy. Distributed training management module It employs distributed parallelism, mixed precision acceleration, gradient pruning, and cosine annealing learning rate scheduling to complete the entire training update process, synchronize global weights, and support breakpoint saving and loading as well as emergency safety shutdown. The system takes in text commands, time-series video, 3D point cloud, environmental audio, and historical action sequences, and processes them sequentially through encoding, fusion, manifold perception, MoE inference, physical modeling, task planning, and joint decoding. Simultaneously, it performs fault repair, cluster collaboration, autonomous evolution, version upgrade, and architecture iteration, and finally outputs text understanding results and robot joint control commands to complete a fully closed-loop autonomous intelligent operation. Beneficial effects Four-modal unified manifold coding fusion enables accurate recognition of the environment across all dimensions, completely solving the problem of fragmented modal features; The heterogeneous MoE expert parallel architecture balances the expressive power of large models with the efficiency of inference computing power. Physical dynamics constraint modeling significantly reduces the discrepancy between virtual cognition and real physical scenes; Layered decision-making combined with trajectory decoding closed loop significantly improves the precision of robot joint control and the smoothness of motion; It has four built-in autonomous engines: self-healing, evolution, upgrade, and architecture development, giving the model the ability to continuously optimize itself. Distributed intelligent agent swarm collaboration, multi-unit linkage operation adapts to complex industrial scenarios; Industrial-grade distributed training management ensures stable numerical values, recoverable faults, and secure and reliable deployment and maintenance. Attached Figure Description Figure 1 Flowchart of the overall architecture of the multimodal manifold autonomous intelligent robot control system This attached diagram illustrates the hierarchical operational architecture of the invention's overall system, which consists of three vertical execution links: a main-link perception, reasoning, and control branch; a parallel autonomous intelligent iterative branch; and a distributed training and management branch. The main link sequentially completes multimodal encoding, fusion, manifold perception, MoE inference, physical modeling, task planning, and joint decoding, outputting robot control commands. The parallel branch simultaneously implements fault repair, agent clustering, autonomous evolution, version upgrades, and architecture development. The management branch completes model training updates and operational support. These three links work together to form a complete autonomous intelligent robot control system. Figure 2 Flowchart of the internal hierarchy for encoding multimodal manifolds This diagram illustrates the internal execution order of the four-modal coding. Independent feature encoding is completed sequentially for text, video, 3D point cloud, and audio, transforming the four types of heterogeneous raw input data into topological features of the same dimension. This provides a standardized feature foundation for subsequent cross-modal fusion, enabling complete extraction of information from the entire scene. Figure 3 Internal flowchart of expert parallel inference for MoE This diagram illustrates the inference process of a heterogeneous expert network. Features are first normalized and preprocessed, then effective expert blocks are dynamically selected and activated by the routing unit. The inference results of multiple experts are aggregated, and the routing balance loss is calculated synchronously. This reduces the consumption of ineffective computing power and improves the efficiency of inference operation while ensuring the expressive power of the model. Figure 4 Flowchart of robot decision-making and control closed loop This diagram illustrates the closed-loop control process from physical state cognition to actual robot operation. The decision-making level is broken down by a hierarchical planning module based on physical world features, and then joint control quantities are generated by a trajectory decoder. After adding physical speed safety constraints, the execution signal is output, completing the entire closed-loop implementation from intelligent decision-making to physical action. Figure 5 Flowchart for Autonomous Intelligent Iterative Operation and Maintenance This diagram illustrates the system's self-optimization and iteration process. Based on model running loss and fitness scores, it sequentially completes fault repair, cluster collaboration, manifold evolution, version upgrade, and architecture self-development, enabling the system to autonomously optimize performance during uninterrupted operation and possess continuous learning and self-improvement capabilities. Figure 6 Distributed training management and operation flowchart This diagram illustrates the entire industrial-grade distributed training process: after batch data completes mixed-precision inference, multi-task loss is calculated, gradient inversion and pruning constraints are executed, weight updates and learning rate adjustments are completed, parameters are globally synchronized across multiple GPUs, training snapshots are stored periodically, and automatic backup and shutdown are implemented in abnormal scenarios to ensure stable and reliable model training. Detailed Implementation The system is built using the PyTorch deep learning framework. During the data input phase, text-based operation instructions for industrial scenarios, monitoring time-series videos, 3D point clouds of equipment, on-site environmental audio, and historical robot motion trajectories are loaded. During the encoding phase, the four types of encoders respectively complete modal feature extraction and output topological features of the same dimension; During the fusion phase, global attention completes cross-modal association aggregation; The manifold perception stage is mapped to a high-dimensional abstract cognitive space; MoE routes dynamically activate expert networks to complete inference computations; Physical model constraints correct environmental states; Layered planning breaks down global tasks into underlying actions; The trajectory decoder generates safe and compliant joint control quantities; During training and operation, fault detection and repair, multi-agent cluster consensus, manifold evolution and iteration, version determination and upgrade, and autonomous architecture search are initiated simultaneously. The distributed module completes multi-card weight synchronization, mixed-precision gradient update, breakpoint snapshot storage, and automatic backup in case of emergency shutdown triggered by an anomaly. The final output is text semantic reasoning results and robot joint servo control signals, which drive the robotic arm to complete industrial operations such as grasping, placing, and transporting, realizing fully unmanned autonomous intelligent control.

Claims

1. A multimodal manifold autonomous intelligent robot control system, characterized in that: It includes a multimodal manifold coding module, a cross-modal topology fusion module, a high-dimensional manifold perception module, a heterogeneous MoE expert parallel reasoning module, a physical world modeling module, a hierarchical task planning module, a robot joint motion decoding module, an autonomous evolution and iteration engine, a fault self-healing and repair module, an automatic version upgrade module, an architecture self-development module, a distributed intelligent agent bee colony cluster module, and a distributed training and management module. The multimodal manifold coding module performs single-modal feature encoding on text, video, point cloud, and audio respectively; The cross-modal topology fusion module splices and fuses four-way features; the high-dimensional manifold perception module extracts high-dimensional abstract cognitive features; the heterogeneous MoE expert parallel reasoning module enables efficient reasoning; the physical world modeling module outputs constrained physical states; the hierarchical task planning module splits global tasks into low-level actions; the robot joint action decoding module generates safe joint control commands; the fault self-healing, autonomous evolution, version upgrade, and architecture development modules constitute an autonomous optimization system; the distributed intelligent agent swarm module enables multi-unit collaboration; and the distributed training and management module completes model training and maintenance.

2. The system according to claim 1, characterized in that: The multimodal manifold coding module includes a text token embedding encoder, a video manifold visual encoder, a 3D point cloud topology encoder, and an audio field encoder. The four types of encoders output topological features of the same dimension.

3. The system according to claim 1, characterized in that: The heterogeneous MoE expert parallel inference module includes an expert routing allocation unit and multiple sets of heterogeneous expert blocks. The routing unit dynamically activates a specified number of expert networks, aggregates inference results, and calculates routing balance loss.

4. The system according to claim 1, characterized in that: The physical world modeling module incorporates a scene graph network and a dynamics prediction network. The output features are truncated by force value limiting to constrain the range of physical state operation.

5. The system according to claim 1, characterized in that: The robot joint motion decoding module adopts a Transformer decoder structure, using historical motion sequences as queries and planning features as memory, to output robot joint control quantities with velocity constraints.

6. The system according to claim 1, characterized in that: The distributed intelligent agent bee colony cluster module can dynamically generate intelligent agent units, build a global topology network, complete collaborative decision-making through cluster consensus voting, and adaptively adjust the scale of cluster collaboration.

7. The system according to claim 1, characterized in that: The distributed training management module adopts distributed parallelism, mixed precision inference, gradient pruning, and cosine annealing learning rate scheduling, and has breakpoint saving, breakpoint loading, and automatic backup functions for emergency shutdown.

8. The system according to claim 1, characterized in that: The system operates three links simultaneously: perception-inference control, autonomous iterative optimization, and distributed training and management. It inputs multimodal industrial scene data and outputs textual semantic results and robot joint control commands to achieve fully closed-loop autonomous intelligent industrial operation.