A robot heterogeneous device unified interconnection architecture based on a CXL switching network
Patent Information
- Application Number
- CN202611309876.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-27
- Publication Date
- 2026-09-25
AI Technical Summary
这样会导致高带宽感知任务和高实时控制任务共用同一传输和调度路径,难以同时兼顾吞吐能力与时延要求
本发明构建了一种不依赖CPU中转的机器人异构设备统一高速互连架构。利用逻辑池化内存空间结合点对点直连避开CPU中转,在消除数据复制开销与主机带宽瓶颈的同时,显著降低传输延迟;通过对观测流与控制流的分流处理,兼顾高带宽感知与低时延控制的双重需求,通过控制桥接模块还实现高频闭环反馈,提升控制稳定性。本发明的机器人异构设备统一高速互连架构消除了多节点间的重复数据搬运,实现了共享数据的高效复用,并支持通过交换网络管理器动态编排资源及平滑扩展节点,大幅提升了系统协同效率、资源利用率与可扩展性。
Smart Images

Figure CN122824699A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of high-speed data transmission technology, specifically relating to a unified interconnection architecture for heterogeneous robot devices based on a CXL switching network. Background Technology
[0002] As robotics technology continues to evolve towards embodied intelligence, autonomous collaboration, and operation in complex environments, the computational, perception, control, and communication resources within robot systems are exhibiting significant heterogeneity and high coupling characteristics. Robots not only need to connect to multimodal perception devices such as cameras, depth cameras, LiDAR, microphone arrays, inertial measurement units, force sensors, and tactile sensors to acquire environmental observation information for feature extraction, fusion analysis, and semantic reasoning, but also need to collect real-time body state data from various execution units, including joint positions, torques, body pose, foot contact states, end effector forces, and battery states. Within a short control cycle, they must integrate this perception and state information to complete motion processing and closed-loop control, thus simultaneously handling a hybrid load of high-bandwidth, low-frequency inference and low-latency, high-frequency control.
[0003] The following are existing solutions for interconnecting and coordinating heterogeneous devices in robot systems: This solution employs a centralized interconnection approach based on a central processing unit (CPU). It utilizes a star topology, connecting peripheral devices such as multiple sensors, actuator controllers, communication interfaces, and accelerators to the main control platform via PCIe, Ethernet, USB, CAN, serial ports, or dedicated buses. The CPU acts as the sole hub of the system, monopolizing the entire process of device management, data aggregation, task scheduling, and result distribution. In AI inference scenarios, heterogeneous accelerators (GPUs / NPUs) or other AI accelerators function merely as passive external computing units: the CPU transports sensor data to the accelerator, processes it, and then sends the results back to the control unit.
[0004] Accelerator access solutions for heterogeneous computing collaboration. With the application of multimodal models in robotics, some solutions are introducing multiple heterogeneous accelerators to handle tasks such as perception and reasoning, path planning, motion generation, or control optimization. These solutions typically use high-speed interconnects such as PCIe as the connection basis between different computing units, and combine shared memory, DMA data transfer, and other methods to achieve data exchange and collaborative processing between heterogeneous computing nodes. At its core, however, the CPU remains the central element: the CPU exclusively handles device enumeration, address mapping, memory management, and task scheduling; perception data, control data, and intermediate computation results must all be transferred or mapped through main memory visible to the CPU.
[0005] The existing solutions for interconnecting and coordinating heterogeneous devices in the above-mentioned robot systems have the following drawbacks: Existing robot systems typically use a mix of USB (Universal Serial Bus), MIPI (Mobile Industry Processor Interface), or Ethernet to access sensors; CAN (Controller Area Network), EtherCAT, or serial ports to access controllers; and PCIe (Peripheral Component Interconnect Express) to access accelerators. The lack of a unified resource organization and access mechanism between different data links results in complex internal interface types, high hardware and software adaptation costs, difficulties in device expansion and architecture migration, and hinders the formation of a unified heterogeneous resource management system.
[0006] Data relies on host relay, resulting in high transmission overhead. In existing solutions, multi-sensor data, control feedback data, and intermediate feature quantities typically need to first enter the central processing unit or main memory before being forwarded by the host to the target accelerator or control node. This method introduces multiple data copying, cache movement, and software intervention processes, which not only consume host bandwidth but also increase transmission latency. When the robot simultaneously processes high-bandwidth sensing flow and high-frequency control flow, the host relay path can easily become a performance bottleneck.
[0007] The efficiency of computing node collaboration is low. Because data exchange between nodes such as GPUs, AI accelerators, and controllers relies on host scheduling, DMA transfer, or software messaging mechanisms, there is a lack of direct data pathways between devices. This results in long data links and high synchronization costs, making it difficult to adapt to the evolving trend of high-level inference and low-level execution collaboration in robotic systems.
[0008] Insufficient resource sharing capabilities. Data such as sensor caches, intermediate features, control descriptors, motion commands, and state variables often need to be repeatedly stored and transported across multiple nodes, which is detrimental to reducing system latency and improving resource utilization. When the robot's task mode changes, existing solutions often struggle to dynamically reorganize and flexibly arrange memory and link resources.
[0009] The perception and control links lack differentiated organization. Observational and control data in robots differ significantly in data volume, latency requirements, priority, and processing frequency. However, existing solutions often only divide tasks at the software level, lacking mechanisms for priority mapping and path optimization of different data streams at the interconnect architecture level. This results in high-bandwidth perception tasks and high-real-time control tasks sharing the same transmission and scheduling path, making it difficult to simultaneously meet throughput and latency requirements.
[0010] Therefore, there is an urgent need in this field for a unified high-speed interconnection system for heterogeneous robot devices that does not rely on CPU relay and effectively shortens the data exchange path, so as to reduce transmission latency and meet the dual requirements of robots for high-bandwidth multimodal perception and strong real-time motion control. Summary of the Invention
[0011] The purpose of this invention is to provide a unified high-speed interconnection system for heterogeneous robot devices that does not rely on CPU relay and effectively shortens the data exchange path. While reducing transmission latency, it meets the dual requirements of robots for high-bandwidth multimodal perception and strong real-time motion control.
[0012] To achieve the above objectives, the technical solution adopted by the present invention is as follows: The first aspect of this invention provides a unified interconnection architecture for heterogeneous robot devices based on a CXL switching network, including a CXL switch, a switching network manager, and an observation bridge module, a control bridge module, an artificial intelligence processor module, and a motion control accelerator module coupled to the CXL switch, wherein: The switching network manager is configured to build a first logical data path, a second logical data path, and a third logical data path that respectively support CXL Peer-to-Peer (CXL P2P) access, and to build and manage a logical pooled memory space that is shared by the observation bridge module, the control bridge module, the artificial intelligence processor module, and the motion control accelerator module. The logical pooled memory space is mapped to a globally unified shared memory address space. The observation bridging module is configured to convert the raw observation data from the robot's multimodal sensing device into standardized observation data and write it into the shared memory address space. The control bridge module is configured to convert the body state data of the robot execution unit into a standardized state representation and write it into the shared memory address space, and to read motion control instructions from the shared memory address space through the second logical data path and send them to the robot execution unit. The artificial intelligence processor module is configured to obtain the standardized observation data from the shared memory address space through the first logical data path, run a multimodal inference model to generate high-level semantic results, and write the high-level semantic results into the shared memory address space; The motion control accelerator module is configured to obtain the standardized state representation from the shared memory address space through the second logical data path and obtain the high-level semantic result from the shared memory address space through the third logical data path, run the motion control model to generate the motion control instruction, and write the motion control instruction into the shared memory address space; The standardized observation data, the high-level semantic results, and the standardized state representation are exchanged between modules in a zero-copy manner through the shared memory address space, without the need for the central processing unit to participate in data transfer.
[0013] This invention uses the CXL Fabric interconnect network as a unified interconnect base for heterogeneous robot devices. An observation bridging module converts raw observation data from multimodal sensing devices into standardized observation data and writes it into shared memory, which can then be directly accessed by the multimodal inference model via CXL P2P. A control bridging module converts the state data of the robot's execution unit into standardized state representations and writes them into shared memory, which can then be directly accessed by the motion control model via CXL P2P. Furthermore, the control bridging module retrieves motion control instructions from the shared memory via CXL P2P and converts them into executable drive control messages for the robot's execution unit. The observation bridging module works in conjunction with the artificial intelligence processor module, which runs the multimodal inference model to achieve high-level perception understanding and semantic reasoning. The control bridging module works in conjunction with the motion control accelerator module, which runs the motion control model to generate actions. By separating the observation and control flows, the dual requirements of large-scale intelligent perception and strong real-time closed-loop control in the robot system are met. Building upon this, and combining the logically pooled memory space and the globally unified shared memory address space, CXL enables direct read and write of shared data between access devices via P2P access, achieving zero-copy data interaction without relying on a central processing unit. Dynamic resource orchestration and unified management are achieved through the switching network manager, enhancing system flexibility and overall operational efficiency.
[0014] In some implementations, the motion control accelerator module is further configured to cache the high-level semantic results obtained from the shared memory address space through the third logical data path as the currently valid high-level semantic results; and to run periodically according to a preset high-frequency control cycle. In each run, based on the currently valid high-level semantic results and the current ontology state data obtained through the second logical data path, the motion control model is run to generate the motion control instructions and write them into the shared memory address space until a new high-level semantic result is obtained through the third logical data path.
[0015] In some implementations, the AI processor module runs high-level semantic results (including task instructions, target poses, safety constraint vectors, etc.) generated by a multimodal inference model (such as a vision-language model, VLM), with a relatively low update frequency (e.g., 10Hz~30Hz); while the motion control accelerator module runs the motion control model at a significantly higher frequency (e.g., 500Hz~2kHz) to meet the robot's requirements for high-frequency closed-loop control. That is, the motion control accelerator module is configured to reuse the same currently valid high-level semantic result across multiple consecutive control cycles between two high-level semantic result updates.
[0016] Furthermore, both the motion control accelerator module and the control bridge module are configured to operate synchronously based on the same preset high-frequency control cycle (e.g., 500Hz~2kHz) to ensure that they exchange data under the same time reference and form a complete closed-loop control.
[0017] When the AI processor module generates and writes new high-level semantic results, the motion control accelerator module learns of this through a doorbell or descriptor update event and refreshes the cached high-level semantic results. Subsequent control cycles will continue to run based on the new semantics.
[0018] Specifically, the control bridge module is also configured to operate periodically according to the high-frequency control cycle, with each operation executing the following sequentially: The robot's execution unit's body state data is converted into a standardized state representation and written into the shared memory address space; Action control instructions are read from the shared memory address space through the second logical data path; The motion control commands are converted into drive control messages and sent to the robot execution unit; The state feedback generated by the robot execution unit is collected to update the body state data for the next high-frequency control cycle.
[0019] The motion control accelerator module is configured to operate synchronously based on the same preset high-frequency control cycle as the control bridge module, so as to ensure that the control bridge module and the motion control accelerator module exchange data on the same time reference and form a complete closed-loop control.
[0020] In some specific embodiments, the motion control accelerator module is configured to write the generated motion control instructions (such as target joint torque, position increment, etc.) into a pre-allocated instruction buffer in the shared memory address space during each high-frequency control cycle, and notify the control bridge module through the doorbell register. The control bridge module is configured to, in response to the doorbell register notification, read the motion control command from the instruction buffer through the second logical data path, convert it into a drive control message (such as EtherCAT PDO or CANopen RPDO) that matches the robot target execution unit, and send it out. After the robot execution unit completes the action output, it collects the actual state feedback generated by the robot execution unit, converts the actual state feedback into the body state data required for the next high-frequency control cycle, and overwrites the original body state data in the shared memory address space for the motion control accelerator module to read in the next control cycle.
[0021] In some implementations, the observation bridging module includes an interface adaptation module, a time synchronization module, and a modal normalization module, wherein: The interface adaptation module is configured to interface with the heterogeneous physical layer and link layer protocols of the multimodal sensing device, and restore the original observation payload to an internal data object; The time synchronization module is configured to apply timestamps to the internal data objects based on a unified reference clock and construct a synchronized observation set according to a fixed time window; The modal normalization module is configured to convert the synchronous observation set into standardized observation data that can be directly called by the multimodal inference model. The standardized observation data includes at least one of the following: image tensor, depth tensor, point cloud tensor, voxel grid, Mel spectrum, audio lexical, and state vector.
[0022] Furthermore, the observation bridging module is also configured to generate an observation descriptor for each of the standardized observation data, and the control bridging module is also configured to generate a control state descriptor for each of the standardized state representations; The observation descriptor and the control state descriptor include at least data type, dimension, load address offset, priority identifier, and validity flag; The observation descriptor and the control state descriptor are stored in a pre-allocated descriptor circular queue in the shared memory address space, and are notified to the corresponding artificial intelligence processor module and motion control accelerator module through the doorbell register.
[0023] Furthermore, the observation bridging module is also configured to perform precoding processing on the standardized observation data. The precoding processing includes at least data precision quantization, normalization, and memory layout alignment to adapt to the cache line access characteristics of the shared memory address space.
[0024] In some specific implementations, the observation bridging module includes a first bridging endpoint and a first CXL memory extension device (i.e., a Type 3 device) that does not contain computational logic. The first CXL memory extension device primarily provides a large-capacity shared memory space to carry the standardized observation data, which is incorporated into the logically pooled memory space as part of a globally unified shared memory address space.
[0025] The first bridging endpoint includes an interface adaptation module, a time synchronization module, a modality normalization module, a preprocessing module, an observation descriptor generation module, and a CXL transaction encapsulation module, wherein: The interface adapter module is configured to be compatible with multiple external bus protocols (including but not limited to MIPI CSI-2, GigE Vision, USB, I2C, SPI, UART, etc.). It performs physical layer and link layer decapsulation on raw signals from cameras, LiDAR, IMU, and force sensors, restoring the raw payload to an internal data object in the format of the first bridge endpoint's internal data bus, and attaching basic metadata such as device ID and original serial number. The time synchronization module is configured to apply a unified timestamp to each modality's data based on the global reference clock provided by the CXL switching network. To address the issue of inconsistent sampling frequencies among sensors, this module performs frame alignment, interpolation, or resampling according to a fixed inference time window, constructing a synchronized observation set to ensure strict alignment of multimodal data in the time dimension. The modal normalization module is configured to convert synchronized internal data objects into a standardized data structure that can be directly consumed by the multimodal inference model. For example, it converts image / depth data into normalized tensors, point clouds into voxel meshes or distance maps, and force data into fixed-dimensional state vectors to eliminate data format heterogeneity. The preprocessing module is configured to quantize, normalize, and memory align the standardized modal tensors or state vectors, mainly to optimize the standardized data before inference. The observation descriptor generation module is configured to generate an observation descriptor for each preprocessed observation data. This descriptor, as a metadata header, includes at least: modality type, sensor identifier, timestamp, tensor dimension, data type, payload address offset in shared memory, and data priority identifier. The CXL transaction encapsulation module is configured to perform a CXL.mem write operation, writing the preprocessed observation data into a designated area in the logical pooled memory space (i.e., the first CXL memory extension device); at the same time, it submits the generated observation descriptors to the descriptor ring buffer located in shared memory, and notifies the artificial intelligence processor module that new data is ready through the doorbell register.
[0026] In some embodiments, the control bridge module includes a state data standardization encoding module, an action decoding module, and an execution unit protocol mapping module, wherein: The state data standardization encoding module is configured to unify the execution units, transform coordinates, and restore the scale of the robot's execution unit's body state data, so as to generate a standardized state representation that can be directly called by the motion control model. The motion decoding module is configured to perform unit conversion of the motion control command (restore the motion control model output from normalized values to physical units of the execution unit), amplitude limiting clipping (constrained by joint angle, speed or torque safety limits), dead zone compensation (compensating for the nonlinear dead zone of the actuator), and smooth interpolation (low-pass filtering or linear / second-order interpolation on commands of adjacent control cycles). The execution unit protocol mapping module is configured to map the processed motion control instructions into data objects of the target bus protocol and generate drive control messages to be sent to the robot execution unit.
[0027] In some implementations, the execution unit protocol mapping module is configured to support at least one of the following industrial control protocols: EtherCAT, CAN, CANopen, Modbus RTU, RS485 / 422 serial communication protocol, and proprietary multi-axis drive protocol; the drive control message includes at least one of position control command, speed control command, torque control command, or current control command.
[0028] In some implementations, the high-level semantic results are represented in the form of structured communication primitives, which include task semantic identifiers, target pose tensors, trajectory parameterization descriptions, implicit feature vectors and their metadata descriptors; The motion control accelerator module is configured to parse the structured communication primitives based on the metadata descriptor, and input the parsed target information and implicit feature vectors into the motion control model.
[0029] In some implementations, the control bridge module includes a second bridge endpoint and a second CXL memory extension device (i.e., a CXL Type 3 device) that does not contain computational logic. The second CXL memory extension device primarily provides a shared memory space carrying the ontology state data, which is incorporated into the logically pooled memory space as part of a globally unified shared memory address space.
[0030] The second bridging endpoint includes a status acquisition and standardization module, a control status descriptor generation module, a CXL transaction encapsulation module, an action decoding and protocol mapping module, and a closed-loop feedback module, wherein: The state acquisition and standardization module is configured to acquire the state data of the robot execution unit in a periodic or event-triggered manner, and perform unit unification, coordinate transformation and scale restoration to generate a standardized state representation that can be directly called by the motion control model. The control state descriptor generation module is configured to generate a control state descriptor for each of the standardized state representations. The control state descriptor includes at least: robot identifier, execution chain identifier, control cycle number, timestamp, state vector dimension, data type, valid bit, fault bit, security level, priority, execution deadline, and write-back address. The CXL transaction encapsulation module is configured to use a combination of fixed-period state updates and abnormal event preemption to write the ontology state data into the second CXL memory expansion device, and submit the control state descriptor to the descriptor circular queue, and notify the action control accelerator through the doorbell register; The motion decoding and protocol mapping module is configured to read the motion control instructions from the motion control accelerator module, perform unit conversion, amplitude limiting and protocol mapping, so as to generate drive control messages that can be executed by the robot execution unit. The closed-loop feedback module is configured to notify the state acquisition and standardization module to acquire the latest state of the execution unit after the execution unit completes the action output, and update the body state data in the second CXL memory expansion device for use by the action control accelerator module in the next control cycle.
[0031] In some implementations, the switching network manager is further configured to prioritize the transmission of the second logical data path over the first logical data path, and to pre-allocate a fast transmission buffer dedicated to body state data and action control instructions in the logical pooled memory space, so that the transmission of body state data upload, action control instruction issuance and abnormal event preemption descriptor meets the preset maximum end-to-end latency and jitter requirements, so as to support closed-loop control running with a preset control cycle.
[0032] Specifically, the abnormal event includes at least one of the following: emergency stop, collision, overcurrent, instability, joint over-limit, or battery failure.
[0033] Specifically, the control bridge module is configured to generate an event descriptor with the highest priority in response to the abnormal event, and send it to the motion control accelerator module in a preemptive manner through the second logical data path to trigger the motion control accelerator module to execute a safety braking or degraded control strategy.
[0034] Specifically, the switching network manager is further configured to configure a shared memory window with a size larger than that of the second logical data path and a descriptor queue with a depth greater than that of the second logical data path for the first logical data path, in order to adapt to the high bandwidth and batch transmission requirements of multimodal sensing data; and to configure a fast transmission buffer dedicated to periodic body state data and action control instructions for the second logical data path, in order to adapt to the high frequency closed-loop control requirements.
[0035] In a specific and preferred embodiment, the switching network manager is configured to perform differentiated Quality of Service (QoS) configurations on the first logical data path and the second logical data path. The switching network manager maps the second logical data path to a higher-priority virtual channel (VC1) or traffic class (TC1) in the CXL switch, while mapping the first logical data path to a lower-priority virtual channel (VC0 or TC0), ensuring that control transactions are prioritized for forwarding during link congestion. Simultaneously, the switching network manager pre-allocates a dedicated Fast Path Buffer in the logical pooled memory space for the ontology state data and the action control instructions. This buffer has a small cache line depth and a short allocation / release path, dedicated to the high-frequency periodic data exchange of the second logical data path, avoiding access contention caused by sharing the same memory area with high-bandwidth observation data.
[0036] Meanwhile, through the combined use of the priority configuration and the fast transmission buffer, the end-to-end transmission delay of the body status data upload, the action control command issuance, and the abnormal event preemption descriptor (such as the emergency stop event descriptor) is guaranteed not to exceed a preset threshold (e.g., ≤100 μs), and the periodic delay fluctuation (Jitter) is also controlled within a preset range (e.g., ≤10 μs), thereby meeting the real-time and deterministic requirements of the closed-loop control task running at the preset high-frequency control cycle (e.g., 1 kHz).
[0037] In some implementations, the switching network manager is further configured to pre-allocate isolated address windows for the observation bridge module, the control bridge module, the artificial intelligence processor module, and the action control accelerator module in the logically pooled memory space, and configure access control lists for each address window to restrict each module to accessing only its corresponding data area. This memory isolation mechanism effectively prevents out-of-bounds writing that could corrupt critical data (such as action control instructions or high-level semantic results) of other modules due to anomalies or firmware defects in any module. It also reduces the reachability of direct memory access (DMA) to improve system security and decreases cache pollution caused by shared access among multiple modules, helping to ensure the time determinism of the high-frequency closed-loop control link.
[0038] In some implementations, the observation bridging module and the control bridging module are further configured to perform integrity verification on the data payload written to the shared memory address space, and set corresponding integrity verification flags in the observation descriptor and the control status descriptor; the artificial intelligence processor module and the motion control accelerator module verify the integrity of the data based on the integrity verification flags before reading the data.
[0039] In a specific implementation, when the AI processor module reads standardized observation data through the first logical data path, and when the motion control accelerator module reads standardized state representations through the second logical data path, both are configured to first parse the integrity check flag in the corresponding descriptor. If the flag indicates that the check passes, the data payload is read and used according to the payload address offset. If the flag indicates that the check fails, the frame data is ignored and the system waits for the next round of new data writing. This mechanism effectively detects data corruption caused by CXL link noise, memory bit flips, or incomplete writes, preventing unsafe control decisions based on error perception or error states.
[0040] In some implementations, the motion control accelerator module is further configured to dynamically adjust the output range of the motion control model or trigger a preset safety response strategy when the high-level semantic result contains a safety constraint vector or an abnormal state. The safety response strategy includes deceleration, force limiting, reverting to a safe pose, or stopping output.
[0041] In some specific implementations, the motion control accelerator module is configured to parse high-level semantic results from structured communication primitives obtained from the shared memory address space; if the high-level semantic results contain safety constraint vectors (such as maximum angles, angular velocities, and torque limits for each joint, or obstacle repulsion weights within the workspace) or abnormal state markers (such as indicators of unusable environmental areas), then: The output of the motion control model is constrained based on the safety constraint vector, for example, by saturation limiting at the policy network output layer, adding obstacle penalty terms, or reducing the motion sampling distribution, so that the generated motion control commands are always within the allowable range of the safety constraints; and / or When the current action is determined to be infeasible based on the safety constraint vector, a preset safety response strategy is directly triggered. The safety response strategy includes at least: reducing the running speed, limiting the maximum output torque, reverting to a predefined safe posture, or stopping the output of action control commands. The triggered safety response strategy is converted into a corresponding drive control message by the control bridging module and sent to the robot execution unit.
[0042] In some implementations, the switching network manager is also configured to monitor the link health and congestion status of the first logical data path, the second logical data path, and the third logical data path, and to reconfigure the logical data path or adjust the quality of service level to maintain the availability of the control link when a link failure or continuous congestion exceeding a threshold is detected.
[0043] In some specific implementations, the switching network manager is configured to periodically or on event triggers query the link registers of each CXL port and logical data path to obtain link health status (such as Link Up / Down, LaneWidth Degrade, CRC Error Count) and congestion proxy indicators (such as incomplete transaction count, Credit exhaustion frequency, descriptor queue occupancy rate, etc.). When a link failure or continuous congestion exceeding a preset threshold (descriptor queue occupancy rate continuously exceeding the threshold) is detected in any logical data path, the switching network manager performs at least one of the following operations: attempts to retrain or switch to a redundant physical path for the faulty link, adjusts the service level weight of the affected logical data path to give the second logical data path priority bandwidth and low queue occupancy, or reallocates the shared memory address window. The protection priority for the second logical data path is higher than that for the first logical data path and the third logical data path, so as to ensure the availability of the high-frequency closed-loop control link.
[0044] In some implementations, the switching network manager is configured to enumerate the observation bridge module and the control bridge module upon system power-on or task startup, and complete the initialization configuration of the logical data paths. Further, the switching network manager is configured to discover and initialize the observation bridge module and the control bridge module after they access the CXL switching network, and configure the shared memory window, descriptor queue, access permissions, and service priority of the first and second logical data paths. The switching network manager is also configured to configure the shared memory window size, descriptor queue depth, and service level of the first and second logical data paths, and to allocate a CXL virtual channel independent of the first logical data path to the second logical data path, thereby completing the initialization configuration of the logical data paths.
[0045] In a specific implementation, the artificial intelligence processor module and the motion control accelerator module are both Type 2 devices, each integrating a dedicated computing core and local private memory for running inference models or motion control models. The local private memory is configured to be included in the logical pooled memory space.
[0046] In some implementations, the architecture further includes a central processing unit (CPU) configured to run a general-purpose operating system (such as Linux or ROS2 distribution) and upper-layer application software. The CPU is primarily responsible for system-level task scheduling, resource management, lifecycle maintenance, and non-real-time business logic processing. It should be specifically noted that the CPU does not participate in the real-time data interaction path between the multimodal observation data, the high-level semantic results, and the action control instructions, nor does it perform read / write operations on core business data in the logically pooled memory space, to avoid the impact of uncertainties and jitter from the general-purpose operating system on the high-frequency closed-loop control link.
[0047] The system also includes a communication device for connecting the robot to an external network. In some preferred embodiments, the communication device is a CXL Type 3 memory expansion device, coupled to the CXL switch via a CXL interface and incorporated into the management scope of the logically pooled memory space via the switch network manager. The communication device is configured as follows: Establish a connection with an external remote server or cloud for uploading robot status logs or downloading updated model weights; Establish cluster communication links with other robot nodes for data synchronization in multi-robot collaborative tasks; Receive manual intervention instructions or task assignment instructions from the remote console.
[0048] Since the communication device is a Type 3 device (possessing only storage capabilities and no computing core), the data it interacts with from the outside world (such as model files, logs, and task instructions) needs to be scheduled to a specific buffer in the logical pooled memory space via the switching network manager, and then loaded into general memory by the CPU for processing, or the AI processor module and the motion control accelerator module are notified to load it. This design ensures that external communication traffic and internal real-time control traffic are isolated from each other in physical memory space and logical data paths, preventing external network interference with the stable operation of the robot.
[0049] A second aspect of this invention provides a robot perception and closed-loop control method based on a CXL switching network, comprising: During system power-on or task startup, the switching network manager enumerates and initializes the CXL switch, logical pooled memory space, observation bridge module, control bridge module, artificial intelligence processor module, and motion control accelerator module. The logical pooled memory space is a globally unified shared memory address space built and managed by the switching network manager for the observation bridge module, control bridge module, artificial intelligence processor module, and motion control accelerator module to access. The switching network manager also builds a first logical data path, a second logical data path, and a third logical data path to support CXL Peer-to-Peer access, respectively. The observation bridging module converts the raw observation data from the robot's multimodal sensing device into standardized observation data and writes it into the shared memory address space. The standardized observation data is obtained from the shared memory address space via the first logical data path through the artificial intelligence processor module, a multimodal inference model is run to generate high-level semantic results, and the results are written into the shared memory address space. The control bridge module collects the body state data of the robot execution unit, converts it into a standardized state representation, and writes it into the shared memory address space. The motion control accelerator module obtains the standardized state representation from the shared memory address space via the second logical data path and the high-level semantic result from the shared memory address space via the third logical data path, runs the motion control model to generate the motion control instruction, and writes the motion control instruction into the shared memory address space. The control bridge module reads the motion control instructions from the shared memory address space via the second logical data path, converts them into drive control messages, and sends them to the robot execution unit; and The control bridge module collects the status feedback generated after the robot execution unit completes the action output, and updates the body status data in the shared memory address space so that the motion control accelerator module can read it in the next control cycle to form closed-loop control.
[0050] In some implementations, both the artificial intelligence processor module and the motion control accelerator module are CXL Type 2 devices, and the logical pooled memory space includes local private memory of the artificial intelligence processor module and the motion control accelerator module, which can be accessed by each module through the CXL Peer-to-Peer method. The high-level semantic results generated by the artificial intelligence processor module and the action control instructions generated by the action control accelerator module are both written into their respective local private memory.
[0051] In some implementations, the step of generating high-level semantic results includes post-processing the output high-level semantic results: The high-level semantic results are encapsulated into structured communication primitives, which include at least task semantic identifiers, target pose tensors, trajectory parameterization descriptions, and latent feature vectors. If the high-level semantic result contains a safety constraint vector or an abnormal state marker, then the safety constraint vector or abnormal state marker is written into the structured communication primitive. The safety constraint vector includes at least the maximum angle, angular velocity and torque limit of each joint, or the repulsive force weight of obstacles in the workspace.
[0052] In some implementations, the step of running the motion control model to generate the motion control commands includes: The motion control accelerator module caches the currently valid high-level semantic results obtained through the third logical data path; and runs periodically according to a preset high-frequency control cycle. In each run, based on the currently valid high-level semantic results and the current ontology state data obtained through the second logical data path, the motion control model is run, the motion control instructions are generated and written into the shared memory address space, until a new high-level semantic result is obtained through the third logical data path. If the high-level semantic result is determined to contain a safety constraint vector or an abnormal state marker, a saturation limit, obstacle penalty term, or reduced sampling distribution is applied to the output of the motion control model based on the safety constraint vector, or a preset safety response strategy is triggered when the current action is determined to be infeasible. The safety response strategy includes at least reducing the running speed, limiting the maximum output torque, returning to a predefined safe pose, or stopping the output.
[0053] In some implementations, the step of converting the raw observation data from the robot's multimodal sensing device into standardized observation data via the observation bridging module includes: By interfacing with the heterogeneous physical layer and link layer protocols of multimodal sensing devices, the original observation payload is restored into an internal data object; The internal data objects are timestamped based on the global reference clock provided by the CXL switching network, and frame alignment or interpolation is performed according to a fixed inference time window to construct a synchronous observation set. The synchronous observation set is converted into standardized observation data that can be directly called by the multimodal inference model, and the unification of dimensions, coordinates and cache lines are completed.
[0054] In some implementations, the steps of writing to the shared memory address space and reading from the shared memory address space are synchronized in the following manner: After writing data, the observation bridging module and the control bridging module generate corresponding observation descriptors or control status descriptors and store them in a descriptor circular queue pre-placed in the shared memory address space. They then notify the corresponding artificial intelligence processing module or motion control accelerator module by writing to the doorbell register. In response to the doorbell register notification, the artificial intelligence processor module or the motion control accelerator module reads the corresponding data from the shared memory address space via CXL point-to-point transmission.
[0055] In some implementations, data integrity assurance operations are also included: Before writing the data payload, the observation bridging module and the control bridging module calculate the integrity check value and write the check result into the integrity check flag in the corresponding descriptor. Before reading data, the artificial intelligence processor module and the motion control accelerator module parse the integrity verification flag. If the verification fails, the current data frame is discarded and the system waits for new data in the next cycle.
[0056] In some implementations, the initialization configuration step of the switching network manager further includes: Enumerate the observation bridge module and the control bridge module, and receive the capability descriptors reported by them. The capability descriptors include at least the supported mode types, maximum data throughput, minimum supported control cycle, and a list of supported external bus protocols. Based on the capability descriptor, the shared memory window size, descriptor circular queue depth, and service level weight of the first logical data path and the second logical data path are determined, and a CXL virtual channel independent of the first logical data path is allocated to the second logical data path.
[0057] In some implementations, a dynamic resource orchestration step is also included: The switching network manager periodically or under event triggering monitors the link health status and congestion status of the first logical data path, the second logical data path, and the third logical data path; When a link failure or continuous congestion exceeding a preset threshold is detected, the faulty link is retrained or switched to a redundant physical path, and the service level weight of the second logical data path or the priority of the virtual channel is increased to maintain the availability of the closed-loop control link.
[0058] In some implementations, the CPU runs the operating system and is responsible for task scheduling. The CPU does not participate in the real-time data interaction path of the standardized observation data, the high-level semantic results, and the action control instructions.
[0059] In some implementations, a communication device (CXL Type 3 device) coupled to the CXL switch accesses an external network, a remote server, or another robot cluster. The storage space of the communication device is incorporated into the logical pooled memory space management and is physically isolated from the first and second logical data paths. This communication device can participate in cross-node data sharing or collaborative processing under the scheduling of the switching network manager, but it does not participate in data interaction of the high-frequency closed-loop control link.
[0060] In this invention, the artificial intelligence processor module can be implemented by a GPU, NPU, dedicated AI chip or other processor suitable for deploying multimodal inference models; the motion control accelerator module can be implemented by a real-time control accelerator, FPGA, ASIC or other computing unit suitable for deploying motion expert models and high-frequency closed-loop control algorithms.
[0061] In this invention, the observation bridging module and the control bridging module can be physically implemented as two independent bridging nodes, or they can be integrated into the same bridging control module. Logical isolation (such as independent address windows, independent virtual channels, or independent descriptor queues) ensures that the observation data stream and the control data stream do not interfere with each other.
[0062] While this invention uses a hierarchical vision-language-action (VLA) architecture as an example, it is not limited to this. The perception and closed-loop control method based on the CXL switching network provided in this invention can be applied to any robot system, embodied intelligence system, industrial automation system, or heterogeneous sensing and control collaborative system that requires simultaneous processing of high-bandwidth observation data and low-latency control data.
[0063] Due to the application of the above-mentioned technical solution, the present invention has the following advantages compared with the prior art: This invention constructs a unified high-speed interconnect architecture for heterogeneous robot devices that does not rely on CPU relay. By utilizing logical pooled memory space combined with point-to-point direct connections to bypass CPU relay, it significantly reduces transmission latency while eliminating data copying overhead and host bandwidth bottlenecks. Through the separate processing of observation and control flows, it addresses the dual requirements of high-bandwidth sensing and low-latency control. Furthermore, a control bridging module enables high-frequency closed-loop feedback, improving control stability. This unified high-speed interconnect architecture for heterogeneous robot devices eliminates redundant data transfer between multiple nodes, achieves efficient reuse of shared data, and supports dynamic resource orchestration and smooth node expansion through a switching network manager, significantly improving system collaboration efficiency, resource utilization, and scalability. Attached Figure Description
[0064] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0065] Figure 1 A schematic diagram of the unified interconnection architecture for heterogeneous robot devices based on the CXL switching network provided in Example 1; Figure 2 This is a schematic diagram of the workflow of the robot perception and closed-loop control method based on the CXL switching network provided in Example 2. Detailed Implementation
[0066] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0067] It should be noted that the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, apparatus, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product or device.
[0068] Example 1: This example provides a unified interconnection architecture for heterogeneous robot devices based on the CXL switching network.
[0069] like Figure 1 As shown, the unified interconnection architecture of heterogeneous robot devices in this embodiment 1 includes a CXL switch, a Fabric Manager, and an Observation Bridge, a Control Bridge, an AI Processor, and an Accelerator coupled to the CXL switch.
[0070] Specifically, the switching network manager is configured to construct a first logical data path, a second logical data path, and a third logical data path that respectively support CXL Peer-to-Peer (CXL P2P) access; construct and manage a logically pooled memory space shared by the observation bridging module, the control bridging module, the artificial intelligence processor module, and the motion control accelerator module; the logically pooled memory space is mapped to a globally unified shared memory address space; the observation bridging module is configured to convert the raw observation data from the robot's multimodal sensing devices into standardized observation data and write it into the shared memory address space; the control bridging module is configured to convert the body state data of the robot's execution unit into a standardized state representation and write it into the shared memory address space, and to connect the two logical data paths via the second logical data path. The logical data path reads motion control instructions from the shared memory address space and sends them to the robot execution unit; the artificial intelligence processor module is configured to obtain standardized observation data from the shared memory address space through the first logical data path, run a multimodal inference model to generate high-level semantic results, and write the high-level semantic results into the shared memory address space; the motion control accelerator module is configured to obtain standardized state representations from the shared memory address space through the second logical data path and obtain high-level semantic results from the shared memory address space through the third logical data path, run a motion control model to generate motion control instructions, and write the motion control instructions into the shared memory address space.
[0071] Specifically, the observation bridging module includes an interface adaptation module, a time synchronization module, and a modality normalization module. The interface adaptation module is configured to interface with the heterogeneous physical layer and link layer protocols of the multimodal sensing device, restoring the original observation payload to internal data objects. The time synchronization module is configured to apply timestamps to the internal data objects based on the global reference clock provided by the CXL switching network, and perform frame alignment or interpolation according to a fixed inference time window to construct a synchronized observation set. The modality normalization module is configured to convert the synchronized observation set into standardized observation data that can be directly called by the multimodal inference model. The standardized observation data includes at least one of the following: image tensor, depth tensor, point cloud tensor, voxel grid, Mel spectrum, audio terms, and state vector. The observation bridging module is also configured to perform pre-coding processing on the standardized observation data (including at least data precision quantization, normalization, and memory layout alignment), and generate observation descriptors (including at least data type, dimension, payload address offset, priority identifier, and validity flag). The observation descriptors are stored in a pre-allocated descriptor circular queue in the shared memory address space and are notified to the corresponding artificial intelligence processor module through a doorbell register.
[0072] In this embodiment, the observation bridging module includes a first bridging endpoint and a first CXL memory extension device (Type 3 device) without computational logic. The first CXL memory extension device provides a large-capacity shared memory space to carry standardized observation data and is incorporated into the logically pooled memory space. The first bridging endpoint includes an interface adaptation module, a time synchronization module, a modality normalization module, a preprocessing module, an observation descriptor generation module, and a CXL transaction encapsulation module, wherein: the interface adaptation module is compatible with multiple external bus protocols (including but not limited to MIPI CSI-2, GigE...). The system uses various communication methods (Vision, USB, I2C, SPI, UART, etc.) to decapsulate raw signals from cameras, LiDAR, IMU, and force sensors, restoring them to internal data objects and attaching basic metadata. The time synchronization module applies timestamps based on a global reference clock, constructing a synchronized observation set that is strictly aligned in the time dimension. The modal normalization module converts the synchronized internal data objects into standardized data structures (such as normalized tensors, voxel grids, or state vectors) that can be directly consumed by the multimodal inference model. The preprocessing module quantizes, normalizes, and aligns the standardized data in memory. The observation descriptor generation module generates a metadata header containing modality type, sensor identifier, timestamp, dimension, data type, and memory address offset. The CXL transaction encapsulation module performs a CXL.mem write operation, writing the preprocessed observation data into the logical pooled memory space and submitting the observation descriptor to the descriptor circular queue, notifying the AI processor module that new data is ready through the doorbell register.
[0073] Specifically, the control bridging module includes a status data standardization encoding module, an action decoding module, and an execution unit protocol mapping module. The state data standardization encoding module is configured to unify the execution units, transform coordinates, and restore the scale of the robot's execution unit's body state data to generate a standardized state representation that can be directly called by the motion control model. The motion decoding module is configured to convert the execution units of motion control commands (restoring the output of the motion control model from normalized values to the physical units of the execution unit), limit clipping (constraining joint angles, speeds, or torque safety limits), compensate for dead zones (compensating for the nonlinear dead zones of the actuators), and smooth interpolation (low-pass filtering or linear / second-order interpolation of commands in adjacent control cycles). The execution unit protocol mapping module is configured to map the processed motion control commands to data objects of the target bus protocol and generate drive control messages to be sent to the robot execution unit. The execution unit protocol mapping module supports at least one of the following industrial control protocols: EtherCAT, CAN, CANopen, Modbus RTU, RS485 / 422 serial communication protocols, and proprietary multi-axis drive protocols. The drive control messages contain at least one of position control commands, speed control commands, torque control commands, or current control commands.
[0074] In this embodiment, the control bridging module includes a second bridging endpoint and a second CXL memory extension device (i.e., a CXL Type 3 device) without computational logic. The second CXL memory extension device mainly provides a shared memory space to carry the body state data and is incorporated into the logical pooled memory space. The second bridging endpoint includes a state acquisition and standardization module, a control state descriptor generation module, a CXL transaction encapsulation module, an action decoding and protocol mapping module, and a closed-loop feedback module. Specifically, the state acquisition and standardization module is configured to acquire the state data of the robot execution unit in a periodic or event-triggered manner, and perform unit unification, coordinate transformation, and scale recovery to generate a standardized state representation. The control state descriptor generation module is configured to generate a control state descriptor for each standardized state representation. The control state descriptor includes at least the robot identifier, execution chain identifier, control cycle number, timestamp, state vector dimension, data type, valid bit, fault bit, security level, priority, execution deadline, and write-back address. The L-transaction encapsulation module is configured to use a combination of fixed-period state updates and abnormal event preemption to write the body state data to the second CXL memory expansion device and submit the control state descriptor to the descriptor circular queue, notifying the motion control accelerator through the doorbell register. The motion decoding and protocol mapping module is configured to read motion control instructions from the motion control accelerator module, perform unit conversion, amplitude limiting, and protocol mapping to generate drive control messages that can be executed by the robot execution unit. The closed-loop feedback module is configured to notify the state acquisition and standardization module to collect the latest state after the execution unit completes the action output, and update the body state data in the second CXL memory expansion device for use by the motion control accelerator module in the next control cycle.
[0075] Specifically, the AI processor module is configured to run a multimodal inference model (such as a vision-language model, VLM) to generate high-level semantic results. These high-level semantic results are represented in the form of structured communication primitives, including task semantic identifiers, target pose tensors, trajectory parameterization descriptions, implicit feature vectors, and their metadata descriptors. The AI processor module writes the generated high-level semantic results into a shared memory address space for use by the motion control accelerator module.
[0076] The motion control accelerator module is configured to run motion control models (such as motion expert models). The AI processor module generates high-level semantic results at a relatively low update frequency (e.g., 20Hz), while the motion control accelerator module needs to run motion control models at a significantly higher frequency (e.g., 1kHz) to meet the robot's requirements for high-frequency closed-loop control.
[0077] In this embodiment, to adapt to the aforementioned frequency differences, the motion control accelerator module is further configured to cache the high-level semantic results obtained from the shared memory address space via the third logical data path, using them as the currently valid high-level semantic results. It also runs periodically according to a preset high-frequency control cycle. In each run, based on the currently valid high-level semantic results and the current ontology state data obtained via the second logical data path, it runs the motion control model, generates motion control instructions, and writes them into the shared memory address space, until a new high-level semantic result is obtained via the third logical data path. That is, the motion control accelerator module is configured to reuse the same currently valid high-level semantic results within multiple consecutive control cycles between two high-level semantic result updates. The control bridge module is configured to run periodically according to the high-frequency control cycle, executing the following sequentially in each run: converting the ontology state data of the robot execution unit into a standardized state representation and writing it into the shared memory address space; reading motion control instructions from the shared memory address space via the second logical data path; converting the motion control instructions into drive control messages and sending them to the robot execution unit; and collecting state feedback generated by the robot execution unit to update the ontology state data for the next high-frequency control cycle. Specifically, the motion control accelerator module is configured to write the generated motion control instructions into a pre-allocated instruction buffer in the shared memory address space during each high-frequency control cycle, and notify the control bridge module via a doorbell register. The control bridge module is configured to, in response to the doorbell register notification, read the motion control instructions from the instruction buffer through a second logical data path and send them to the robot execution unit. After the robot execution unit completes the action output, it collects the actual state feedback generated by the robot execution unit, converts the actual state feedback into the ontology state data required for the next high-frequency control cycle, and overwrites the original ontology state data in the shared memory address space for the motion control accelerator module to read in the next high-frequency control cycle. Specifically, the motion control accelerator module is configured to parse the high-level semantic results based on the metadata descriptor of the structured communication primitive, and use the parsed target information and implicit feature vectors as input to the motion control model. When the artificial intelligence processor module generates and writes new high-level semantic results, the motion control accelerator module detects this through a doorbell or descriptor update event, refreshes the cached high-level semantic results, and subsequent control cycles will continue to run based on the new semantics. Both the motion control accelerator module and the control bridge module are configured to operate synchronously based on the same preset high-frequency control cycle to ensure that they exchange data under the same time reference, forming a complete closed-loop control. In this embodiment, the artificial intelligence processor module and the motion control accelerator module are both Type 2 devices, each integrating a dedicated computing core and local private memory for running inference models or motion control models. The local private memory is configured to be included in the logical pooled memory space.
[0078] In this implementation, the observation bridging module and the control bridging module are also configured to perform integrity verification on the data payload written to the shared memory address space, and set corresponding integrity verification flags in the observation descriptor and the control state descriptor; the artificial intelligence processor module and the motion control accelerator module verify the data integrity based on the verification flags before reading the data. The motion control accelerator module is also configured to dynamically adjust the output range of the motion control model or trigger a preset safety response strategy when it detects that the high-level semantic results contain safety constraint vectors or abnormal states. The safety response strategy includes deceleration, force limiting, regressing to a safe pose, or stopping output.
[0079] Specifically, the switching network manager is configured to enumerate and initialize the CXL switch, logical pooled memory space, and each access module during system power-on or task startup. Specifically, the observation bridge module and control bridge module are also configured to report their own capability descriptors to the switching network manager in response to enumeration requests; the capability descriptors include at least the supported modal types (such as RGB, Depth, LiDAR, IMU, JointState), maximum data throughput, minimum supported control cycle, and a list of supported external bus protocols (such as EtherCAT, CAN, CANopen, RS485). Based on the capability descriptors, the switching network manager determines the shared memory window size, descriptor ring queue depth, and service level weights for the first and second logical data paths, and allocates a CXL virtual channel independent of the first logical data path for the second logical data path to complete the initialization configuration of the logical data path. Specifically, the switching network manager is configured to configure a shared memory window with a size larger than that of the second logical data path and a descriptor queue with a depth greater than that of the second logical data path for the first logical data path, in order to adapt to the high bandwidth and batch transmission requirements of multimodal sensing data; and to configure a fast transmission buffer dedicated to periodic body state data and motion control commands for the second logical data path, in order to adapt to the high frequency closed-loop control requirements.
[0080] The switching network manager is also configured to perform Quality of Service (QoS) differentiation on the first and second logical data paths to support high-frequency closed-loop control. The switching network manager maps the second logical data path to a higher-priority virtual channel (VC1) or traffic class (TC1) in the CXL switch, while mapping the first logical data path to a lower-priority virtual channel (VC0 or TC0), ensuring that control transactions are prioritized during link congestion. Simultaneously, the switching network manager pre-allocates a dedicated fast transmission buffer in the logically pooled memory space for ontology state data and action control instructions. This buffer has a small cache line depth and a short allocation / release path, dedicated to the high-frequency periodic data exchange of the second logical data path, avoiding access contention caused by sharing the same memory area with high-bandwidth observation data. By employing priority configuration and a fast transmission buffer, the switching network manager ensures that the end-to-end transmission latency for body status data upload, motion control command issuance, and abnormal event preemption descriptor does not exceed a preset threshold (e.g., ≤100μs), and that periodic latency jitter is also controlled within a preset range (e.g., ≤10 μs). This satisfies the real-time and deterministic requirements of closed-loop control tasks operating at a preset high-frequency control cycle (e.g., 1 kHz). Abnormal events include at least one of the following: emergency stop, collision, overcurrent, instability, joint over-limit, or battery failure. The control bridge module is configured to generate an event descriptor with the highest priority in response to an abnormal event and send it to the motion control accelerator module via a second logical data path in a preemptive manner to trigger the motion control accelerator module to execute a safety braking or degraded control strategy.
[0081] The switching network manager is also configured to pre-allocate isolated address windows for each module within the logically pooled memory space, and configure access control lists (ACLs) for each address window to restrict each module to accessing only its corresponding data area. This memory isolation mechanism effectively prevents out-of-bounds writes that could corrupt critical data in other modules due to anomalies or firmware defects in any module. It also narrows the reachable range of DMA to improve system security and reduces cache pollution caused by shared access from multiple modules, helping to ensure the time determinism of the high-frequency closed-loop control link.
[0082] The switching network manager is also configured to monitor the link health and congestion status of each logical data path, and perform dynamic orchestration to maintain the availability of control links when anomalies are detected. The switching network manager periodically or on event triggers queries the link registers of each CXL port to obtain link health status (such as Link Up / Down, Lane WidthDegrade, CRC Error Count) and congestion proxy metrics (such as incomplete transaction count, Credit exhaustion frequency, descriptor queue occupancy, etc.). When a link failure or persistent congestion exceeding a preset threshold is detected, the switching network manager performs at least one of the following operations: attempts to retrain or switch to a redundant physical path for the failed link, adjusts the service class weight of the affected logical data path to give priority bandwidth to the second logical data path, or reallocates shared memory address windows; wherein, the protection priority for the second logical data path is higher than that for other paths to ensure the availability of high-frequency closed-loop control links.
[0083] In this embodiment, the unified interconnect architecture for heterogeneous robot devices may further include a central processing unit (CPU) and a communication device (CXL Type 3 memory extension device). The CPU is configured to run a general-purpose operating system (such as Linux or ROS2 distribution) and upper-layer application software, primarily responsible for system-level task scheduling, resource management, lifecycle maintenance, and non-real-time business logic processing. The CPU does not participate in the real-time data interaction path of multimodal observation data, high-level semantic results, and motion control instructions, nor does it perform read / write operations on core business data in the logical pooled memory space, to avoid the impact of uncertainties and jitters of the general-purpose operating system on the high-frequency closed-loop control link. The communication device is used to connect the robot to an external network. In some preferred embodiments, the communication device is a CXL Type 3 memory extension device, which is coupled to a CXL switch through a CXL interface and included in the management scope of the logical pooled memory space via the switch network manager. The communication equipment is configured to: establish a connection with an external remote server or cloud for uploading robot status logs or downloading updated model weights; establish a cluster communication link with other robot nodes for data synchronization in multi-robot collaborative tasks; and receive manual intervention instructions or task assignment instructions from a remote control console.
[0084] Example 2: This example provides a robot perception and closed-loop control method based on a CXL switching network. This method is implemented on the unified interconnection architecture of heterogeneous robot devices in Example 1.
[0085] like Figure 2 As shown, the robot perception and closed-loop control method in this embodiment includes: Initialization: During system power-on or task startup, the switching network manager enumerates and initializes the CXL switch, logical pooled memory space, observation bridge module, control bridge module, artificial intelligence processor module, and motion control accelerator module. The logical pooled memory space is a globally unified shared memory address space built and managed by the switching network manager for the observation bridge module, control bridge module, artificial intelligence processor module, and motion control accelerator module to access. The switching network manager also builds the first, second, and third logical data paths to support CXL Peer-to-Peer access, respectively. The observation bridging module converts the raw observation data into standardized observation data and writes it into shared memory: The observation bridging module converts the raw observation data from the robot's multimodal sensing devices into standardized observation data and writes it into the shared memory address space. The artificial intelligence processor module acquires observation data and generates high-level semantic results: The artificial intelligence processor module acquires standardized observation data from the shared memory address space via the first logical data path, runs a multimodal inference model to generate high-level semantic results, and writes them into the shared memory address space. The motion control accelerator module generates motion control commands and writes them to shared memory; the control bridging module collects body state data and writes it to shared memory; the control bridging module collects the body state data of the robot execution unit, converts it into a standardized state representation, and writes it to the shared memory address space. The control bridge module reads motion control instructions and sends them to the execution unit: the motion control accelerator module obtains standardized state representations from the shared memory address space via the second logical data path and high-level semantic results from the shared memory address space via the third logical data path, runs the motion control model to generate motion control instructions, and writes the motion control instructions into the shared memory address space. The control bridging module collects status feedback and forms closed-loop control: the control bridging module reads motion control instructions from the shared memory address space via the second logical data path, converts them into drive control messages and sends them to the robot execution unit; and the control bridging module collects the status feedback generated after the robot execution unit completes the motion output, overwrites and updates the body status data in the shared memory address space, so that the motion control accelerator module can read it in the next control cycle, thus forming closed-loop control.
[0086] The specific initialization operation is as follows: During system power-on or task startup, the CXL switch, logical pooled memory space and observation bridge module, control bridge module, artificial intelligence processor module and motion control accelerator module are enumerated and initialized through the switching network manager.
[0087] The switching network manager discovers and enumerates ports of each node in the CXL switching network, negotiates link capabilities, allocates address windows and buffers, and configures target routes, priority policies, access permissions, and doorbell mechanisms for observation and control paths. It also initializes the buffer descriptor queues of each module. Based on this, the switching network manager constructs a logically pooled memory space, performs unified addressing and mapping of memory expansion devices connected to the CXL switching network, and establishes a globally unified shared memory address space.
[0088] Among them, the test bridging module and the control bridging module are connected to the CXL switching network as bridging nodes, responsible for bridging, buffering, descriptor management and transaction initiation; the artificial intelligence processor module and the motion control accelerator module are CXL Type 2 devices, and their local private memory is included in the logical pooled memory space for each module to access through CXL P2P.
[0089] The switching network manager builds three logical data paths supporting CXL P2P access based on the CXL switch: First logical data path: connects the observation bridging module and the artificial intelligence processor module, carrying high-bandwidth, multimodal, low-frequency sensing data; Second logical data path: connects the control bridge module and the motion control accelerator module, carrying low-latency, high-frequency closed-loop control data; The third logical data path is used by the AI processor module to transmit high-level semantic results to the motion control accelerator module.
[0090] Specifically, the observation bridging module performs the following operations to convert raw observation data into standardized observation data and write it into shared memory: The observation bridging module receives raw data from cameras, depth cameras, LiDAR, microphone arrays, IMUs, force and tactile sensors, and accesses them via heterogeneous protocols such as MIPI CSI-2, GigE Vision, USB, I2C, SPI, and UART.
[0091] The observation bridging module first performs physical layer and link layer adaptation, restoring the original payload to an internal data object and adding metadata such as sensor identifier, mode identifier, sampling time, and sequence number. Subsequently, the time synchronization module applies a unified timestamp based on the global reference clock provided by the CXL switching network and performs frame alignment, interpolation, or resampling according to a fixed inference time window to construct a synchronized observation set.
[0092] The modal normalization module converts the synchronous observation set into standardized observation data that can be directly called by the multimodal inference model, including image tensors, depth tensors, point cloud tensors, voxel grids, Mel spectra, state vectors, etc., and completes the unification of dimensions, coordinates, and cache line alignment.
[0093] The observation bridging module generates an observation descriptor (including modality type, tensor dimension, payload address offset, priority identifier, etc.) for each standardized observation data. It adopts an organization method that separates the payload and the descriptor, writes the observation payload into the shared memory address space, stores the observation descriptor in the descriptor circular queue, and notifies the artificial intelligence processor module by writing to the doorbell register.
[0094] Specifically, the AI processor module acquires observation data and generates high-level semantic results through the following operations: In response to a doorbell notification, the AI processor module reads standardized observation data from the shared memory address space via the first logical data path and runs a visual-language model (VLM) or a multimodal reasoning model to perform fusion analysis on visual, language, depth map, point cloud and state information, and outputs high-level semantic results.
[0095] High-level semantic results are represented in the form of structured communication primitives, including at least task semantic identifiers, target pose tensors, trajectory parameterization descriptions, and implicit feature vectors. If the inference results contain safety constraint vectors (such as maximum angles, angular velocities, and torque limits for each joint) or abnormal state markers, the AI processor module encapsulates them together in the structured communication primitives and writes them to the shared memory address space via CXL P2P for use by the motion control accelerator module.
[0096] Specifically, the control bridging module collects the body status data and writes it to the shared memory in the following ways: The control bridging module collects body status data such as joint position, speed, torque, body posture, linear velocity, angular velocity, foot contact state, end force, and battery status through industrial protocols such as EtherCAT, CAN, CANopen, and serial communication interfaces.
[0097] Because the data formats, unit systems, and coordinate systems of different execution units are not uniform, the control bridging module performs unit unification, coordinate transformation, and scale restoration (such as converting encoder counts into joint positions), and organizes them into standardized state representations that can be directly called by the motion expert model, including single-moment state vectors, fixed-time-window state tensors, and safety constraint vectors.
[0098] The control bridging module generates a control state descriptor for each state object, which includes at least the robot identifier, control cycle number, timestamp, state vector dimension, valid bit, fault bit, safety level, priority, and execution deadline. The state payload and control descriptor are written into the shared memory address space and the descriptor circular queue, respectively, and the motion control accelerator module is notified through a doorbell.
[0099] The specific operations of the motion control accelerator module in generating motion control instructions and writing them to shared memory are as follows: The motion control accelerator module obtains the standardized state representation via the second logical data path and the high-level semantic results written by the artificial intelligence processor module via the third logical data path.
[0100] The motion control accelerator module caches the currently valid high-level semantic results. Since the update frequency of the high-level semantic results (e.g., 20 Hz) is much lower than the running frequency of the motion control model (e.g., 1 kHz), the motion control accelerator module reuses the same currently valid high-level semantic results between two high-level semantic result updates, runs the motion control model multiple times consecutively according to the preset high-frequency control cycle, generates motion control instructions frame by frame, and writes them to the shared memory address space.
[0101] If the high-level semantic result contains a safety constraint vector or an abnormal state marker, the motion control accelerator module applies a saturation limit, obstacle penalty term, or reduces the sampling distribution to the model output based on the vector, or triggers a preset safety response strategy when the current action is determined to be infeasible, including reducing the running speed, limiting the maximum output torque, reverting to a predefined safe pose, or stopping the output.
[0102] Specifically, the control bridging module reads the action control command and sends it to the execution unit as follows: In response to the doorbell notification written by the motion control accelerator module, the control bridge module reads the motion control instruction from the shared memory address space via the second logical data path and enters the reverse semantic decoding stage.
[0103] The control bridging module determines whether the motion result should be mapped to position control, speed control, torque control, or current control based on the actuator type. It then performs unit conversion, proportional recovery, amplitude limiting, dead zone compensation, filtering, and smooth interpolation on the model output. Finally, it maps the processed control quantity to the target bus protocol data object (such as CANopen PDO or private serial port driver control message) and sends it to the robot execution unit.
[0104] Specifically, the operation of the control bridging module to collect status feedback and form closed-loop control is as follows: After the actuator completes its action output, the control bridging module collects new position, speed, torque, fault and status feedback, re-encodes it into a standard status representation, and overwrites and updates the body status data in the shared memory address space so that the motion control accelerator module can read it in the next control cycle to form a complete closed-loop control.
[0105] In this embodiment, between the above-mentioned write and read steps, the observation bridging module and the control bridging module calculate the integrity check value (such as CRC) before writing the data payload, and write the check result into the integrity check flag in the descriptor; the artificial intelligence processor module and the motion control accelerator module parse the integrity check flag before reading the data, and if the check fails, discard the current data frame and wait for the new data in the next cycle.
[0106] The switching network manager periodically monitors the link health and congestion status of each logical data path, either directly or in response to events. When a link failure or persistent congestion exceeding a threshold is detected, the manager attempts to retrain the faulty link or switch it to a redundant physical path. It also increases the service level weight or virtual channel priority of the second logical data path to maintain the availability of the closed-loop control link. Simultaneously, the switching network manager restricts each module to accessing only its corresponding data area by pre-allocating isolated address windows and configuring access control lists (ACLs), preventing out-of-bounds writing that could corrupt critical data. The switching network manager configures a large shared memory window and a deep descriptor queue for the first logical data path to accommodate large data transfers; and configures a low-latency, small-granularity periodic buffer and a high-priority descriptor queue for the second logical data path, combined with an abnormal event preemption strategy, to ensure the real-time performance and determinism of the control link.
[0107] The above embodiments have the following advantages: Enhance unified access capabilities: Integrate the observation bridging module, control bridging module, artificial intelligence processor module, and motion control accelerator module into the CXL switching network to avoid the problem of scattered device access and significantly improve system integration.
[0108] Reduce data transmission latency: CXL P2P enables direct data exchange between devices, completely bypassing the CPU relay and data copying process, effectively shortening the data exchange path.
[0109] Alleviating host bandwidth bottlenecks: Reducing the dependence of multi-sensor observation data and control and status data on the CPU relay path, significantly reducing host bandwidth usage, and improving the overall system throughput.
[0110] Improve the efficiency of heterogeneous node collaboration: Achieve zero-copy high-speed collaboration between the artificial intelligence processor module and the motion control accelerator module, and significantly improve the collaborative processing capabilities of the perception, decision-making and control links.
[0111] Balancing high-bandwidth perception with low-latency control: By splitting the observation stream and control stream and implementing differentiated QoS scheduling, the system meets the dual requirements of large-scale intelligent perception and strong real-time closed-loop control in robot systems.
[0112] Enhanced control closed-loop performance: By using the control bridging module and CXL P2P to achieve uplink control information and downlink feedback of execution results, the real-time performance and stability of the control closed loop are improved.
[0113] Improve the efficiency of shared data reuse: By forming a logical pooled memory space that can be shared and accessed by all nodes through memory pooling and unified addressing, the redundant storage and transportation are reduced, thereby improving the efficiency of data sharing.
[0114] Improve resource utilization and scalability: Dynamic orchestration and unified management of resources are achieved through the switching network manager, enhancing system flexibility and overall operating efficiency; when adding nodes, flexible expansion can be completed within the same CXL switching network, improving the system's adaptability and scalability.
[0115] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A unified interconnection architecture for heterogeneous robot devices based on a CXL switching network, characterized in that, It includes a CXL switch, a switching network manager, and observation bridge modules, control bridge modules, artificial intelligence processor modules, and motion control accelerator modules coupled to the CXL switch, wherein: The switching network manager is configured to build a first logical data path, a second logical data path, and a third logical data path that respectively support CXL Peer-to-Peer access, and to build and manage a logically pooled memory space that is shared by the observation bridge module, the control bridge module, the artificial intelligence processor module, and the motion control accelerator module. The logically pooled memory space is mapped to a globally unified shared memory address space. The observation bridging module is configured to convert the raw observation data from the robot's multimodal sensing device into standardized observation data and write it into the shared memory address space. The control bridge module is configured to convert the body state data of the robot execution unit into a standardized state representation and write it into the shared memory address space, and to read motion control instructions from the shared memory address space through the second logical data path and send them to the robot execution unit. The artificial intelligence processor module is configured to obtain the standardized observation data from the shared memory address space through the first logical data path, run a multimodal inference model to generate high-level semantic results, and write the high-level semantic results into the shared memory address space; The motion control accelerator module is configured to obtain the standardized state representation from the shared memory address space through the second logical data path and obtain the high-level semantic result from the shared memory address space through the third logical data path, run the motion control model to generate the motion control instruction, and write the motion control instruction into the shared memory address space; The standardized observation data, the high-level semantic results, and the standardized state representation are exchanged between modules in a zero-copy manner through the shared memory address space, without the need for the central processing unit to participate in data transfer.
2. The unified interconnection architecture for heterogeneous robot devices according to claim 1, characterized in that, The motion control accelerator module is also configured to cache the high-level semantic results obtained from the shared memory address space through the third logical data path as the currently valid high-level semantic results; It runs periodically according to a preset high-frequency control cycle. In each run, based on the current valid high-level semantic results and the current ontology state data obtained through the second logical data path, the motion control model is run, the action control instructions are generated and written into the shared memory address space, until a new high-level semantic result is obtained through the third logical data path.
3. The unified interconnection architecture for heterogeneous robot devices according to claim 2, characterized in that, The control bridge module is also configured to operate periodically according to the high-frequency control cycle, with each operation executing the following sequentially: The robot's execution unit's body state data is converted into a standardized state representation and written into the shared memory address space; Action control instructions are read from the shared memory address space through the second logical data path; The motion control commands are converted into drive control messages and sent to the robot execution unit; The state feedback generated by the robot execution unit is collected to update the body state data for the next high-frequency control cycle.
4. The unified interconnection architecture for heterogeneous robot devices according to claim 3, characterized in that, The motion control accelerator module is configured to write the generated motion control instruction into a pre-allocated instruction buffer in the shared memory address space during each high-frequency control cycle, and notify the control bridge module through the doorbell register. The control bridge module is configured to respond to the doorbell register notification by reading the motion control command from the instruction buffer through the second logical data path and sending it to the robot execution unit. After the robot execution unit completes the motion output, the module collects the actual state feedback generated by the robot execution unit, converts the actual state feedback into the body state data required for the next high-frequency control cycle, and overwrites the original body state data in the shared memory address space for the motion control accelerator module to read in the next high-frequency control cycle.
5. The unified interconnection architecture for heterogeneous robot devices according to claim 1, characterized in that, The observation bridging module includes an interface adaptation module, a time synchronization module, and a modal standardization module, wherein: The interface adaptation module is configured to interface with the heterogeneous physical layer and link layer protocols of the multimodal sensing device, and restore the original observation payload to an internal data object; The time synchronization module is configured to apply timestamps to the internal data objects based on a unified reference clock and construct a synchronized observation set according to a fixed time window; The modal normalization module is configured to convert the synchronous observation set into standardized observation data that can be directly called by the multimodal inference model. The standardized observation data includes at least one of the following: image tensor, depth tensor, point cloud tensor, voxel grid, Mel spectrum, audio lexical, and state vector.
6. The unified interconnection architecture for heterogeneous robot devices according to claim 1, characterized in that, The observation bridging module is further configured to generate an observation descriptor for each of the standardized observation data, and the control bridging module is further configured to generate a control state descriptor for each of the standardized state representations; The observation descriptor and the control state descriptor include at least data type, dimension, load address offset, priority identifier, and validity flag; The observation descriptor and the control state descriptor are stored in a pre-allocated descriptor circular queue in the shared memory address space, and are notified to the corresponding artificial intelligence processor module and motion control accelerator module through the doorbell register.
7. The unified interconnection architecture for heterogeneous robot devices according to claim 5, characterized in that, The observation bridging module is also configured to perform precoding processing on the standardized observation data. The precoding processing includes at least data precision quantization, normalization, and memory layout alignment to adapt to the cache line access characteristics of the shared memory address space.
8. The unified interconnection architecture for heterogeneous robot devices according to claim 1, characterized in that, The control bridge module includes a state data standardization encoding module, an action decoding module, and an execution unit protocol mapping module, wherein: The state data standardization encoding module is configured to unify the execution units, transform coordinates, and restore the scale of the robot's execution unit's body state data, so as to generate a standardized state representation that can be directly called by the motion control model. The motion decoding module is configured to perform unit conversion, amplitude clipping, dead zone compensation, and smooth interpolation on the motion control command. The execution unit protocol mapping module is configured to map the processed motion control instructions into data objects of the target bus protocol and generate drive control messages to be sent to the robot execution unit.
9. The unified interconnection architecture for heterogeneous robot devices according to claim 8, characterized in that, The execution unit protocol mapping module is configured to support at least one of the following industrial control protocols: EtherCAT, CAN, CANopen, ModbusRTU, RS485 / 422 serial communication protocol, and proprietary multi-axis drive protocol; the drive control message includes at least one of position control command, speed control command, torque control command, or current control command.
10. The unified interconnection architecture for heterogeneous robot devices according to claim 1, characterized in that, The high-level semantic results are represented in the form of structured communication primitives, which include task semantic identifiers, target pose tensors, trajectory parameterization descriptions, implicit feature vectors and their metadata descriptors; The motion control accelerator module is configured to parse the structured communication primitives based on the metadata descriptor, and input the parsed target information and implicit feature vectors into the motion control model.
11. The unified interconnection architecture for heterogeneous robot devices according to claim 1, characterized in that, The switching network manager is also configured to prioritize the transmission of the second logical data path over the first logical data path, and to pre-allocate a fast transmission buffer dedicated to the body state data and action control instructions in the logical pooled memory space, so that the transmission of body state data upload, action control instruction issuance and abnormal event preemption descriptor meets the preset maximum end-to-end latency and jitter requirements, so as to support closed-loop control running with a preset control cycle.
12. The unified interconnection architecture for heterogeneous robot devices according to claim 11, characterized in that, The abnormal events include at least one of the following: emergency stop, collision, overcurrent, instability, joint over-limit, or battery failure; The control bridge module is configured to generate an event descriptor with the highest priority in response to the abnormal event, and send it to the motion control accelerator module in a preemptive manner through the second logical data path to trigger the motion control accelerator module to execute a safety braking or degraded control strategy.
13. The unified interconnection architecture for heterogeneous robot devices according to claim 1, characterized in that, The switching network manager is also configured to configure a shared memory window with a size larger than that of the second logical data path and a descriptor queue with a depth greater than that of the second logical data path for the first logical data path, in order to adapt to the high bandwidth and batch transmission requirements of multimodal sensing data; and to configure a fast transmission buffer dedicated to periodic body state data and action control commands for the second logical data path, in order to adapt to the high frequency closed-loop control requirements.
14. The unified interconnection architecture for heterogeneous robot devices according to claim 1, characterized in that, The switching network manager is also configured to pre-allocate isolated address windows for the observation bridge module, the control bridge module, the artificial intelligence processor module, and the motion control accelerator module in the logical pooled memory space, and configure an access control list for each address window to restrict each module to access only its corresponding data area.
15. The unified interconnection architecture for heterogeneous robot devices according to claim 1, characterized in that, The observation bridging module and the control bridging module are further configured to perform integrity verification on the data payload written to the shared memory address space, and set corresponding integrity verification flags in the observation descriptor and the control status descriptor; the artificial intelligence processor module and the motion control accelerator module verify the data integrity based on the verification flags before reading the data.
16. The unified interconnection architecture for heterogeneous robot devices according to claim 1, characterized in that, The motion control accelerator module is also configured to dynamically adjust the output range of the motion control model or trigger a preset safety response strategy when the high-level semantic result contains a safety constraint vector or an abnormal state. The safety response strategy includes deceleration, force limitation, regression to a safe pose, or stopping output.