Special chip architecture for sensing-computing-control integrated robot controller
By adopting a multi-core heterogeneous system architecture and a custom instruction integrated accelerator in the robot controller, the problems of insufficient real-time performance, computing efficiency and flexibility in the existing technology are solved, and an efficient and low-latency robot controller design is achieved.
Patent Information
- Application Number
- CN202510940864.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-17
AI Technical Summary
Existing robot controller chips have deficiencies in real-time performance, computing efficiency, energy efficiency, and flexibility, making it difficult to meet the requirements of high-precision, low-latency real-time control. In addition, the system hardware is complex and costly, and inter-module communication delays become a performance bottleneck.
It adopts a multi-core heterogeneous system architecture, integrates the control core, computing core and perception core into a single SoC chip, integrates a dedicated coprocessor interface and accelerator, realizes the integration of sensing, computing and control, and directly interacts with the accelerator through custom instructions, reducing control latency and improving computing performance and system integration.
It significantly improves the real-time response capability and computing efficiency of the robot control system, reduces hardware complexity and cost, meets the requirements of microsecond-level real-time closed-loop control, and realizes high flexibility and low power consumption system design.
Smart Images

Figure CN120804022A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of integrated circuit design and robot technology, in particular to a high-performance, low-power, high-flexibility System on Chip (SoC) chip for robot controller, and especially to a sensor-computation-control integrated robot controller special chip architecture. BACKGROUND
[0002] Currently, the technical solutions of robot controller chips mainly have the following problems and defects: 1. Performance and real-time bottleneck: Traditional robot controllers mostly use general-purpose microprocessors (MCU) or general-purpose processors, whose architecture design is not for complex calculations in the field of robots. When dealing with high-frequency and high-throughput computing tasks such as kinematics and dynamics forward and inverse solutions, multi-axis collaborative interpolation, artificial intelligence model reasoning, etc., general-purpose processors are limited by their serial execution mode and limited parallel capabilities, making it difficult to meet the real-time control requirements of modern robots for high precision and low latency.
[0003] 2. Low efficiency of hardware accelerator integration: In existing technologies, to improve the computing performance of specific tasks, a special hardware accelerator (such as an AI processor or a matrix operation unit) is usually mounted as a peripheral on the system bus (such as AXI bus). In this way, the processor sends control words and data to the accelerator through memory instructions, which has significant control delay and data transfer overhead. Instructions need to go through multiple links such as processor pipeline and bus arbitration before reaching the accelerator, which seriously restricts the response speed of the acceleration unit and cannot meet the microsecond-level real-time closed-loop control requirements.
[0004] 3. System bottleneck caused by separation of "sensing-computing-control": Existing robot systems usually distribute sensing, computing, and control functions on different hardware modules or chips, and exchange data between modules through buses or communication interfaces. This separated architecture results in complex and costly system hardware, and the communication delay between modules becomes a performance bottleneck of the entire system, making it difficult to achieve close and low-latency collaboration of multi-modal sensing data and control loops.
[0005] Flexibility and energy efficiency: Existing solutions cannot balance performance, flexibility, and energy efficiency. General-purpose processors are flexible but have low energy efficiency; application-specific integrated circuits (ASICs) are efficient but have fixed functions and cannot adapt to rapidly iterating robot algorithms. There is a lack of a domain-specific architecture (DSA) that has both software programmable flexibility and ASIC-level performance and energy efficiency for robot core tasks.
[0006] Therefore, there is an urgent need for a new technical solution to solve the above technical problems. SUMMARY
[0007] The present application aims to overcome the problems of the prior art, and provides a sensor-computation-control integrated robot controller special chip architecture to solve the technical problems of the real-time performance, computational efficiency, energy efficiency and flexibility of the robot controller in the prior art.
[0008] The above-mentioned object is achieved by the following technical solutions: A sensor-computation-control integrated robot controller special chip architecture adopts a multi-core heterogeneous system architecture, including a control core, a computation core and a perception core, the control core is a high-performance real-time processor core cluster based on the RISC-V instruction set; the computation core is used for accelerating AI task inference computation of visual language models, reinforcement learning and target detection; the perception core is used for gathering and preprocessing data from various sensors; the control core, the computation core and the perception core are integrated in a single SoC chip, realizing integrated fusion of multi-modal perception, high-throughput computation and high real-time control functions.
[0009] Further, the control core includes an application core and a real-time core, the application core runs a Linux operating system and processes non-real-time but computationally intensive upper-layer applications; the real-time core runs a real-time operating system and is used for processing trajectory planning and kinematics / dynamics solving tasks, and the real-time core is equipped with instruction local storage and data local storage.
[0010] Further, the computation core is a tensor processing unit with 16-32TOPS parallel computing capability.
[0011] Further, the perception core is a SensorHub.
[0012] Further, an accelerator integration scheme based on custom instruction extension is adopted, including a special coprocessor interface, embedding a domain-specific acceleration unit into the execution pipeline of the real-time core, the decoding dispatch unit of the processor identifies the extended instruction, transmits the instruction and operand to the acceleration unit through the special coprocessor interface, and the acceleration unit writes the calculation result back to the register file of the processor through the special coprocessor interface.
[0013] Further, the domain-specific acceleration unit includes a matrix operation accelerator, the matrix operation accelerator contains a parallel floating-point calculation array and a matrix register, used for accelerating matrix operations in robot kinematics and dynamics; a multi-axis coordination and PID control acceleration engine, the multi-axis coordination and PID control acceleration engine contains a parallel PID processing unit, driven by a finite state machine, supporting position-velocity cascade control, independent position loop mode, and cooperates with an external FPGA to realize closed-loop control of nine-axis motors.
[0014] Further, the parallel floating-point calculation array of the matrix operation accelerator includes 9 multipliers, 9 adders and 1 reciprocal device, the matrix register of the matrix operation accelerator is 4 blocks, the maximum support is 9*9 matrix, and matrix multiplication, matrix inversion and matrix transpose operation are realized through a custom matrix operation instruction.
[0015] Further, the PID closed-loop control response time of the multi-axis coordination and PID control acceleration engine is 125 mu s, the nine-axis synchronization error is less than 5 ns, the eight-line parallel interface is connected with an external FPGA, the FPGA is responsible for current loop control, and the current loop update frequency is 80 kHz.
[0016] Further, the peripherals and storage system special for robots are integrated; the peripherals include CANFD, EtherCAT master station, SPI, I2C, UART, Ethernet and USB communication interfaces; the storage system adopts a multi-level storage architecture, including a core independent L1 instruction / data cache, a core and peripheral shared L2 system cache, and a real-time core dedicated instruction / data local storage, and supports DDR4 and SD card external storage.
[0017] Further, in the multi-level storage architecture, the L2 system cache is 256 kB, and the instruction local storage and data local storage of the real-time core realize zero-wait access of key task code and data.
[0018] The application provides a kind of sense-computing-control integrated robot controller special chip architecture, through software and hardware collaborative design, multi-modal perception, complex calculation and real-time control are deeply integrated in single SoC chip, to significantly improve the comprehensive performance of robot control system, realize the unity of high real-time response, high computing efficiency and high design flexibility. Specific advantages are as follows: 1. Computing performance and efficiency: through special hardware accelerator and custom instruction, the performance of core computing task is greatly improved. For example, the matrix operation accelerator realizes up to 143 times and 42 times acceleration in executing matrix multiplication and inversion, compared with general RISC-V processor under the same process, and more than 10 times acceleration compared with vector processor (VPU), and the chip area is only 1 / 5 of CPU.
[0019] 2. Real-time response capability: the accelerator is integrated into the processor pipeline, which greatly reduces the control delay. The innovative multi-axis coordination control architecture stabilizes the PID closed-loop control response time at 125 mu s, and the nine-axis synchronization error is less than 5 ns. The actual measurement of the real-time kernel prototype chip shows that the interrupt response delay is as low as 435 ns, with very small jitter, fully meeting the hard real-time control requirements.
[0020] 3. High system integration and low cost: Through the "sensing-calculation-control" integrated design, the functions realized by multiple chips in the past are integrated in a single SoC, significantly simplifying the hardware system design of the robot controller, reducing the material cost (BOM), power consumption and physical size, while eliminating the inter-chip communication bottleneck, improving the stability and reliability of the system. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 A functional architecture diagram of the robot controller special chip architecture of the sensing-calculation-control integration of the application; Figure 2 An instruction set extension and accelerator integration architecture diagram in the robot controller special chip architecture of the sensing-calculation-control integration of the application; Figure 3 A matrix operation accelerator hardware architecture diagram in the robot controller special chip architecture of the sensing-calculation-control integration of the application; Figure 4 A multi-axis coordination and control accelerator system architecture diagram in the robot controller special chip architecture of the sensing-calculation-control integration of the application. DETAILED DESCRIPTION
[0022] The application will be further described in detail below according to the drawings and embodiments. The described embodiments are only a part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0023] As shown in Figure 1 A robot controller special chip architecture of sensing-calculation-control integration, the chip adopts a multi-core heterogeneous system oriented to the task requirements of the robot, which is divided into three core modules of control core, calculation core and perception core in function, and is physically mapped to a dedicated hardware unit, realizing the deep integration of "sensing-calculation-control", including: Control core (motion control): a high-performance real-time processor core cluster based on RISC-V instruction set is adopted, which is divided into "application core" and "real-time core". The application core runs a complex operating system such as Linux, processing non-real-time but computation-intensive upper-layer applications; the real-time core runs a real-time operating system (RTOS), which is dedicated to processing tasks such as trajectory planning, kinematics / dynamics solving, etc. with strict determinacy requirements. The real-time core is equipped with instruction local storage (ILM) and data local storage (DLM) to ensure zero-wait access of critical task code and data, and to guarantee hard real-time performance.
[0024] Compute Core (Planning Decision): Integrates a high-performance Tensor Processing Unit (TPU) as the core of AI computing. The TPU has strong parallel computing capabilities (such as 16-32 TOPS) and is dedicated to accelerating the inference computing of AI tasks such as visual language models, reinforcement learning, and object detection.
[0025] Perception Core (Multi-modal Perception): Integrates a SensorHub for efficient gathering and preprocessing of data from various sensors (such as vision, force, IMU, encoder), reducing the burden of the control core, and providing high-quality perception information for upper-level decision-making.
[0026] The control core, the compute core, and the perception core are integrated into a single SoC chip, realizing the integration of multi-modal perception, high-throughput computing, and high real-time control functions.
[0027] The chip architecture of the present embodiment is shown in Figure 1 The chip center is interconnected by AXI bus to form a distributed system. Among them: The "application core cluster" runs the Linux operating system, handles human-computer interaction, network communication, and high-level applications, includes Core0, Core1, and other processor cores, configures I-Cache, D-Cache, Mem, PPI, and other modules, and connects SystemCache (256kB) and USART*3, USB, QSPController, DDR, SDIO, and other peripherals through BusFabric.
[0028] The "real-time core cluster" runs FreeRTOS, and each real-time core is responsible for high-deterministic tasks such as kinematics solving, equipped with I-Cache, ILM, DLM, etc., connected to the bus through PPI. In addition, the chip also integrates TPU computing core and SensorHub perception core, TPU for AI computing, and SensorHub for sensor data gathering and preprocessing.
[0029] To solve the delay bottleneck of traditional bus-mounted accelerators, the present scheme proposes an innovative accelerator integration scheme based on custom instruction extension within the processor pipeline, including: 1. Special co-processor interface (NICE): Designs and adopts a Nuclei Instruction Co-unit Extension (NICE) instruction extension co-processor interface, which embeds domain-specific acceleration units (such as matrix accelerator, PID engine) as functional units directly into the execution pipeline of RISC-V real-time core.
[0030] 2. Direct instruction dispatch: the decode / dispatch unit of the processor can directly recognize the custom extension instructions and send the instructions and operands (source register values) to the corresponding accelerator units through the NICE interface, bypassing the traditional memory access and system bus. After the accelerator completes the calculation, the result is written back to the register file of the processor directly through the NICE interface. This scheme shortens the control delay of the accelerator from tens or even hundreds of cycles to a few cycles.
[0031] As shown in Figure 2 , it is a schematic diagram of the instruction set extension and accelerator integration architecture in this scheme, showing how standard instructions and extension instructions are sent to standard execution units and domain accelerator units through the decode / dispatch unit, and how the domain accelerator unit interacts with the processor pipeline through the NICE special interface. After the decode / dispatch stage of the processor, add judgment logic based on the instruction opcode. If it is a standard RISC-V instruction, send it to the standard execution unit such as ALU, LSU, etc.; if it is a custom-1 opcode, send it to the "domain accelerator unit" through the NICE interface. The NICE interface consists of multiple channels such as request, response, memory access, etc. For example, when the CPU executes the custom matrix multiplication instruction mmulmrd, mrs0, mrs1, the decoder sends the opcode, source register number, and destination register number to the matrix accelerator through the NICE request channel. After the accelerator completes the execution, it sends the completion signal and result back to the CPU through the response channel, and the CPU writes to the destination register file.
[0032] This embodiment designs and integrates multiple special hardware accelerators for the most frequent and time-consuming core tasks in robot algorithms, including: 1. Matrix operation accelerator: used to accelerate the large number of matrix operations in robot kinematics and dynamics. It contains a parallel floating-point calculation array (9 multipliers, 9 adders, 1 reciprocal) and 4 special matrix registers (maximum support 9x9 matrix) inside. Through custom matrix operation instructions (such as matrix multiplication, matrix inversion, matrix transpose), the performance can be improved by tens to hundreds of times compared to general-purpose CPUs, and the hardware area is much smaller than general-purpose processors.
[0033] As shown in Figure 3Figure 2 shows the hardware architecture of the matrix operation accelerator in this solution, illustrating its internal parallel floating-point computation array, matrix operation controller, matrix configuration registers, and interface with the external matrix register array. The matrix accelerator contains a parallel computation array consisting of nine 32-bit floating-point multipliers and nine 32-bit floating-point adders, connected to four matrix register arrays, each capable of storing a 9×9 single-precision floating-point matrix. When performing a matrix inversion operation, the software configures the matrix dimensions using the custom_msetmnp(n,n) instruction. The custom_mload0(addr) instruction loads the matrix to be inverted from DDR memory into matrix register 0. After executing the custom_minv() instruction, the matrix inversion controller is activated, using the Gauss-Jordan elimination method. Controlled by a hardware state machine, the matrix reads rows from the matrix registers and transforms them into rows in the computation array. Intermediate results are written back. After the operation is completed, the result matrix is stored in matrix register 1 and can be written back to memory using the custom_mstore1(addr) instruction.
[0034] 2. Multi-Axis Collaboration and PID Control Acceleration Engine: To achieve high-precision multi-axis motor servo control, a reconfigurable PID calculation engine was designed. This engine comprises multiple parallel PID processing units (PID_PE), driven by a finite state machine (FSM). It supports various modes, including position-velocity cascade control and independent position loops. Working in conjunction with an external FPGA (current loop), it achieves high-speed (125μs response period) and highly synchronized (inter-axis error <5ns) closed-loop control of nine motor axes.
[0035] like Figure 4 Figure 2 shows the architecture of the multi-axis coordination and control accelerator system in this solution. It illustrates how the PID array within the SoC implements hierarchical coordinated control with the external FPGA responsible for the current loop and motors via the EPPI interface. This embodiment is used to drive a nine-axis robotic arm. The SoC chip serves as the master controller, with its internal PID calculation engine responsible for position and velocity loop calculations. Three external FPGA development boards serve as slave controllers, each driving three motors and responsible for the current loop with an 80kHz update frequency. The SoC and FPGAs are connected via a customized eight-wire parallel interface (EPPI). The PID engine within the SoC performs position / velocity loop calculations for the nine axes in a 125μs cycle. Its three internal PID_PE processing units process three sets of motor data in parallel in a pipelined manner, with the calculation results written to the EPPI interface's transmit FIFO via the AXI-Lite bus. The EPPI master controller sequentially selects the three FPGAs and issues current setpoints. Simultaneously, the master FPGA generates an 8kHz synchronization pulse signal, which is distributed to all nodes to latch the motor encoder position data, ensuring data sampling synchronization and achieving high dynamic response and high-precision trajectory tracking for the nine-axis system.
[0036] In addition, the embodiment also provides a robot-specific peripheral and storage system, which comprises: 1. Interface integration: The SoC integrates various communication interfaces necessary for the robot system, such as CAN FD, EtherCAT master station, SPI, I2C, UART, Ethernet, USB, etc., to meet the connection requirements with various sensors, actuators, and external devices.
[0037] 2. Storage hierarchy: A multi-level storage architecture is adopted, including L1 instruction / data cache (I / D-Cache) for each core, L2 system cache (System Cache) shared by all cores and peripherals, and instruction / data local storage (ILM / DLM) dedicated to real-time cores. This system balances the needs of high throughput and low latency deterministic access. It also supports external mass storage such as DDR4 and SD cards.
[0038] As a further description of the present solution, the following is provided: 1. Overall architecture implementation: As shown in Figure 1 and Figure 2 , the chip adopts a multi-core heterogeneous architecture. The center of the chip is a distributed system interconnected by AXI buses. Among them, one cluster acts as an "application core cluster" running Linux operating system, which is used to handle human-computer interaction, network communication and high-level applications. Another cluster acts as a "real-time core cluster" running FreeRTOS, and each real-time core is responsible for one or more high-deterministic tasks, such as kinematics solving. In addition, a TPU computing core and a Sensor Hub sensing core are integrated on the chip.
[0039] 2. Real-time core and accelerator integration implementation: The real-time core selects a 64-bit RISC-V processor, whose pipeline is expanded. As shown in Figure 2 , after the decoding dispatch stage of the processor, a judgment logic based on instruction operation code is added. If the instruction is a standard RISC-V instruction, it is sent to the standard execution units such as ALU and LSU. If the instruction is a custom-1 operation code, it is transmitted to the "domain accelerator" through the NICE interface. The NICE interface is composed of multiple channels such as request, response, and memory access. For example, when the CPU executes a custom matrix multiplication instruction mmul mrd, mrs0, mrs1, the decoder sends the operation code, source register number (mrs0, mrs1), and destination register number (mrd) of the instruction to the matrix accelerator through the NICE request channel. After the accelerator completes the execution, it sends the completion signal and the result back to the CPU through the NICE response channel, and the CPU writes the result to the destination register file.
[0040] 3. Matrix operation accelerator implementation: As shown in Figure 3As shown, the matrix accelerator is implemented as a specific hardware module. The module contains a parallel computing array composed of 9 32-bit floating point multipliers and 9 32-bit floating point adders. The module is surrounded by 4 matrix register arrays, each of which can store a complete 9x9 single-precision floating point matrix. When performing the matrix inversion operation, the software first configures the matrix dimension through the custom_msetmnp(n, n) instruction, and then loads the matrix to be inverted from the DDR memory into the 0th matrix register using the custom_mload0(addr) instruction. Subsequently, the custom_minv() instruction is executed. At this time, the matrix inversion controller is started, and the Gaussian-Jordan elimination method is used to iteratively read data rows from the matrix register, send them to the parallel computing array for row transformation, and write the intermediate results back over multiple cycles. The entire process is controlled by a hardware state machine without CPU intervention. After the operation is completed, the result matrix is stored in the 1st matrix register and can be written back to memory using the custom_mstorel(addr) instruction.
[0041] 4. Multi-axis cooperative control embodiment: As shown in Figure 4 , this embodiment is used to drive a nine-axis robot arm. The SoC chip serves as the master controller, and its internal PID calculation engine is responsible for position loop and speed loop calculations. Three external FPGA development boards serve as slave controllers, each driving three motors and responsible for the current loop with the highest real-time requirement (80 kHz update frequency). The SoC and FPGA are connected through a custom eight-wire parallel interface (EPPI). The PID engine inside the SoC polls the position / speed loop of the nine axes at a period of 125 µs. The three PID_PE processing units inside it handle the data of three groups of motors in a pipeline manner. The calculation results are written to the transmit FIFO of the EPPI interface through the AXI-Lite bus. The EPPI master controller selects the three FPGAs in turn and issues the calculated current set value. At the same time, an 8 kHz synchronization pulse signal generated by the master FPGA is distributed to all nodes to latch the position data of all motor encoders, ensuring the synchronization of data sampling. This layered and cooperative architecture ensures high dynamic response and high-precision trajectory tracking of the entire nine-axis system.
[0042] The above merely describes the embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made by those skilled in the art within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A dedicated chip architecture for a robot controller that integrates sensing, computing, and control, characterized by ,Adopting multi-core heterogeneous system architecture, including: A control core, which is a high-performance real-time processor core cluster based on the RISC-V instruction set; Computing cores, which are used to accelerate reasoning and computation for AI tasks such as visual language models, reinforcement learning, and object detection; A perception core, which is used to aggregate and pre-process data from various sensors; The control core, the computing core and the perception core are integrated into a single SoC chip, realizing the integrated fusion of multimodal perception, high-throughput computing and high real-time control functions.
2. The dedicated chip architecture for a robot controller integrating sensing, computing and control according to claim 1, characterized in that , the control core includes: An application core, which runs a Linux operating system and processes non-real-time but computationally intensive upper-layer applications; A real-time core runs a real-time operating system and is used to process trajectory planning and kinematics / dynamics solving tasks. The real-time core is equipped with local instruction storage and local data storage.
3. The dedicated chip architecture for a robot controller integrating sensing, calculation and control according to claim 1, characterized in that ,The computing core is a tensor processing unit with 16-32TOPS parallel computing ,capacity.
4. The dedicated chip architecture for a robot controller integrating sensing, computing and control according to claim 1, characterized in that ,The perception core is SensorHub.
5. The dedicated chip architecture for a robot controller integrating sensing, computing and control according to claim 1, characterized in that: An accelerator integration solution based on custom instruction extension is adopted, including a dedicated coprocessor interface, and a domain-specific acceleration unit is embedded in the execution pipeline of the real-time core. The processor's decoding and dispatching unit recognizes the extended instructions and transmits the instructions and operands to the acceleration unit through the dedicated coprocessor interface. The calculation results of the acceleration unit are written back to the processor's register file through the dedicated coprocessor interface.
6. The dedicated chip architecture for a robot controller integrating sensing, computing and control according to claim 5, characterized in that: The domain-specific acceleration units include: A matrix operation accelerator, comprising a parallel floating-point calculation array and matrix registers, for accelerating matrix operations in robot kinematics and dynamics; A multi-axis collaboration and PID control acceleration engine includes a parallel PID processing unit, is driven by a finite state machine, supports position-speed cascade control and independent position loop mode, and collaborates with an external FPGA to achieve closed-loop control of nine-axis motors.
7. The dedicated chip architecture for a robot controller integrating sensing, computing and control according to claim 6, characterized in that The parallel floating-point calculation array of the matrix operation accelerator includes 9 multipliers, 9 adders, and 1 invertor. The matrix registers of the matrix operation accelerator are 4 blocks, which support a maximum of 9×9 matrices. Matrix multiplication, matrix inversion, and matrix transposition operations are realized through custom matrix operation instructions.
8. The dedicated chip architecture for a robot controller integrating sensing, computing and control according to claim 6, characterized in that The PID closed-loop control response time of the multi-axis collaboration and PID control acceleration engine is 125μs, and the nine-axis synchronization error is less than 5ns. It is connected to the external FPGA through an eight-wire parallel interface. The FPGA is responsible for current loop control, and the current loop update frequency is 80kHz.
9. The dedicated chip architecture for a robot controller integrating sensing, computing and control according to claim 1, characterized in that , integrating robot-specific peripherals and storage systems; the peripherals include CANFD, EtherCAT master station, SPI, I2C, UART, Ethernet, and USB communication interfaces; the storage system adopts a multi-level storage architecture, including core-independent L1 instruction / data cache, L2 system cache shared by the core and peripherals, and real-time core-specific instruction / data local storage, supporting DDR4 and SD card external storage.
10. The dedicated chip architecture for a robot controller integrating sensing, calculation and control according to claim 9, characterized in that In the multi-level storage architecture, the L2 system cache is 256kB, and the instruction local storage and data local storage of the real-time core realize zero-wait access to critical mission codes and data.
Citation Information
Cited By
Disposable temperature-salinity-depth sensor and acquisition method and system using same
CN121089824A