A three-dimensional reconfigurable hardware acceleration core chip

By designing a three-dimensional reconfigurable hardware acceleration core chip and employing a reconfigurable computing array and controller set, the problem of low resource utilization of DSP chips under diverse computing needs is solved, achieving efficient DSP algorithm processing and flexible computing capabilities.

CN119441130BActive Publication Date: 2025-10-24NANJING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411555385.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-10-24
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

The fixed hardware architecture of existing DSP chips is difficult to adapt to diverse and ever-changing computing needs, resulting in low resource utilization and low design efficiency, especially when processing complex digital signal processing algorithms.

Method used

Design a three-dimensional reconfigurable hardware acceleration core chip, which adopts a reconfigurable computing array, storage array and controller set, and realizes efficient mapping and deployment of signal processing algorithms through data flow driven method, reduces the design complexity of data and control paths, and supports multi-dimensional reconfiguration in spatial, temporal and resource dimensions.

Benefits of technology

It achieves the flexibility of computing arrays to adapt to different digital signal processing tasks, improves resource utilization and computing efficiency, supports efficient processing of various DSP algorithms, and reduces the difficulty of physical implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441130B_ABST
    Figure CN119441130B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional reconfigurable hardware acceleration core chip, and belongs to the technical field of chips.The technical scheme of the three-dimensional reconfigurable hardware acceleration core chip comprises the following: a reconfigurable operation array is used for providing at least one unit-level computing unit and at least one algorithm-level computing unit; a storage array is used for storing operation data input through an AXI bus and output from the reconfigurable operation array; and a controller set is used for controlling the at least one unit-level computing unit and the at least one algorithm-level computing unit to respectively realize unit-level computing operation and algorithm-level computing operation, and controlling operation data storage of the storage array. The application manages scheduling functions such as configuration decoding, reconfiguration control, computing control, data distribution and storage control through an independent control system, constructs a three-dimensional reconfigurable hardware acceleration core chip based on a static scheduling and static data flow model, and realizes multi-dimensional reconfiguration in space dimension, time dimension and resource dimension through storage-computing decoupling.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of chips, in particular to a three-dimensional reconfigurable hardware acceleration core chip. BACKGROUND

[0002] DSP algorithms are used in communication systems for modulation and demodulation, signal encoding and decoding, and error correction, improving the efficiency and reliability of communication. In audio processing, they are used for noise reduction, echo cancellation, and sound enhancement, widely used in sound systems and speech recognition devices. In image processing, these algorithms support image filtering, edge detection, and image compression, used in digital cameras, video surveillance, and medical image analysis. Radar and sonar systems rely on DSP algorithms for target detection and tracking to improve accuracy and resolution. In medical devices, DSP algorithms are used for medical imaging and electrocardiogram analysis to achieve high-precision diagnosis and real-time monitoring. In automatic control systems, they help achieve real-time control and signal analysis. DSP chips play a crucial role in these applications, designed for efficient processing of digital signals, capable of executing complex operations at extremely high speeds, with dedicated computing units such as multipliers and circular shifters, optimized for power consumption, and usually programmable to allow adjustment of processing algorithms as needed. This integrated design simplifies system design, reduces cost and complexity, making DSP chips a core component in modern electronic systems.

[0003] For example, a computing system based on a DSP chip array is provided in Chinese patent CN114185599A, which connects the DSP chip array and array support unit through a hardware link, supports parallel computing, greatly improves the module-level computing unit computing power, starts faster, and effectively reduces hardware costs.

[0004] However, the above-mentioned patent is similar to traditional DSP chips, which usually use fixed hardware architecture and preset operation units, optimized for specific types of computing tasks. Although this specialized architecture can provide high performance when processing specific algorithms, it is difficult to adapt to changing application requirements and algorithm upgrades. As application scenarios diversify and computing needs continue to evolve, fixed hardware architectures often cannot meet new algorithm requirements or performance optimization needs. Especially when dealing with complex and diverse digital signal processing algorithms such as Fast Fourier Transform (FFT), Finite Impulse Response (FIR) filters, autocorrelation, cross-correlation, and various matrix operations, the resource utilization rate and design efficiency of existing DSP chips are low, so the existing technology has deficiencies. SUMMARY

[0005] In view of the deficiencies in the prior art, the purpose of the present application is to provide a three-dimensional reconfigurable hardware acceleration core chip, which can realize multi-dimensional reconfiguration on a chip in space dimension, time dimension and resource dimension, complete efficient mapping and deployment of signal processing algorithms in a data flow driven manner, reduce the design complexity of data and control paths, and effectively reduce the difficulty of physical implementation.

[0006] To achieve the above object, the present application provides the following technical solutions.

[0007] A three-dimensional reconfigurable hardware acceleration core chip comprises a reconfigurable operation array, a storage array and a controller set.

[0008] The reconfigurable operation array is configured to provide at least one unit-level computing unit and at least one algorithm-level computing unit.

[0009] The storage array is configured to store operation data input via an AXI bus and output by the reconfigurable operation array.

[0010] The controller set is configured to control the at least one unit-level computing unit and the at least one algorithm-level computing unit to respectively implement unit-level computing operation and algorithm-level computing operation, and control operation data storage of the storage array.

[0011] As a further improvement of the present application, the controller set comprises a DSP scheduler, a computing resource controller, a memory access resource controller and a DMA resource controller.

[0012] The computing resource controller is configured to schedule and control the reconfigurable operation array.

[0013] The memory access resource controller is configured to control data scheduling between the storage array and the reconfigurable operation array.

[0014] The DMA resource controller is configured to control data transfer between the storage array and the AXI bus.

[0015] The DSP scheduler is configured to control the reconfiguration mode and running state of the computing resource controller, the memory access resource controller and the DMA resource controller.

[0016] As a further improvement of the present application, the computing resource controller comprises an input / output buffer, a computing resource decoder, a first register module and a first reconfiguration state machine module.

[0017] The input / output buffer is configured to buffer output data of the memory access resource controller and the reconfigurable operation array.

[0018] The computing resource decoder is configured to decode the control instructions sent by the DSP scheduler.

[0019] The first register module is configured to register the reconfiguration information of the computing resource controller.

[0020] The first reconfiguration state machine module is configured to control the operation process of the reconfigurable operation array.

[0021] As a further improvement of the present application, the memory access resource controller comprises a data buffer module, a second reconfiguration state machine module, a second register module and a sub-controller.

[0022] The data buffer module is configured to buffer the output data of the DMA resource controller and the computing resource controller.

[0023] The second reconfiguration state machine module is configured to control the reconfiguration process of the sub-controller.

[0024] The second register module is configured to register the reconfiguration information of the memory access resource controller.

[0025] The sub-controller comprises a plurality of unit-level memory access controllers and a plurality of algorithm-level memory access controllers, each of the unit-level memory access controllers corresponds to a unit-level memory access mode, and each of the algorithm-level memory access controllers corresponds to an algorithm-level memory access mode.

[0026] As a further improvement of the present application, the DMA resource controller comprises a DMA module, a logic control module, a data configuration module and a DMA_Port.

[0027] The DMA module comprises a first status register configured to register the working status of the DMA module.

[0028] The logic control module is configured to control the state machine to jump, the state machine is configured to control the data transfer process of the DMA module, and the state machine is composed of the first status register and a combinational logic circuit.

[0029] The data configuration module is configured to register the configuration information of the DMA module and use the configuration information to configure the DMA module.

[0030] The DMA_Port is configured to perform data transmission between the DMA module and the built-in memory of the device.

[0031] As a further improvement of the present application, the DSP controller comprises a buffer area, a decoder, a configuration circuit module, a register group, an FSM state machine and a transmitter.

[0032] The buffer area is configured to buffer the instructions sent on the AXI bus.

[0033] The coder is used for decoding instructions received by the three-dimensional reconfigurable hardware acceleration core;

[0034] The configuration circuit module comprises a read configuration circuit, a write configuration circuit and an interface circuit.

[0035] As a further improvement of the present application, the register group comprises a device register, a configuration register and a second state register;

[0036] The device register is used for storing response mode information and working mode information of the three-dimensional reconfigurable hardware acceleration core chip;

[0037] The configuration register is used for storing configuration information of the memory resource controller and the DMA resource controller;

[0038] The second state register is used for storing running state information of the controller set.

[0039] As a further improvement of the present application, the FSM state machine is used for managing the register group, managing storage permissions, managing calculation sequences and detecting data correlation.

[0040] As a further improvement of the present application, the transmitter comprises a LEN instruction counter unit and a path selection unit;

[0041] The LEN instruction counter unit is used for recording the number of instructions sent by the device register, identifying the address of the instructions, and sending the address to the DSP scheduler;

[0042] The path selection unit is used for selecting a distribution path of configuration information of the configuration register, or a receiving path of state information of the state register.

[0043] The present application provides an algorithm task reconstruction method applied to the three-dimensional reconfigurable hardware acceleration core chip, and the algorithm task reconstruction method comprises:

[0044] The three-dimensional reconfigurable hardware acceleration core chip receives device configuration information sent by the management core chip, and writes the device configuration information into a configuration register;

[0045] Each controller in the controller set completes configuration according to the device configuration information, and realizes reconstruction of the algorithm task.

[0046] As a further improvement of the present invention, if the three-dimensional reconfigurable hardware acceleration core chip is the main processing device of the algorithm task, the three-dimensional reconfigurable hardware acceleration core chip receives the device configuration information sent by the management core chip and writes the device configuration information into the configuration register, including:

[0047] The three-dimensional reconfigurable hardware acceleration core chip obtains the device configuration information by polling the control command area in the double-speed synchronous dynamic random access memory and writes the device configuration information into the configuration register. The device configuration information is written into the control command area of ​​the double-speed synchronous dynamic random access memory by the management core chip through the crossbar switch matrix.

[0048] As a further improvement of the present invention, if the three-dimensional reconfigurable hardware acceleration core chip is a slave processing device for the algorithm task, the three-dimensional reconfigurable hardware acceleration core chip receives device configuration information sent by the management core chip and writes the device configuration information into a configuration register, including:

[0049] The three-dimensional reconfigurable hardware acceleration core chip receives the device configuration information sent by the management core chip through the crossbar switch matrix, and writes the device configuration information into the configuration register in sequence.

[0050] The present invention provides a processing device equipped with the above-mentioned management core chip and the above-mentioned three-dimensional reconfigurable hardware acceleration core chip.

[0051] The present invention dynamically configures multiple computational units through a reconfigurable computational array, enabling the array to adapt to diverse digital signal processing tasks without requiring dedicated hardware for each task. Users can also configure computational units based on specific application requirements, optimizing computational performance and improving resource utilization, thereby enabling efficient processing of a variety of DSP algorithms. This overcomes the flexibility limitations of traditional DSP chips, enabling a wider range of applications and greater computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is the overall architecture diagram of the 3D reconfigurable hardware acceleration core chip;

[0053] Figure 2 This is the structural diagram of the DSP scheduling controller;

[0054] Figure 3 It is the state jump diagram of the main state machine;

[0055] Figure 4 Reconstruct implementation flow charts for algorithmic tasks. DETAILED DESCRIPTION

[0056] The technical solutions of the present application will be described in detail below with reference to the drawings and specific embodiments. It should be understood that the embodiments and specific features in the embodiments are detailed descriptions of the technical solutions of the present application, but not limitations of the technical solutions of the present application.

[0057] The term "and / or" in the following description only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent three cases: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " generally represents that the associated objects before and after it are in an "or" relationship.

[0058] As shown in Figure 1 The three-dimensional reconfigurable hardware acceleration core chip of the embodiment of the present application comprises a reconfigurable operation array, a storage array and a controller set.

[0059] The reconfigurable operation array is used to provide at least one unit-level computing unit and at least one algorithm-level computing unit.

[0060] The storage array is used to store the operation data input through the AXI bus and output by the reconfigurable operation array.

[0061] The controller set is used to control the at least one unit-level computing unit and the at least one algorithm-level computing unit to respectively implement unit-level computing operation and algorithm-level computing operation, and control the storage of the operation data of the storage array.

[0062] Specifically, the reconfigurable computing array, i.e. the PE array, is divided into single-precision PE array and double-precision PE array. The single-precision PE array is two multifunctional arithmetic computing units (1 / 2PE) containing independent data paths and one PE outside multiplication-addition bypass array. The 1 / 2PE can realize intra-PE reconfiguration and inter-PE reconfiguration through the multiplication-addition bypass outside the PE. Each independent 1 / 2PE computing unit includes 8 multiplication-accumulation structures, 1 comparator, wherein the multiplication-accumulation structure includes 1 complex multiplier and 2 complex adders, the comparator array includes 8 comparators, and the multiplication-addition bypass outside the PE is composed of 4 complex adders, 2 complex multipliers, 1 SQRT and a register group. That is, the final operation resources of the single-precision PE array are composed of 18 complex multipliers, 36 complex adders, 16 comparators and 1 sqrt computing unit. The double-precision PE array is an algorithm-level computing unit, which is used for double-precision matrix inversion algorithm through algorithm-level reconfiguration. The final operation resources thereof are composed of 8 double-precision complex adders, 2 double-precision complex multipliers, 1 double-precision comparator and 1 double-precision division unit.

[0063] The storage array is composed of an sram array and a rom array, and functions to store data and results required by each algorithm operation. The sram array is composed of 64 bank sub-modules, and has a total storage capacity of 2MB, a bit width capable of supporting 8 / 16 / 32 / 64 / 128 / 256 bits, two independent ports, one read-only port and one write-only port. The ROM array is composed of 16 ROMs with a capacity of 8KB, and has a fixed interface bit width of 64 bits (ROM frequency 1Ghz). Each bank module is internally composed of 4 2-port SRAM_1024x64bit modules to form 32KB storage, and through internal logic control, the reading and writing of different data bit width data are realized. Each group of signals connected to the bank includes a clock signal clk, a chip selection signal bce, a write mode configuration bwmod, a read / write signal bwren, a write address signal bwaddr, a write data signal bwdata, a read mode configuration brmod, a read address signal braddr, a read data signal brdata, and the like.

[0064] The embodiment of the application dynamically configures a plurality of operation units through the reconfigurable computing array, so that the computing array can adapt to different digital signal processing tasks, and thus the hardware chip can be adjusted in real time according to different algorithm requirements, without the need to design a dedicated hardware for each task.

[0065] Further, the controller set includes a DSP scheduler, a computing resource controller, a memory access resource controller, and a DMA resource controller.

[0066] The computing resource controller is configured to schedule and control the reconfigurable operation array.

[0067] The memory access resource controller is configured to control data scheduling between the storage array and the reconfigurable operation array.

[0068] The DMA resource controller is configured to control data transfer between the storage array and the AXI bus.

[0069] The DSP scheduler is configured to control the reconfiguration mode and the running state of the computing resource controller, the memory access resource controller, and the DMA resource controller.

[0070] The computing resource controller includes an input / output buffer, a computing resource decoder, a first register module, and a first reconfiguration state machine module.

[0071] The input / output buffer is configured to buffer output data of the memory access resource controller and the reconfigurable operation array.

[0072] The computing resource decoder is configured to decode a control instruction sent by the DSP scheduler.

[0073] The first register module is configured to register reconfiguration information of the computing resource controller.

[0074] The first reconfiguration state machine module is configured to control an operation process of the reconfigurable operation array.

[0075] The output data of the DMA resource controller is used as the input data of the memory resource controller, and the output data of the computing resource controller is used as the output data of the memory resource controller.

[0076] Specifically, the computing resource controller is responsible for allocating computing resources, controlling the execution order and timing of the computing module, and managing and scheduling the reconfigurable computing array. The reconfiguration and interconnection mode of multiple operation resources can be controlled. The three-dimensional reconfigurable acceleration core includes more than ten unit-level computing modes, including: (real / complex) multiplication operation, (real / complex) addition operation, (real / complex) subtraction operation, (complex) conjugate operation, (complex) conjugate accumulation tree, butterfly algorithm unit, (real / complex) multiplication accumulation tree, (real / complex) accumulation tree, (real / complex) mean accumulation tree, (complex) conjugate multiplication operation, (real) comparison operation, (real) comparison tree operation, (complex) multiplication accumulation operation, (complex) conjugate multiplication addition tree, (complex) multiplication addition operation, and four algorithm-level computing modes, which are: butterfly operation, double-precision multiplication addition operation, double-precision division operation, and double-precision comparison operation.

[0077] The memory resource controller includes a data buffer module, a second reconfiguration state machine module, a second register module, and a sub-controller.

[0078] The data buffer module is configured to buffer the output data of the DMA resource controller and the computing resource controller.

[0079] The second reconfiguration state machine module is configured to control the reconfiguration process of the sub-controller.

[0080] The second register module is configured to register reconfiguration information of the memory resource controller.

[0081] The sub-controller includes a plurality of unit-level memory controllers and a plurality of algorithm-level memory controllers. Each unit-level memory controller corresponds to a unit-level memory mode, and each algorithm-level memory controller corresponds to an algorithm-level memory mode.

[0082] The output data of the DMA resource controller, i.e., the output data of the DMA_Port, is used as the input data of the memory resource controller, and the output data of the computing resource controller is used as the output data of the memory resource controller. The second reconfiguration state machine module is further configured to allocate the interface of the bank cluster space and the data buffer. The reconfiguration information of the memory resource controller includes the memory resource type, the bank cluster division strategy, and the ping-pong strategy in the bank cluster.

[0083] Specifically, the memory access resource controller is responsible for efficiently distributing data, managing read-write timing control, and completing data scheduling between the storage array and the reconfigurable computing array to ensure efficient data transmission. It contains four unit-level memory access controllers, including a sliding window memory access controller supporting DSP algorithms, a vector memory access controller, a two-dimensional memory access controller, and a sorting memory access controller. In addition, there are two algorithm-level memory access controllers, including a butterfly memory access controller, a single-precision matrix inversion memory access controller, and a double-precision matrix inversion memory access controller.

[0084] The DMA resource controller includes a DMA module, a logic control module, a data configuration module, and a DMA_Port;

[0085] The DMA module includes a first status register for storing the working status of the DMA module;

[0086] The logic control module is used to control the state machine jump, and the state machine is used to control the data transfer process of the DMA module. The state machine is composed of the first status register and a combination logic circuit;

[0087] The data configuration module is used to store the configuration information of the DMA module and configure the DMA module using the configuration information;

[0088] The DMA_Port is used for data transmission between the DMA module and the built-in memory of the device.

[0089] The data transfer process of the DMA module includes dividing data blocks, setting data transfer dimensions, controlling outstanding, and controlling burst length. The configuration information of the DMA module includes data address segments, data dimensions, and data lengths. In the three-dimensional reconfigurable hardware acceleration core architecture, the built-in memory of the device is an SRAM (static random access memory).

[0090] Specifically, the DMA resource controller contains a DMA module that performs data transmission through an AXI bus interface, and realizes the following main functions: each DMA supports two channels; configuration register request transmission and hardware request transmission; supports reading, writing, reading and writing, frame transmission, and multi-block data transmission; supports interrupts and error interrupts; supports pausing and canceling DMA transmission; each DMA channel supports at least 16 outstanding numbers; uses physical addresses; each channel contains a 16KB buffer, and a total of 32KB buffer; supports more than ten data transfer modes; supports configuration and data transfer.

[0091] As Figure 2The DSP scheduler includes a buffer area, a decoder, a configuration circuit module, a register group, an FSM state machine and a transmitter.

[0092] The buffer area is used for buffering instructions sent on the AXI bus.

[0093] The decoder is used for decoding instructions received by the three-dimensional reconfigurable hardware acceleration core.

[0094] The configuration circuit module includes a read configuration circuit, a write configuration circuit and an interface circuit.

[0095] The register group includes a device register, a configuration register and a second state register.

[0096] The device register is used for storing response mode information and working mode information of the three-dimensional reconfigurable hardware acceleration core chip.

[0097] The configuration register is used for storing configuration information of the memory access resource controller and the DMA resource controller.

[0098] The second state register is used for storing running state information of the controller set.

[0099] The FSM state machine is used for managing the register group, managing storage permissions, managing calculation sequences and detecting data correlation.

[0100] The transmitter includes a LEN instruction counter unit and a path selection unit.

[0101] The LEN instruction counter unit is used for recording the number of instructions sent by the device register, identifying the address of the instructions, and sending the address to the DSP scheduler.

[0102] The path selection unit is used for selecting a distribution path of configuration information of the configuration register, or a receiving path of state information of the state register.

[0103] Specifically, the DSP scheduler is responsible for configuration information decoding, supervising and controlling the reconfiguration mode and the running state of each sub-module. It internally contains a register group, a finite state machine, a program counter, a decoder and a buffer area. The register group includes: a device register responsible for controlling the response mode and the working mode of the acceleration core; a configuration register responsible for storing parameters such as start address, memory access type, operation type, data length and operation point number; a main mode register responsible for controlling program running in the main mode; a state register responsible for representing the running state of sub-module controllers such as the memory access resource controller and the calculation resource controller; and an exception state register indicating calculation exceptions.

[0104] Further, the DSP scheduling controller is also used to perform operation control of the acceleration core. If the acceleration core is idle at present, the scheduling controller reads out the relevant information in the configuration register, such as the start address, data length, operation point number, etc., and writes them into the internal register of the DMA and the internal register of the reconfigurable control unit respectively, starts the DMA to perform data transfer, and informs the operation array to perform the specified operation. After the operation array finishes the operation, it informs the scheduling controller, and the scheduling controller performs the next control scheduling.

[0105] The FSM main state machine is used to control the signal processing algorithm. When the acceleration core is in the starting state, the corresponding state machine is selected and started according to the executed algorithm. When the acceleration core is started, the main state machine is in the IDLE state. The state machine jump flow is as shown in Figure 3 The meanings of the states are as follows:

[0106] IDLE: If the configuration information is valid, the configuration information is read in, and the READY state is entered. If the configuration register is invalid, the DSP is informed by the master to send the acceleration core idle flag, and the IDLE state is kept. The sub-module enable signals are all in the reset state.

[0107] READY: The operation type in the configuration information is analyzed. If it is an illegal operation type (i.e. an undefined operation type), an error is reported, the err of the state register becomes 4’b1111, and the state machine returns to the IDLE state. If the operation type is legal, it is decoded and the CONFIG_READY state is entered.

[0108] CONFIG_READY: The master pulls up the DMA resource controller enable (only indicating that the data transfer is started), the memory controller enable and the calculation resource controller enable, waits for the configuration completion signal to return, and enters the SRC_TRANS state. Otherwise, the CONFIG_READY state is kept.

[0109] SRC_TRANS: If the source data does not need to be continuously transferred (the source data transfer completion signal is pulled up, or the source data transfer completion signal is pulled up by the DMA resource controller by skipping the source data transfer), the CALCULATE_READY state is jumped to. Otherwise, the SRC_TRANS state is kept.

[0110] CALCULATE_READY: The master queries the memory resource state register. If the memory resource controller is idle, the CALCULATE state is jumped to, and the cal_strat signal is generated. Otherwise, the CALCULATE_READY state is kept.

[0111] CALCULATE: The master controls the query access resource state register, if the access resource controller is idle and the calculation completion flag is pulled high, then jump into the RESULT_TRANS state, otherwise, keep the CALCULATE state.

[0112] RESULT_TRANS: If there is no need to continue to carry the result data (the result data carrying is completed, or the result data carrying completion signal is pulled high by the DMA resource controller to skip the result data carrying), jump to the RESULT_TRANS_FINSH state, otherwise keep the CALCULATE_FINISH.

[0113] RESULT_TRANS_FINSH: The master queries whether the next configuration register is valid in the slave mode or whether the master mode is the last one, if valid, jump into the MASTER_TRANS state to continue the next batch of configuration decoding; if invalid, write the DONE signal to the master state register, and the master returns to the IDLE state and generates an interrupt signal to send to the CPU.

[0114] Further, the three-dimensional reconfigurable hardware acceleration core chip provided by the embodiment further comprises an AXI interface, the AXI interface uses DW_axi_gs IP in the DW_axi_gs IP of the DesignWare synthesizable device group of the Synopsis company. This IP provides a simple method for third parties to connect to the AMBA AXI bus. They use a simplified interface to reduce the design complexity of the third-party controller required when the third party is connected to the AXI bus.

[0115] The three-dimensional reconfigurable hardware acceleration core chip provided by the embodiment forms a reconfigurable acceleration architecture for realizing a static scheduling and static data flow model through an independent scheduling system of configuration decoding, reconfiguration control, calculation control, data distribution and storage control. The architecture realizes multi-dimensional reconfiguration technology in the spatial dimension, the time dimension and the resource dimension through storage and calculation decoupling.

[0116] Further, as shown in Figure 4 the algorithm task reconfiguration method provided by the embodiment is applied to the three-dimensional reconfigurable hardware acceleration core chip, and comprises the following steps:

[0117] The three-dimensional reconfigurable hardware acceleration core chip receives the device configuration information sent by the management core chip, and writes the device configuration information into a configuration register;

[0118] Each controller in the controller set completes the configuration according to the device configuration information, and realizes reconfiguration of the algorithm task.

[0119] If the three-dimensional reconfigurable hardware acceleration core chip is a master processing device of an algorithm task, the three-dimensional reconfigurable hardware acceleration core chip acquires device configuration information in a control command area of a double-rate synchronous dynamic random access memory through polling, and writes the device configuration information into a configuration register, wherein the device configuration information is written into the control command area of the double-rate synchronous dynamic random access memory by a management core chip through a crossbar matrix.

[0120] If the three-dimensional reconfigurable hardware acceleration core chip is a slave processing device of an algorithm task, the three-dimensional reconfigurable hardware acceleration core chip receives device configuration information sent by a management core chip through a crossbar matrix, and sequentially writes the device configuration information into a configuration register.

[0121] Specifically, when the acceleration core is a master device, the management core needs to compile and generate configuration information in advance, and write the configuration information into a control command area of a DDR through an L3 Cache by the management core through a CrossBar (crossbar matrix). The acceleration core acquires the configuration information in the control command area of the DDR through polling, and ping-pong writes a batch / multiple batches of configuration information into a configuration register in the core. The three-dimensional reconfigurable acceleration core has one configuration register in the master mode, which contains 512 bits of configuration information. When the configuration information is transported, the DMA resource controller waits for the next command of the master. At this time, the in-core decoder decodes the configuration information and starts the DMA resource controller, the memory access resource controller, and the calculation resource controller configuration. When all the configuration completion signals in the status register are pulled high, the acceleration core starts operation. The DMA resource controller starts data transport according to the configuration information. After the first batch of data transport is completed, the DMA resource controller writes a (first batch) data transport completion signal into the status register. The calculation resource controller and the memory access resource controller are started, and the calculation is started. When the unit-level operation is completed, the unit-level memory access controller writes a result data transport signal into the status register. The DMA resource controller queries the status register and transports the operation result to the DDR. After the current result data transport is completed, if it is not the last batch of master mode configuration information, the master sends a configuration transport task to the DMA resource controller again. The DMA continues to acquire the next batch of configuration information in the control command area of the DDR through polling, until the last batch of configuration operation in the master mode is completed, and the master sends an interrupt request to the CPU.

[0122] When the acceleration core is a slave device, the management core can sequentially write the generated configuration information into the configuration register in the core through the CrossBar. In the slave mode, the three-dimensional reconfigurable acceleration core has four configuration registers, each of which contains 512 bits of configuration information. The acceleration core master will sequentially execute valid configurations until all configurations are executed, thereby realizing the configuration and calculation of the acceleration core in the slave mode.

[0123] The three-dimensional reconfigurable hardware acceleration core chip provided by the embodiment adopts a three-dimensional reconfigurable hardware acceleration core architecture, supports a unit-level reconfiguration mode when different algorithms are executed, and realizes rapid reconfiguration of an algorithm-level task through a reconfigurable operation and memory access mode combination.

[0124] The existing RASP architecture adopts an algorithm reconfiguration controller control mode, and different algorithm controllers are selected to realize specific algorithm functions and uses. Different from the prior art, the three-dimensional reconfigurable chip provided by the embodiment adopts a three-dimensional reconfigurable hardware acceleration core architecture, and through a storage and calculation decoupling control mode, the calculation reconfiguration control and data memory access control are decoupled, and different calculation resources and memory access resources are cross combined to realize specific algorithm functions and uses.

[0125] Specifically, the DSP scheduling controller offloads past data flow control tasks to the sub-controllers, and the sub-controllers are divided into: 1) a calculation resource controller responsible for calculation resource allocation and calculation timing control; 2) a memory access resource controller responsible for efficient data distribution and read-write timing control; and 3) a DMA resource controller responsible for data storage location and moving strategy control.

[0126] All modules complete unified information interaction through a register group, and the register group includes device registers, configuration registers, state registers and the like. The configuration registers are used to realize reconfiguration control, and the maintenance work thereof is completed by the scheduling controller (writable), and the remaining modules are read-only. The state registers are used to realize unified information interaction, and the maintenance work thereof is completed by the sub-resource controllers (writable), and the remaining modules are read-only except the scheduling controller. Based on the above design, the data flow and the control flow are logically isolated, each sub-resource controller has the read permission of all configuration and state registers and the write permission of the state register maintained by itself, except the scheduling controller, the data flow conflict risk is reduced, the hardware maintenance efficiency is improved, and the subsequent testing and use are facilitated.

[0127] In addition, the existing RASP architecture also has a complex wiring complexity, and the main reasons for this situation are: 1) a large number of BANKs: the increase in the number of BANKs will bring a proportional increase in the number of wirings; 2) each of the reconfigurable calculation resources and the BANKs is separately interconnected: there is no unified connection mode between the calculation resources and the BANKs, the reconfiguration process of each calculation resource is performed in the reconfigurable controller, each operation unit is interconnected with all the BANK areas, and the wiring quantity is doubled with the increase of each reconfigurable controller; 3) redundant data control logic: for the same BANK area, the address generation and address distribution logic between different algorithm reconfigurable controllers cannot be reused even if they are the same, so that each increase in an algorithm must increase a reconfigurable controller, and the wiring quantity increases exponentially.

[0128] To solve the above problems, the embodiments of the present application provide three methods for reducing the complexity of wiring, which are as follows:

[0129] 1. Reducing the number of BANKs, which is the most direct method for reducing the number of wirings, optimizing the architecture design scheme, and reducing the number of wirings by reducing 128 BANKs to 72 BANKs.

[0130] 2. Modifying the calculation resource reconstruction method: by configuring the calculation resource path, reconstructing the calculation structure required by the algorithm, isolating the calculation resources inside the PE array from the outside interaction, and interconnecting with the BANK through a unified input / output buffer, the number of interconnection lines is reduced.

[0131] 3. Storage and calculation decoupling: for the same bank area, one address generation and address distribution logic corresponds to one memory access mode, even if the algorithm is different, the same memory access mode can be reused. At the same time, the calculation resource and the storage resource are decoupled, and are interconnected through a unified data buffer, so that multiple algorithm modes can be realized by mutual combination, and the number of wirings can be reduced.

[0132] Further, the embodiments of the present application provide two configuration time optimization strategies for the three-dimensional reconfigurable hardware acceleration core, which are as follows:

[0133] 1. Compressing redundant configuration information to improve configuration efficiency: the total length of the configuration information of the three-dimensional reconfigurable acceleration core is 512 bits, which is shortened by 39% compared with the RISC-V core of 832 bits.

[0134] 2. Expanding the size of a single configuration register to 256 bits to reduce the number of configuration registers and save configuration time. The length of 256 bits is selected to maximize the utilization of AXI bus bandwidth.

[0135] After the above two optimizations, the configuration time of the FFT algorithm of the three-dimensional reconfigurable hardware acceleration core is shortened from the original 7.857us to 0.657us, with a reduction of 91.6%.

[0136] Through the above analysis, it can be seen that the three-dimensional reconfigurable acceleration core chip provided by the embodiments of the present application is designed based on the static scheduling and static data flow execution (SSD) model of storage and calculation decoupling. Its design goal is to provide high-performance, low-power signal processing acceleration capability, which has the following characteristics:

[0137] 1. Scalability: using the storage and calculation decoupling method, the fixed storage mode and calculation mode are combined to realize the expansion of the algorithm support range, which can adapt to different application scenarios and algorithm requirements and support part of unknown algorithms. For example, the chip supports different data types and different calculation unit and storage unit configurations, rather than different algorithm configuration modes.

[0138] Specifically, the scalability is embodied in the following aspects:

[0139] 1) Based on the SSD model, the design idea of decoupling storage and calculation is added to further split the static data stream, and the data transmission path is divided into storage, calculation and distribution, which are encapsulated. While taking advantage of the high computational efficiency of static data flow, the coupling mode is increased and the application function is expanded.

[0140] 2) A unit-level scheduling mechanism is designed, which realizes a variety of algorithms through the combination of different storage control logic and dozens of calculation control logic, and supports variable parameters. This design provides scalable algorithm support for future users.

[0141] 3) For algorithms with typical and determined requirements, API function library support is used to realize them, and the reconstruction path, algorithm type and execution order are determined to ensure their performance. For algorithms with uncertain requirements, unit-level configuration is used to realize them, ensuring that new algorithms proposed by future applications can be deployed on the acceleration core, realizing scalability.

[0142] 4) The operation resources and memory resources are divided into independent data flow units, which perform calculation operations in parallel under the control of the clock signal. Static data flow will solidify the data transmission path for different algorithms inside the hardware, and encapsulate it as a black box. The operation state in the black box is controlled by configuration registers. This design method can reduce the complexity of the data path, reduce or avoid the problems caused by data conflicts and resource competition, thereby providing scalability while improving computational efficiency.

[0143] 5) In addition, with the widespread application of digital images, there is an urgent need for high-precision image data processing and analysis capabilities in new fields such as computer vision, medical image processing, and unmanned driving. Image processing algorithms play a crucial role in these applications, such as image enhancement, object recognition, and motion tracking. High-performance DSP processors can execute image processing algorithms more efficiently, providing real-time processing speed and lower power consumption. This provides a technical foundation for the development of image processing algorithms, enabling more complex algorithms to be implemented in practical applications. Therefore, to achieve algorithm scalability, a two-dimensional sliding window memory controller for image processing is added to the three-dimensional reconfigurable hardware acceleration chip, which can implement a two-dimensional sliding window process and implement image enhancement algorithms such as neighborhood averaging, weighted averaging, Gaussian filtering, bilateral filtering, Laplace operator and Sobel operator, and edge detection.

[0144] 2. Universality, specifically embodied in the following aspects:

[0145] 1) It has the characteristics of software-defined hardware, and its design idea is also to abstract hardware resources into programmable modules, so that they can be reconfigured and reallocated as needed. Through software-defined hardware, the function of the hardware can be changed by reprogramming and reconfiguration without replacing the hardware. Such functions make it suitable for multiple application fields and scenarios, such as graphics processing, machine learning inference, signal processing, etc.

[0146] 2) The resources of the computing unit can be reallocated according to the needs of the task to achieve better performance and efficiency. In addition, since static scheduling is performed at runtime, it can adapt to various different computing loads and needs, providing better versatility.

[0147] 3) It supports the decomposition of computing tasks into many small operations, each corresponding to a node in the computing graph. These operation steps can be executed in different orders, thus realizing the decomposition and reorganization of computing tasks.

[0148] 4) Fixed-length configuration registers are used to control the configuration method and provide corresponding API configuration functions, so that the chip architecture can be easily integrated into other compilation frameworks, thus further improving the versatility of the chip architecture.

[0149] 3、Parallelism: Static data flow execution can fully utilize the advantages of parallel data flow and hardware resources, achieving higher computing parallelism.

[0150] 4、Flexibility: It can support real-time processing needs, and through the combination of CPU scheduling and static data flow execution, it can be reconfigured in a short time to meet the real-time processing needs of various algorithms or data, providing flexible acceleration solutions for various applications. For example, it can support different data flow processing methods as needed, such as high-parallelism low-correlation data flow processing, low-parallelism high-correlation data flow processing, and full-flow high-speed data flow processing; support different operation sequence combinations, such as free combination of memory, computation, and storage modes.

[0151] 5、Ease of use: It has better ease of use, and can provide users with more friendly function libraries and programming interfaces, making it easier for users to use the acceleration core to accelerate signal processing tasks.

[0152] Those skilled in the art will appreciate that embodiments of the application can be devised for a method, a system, or a computer program product. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.

[0153] The present application is described in reference to the flowchart and / or block diagrams of the method, apparatus (system) and computer program product according to embodiments of the application. It should be understood that each flow and / or block in the flowchart and / or block diagrams, and a combination of flows and / or blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, a special purpose computer, an embedded processor or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagrams block or blocks. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the function specified in the flowchart and / or block diagram block or blocks.

[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions means which implement the function specified in the flowchart and / or block diagrams block or blocks. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the function specified in the flowchart and / or block diagram block or blocks.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the function specified in the flowchart and / or block diagrams block or blocks. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the function specified in the flowchart and / or block diagram block or blocks.

[0156] The above description is only preferred embodiments of the present application, the protection scope of the present application is not limited to the above-mentioned embodiments, any technical scheme falling within the concept of the present application shall be within the protection scope of the present application. It should be noted that, for those skilled in the art, some improvements and refinements without departing from the principle of the present application shall be considered as the protection scope of the present application.

Claims

1. A three-dimensional reconfigurable hardware acceleration core chip, characterized by, The three-dimensional reconfigurable hardware acceleration core chip comprises a reconfigurable operation array, a storage array and a controller set; The reconfigurable operation array is configured to provide at least one unit-level computing unit and at least one algorithm-level computing unit; The storage array is configured to store operation data input via an AXI bus and output by the reconfigurable operation array; The controller set is configured to control the at least one unit-level computing unit and the at least one algorithm-level computing unit to respectively implement unit-level computing operation and algorithm-level computing operation, and control operation data storage of the storage array; The controller set comprises a DSP scheduler, a computing resource controller, a memory access resource controller and a DMA resource controller; The computing resource controller is configured to schedule and control the reconfigurable operation array; The memory access resource controller is configured to control data scheduling between the storage array and the reconfigurable operation array; The DMA resource controller is configured to control data transfer between the storage array and the AXI bus; The DSP scheduler is configured to control reconfiguration mode and running state of the computing resource controller, the memory access resource controller and the DMA resource controller; The DMA resource controller comprises a DMA module, a logic control module, a data configuration module and a DMA_Port; The DMA module comprises a first state register configured to register working state of the DMA module; The logic control module is configured to control state machine jumping, the state machine being configured to control data transfer process of the DMA module, and the state machine being composed of the first state register and a combination logic circuit; The data configuration module is configured to register configuration information of the DMA module and use the configuration information to configure the DMA module; The DMA_Port is configured to perform data transmission between the DMA module and a device built-in memory.

2. The three-dimensional reconfigurable hardware acceleration core chip according to claim 1, wherein, The computing resource controller comprises an input / output buffer, a computing resource decoder, a first register module and a first reconfiguration state machine module; The input / output buffer is configured to buffer output data of the memory access resource controller and the reconfigurable operation array; The computing resource decoder is configured to decode a control instruction sent by the DSP scheduler; The first register module is configured to register reconfiguration information of the computing resource controller; The first reconfiguration state machine module is configured to control operation process of the reconfigurable operation array.

3. The 3-D reconfigurable hardware acceleration core chip of claim 1, wherein, The memory access resource controller comprises a data buffer module, a second reconfiguration state machine module, a second register module and a sub-controller; The data buffer module is configured to buffer output data of the DMA resource controller and the computing resource controller; The second reconfiguration state machine module is configured to control reconfiguration process of the sub-controller; The second register module is configured to register reconfiguration information of the memory access resource controller; The sub-controller comprises a plurality of unit-level memory access controllers and a plurality of algorithm-level memory access controllers, each unit-level memory access controller corresponding to a unit-level memory access mode, and each algorithm-level memory access controller corresponding to an algorithm-level memory access mode.

4. The 3-D reconfigurable hardware acceleration core chip of claim 1, wherein, The DSP scheduler comprises a buffer, a decoder, a configuration circuit module, a register set, an FSM state machine and a transmitter. The buffer is used for buffering instructions sent on an AXI bus. The decoder is used for decoding instructions received by the three-dimensional reconfigurable hardware acceleration core. The configuration circuit module comprises a read configuration circuit, a write configuration circuit and an interface circuit.

5. The three-dimensional reconfigurable hardware acceleration core chip according to claim 4, wherein, The register set comprises a device register, a configuration register and a second state register. The device register is used for storing response mode information and working mode information of the three-dimensional reconfigurable hardware acceleration core chip. The configuration register is used for storing configuration information of the memory access resource controller and the DMA resource controller. The second state register is used for storing running state information of the controller set.

6. The three-dimensional reconfigurable hardware acceleration core chip according to claim 4, wherein, The FSM state machine is used for managing the register set, managing storage permissions, managing calculation sequences and detecting data correlation.

7. The three-dimensional reconfigurable hardware acceleration core chip according to claim 5, wherein, The transmitter comprises a LEN instruction counter unit and a path selection unit. The LEN instruction counter unit is used for recording the number of instructions sent by the device register, identifying the address of the instructions and sending the address to the DSP scheduler. The path selection unit is used for selecting a distribution path of configuration information of the configuration register or a receiving path of state information of the second state register.

8. An algorithm task reconfiguration method applied to the three-dimensional reconfigurable hardware acceleration core chip according to claim 1, characterized in that, The algorithm task reconstruction method comprises: The three-dimensional reconfigurable hardware acceleration core chip receives device configuration information sent by a management core chip and writes the device configuration information into a configuration register. Each controller in the controller set completes configuration according to the device configuration information and realizes reconstruction of the algorithm task.

9. The method of claim 8, wherein, If the three-dimensional reconfigurable hardware acceleration core chip is a master processing device of the algorithm task, the three-dimensional reconfigurable hardware acceleration core chip receives device configuration information sent by a management core chip and writes the device configuration information into a configuration register, comprising: The three-dimensional reconfigurable hardware acceleration core chip acquires the device configuration information in a control command area of a double-rate synchronous dynamic random access memory by polling and writes the device configuration information into the configuration register, wherein the device configuration information is written into the control command area of the double-rate synchronous dynamic random access memory by the management core chip through a crossbar matrix.

10. The method of claim 8, wherein, If the three-dimensional reconfigurable hardware acceleration core chip is a slave processing device of the algorithm task, the three-dimensional reconfigurable hardware acceleration core chip receives device configuration information sent by a management core chip and writes the device configuration information into a configuration register, comprising: The three-dimensional reconfigurable hardware acceleration core chip receives the device configuration information sent by the management core chip through a crossbar matrix and sequentially writes the device configuration information into the configuration register.

11. A processing device, characterized by The three-dimensional reconfigurable hardware acceleration core chip is loaded with a management core chip and the three-dimensional reconfigurable hardware acceleration core chip as claimed in any one of claims 1-7.

Citation Information

Patent Citations

  • Operational system based on DSP (Digital Signal Processor) chip array

    CN114185599A

  • Coarse granularity dynamic reconfigurable data integration and control unit structure

    CN103761075A

  • High-efficient controller and control method of configurable water flow signal processing core

    CN105955923A