Neural network acceleration system and method, device, storage medium, and program

CN122311317BActive Publication Date: 2026-09-18INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610761632.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-09-18
Estimated Expiration
2046-05-29

AI Technical Summary

Technical Problem

而这种通过上位机中转的方式,使得上位机处理器成为性能瓶颈,导致输入/输出(Input/Output,I/O)吞吐率与计算资源利用率低下

Benefits of technology

[0011] The neural network acceleration system, method, device, storage medium, and program provided in this application establish a direct data path from storage to computation by directly connecting the storage module to the accelerator and having the operating system and file system running within the accelerator directly manage the storage module. This reduces data path redundancy and data transport overhead caused by the host computer intermediary in traditional architectures. The processing system and programmable logic module collaborate through an on-chip bus interface, enabling autonomous execution of data reading, task scheduling, and neural network computation within the accelerator. This allows for a full overlap between data loading and neural network computation processes, thereby improving input/output throughput and computational resource utilization, while preventing the host computer processor from becoming a performance bottleneck, achieving efficient integration of storage and computation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122311317B_ABST
    Figure CN122311317B_ABST
Patent Text Reader

Abstract

The application discloses a neural network acceleration system and method, equipment, a storage medium and a program, relates to the technical field of artificial intelligence, and constructs a direct data path from storage to calculation by directly connecting a storage module with an accelerator and directly managing the storage module by an operating system and a file system running in a processing system in the accelerator, reduces data path redundancy and data carrying overhead caused by transfer in a host computer in a traditional architecture. The processing system and the programmable logic module are cooperated through an on-chip bus interface, autonomous execution of data reading, task scheduling and neural network calculation in the accelerator is realized. Data loading and the neural network calculation process can be fully overlapped, so that the input / output throughput and the utilization rate of the computing resources are improved, and meanwhile, the host processor is avoided from becoming a performance bottleneck, efficient fusion of storage and calculation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a neural network acceleration system, method, device, storage medium and program. Background Technology

[0002] During large-scale neural network training, data needs to be frequently transferred between storage devices and computing units. In related technologies, data transfer relies on the involvement of a host computer. Data is first loaded from the storage device into the host computer's memory, and then transferred to a dedicated computing unit for neural network computation. This method of relaying data through the host computer makes the host computer processor a performance bottleneck, resulting in low input / output (I / O) throughput and low utilization of computing resources.

[0003] Therefore, how to achieve efficient integration of storage access and neural network computing without the intervention of a host computer, in order to eliminate data path redundancy and host computer processor bottlenecks, and improve I / O throughput and computing resource utilization, is an urgent problem to be solved. Summary of the Invention

[0004] This application provides a neural network acceleration system, method, device, storage medium, and program to at least solve the problem in related technologies of how to achieve efficient integration of storage access and neural network computation without the intervention of a host computer.

[0005] This application provides a neural network acceleration system, including: an accelerator and at least one storage module directly connected to the accelerator. The accelerator includes a processing system and a programmable logic module. The processing system and the programmable logic module exchange data and control interaction through multiple on-chip bus interfaces. The processing system includes a processor, memory and an input / output controller. The programmable logic module includes on-chip memory and an acceleration module.

[0006] The processor runs an operating system and file system, which are used to directly manage the data to be processed in the storage module, perform task scheduling, and control and configure the acceleration module through the input / output controller. The memory is used to store the data to be processed that the processing system reads from the storage module. The on-chip memory is used to cache data to be processed from the processing system, the acceleration module is used to perform neural network calculations based on the data to be processed in the on-chip memory, and the programmable logic module writes the calculation results of the neural network calculations back to memory. The processor is also used to process data based on the calculation results and write the data processing results or calculation results back to the storage module. The process from reading the data to be processed to writing back the data processing results or calculation results is autonomously scheduled by the accelerator.

[0007] This application also provides a neural network acceleration method, which is applied to the processor in the aforementioned neural network acceleration system, comprising: The system reads the data to be processed directly from the storage module into memory through the file system, configures the calculation parameters of the acceleration module, and sends control commands to the acceleration module. The process data in memory is transferred to on-chip memory via the direct memory access engine. The acceleration module is controlled by control commands to perform neural network calculations based on the data to be processed in the on-chip memory, and the calculation results are obtained. Data processing is performed based on the calculation results in memory, and the data processing results or calculation results are written back to the storage module. The calculation results are written back to memory from the on-chip memory by the direct memory access engine.

[0008] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described neural network acceleration methods when executing the computer program.

[0009] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described neural network acceleration methods.

[0010] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described neural network acceleration methods.

[0011] The neural network acceleration system, method, device, storage medium, and program provided in this application establish a direct data path from storage to computation by directly connecting the storage module to the accelerator and having the operating system and file system running within the accelerator directly manage the storage module. This reduces data path redundancy and data transport overhead caused by the host computer intermediary in traditional architectures. The processing system and programmable logic module collaborate through an on-chip bus interface, enabling autonomous execution of data reading, task scheduling, and neural network computation within the accelerator. This allows for a full overlap between data loading and neural network computation processes, thereby improving input / output throughput and computational resource utilization, while preventing the host computer processor from becoming a performance bottleneck, achieving efficient integration of storage and computation. Attached Figure Description

[0012] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a schematic diagram of the structure of a neural network acceleration system provided in an embodiment of this application; Figure 2 This is a schematic diagram of another neural network acceleration system provided in an embodiment of this application; Figure 3 This application provides a schematic diagram of the workflow of a neural network acceleration system according to an embodiment of the present application. Figure 4 This is a flowchart illustrating a neural network acceleration method provided in an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0015] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0016] The neural network acceleration system of this application is a hardware acceleration system that integrates storage access, control interaction and neural network computation into a single accelerator, thereby enabling efficient and autonomous processing of large-scale data to be processed. It can be applied to the fields of machine learning and artificial intelligence. In the training and inference tasks of neural network models, it can effectively process ultra-large-scale datasets that exceed the memory capacity of traditional processors.

[0017] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] Figure 1 This document provides a schematic diagram of a neural network acceleration system according to an embodiment of this application. A detailed description of the structure of the neural network acceleration system will follow.

[0019] like Figure 1As shown, the neural network acceleration system includes: an accelerator and at least one storage module directly connected to the accelerator. The accelerator includes a processing system and a programmable logic module. The processing system and the programmable logic module exchange data and interact with each other through multiple on-chip bus interfaces. The processing system includes a processor, memory and an input / output controller. The programmable logic module includes on-chip memory and an acceleration module. The processor runs an operating system and file system, which are used to directly manage the data to be processed in the storage module, perform task scheduling, and control and configure the acceleration module through the input / output controller. The memory is used to store the data to be processed that the processing system reads from the storage module. The on-chip memory is used to cache data to be processed from the processing system, the acceleration module is used to perform neural network calculations based on the data to be processed in the on-chip memory, and the programmable logic module writes the calculation results of the neural network calculations back to memory. The processor is also used to process data based on the calculation results and write the data processing results or calculation results back to the storage module. The process from reading the data to be processed to writing back the data processing results or calculation results is autonomously scheduled by the accelerator.

[0020] In the embodiments of this application, to facilitate understanding of the neural network acceleration system of the embodiments of this application, this application also provides a schematic diagram of another neural network acceleration system, such as... Figure 2 As shown.

[0021] The accelerator, acting as a computing and scheduling unit, integrates heterogeneous acceleration cores (programmable logic modules) and processing cores (processing systems). At least one storage module directly connected to the accelerator is used for persistent storage of data to be processed. This storage module can be configured as a Non-Volatile Memory Host Controller Interface Solid State Drive (NVMe SSD). The number of storage modules can be configured according to data volume and throughput requirements to achieve storage bandwidth aggregation. The storage modules and accelerator can be directly connected via a Mini Cool Edge Input / Output (MCIO) interface. By directly connecting the storage modules to the accelerator, data can be transferred between the accelerator and storage modules without passing through a host computer.

[0022] The processing system (PS) and programmable logic (PL) modules integrated within the accelerator are packaged together within the same accelerator chip (such as a field-programmable gate array (FPGA) or system-on-chip (SoC)). The processing system is a computing subsystem based on a general-purpose processor, which includes a processor (e.g., an ARM Cortex series multi-core processor), memory (typically Double Data Rate Synchronous Dynamic Random Access Memory (DDR) or High Bandwidth Memory (HBM)), and input / output controllers (hardware controllers used to connect and manage external storage modules and other peripherals, such as Peripheral Component Interconnect Express (PCIe) / Non-Volatile Memory Express (NVMe) controllers).

[0023] Programmable logic modules (PLMs) are hardware logic resource areas based on FPGA architecture. They include on-chip memory (e.g., Block Random-Access Memory (BRAM) / Ultra Random-Access Memory (URAM), and high-speed, low-latency static random-access memory embedded within the PLM) and acceleration modules (i.e., Graph Neural Network (GNN) acceleration cores). The processing system and PLMs are interconnected through multiple on-chip bus interfaces. These on-chip bus interfaces are hardware channels that conform to preset high-speed interconnect protocols (e.g., Advanced eXtensible Interface (AXI)). The on-chip bus interfaces provide pathways for sending control commands, transmitting status information, and moving large amounts of data between memory and the PLMs.

[0024] On the processing system side, the processor runs a complete operating system (e.g., embedded Linux operating system) and file system (e.g., fourth-generation extended file system (ext4), flash-friendly file system (F2FS) etc.), which enables the processor to directly manage and access the data to be processed in the storage module through the running file system based on the input / output controller (PCIe / NVMe controller), without relying on the host computer for device management and data preprocessing.

[0025] The data to be processed is data that requires neural network computation, such as graph data in graph neural network computation. Furthermore, for ease of understanding of the embodiments of this application, the data to be processed will be described using graph data as an example, and the neural network computation process will be described using graph neural network computation as an example.

[0026] The processor can plan the data flow according to the needs of training or inference tasks, including reading data from the file system, configuring Direct Memory Access (DMA) engine parameters, initiating data transfer, waiting for the PL side to complete computation and receive interrupt notifications, and finally reading the result data. Simultaneously, the processor is also responsible for controlling and configuring the acceleration modules in the programmable logic module (PLMB), such as setting the number of layers, dimensions, sampling parameters, and computation start instructions of the neural network model by writing to configuration registers. The processing system's memory is used to receive and temporarily store raw or preprocessed data to be processed read from external storage modules. This data includes the graph's topology (such as adjacency lists) and vertex feature vectors.

[0027] On the programmable logic module (PLM) side, the on-chip memory receives and buffers data to be processed transmitted from the processing system memory via the on-chip bus interface. Due to its low access latency, the PLM effectively meets the high-speed data access requirements of the acceleration module. The acceleration module is a hardware circuit specifically designed for neural network computations (such as neighbor sampling, feature aggregation, and multilayer perceptron computation). The acceleration module reads the required data from the PLM and performs highly parallelized and pipelined computations. After computation, the PLM efficiently writes the neural network computation results (such as forward propagation output features or backpropagation gradient data) back to the processing system memory via the on-chip bus interface.

[0028] Subsequently, the processor of the processing system performs data processing based on the calculation results. For example, in a training scenario, the processor reads gradient data from memory and executes optimizer algorithms (such as Stochastic Gradient Descent (SGD)) to update the model parameters; in an inference scenario, it may perform subsequent logical judgments or formatting on the output features. Finally, the processor writes the data processing results (such as updated model parameters) or direct calculation results (such as inference labels) back to the external storage module through the input / output controller for persistent storage or later use.

[0029] It's important to note that the entire process, from reading the data to be processed to writing back the processing or calculation results, is autonomously scheduled by the accelerator. This means the entire process—from reading the data from the storage module, executing neural network calculations, to writing the results back to the storage module—is completed internally by the processor and programmable logic modules, without the need for data transfer or process scheduling via a host computer's central processing unit. In other words, from reading the raw data from the storage module to writing the processing or calculation results back, the entire process is autonomously scheduled and coordinated by the accelerator's internal processor and programmable logic modules. The host computer only needs to initiate a task request and receive the final result; it does not need to participate in intermediate steps such as data transfer or computation scheduling.

[0030] This application integrates the file system, task scheduling, and computation acceleration functions entirely within the accelerator, avoiding the overhead of intermediary work with a host computer. The heterogeneous collaboration between the processing system and the programmable logic module achieves the separation and efficient coordination of control and data flows. The processing system is dedicated to file management and task scheduling, while the programmable logic module is dedicated to neural network computation, improving the overall system's resource utilization and achieving efficient integration of storage and computation.

[0031] In one possible implementation of this application embodiment, the programmable logic module further includes a direct memory access engine, which supports distributed aggregation operation function to support the transmission of discontinuous pending data and the management of descriptor queues. The direct memory access engine and multiple on-chip bus interfaces enable data exchange and control interaction between the processing system and the programmable logic module via direct memory access.

[0032] In the embodiments of this application, the Direct Memory Access (DMA) engine is used to perform high-speed, background data transfer between the memory of the processing system and the on-chip memory of the programmable logic module. Its operating mode is direct memory access, in which data movement does not require direct processor intervention. The processor only needs to initialize and start the data transfer task; subsequent data transfer is completed independently by the engine, allowing the processor to focus on higher-level logic such as task scheduling and file system management.

[0033] This application utilizes a direct memory access engine that supports distributed aggregation operations, enabling efficient processing of discontinuous memory access patterns of data to be processed through hardware acceleration. The data exchange and control interaction mechanism based on direct memory access improves the collaboration between the processing system and the programmable logic module, effectively enhancing the overall system throughput.

[0034] In one possible implementation of this application embodiment, the memory includes a file system cache area, a direct memory access buffer, and a shared memory area with the programmable logic module; The file system cache is used to cache data to be processed read from the storage module; the direct memory access buffer is used to temporarily store data to be processed that is to be transferred to the programmable logic module or the calculation results returned from the programmable logic module; the shared memory area is used to store control interaction information between the processor and the programmable logic module.

[0035] In the embodiments of this application, the file system cache is a region serving the file system on the processor. When the processor reads data files to be processed from the external storage module through the input / output controller, these data first enter the file system cache. The direct memory access buffer is a memory region specifically established for direct memory access data transfer operations. When it is necessary to send data from the file system cache or processed data to the programmable logic module for computation, this data is copied into the direct memory access buffer to form data blocks suitable for DMA engine transfer. The shared memory region is a memory region that can be directly accessed by both the processor and the programmable logic module of the processing system, enabling asynchronous and efficient bidirectional communication between the processor and the acceleration module.

[0036] Control interaction information includes, but is not limited to: command descriptors issued by the processor to the programmable logic module, calculation task parameter blocks, small-scale configuration data that needs to be updated in real time; and status words, completion notifications, interrupt request-related information, or lightweight result data that the programmable logic module feeds back to the processor.

[0037] This application achieves high-efficiency data transmission through memory partitioning management. The file system cache leverages the capabilities of the file system to improve storage access efficiency; the direct memory access buffer provides a standardized and efficient transit area for high-speed data transmission, ensuring stable and smooth data transfer; and the shared memory area provides a low-latency channel for flexible control interaction. These three memory areas work together to eliminate potential conflicts and bottlenecks between data flow and control flow within the system, enabling the process of reading data from storage and writing back the final result to be completed autonomously and efficiently by the accelerator, thus improving the overall throughput and responsiveness of the system.

[0038] Furthermore, regarding the data storage and access aspects of the neural network acceleration system in this application, the following explanation can be provided: The storage subsystem connects the embedded processor and the GNN acceleration core via a DMA mechanism, serving as the system's data exchange hub. The storage subsystem includes memory, a DMA engine, on-chip memory, and a PCIe / NVMe controller. The storage subsystem has bidirectional communication capabilities: in the download direction, it is responsible for transmitting the graph structure data, vertex feature vectors, training labels, and neural network weight parameters prepared by the PS side to various functional modules within the GNN acceleration core; in the upload direction, it undertakes two different responsibilities: in training mode, it uploads the calculated gradient data for parameter updates by the PS side; in inference mode, it directly uploads the final prediction results.

[0039] This centralized data interface design allows the GNN acceleration core to focus on computationally intensive tasks, while the flexible control logic is handled by the embedded processor on the PS side. The PS-side storage subsystem integrates a complete PCIe / NVMe controller and driver stack, responsible for establishing high-speed connections with external NVMe SSDs. The NVMe driver runs in the operating system kernel space, providing a unified I / O interface upwards through the standard block device layer, and works with the I / O scheduler to optimize the ordering and merging strategies of read and write requests to improve the access efficiency of NVMe SSDs. Data to be processed read from the NVMe SSD is first loaded into the PS-side DDR memory. This memory is used not only to store file system caches but also to allocate a dedicated DMA buffer and a shared memory area with the PL side, preparing for subsequent data transfers.

[0040] In one possible implementation of this application embodiment, the processor also runs a driver corresponding to the programmable logic module. The driver encapsulates the configuration of the acceleration module, the configuration of the direct memory access engine, and the processing of interrupt signals.

[0041] In the embodiments of this application, in addition to the basic operating system and file system services, the processor's operating system also runs a driver specifically developed for programmable logic modules. This driver is a software module in the operating system kernel space, which acts as a bridge between user-space applications and underlying hardware, providing a concise application programming interface (API) to user space while efficiently managing and driving hardware resources.

[0042] This application enables efficient configuration of acceleration module parameters by encapsulating the configuration of the acceleration module within the driver program. Encapsulating the configuration of the direct memory access engine ensures efficient data transfer. Encapsulating interrupt signal handling allows for timely signal response, guaranteeing efficient collaboration between the processor and the acceleration module.

[0043] In one possible implementation of this application embodiment, the processor is further configured to control the neural network computation process; Controlling the neural network computation process includes at least the following: reading the data to be processed from the storage module into memory via the file system; configuring the parameters of the direct memory access engine and initiating data transfer; waiting for the programmable logic module to complete the neural network computation and receiving an interrupt notification; and reading the computation results from memory.

[0044] In the embodiments of this application, the processor acts as a command center, executing key operation sequences during the accelerator's autonomous completion of neural network computation tasks. The processor does not passively wait, but actively controls the neural network computation process. This control process is manifested as a series of carefully choreographed, repeatable steps, ensuring that the entire process from data reading to result retrieval is orderly, efficient, and reliable.

[0045] The control process begins at least with the data loading phase, where the processor reads the data to be processed from the storage module into memory via the file system. At this point, the operating system and its file system running on the processor initiate a read operation to the directly connected storage module through the input / output controller. The raw data to be processed is read into the file system cache, and subsequently copied or mapped to the direct memory access buffer.

[0046] Next, the processor enters the data transfer configuration phase, which involves configuring the parameters of the direct memory access engine and initiating data transfer. When it is necessary to send data to be processed in memory to the programmable logic module for computation, the processor configures a series of key parameters that define the details of the data transfer.

[0047] Subsequently, the processor enters a synchronous waiting phase, waiting for the programmable logic module to complete the neural network calculation and receive an interrupt notification. After the acceleration module inside the programmable logic module completes the neural network calculation, it generates an interrupt notification signal, and the processor responds to this interrupt to confirm the completion of the calculation.

[0048] Finally, the processor enters the result write-back phase. After confirming the completion of the computation, the processor determines that the computation result generated by the programmable logic module and written back through the direct memory access engine is now stored in memory. The processor then accesses the computation result. Depending on the task requirements, the processor may directly use the computation result (e.g., classifying the output in an inference task), or it may perform further data processing based on the computation result (e.g., calling the optimizer algorithm to update model parameters using gradients in a training task).

[0049] This application defines the processor's function in the neural network computation process. By decomposing the control process into data reading, transmission parameter configuration and startup, waiting for notification, and result reading, and utilizing file systems, direct memory access, interrupt notifications, etc., efficient control is achieved. This enables the processor to control the accelerator's operation in a highly efficient manner.

[0050] Furthermore, the functionality of the processing system in this application can be described as follows: The PS (Power Supply Unit) is the software control center of the entire system, undertaking the core responsibilities of operating system operation, file system management, peripheral device control, and collaboration with the PL (Power Processor). The hardware foundation of the PS is an embedded processor. This embedded processor runs a complete operating system, providing the system with a rich software ecosystem and a standard development environment. Above the operating system, the PS manages mature file systems such as ext4 or F2FS, which allows the storage and access of data to be processed to use standard file I / O interfaces, simplifying the development complexity of applications. Data management can be performed like operating a regular operating system, without needing to handle the details of underlying block devices.

[0051] Another key responsibility of the PS side is to act as the master controller of the GNN acceleration core in the PL side. The embedded processor accesses the control registers on the PL side through the AXI General Purpose (AXI GP) port to configure parameters, start tasks, and monitor the status of the GNN acceleration core. Simultaneously, the PS side also runs a specially developed GNN acceleration core driver, which encapsulates low-level details such as register operations, DMA transfer configuration, and interrupt handling, providing a concise API interface to user-space applications.

[0052] When GNN calculations are required, the PS side is responsible for orchestrating the entire process: reading data from the file system, configuring DMA engine parameters, initiating data transfer, waiting for the PL side to complete the calculation and receive interrupt notifications, and finally reading the result data. This makes the PS side act as the command center, coordinating various aspects such as storage access, data migration, and accelerated calculations to ensure the system operates efficiently and orderly.

[0053] In one possible implementation of this application embodiment, a model parameter updater also runs in the processor; In response to the neural network computation process for model training, the processor retrieves the computation results from memory and performs data processing based on the computation results through the model parameter updater to obtain the data processing results; In response to the neural network computation process for model inference, the processor retrieves the computation results from memory; The processor writes the data processing results or calculation results into memory; If the neural network calculation process is model training, the calculation result is gradient data used for model parameter optimization, and the data processing result is the updated model parameters; if the neural network calculation process is model inference, the calculation result is the model inference result.

[0054] In the embodiments of this application, the model parameter updater is a software functional module running on the processor, which encapsulates the strategies and methods for adjusting the learnable parameters inside the neural network model according to the optimization objective.

[0055] Neural network acceleration systems typically operate in at least two modes: model training and model inference. The processor needs to adapt the computation results returned by the programmable logic module according to the current task mode. When the neural network computation process is for model training, the computation results are essentially gradient data used for model parameter optimization. The processor processes the data based on the computation results through a model parameter updater. The model parameter updater applies a preset optimization algorithm, using the current model parameters, the acquired gradient data, and hyperparameters such as learning rate and momentum as input, to perform a series of mathematical operations, ultimately obtaining the data processing result, i.e., the updated model parameters.

[0056] Conversely, when the neural network computation process involves model inference, the resulting computation is directly the model inference result. For inference tasks, the processor can retrieve the computation result from memory after completion, typically without needing to invoke complex parameter update algorithms. The computation result at this point is the final data usable by upper-layer applications, requiring no further conversion into data processing results.

[0057] In both training and inference modes, after the processor completes the processing of the computation results, it needs to write the results back to memory for subsequent persistence. Therefore, the processor writes the data processing results or computation results to memory. In training scenarios, the updated model parameters are written, which are typically used in the next iteration of computation. In inference scenarios, the original model inference results are written, for the application to read and output.

[0058] This application implements the model training iteration process and the model inference process by using a model parameter updater and differentiating the processing logic by mode. This enables the accelerator to function as a module supporting model training and inference. The processor enhances the system's versatility and practicality by dynamically invoking different processing methods (model parameter updates or directly writing back the calculation results).

[0059] Furthermore, regarding the function of the model parameter updater in this application, it can be described as follows: The model parameter updater runs on the PS side and is responsible for the parameter optimization phase throughout the training process. After the FPGA completes a small batch of forward and backward computations, the embedded processor on the PS side reads all uploaded gradient data from DDR and then calls the selected optimization algorithm to update the parameters. Specifically, the model parameter updater executes optimization algorithms based on gradients, such as: basic SGD, momentum SGD, etc., using the current model weights, such as: Gradient data, such as: Hyperparameters, such as learning rate and momentum coefficient, are used as inputs to update the model weight parameters. The updated weights are then used as follows: Write back to the weighted storage area of ​​DDR and send a signal to notify the FPGA that the next small batch can begin.

[0060] After the parameter update is complete, the embedded processor on the PS side writes the new weight parameters back to DDR and triggers a transfer to the FPGA weight cache via the DMA engine. Due to the system's double-buffering mechanism, while the new parameters are being loaded in the background, the FPGA can continue processing the next small batch using the old parameters, achieving overlap between computation and communication and maximizing hardware utilization. The embedded processor and GNN accelerator are designed collaboratively to fully leverage their respective strengths: the GNN accelerator handles parallel computation of rules, while the embedded processor handles flexible algorithm logic.

[0061] In one possible implementation of this application embodiment, the plurality of on-chip bus interfaces include a first interface for data transmission, a second interface for control register access, and a third interface for cache coherency access. The first interface connects to the memory and on-chip memory, the second interface connects to the processor and acceleration module, and the third interface connects to the cache coherence unit and programmable logic module of the processing system.

[0062] In the embodiments of this application, multiple on-chip bus interfaces (AXI bus interfaces) serve as the communication infrastructure connecting the processing system (PS) and the FPGA programmable logic (PL) module. These interfaces include various interface channels with different characteristics to meet diverse communication needs such as control configuration, high-speed data transmission, and cache-coherent access. The multiple on-chip bus interfaces can be based on the Advanced Microcontroller Bus Architecture 4 (AMBAAXI4) protocol standard, supporting high-bandwidth, low-latency burst transmissions and achieving full-duplex communication through independent read / write channels, making them suitable for data interaction between processors and hardware accelerators in heterogeneous systems.

[0063] Multiple on-chip bus interfaces are pre-configured between the PS and PL. These on-chip bus interfaces have completed timing convergence and interconnection optimization at the physical level, and can directly connect custom IP cores on the PL side without having to deal with complex cross-clock domain and physical interface issues.

[0064] The on-chip bus interfaces specifically include three core types of interfaces: a first interface for data transmission, such as the AXI High Performance (AXI HP) port; a second interface for control register access, such as the AXI GP port; and a third interface for cache coherency access, such as the AXI Accelerator Coherency Port (AXI ACP). By differentiating these interfaces, different types of data and signal streams can be isolated onto dedicated physical channels, thereby avoiding resource contention, improving communication bandwidth and efficiency, and meeting the latency, bandwidth, and consistency requirements of different interaction modes.

[0065] This application enables efficient communication by dividing the on-chip bus interface. The first interface ensures high-speed data transmission; the second interface provides a low-latency channel for control and status interaction; and the third interface, through cache coherency, enables memory access of heterogeneous modules, thereby improving system performance.

[0066] In one possible implementation of this application embodiment, the processing system and the programmable logic module implement data transmission based on the first interface through a direct memory access engine; Data transmission includes at least: transferring the data to be processed stored in memory to the on-chip memory, and transferring the calculation results stored in the on-chip memory to memory.

[0067] In the embodiments of this application, the first interface, such as the AXI HP interface, is a set of ports specifically designed for high-speed data transmission, typically with four channels available: HP0 to HP3. The AXI HP interface is directly connected to the DDR memory controller on the PS side. In GNN acceleration scenarios, the DMA engine uses the AXI HP interface to transfer the data to be processed stored in DDR to the on-chip memory on the PL side at high speed, or to write the calculated results back to DDR. The entire process does not require the embedded processor to participate in the movement of each data byte, thereby significantly reducing the burden on the embedded processor and improving data throughput. Since GNN processing often involves large-scale vertex feature matrices and adjacency list data, using DMA in conjunction with the HP interface can effectively reduce the risk of data transmission becoming a system bottleneck.

[0068] This application uses a high-bandwidth first interface and an automated direct memory access engine to exchange data between the processing system's memory and the on-chip memory of the programmable logic module, reducing latency and overhead during data transmission and enabling efficient collaboration between the processing system and the programmable logic module.

[0069] In one possible implementation of this application embodiment, the processor accesses the configuration register of the acceleration module through the second interface to set parameters of the acceleration module, issue control commands, and query the status during the neural network calculation process.

[0070] In the embodiments of this application, the second interface, such as the AXI GP interface, is primarily used for lightweight control plane communication. Unlike the HP interface, which focuses on data transmission, the GP interface is designed to provide flexible register-level access capabilities, with relatively low bandwidth but controllable latency. The embedded processor accesses the configuration registers and status registers of the PL-side GNN acceleration core through the AXI GP interface to set parameters (such as learning rate, number of layers, activation function type, etc.), issue task start commands, and query computation progress and error status. This design of separating control and data allows the software to monitor and adjust the behavior of the hardware accelerator at any time without interfering with the data flow. In addition, the AXI GP interface also supports interrupt signal transmission, allowing the PL side to notify the PS-side driver via an interrupt when computation is complete or an exception occurs, triggering the corresponding subsequent processing flow.

[0071] This application establishes a control and status monitoring channel between the processor and the acceleration module through a second interface and a configuration register, thereby separating the control flow from the data flow. This allows the processor to configure, command-control, and monitor the status of the acceleration module without interfering with the data flow, thus improving the accuracy of system control.

[0072] In one possible implementation of this application embodiment, the third interface is used to enable the programmable logic module to access memory in a cache-consistent manner.

[0073] In the embodiments of this application, the third interface, such as the AXI ACP interface, is an optional high-level feature interface. The AXI ACP interface is directly connected to the cache coherency unit of the embedded processor, allowing the acceleration module on the PL side to access the DDR memory on the PS side in a cache-coherent manner. When accessing memory via the AXI HP interface, it bypasses the cache hierarchy of the embedded processor. If some data in the DDR also exists in the embedded processor cache, explicit cache refresh or invalidation operations need to be performed; otherwise, data inconsistency may occur. However, through the AXI ACP interface, access on the PL side automatically triggers the cache coherency protocol, and the hardware is responsible for checking and updating the cache status, thereby avoiding cumbersome cache management operations at the software level and reducing unnecessary memory overhead. If the PS side frequently preprocesses or postprocesses the data to be processed, using the AXI ACP interface allows the PL side to directly read the latest data in the embedded processor cache or write the results to the cache, improving overall data consistency and access efficiency. However, it should be noted that the bandwidth of the AXI ACP interface is usually lower than that of the AXI HP interface, and it introduces additional cache coherency maintenance overhead. Therefore, its use needs to be weighed according to the specific data access mode.

[0074] Both the AXI ACP and AXI HP interfaces can achieve data transfer between the processing system and programmable logic modules through the direct memory access engine. However, the choice between the AXI ACP and AXI HP interfaces can be dynamically selected based on the specific data access mode or a custom switching algorithm.

[0075] This application simplifies the management of shared data between the processing system and the programmable logic module and improves system performance and reliability by using a cache-consistent access path through a third interface.

[0076] In one possible implementation of this application embodiment, the acceleration module includes: The sampling module is used to perform parallel sampling processing on the data to be processed, generating multiple sub-data; The aggregation calculation module is used to aggregate data features from multiple sub-data sets; The feature transformation module is used to perform feature transformation calculations during the neural network computation process.

[0077] In the embodiments of this application, the PL side is the computational core of the entire system, carrying all the hardware logic implementation for GNN acceleration. Unlike the instruction-driven execution mode of a processor, the PL side implements data paths and control logic optimized for GNN computational characteristics through dedicated hardware circuits, enabling efficient processing of aggregation, transformation, and propagation operations of graph-structured data in a pipelined and parallelized manner. The entire PL side design revolves around the computational model of GNN. The GNN acceleration core (acceleration module) includes: The sampling module has the function of transforming a large graph into smaller subgraphs through sampling, thus addressing a problem in the training of large-scale neural networks.

[0078] The aggregation computing module is one of the core modules of the GNN accelerator, responsible for implementing neighbor feature fusion on the graph structure.

[0079] The feature transformation module is responsible for performing feature transformation operations in GNN. This module achieves unified processing of forward propagation and backward propagation on the same hardware structure through a bidirectional reconfigurable systolic array architecture.

[0080] This application optimizes different computational characteristics by decomposing the computational process of a neural network and processing it through three dedicated modules. These three modules, within the acceleration module, form a highly efficient computational pipeline, allowing sampling, aggregation, and transformation stages to partially overlap, thereby reducing processing latency at each stage and improving overall computational throughput and hardware resource utilization.

[0081] In one possible implementation of this application embodiment, the sampling module adopts a multi-level pipelined parallel architecture, which includes multiple parallel sampling processing units, each of which independently performs neighbor reading, random sampling, and result caching operations; The feature transformation module adopts a reconfigurable bidirectional systolic array architecture to support forward and backward propagation calculations in the neural network computation process on the same hardware structure; During the acceleration module's execution of the current batch of neural network calculations, the processor performs parallel operations to pre-fetch the data to be processed for the next batch of neural network calculations from the storage module.

[0082] In the embodiments of this application, the sampling module adopts a multi-stage pipelined parallel architecture to maximize sampling throughput. Multi-stage pipelined architecture refers to decomposing the sampling task into multiple sequentially executed stages, with these stages overlapping in time. The sampling module contains multiple parallel sampling processing units. These parallel sampling processing units are physically independent hardware logical copies, each capable of fully executing a sampling task. Each parallel sampling processing unit independently performs a series of standardized operations at runtime: neighbor reads, random sampling, and result caching.

[0083] The feature transformation module is a reconfigurable bidirectional systolic array architecture. It features both reconfigurability and bidirectionality. Reconfigurability means that the data paths and computational units within the array can be dynamically adjusted at runtime using configuration signals. Bidirectionality refers to the array's design to support both forward and backward propagation computations in neural network computation on the same hardware architecture.

[0084] Meanwhile, while the acceleration module is performing neural network calculations for the current batch, the processor is performing operations in parallel to pre-fetch the data to be processed for the next batch of neural network calculations from the storage module.

[0085] This application optimizes neural network computation and task scheduling through a parallel pipeline architecture for the sampling module, a reconfigurable bidirectional computing array for the feature transformation module, and parallelization of processor prefetching and accelerator computation. This results in overall low-latency and high-throughput neural network processing capabilities.

[0086] In one possible implementation of this application embodiment, at least one storage module is configured in striped RAID0 mode and uses a hash partitioning or graph partitioning strategy to distribute the data to be processed.

[0087] In the embodiments of this application, the storage module (NVMe SSD) is configured as a RAID 0 array, and the data to be processed is distributed through hash partitioning or graph partitioning strategies, and the aggregated storage bandwidth is matched with the throughput calculated on the PL side.

[0088] This application enables high-speed data transfer by configuring at least one storage module in striped RAID 0 mode. Furthermore, by combining hash partitioning or graph partitioning strategies to distribute the data to be processed, efficient bandwidth utilization is ensured.

[0089] In one possible implementation of this application embodiment, the system further includes a host computer, which is connected to the accelerator via a peripheral interface, and is configured to provide control information to the processor to control the neural network computation process.

[0090] In the embodiments of this application, the host computer refers to a computing device independent of the accelerator, such as a server, workstation, or personal computer. It serves as the entry point for the management, scheduling, and user interaction of the entire acceleration system. The host computer and the accelerator are connected via a peripheral interface. This peripheral interface can be a PCIe interface.

[0091] Control information includes, but is not limited to: the type and structure definition of the neural network model to be used, the dataset used for the training task, the total number of training iterations, batch size, hyperparameters such as optimizer type and learning rate, and descriptions of the input data required for the inference task.

[0092] Specifically, the neural network acceleration system in this application embodiment can be further described through, but is not limited to, the following example: The system consists of a host computer, an FPGA accelerator, and multiple interconnected NVMe SSDs. The FPGA accelerator is connected to the host computer via a PCIe slot; the NVMe SSDs are connected to the FPGA accelerator via MCIO interfaces. The FPGA accelerator has an FPGA MPSoC chip, which integrates a Processing System (PS) and Programmable Logic (PL), which communicate at high speed through multiple on-chip AXI interconnect interfaces. On the PS side, the embedded processor runs an operating system and manages a complete file system, accesses external NVMe SSD storage through a PCIe / NVMe controller, and is responsible for reading and writing data to be processed, task scheduling, and controlling and configuring the accelerator on the PL side. On the PL side, a GNN acceleration core is deployed to perform GNN calculations, equipped with a DMA engine to achieve high-speed data transfer with the PS-side DDR memory, and utilizes on-chip BRAM / URAM cache graph structure and vertex feature data to accelerate access.

[0093] Furthermore, regarding the data flow of the neural network acceleration system, embodiments of this application also provide a schematic diagram of the neural network acceleration system's workflow, such as... Figure 3 As shown, the entire system's data flow follows the path of "storage, processing, acceleration, and write-back": The PS-side embedded processor reads the data to be processed from the SSD into the PS's memory, configures the parameter register of the GNN acceleration core through the AXI GP interface, and then the DMA engine transmits the data to the PL's on-chip storage through the AXI HP interface. After the GNN acceleration core completes the calculation, the result is written back to the PS-side memory through DMA. Finally, the PS-side embedded processor processes the data and writes it back to the NVMe SSD as needed.

[0094] Figure 4 The following is a flowchart illustrating a neural network acceleration method provided for an embodiment of this application. like Figure 4 As shown, this neural network acceleration method is applied to Figure 1 The processor in the neural network acceleration system shown includes: Step 401: Read the data to be processed directly from the storage module into memory through the file system, configure the calculation parameters of the acceleration module, and send control commands to the acceleration module.

[0095] Step 402: Transfer the data to be processed in memory to the on-chip memory through the direct memory access engine.

[0096] Step 403: Control the acceleration module to perform neural network calculations based on the data to be processed in the on-chip memory through control commands to obtain the calculation results.

[0097] Step 404: Perform data processing based on the calculation results in memory, and write the data processing results or calculation results back to the storage module. The calculation results are written back to memory from the on-chip memory by the direct memory access engine.

[0098] In one possible implementation of this application embodiment, the method further includes: In response to the neural network computation process for model training, the computation results are retrieved from memory, and the model parameter updater performs data processing based on the computation results to obtain the data processing results; In response to the neural network computation process for model inference, the computation results are retrieved from memory; Write the data processing results or calculation results into memory; If the neural network calculation process is model training, the calculation result is gradient data used for model parameter optimization, and the data processing result is the updated model parameters; if the neural network calculation process is model inference, the calculation result is the model inference result.

[0099] In one possible implementation of this application embodiment, the method further includes: After the model parameter updater processes the data based on the calculation results and obtains the data processing results, the data processing results are transmitted to the programmable logic module for the next batch of neural network calculations. Specifically, while the control acceleration module is executing the neural network calculations for the current batch, it is simultaneously executing the operation of pre-reading the data to be processed for the next batch of neural network calculations from the storage module.

[0100] For a description of the features in the embodiments corresponding to the neural network acceleration method, please refer to the relevant descriptions in the embodiments corresponding to the neural network acceleration system, which will not be repeated here.

[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0102] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described neural network acceleration method embodiments.

[0103] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described neural network acceleration method embodiments at runtime.

[0104] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0105] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described neural network acceleration method embodiments.

[0106] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described neural network acceleration method embodiments.

[0107] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0108] The foregoing has provided a detailed description of a neural network acceleration system, method, device, storage medium, and program provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A neural network acceleration system, characterized in that, include: An accelerator and at least one storage module directly connected to the accelerator, the accelerator including a processing system and a programmable logic module, the processing system and the programmable logic module exchanging data and interacting with each other through multiple on-chip bus interfaces, the processing system including a processor, memory and input / output controller, the programmable logic module including on-chip memory and an acceleration module; The processor runs an operating system and a file system, which are used to directly manage the data to be processed in the storage module, perform task scheduling, and control and configure the acceleration module through the input / output controller. The memory is used to store the data to be processed read by the processing system from the storage module. The on-chip memory is used to cache the data to be processed from the processing system, the acceleration module is used to perform neural network calculations based on the data to be processed in the on-chip memory, and the programmable logic module writes the calculation results of the neural network calculations back to the memory. The processor is also configured to perform data processing based on the calculation result, and write the data processing result or the calculation result back to the storage module; The processor also runs a model parameter updater; In response to the neural network computation process for model training, the processor retrieves the computation result from the memory and performs data processing based on the computation result through the model parameter updater to obtain the data processing result; the processor writes the data processing result into the memory. In response to the neural network computation process for model inference, the processor retrieves the computation result from the memory; the processor writes the computation result back into the memory. Wherein, if the neural network calculation process is model training, the calculation result is gradient data used for model parameter optimization, and the data processing result is the updated model parameters; if the neural network calculation process is model inference, the calculation result is the model inference result.

2. The neural network acceleration system according to claim 1, characterized in that, The programmable logic module also includes a direct memory access engine, which supports distributed aggregation operation functions to support the transmission of discontinuous pending data and the management of descriptor queues. The direct memory access engine and the plurality of on-chip bus interfaces are used to enable the processing system and the programmable logic module to exchange data and interact with each other via direct memory access.

3. The neural network acceleration system according to claim 1, characterized in that, The memory includes a file system cache area, a direct memory access buffer, and a shared memory area with the programmable logic module; The file system cache area is used to cache the data to be processed read from the storage module; The direct memory access buffer is used to temporarily store the data to be processed that is to be transmitted to the programmable logic module or the calculation results returned from the programmable logic module; the shared memory area is used to store control interaction information between the processor and the programmable logic module.

4. The neural network acceleration system according to claim 1, characterized in that, The processor also runs a driver program corresponding to the programmable logic module, which encapsulates the configuration of the acceleration module, the configuration of the direct memory access engine, and the handling of interrupt signals.

5. The neural network acceleration system according to claim 1, characterized in that, The processor is also used to control the neural network computation process; The controlled neural network computation process includes at least the following steps: reading the data to be processed from the storage module to the memory through the file system; configuring the parameters of the direct memory access engine and starting data transmission; waiting for the programmable logic module to complete the neural network computation and receiving an interrupt notification; and reading the computation result from the memory.

6. The neural network acceleration system according to claim 1, characterized in that, The plurality of on-chip bus interfaces include a first interface for data transmission, a second interface for control register access, and a third interface for cache coherency access. The first interface connects to the memory and the on-chip memory, the second interface connects to the processor and the acceleration module, and the third interface connects to the cache coherence unit of the processing system and the programmable logic module.

7. The neural network acceleration system according to claim 6, characterized in that, The processing system and the programmable logic module implement data transmission based on the first interface through a direct memory access engine. The data transmission includes at least: transferring the data to be processed stored in the memory to the on-chip memory, and transferring the calculation results stored in the on-chip memory to the memory.

8. The neural network acceleration system according to claim 6, characterized in that, The processor accesses the configuration register of the acceleration module through the second interface to set parameters of the acceleration module, issue control commands, and query the status during the neural network calculation process.

9. The neural network acceleration system according to claim 6, characterized in that, The third interface is used to enable the programmable logic module to access the memory in a cache-consistent manner.

10. The neural network acceleration system according to claim 1, characterized in that, The acceleration module includes: The sampling module is used to perform parallel sampling processing on the data to be processed to generate multiple sub-data; The aggregation calculation module is used to perform aggregation operations on the data features of the multiple sub-data; The feature transformation module is used to perform feature transformation calculations during the neural network computation process.

11. The neural network acceleration system according to claim 10, characterized in that, The sampling module adopts a multi-stage pipelined parallel architecture, which includes multiple parallel sampling processing units. Each of the parallel sampling processing units independently performs neighbor reading, random sampling, and result caching operations. The feature transformation module adopts a reconfigurable bidirectional pulsating array architecture to support forward and backward propagation calculations in the neural network computation process on the same hardware structure. During the acceleration module's execution of the current batch of neural network calculations, the processor concurrently performs the operation of pre-reading the data to be processed for the next batch of neural network calculations from the storage module.

12. The neural network acceleration system according to claim 1, characterized in that, The at least one storage module is configured in striped RAID 0 mode and uses hash partitioning or graph partitioning strategies to distribute the data to be processed.

13. The neural network acceleration system according to claim 1, characterized in that, The system further includes a host computer, which is connected to the accelerator via a peripheral interface, and is configured to provide the processor with control information for controlling the neural network computation process.

14. A neural network acceleration method, characterized in that, The method is applied to a processor in a neural network acceleration system as described in any one of claims 1-13, comprising: The system reads the data to be processed directly from the storage module into memory through the file system, configures the calculation parameters of the acceleration module, and sends control commands to the acceleration module. The data to be processed in the memory is transferred to the on-chip memory via the direct memory access engine; The control command controls the acceleration module to perform neural network calculations based on the data to be processed in the on-chip memory, and obtains the calculation results. Data processing is performed based on the calculation results in the memory, and the data processing results or the calculation results are written back to the storage module. The calculation results are written back to the memory from the on-chip memory by the direct memory access engine. The method further includes: In response to the neural network computation process for model training, the processor retrieves the computation results from the memory and performs data processing based on the computation results through a model parameter updater to obtain data processing results; the processor writes the data processing results into the memory. In response to the neural network computation process for model inference, the processor retrieves the computation result from the memory; the processor then writes the computation result into the memory. Wherein, if the neural network calculation process is model training, the calculation result is gradient data used for model parameter optimization, and the data processing result is the updated model parameters; if the neural network calculation process is model inference, the calculation result is the model inference result.

15. The neural network acceleration method according to claim 14, characterized in that, The method further includes: After the model parameter updater processes the data based on the calculation results and obtains the data processing results, the data processing results are transmitted to the programmable logic module for the next batch of neural network calculations. In the process of controlling the acceleration module to perform the current batch of neural network calculations, the operation of pre-reading the data to be processed for the next batch of neural network calculations from the storage module is performed in parallel.

16. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the neural network acceleration method as described in any one of claims 14 to 15.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the neural network acceleration method as described in any one of claims 14 to 15.

18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the neural network acceleration method as described in any one of claims 14 to 15.

Citation Information

Patent Citations

  • Neural network forward prediction hardware accelerator based on programmable logic device

    CN119578475A

  • Processor and accelerator direct connection interface embedded in pipeline, and implementation method therefor

    WO2025194562A1