Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6 results about "Fpga acceleration" patented technology

The Intel FPGA Acceleration Stack. The Acceleration Stack for Intel Xeon CPU with FPGAs is a robust collection of software, firmware, and tools designed and distributed by Intel to make it easier to develop and deploy Intel FPGAs for workload optimization in the data center.

PCIe link monitoring and self-repair system for FPGA acceleration card

The application provides a PCIe link monitoring and self-repair system for an FPGA acceleration card, and the system comprises: a signal acquisition module having a main sampler anchored at a central sampling point and an offset sampler capable of biased sampling; a monitoring point calibration module for controlling the offset sampler to scan along a voltage axis and a time axis after the link is ready, determining an eye diagram boundary and calculating a monitoring sampling point; an online monitoring module for polling each monitoring point and interrupting when an error code exceeds a limit, and outputting a repair trigger signal; a repair decision module for determining a distortion type according to an abnormal point position and outputting a coding signal; and a parameter adjustment module for selecting a target equalizer according to the coding signal, and performing adaptive iterative adjustment with the monitoring result as feedback until the link returns to normal. The application realizes real-time monitoring, rapid diagnosis and closed-loop self-repair of the PCIe link signal under the premise of uninterrupted service, and significantly improves the stability and reliability of the system.
Owner:SHANGHAI XINLIJI SEMICON CO LTD

Vision transformer neural network acceleration system and method based on cpu and fpga

This invention proposes a visual Transformer neural network acceleration system and method based on CPU and FPGA. The method's implementation steps are as follows: an embedding module constructs a feature matrix; a preprocessing module preprocesses the feature matrix and weight matrix; the preprocessed feature matrix and weight matrix are moved and cached; an acceleration unit constructs a self-attention matrix; the self-attention matrix is ​​cached and moved; a post-processing unit constructs a global feature matrix; and an MLP Head module obtains the classification result. The preprocessing module in the CPU reduces the additional data processing time of the FPGA acceleration unit by preprocessing the feature matrix and weight matrix, effectively improving the neural network's computation speed. Simultaneously, the normalization operation module in the acceleration unit utilizes exponential and logarithmic approximation results during approximation calculations, eliminating floating-point exponentiation and division operations, effectively reducing hardware resource consumption.
Owner:XIDIAN UNIV

Fpga acceleration card power consumption test method, device and electronic equipment

ActiveCN115114098BFaulty hardware testing methodsEnergy efficient computingTest sceneVoltage
The application discloses a kind of FPGA acceleration card power consumption test method, device and electronic equipment, the method includes: based on CPLD receiving the first instruction sent by BMC, first instruction at least includes modification instruction and target voltage value;According to modification instruction and target voltage value, modify the on-board voltage value of FPGA acceleration card;Based on the second instruction sent by BMC, second instruction at least includes save effective instruction;According to second instruction, save current on-board voltage value and make it effective;Based on current on-board voltage value, upgrade the firmware of FPGA acceleration card and execute power consumption stress test;By BMC and the CPLD of FPGA acceleration card are interacted to realize the modification on-board voltage value of FPGA acceleration card, then high-power consumption version of FW corresponding voltage value is burned again to carry out high-power consumption test, and the test scene of FPGA acceleration card is enriched, and the stability of FPGA acceleration card is ensured.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Method and system for virtual-real combination simulation verification of CPU and FPGA

The application provides a CPU and FPGA virtual-real combination simulation verification method and system, comprising: creating a CPU simulation model and loading a CPU-side binary executable file; deploying FPGA code through FPGA acceleration hardware based on the CPU simulation model; dynamically configuring the connection relationship and connection interface of the CPU simulation model and the FPGA hardware, and converting the calling interface of the FPGA code into a network interface transceiver; matching the running speed of the CPU simulation model and the FPGA acceleration hardware, and then running an external test device to obtain a test result. Through software and hardware collaborative adjustment and data caching technology, the application ensures the relative uniformity of the simulation timing, improves the reliability and correctness of the joint verification, and effectively supports the simulation deduction of the system.
Owner:VISION MICROSYST (SHANGHAI) CO LTD

FPGA-based MTLA-Transformer hardware accelerator

PendingCN122154768AInference methodsPhysical realisationFeed forward networkParallel computing
The application discloses an MTLA-Transformer hardware accelerator based on FPGA and relates to the technical field of FPGA acceleration and deep learning inference optimization. The accelerator comprises a controller module, an attention calculation module and a feedforward network module. The controller module is used for completing data scheduling of on-chip cache and off-chip memory and KV cache update management. The attention calculation module comprises a reusable matrix multiplication and addition calculation array, a position coding submodule, a HyperNet calculation submodule, a KV cache management submodule, a Softmax submodule and a residual normalization submodule. Projection calculation, fractional calculation and output projection and other matrix multiplication and addition operation multiplexing are realized through a systolic array. Position coding rotation factor precalculation lookup table, HyperNet position correlation result offline precalculation and Softmax scaling coefficient fusion are adopted to reduce online calculation amount and memory access overhead. The feedforward network module completes two-layer linear transformation and activation operation and realizes residual connection normalization processing. The application can improve incremental inference throughput and reduce hardware resource overhead under the premise of ensuring inference correctness and is suitable for edge end low-power real-time inference scenarios.
Owner:SUN YAT SEN UNIV +1

A low-latency FPGA acceleration method for graph neural network inference

The application relates to the field of artificial intelligence hardware acceleration and wireless communication technology, and particularly discloses a low-delay FPGA acceleration method for graph neural network inference, which comprises the following steps: initializing FPGA hardware resources, loading satellite communication channel data and weight and bias parameters of a graph neural network model from an off-chip memory to an on-chip memory, and initializing each calculation module in a calculation engine; performing matrix operation in the graph neural network by using a parallel calculation engine in the FPGA, wherein the parallel calculation engine comprises a plurality of contraction arrays, each contraction array is composed of a plurality of processing elements, and is used for performing matrix multiplication calculation of a full connection layer in parallel; parallel operation is realized between data processing and data transmission by adopting a double-buffering technology, and a plurality of calculation layers are merged into a calculation group by using a layer fusion technology, so that the storage and transmission of intermediate data are reduced; and the calculated beamforming data is output to the off-chip memory, so that the acceleration task is completed.
Owner:PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV