Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.
17 results about "Fpga acceleration" patented technology
Filter
Efficacy Topic
Property
Owner
Technical Advancement
Application Domain
Technology Topic
Technology Field Word
Patent Country/Region
Patent Type
Patent Status
Application Year
Inventor
The Intel FPGA Acceleration Stack. The Acceleration Stack for Intel Xeon CPU with FPGAs is a robust collection of software, firmware, and tools designed and distributed by Intel to make it easier to develop and deploy Intel FPGAs for workload optimization in the data center.
The invention discloses an internal combustion engine in-cylinder transient self-adaptive exposure control method based on FPGA acceleration, and belongs to the technical field of internal combustion engine endoscopic visualization. In order to solve the technical problem of image overexposure or underexposure caused by violent illumination change in the combustion process in an internal combustion engine cylinder, a real-time exposure detection unit is fused with a camera built-in sensor and an endoscope added sensor to monitor light intensity data; the adaptive exposurecontrol unit accelerated by the FPGA adopts adaptive PID control and fuzzy logic control to generate a parameter adjustment direction, the adaptive PID dynamically adjusts gain based on a machine learning model, and the fuzzy logic outputs a decision according to a fuzzy rule base of brightness, light intensity and errors; the parameter adjusting unit adjusts the aperture size, the shutter speed and the light sensitivity in a linkage mode according to the target and actual gray scale difference value, and millisecond-level multi-parameter exposure control is achieved. Overexposure of a highlight area and loss of details of a low-light area are effectively avoided, and stable and clear images are obtained in the whole combustion process.
A method and device for interpreting a graph neural network based on FPGA acceleration propose to use FPGA hardware to accelerate interpretation process of the graph neural network oriented to node classification in parallel, and improve node traversal and shortest path search of BFS, thereby optimizing requirements of algorithm calculation and storage, and accelerating generation of interpretation results. During calculating HN values, the present disclosure optimizes multiplication operation using the matrix characteristics, transforms the dense matrix multiplication into sparse-dense matrix multiplication, and optimizes the resource occupation using multi-PE parallel processing, greatly improving performance of graph neural network interpretation acceleration. Moreover, an overall architecture based on FIFO storage calculation task distribution is designed to reduce calculation difference between nodes, solving the difficulty of high time complexity of the graph neural network interpretation method based on node classification in actual data applications and improving the time efficiency of interpretation.
The application discloses a quantum computing data processing method and device and a storage medium, comprising: collecting real-time data of an application scene by a terminal device, and transmitting the real-time data to an edge processing node cluster; using an FPGA acceleration module integrated in the edge processing node cluster to perform a processing operation on the real-time data to generate a quantum data packet, wherein the processing operation comprises preprocessing, quantumization coding and data routing, and the edge processing node cluster sends the quantum data packet and a corresponding task request to a quantum computing center; constructing a quantum service blockchain network for analyzing and identifying the task request, and dynamically calling, combining and deploying the corresponding quantum service; constructing a quantum computing model, using the quantum service to realize optimized calculation on the task request; and transmitting the optimized calculation result generated by the quantum computing center to the edge processing node cluster through a quantum encryptiontransmission channel, and driving the terminal device to execute by the edge processing node cluster.
The application provides a PCIe link monitoring and self-repair system for an FPGA acceleration card, and the system comprises: a signal acquisition module having a main sampler anchored at a central sampling point and an offset sampler capable of biased sampling; a monitoring point calibration module for controlling the offset sampler to scan along a voltage axis and a time axis after the link is ready, determining an eye diagram boundary and calculating a monitoring sampling point; an online monitoring module for polling each monitoring point and interrupting when an error code exceeds a limit, and outputting a repair trigger signal; a repair decision module for determining a distortion type according to an abnormal point position and outputting a coding signal; and a parameter adjustment module for selecting a target equalizer according to the coding signal, and performing adaptive iterative adjustment with the monitoring result as feedback until the link returns to normal. The application realizes real-time monitoring, rapid diagnosis and closed-loop self-repair of the PCIe link signal under the premise of uninterrupted service, and significantly improves the stability and reliability of the system.
This invention proposes a visual Transformer neural network acceleration system and method based on CPU and FPGA. The method's implementation steps are as follows: an embedding module constructs a feature matrix; a preprocessing module preprocesses the feature matrix and weight matrix; the preprocessed feature matrix and weight matrix are moved and cached; an acceleration unit constructs a self-attention matrix; the self-attention matrix is cached and moved; a post-processing unit constructs a global feature matrix; and an MLP Head module obtains the classification result. The preprocessing module in the CPU reduces the additional data processing time of the FPGA acceleration unit by preprocessing the feature matrix and weight matrix, effectively improving the neural network's computation speed. Simultaneously, the normalization operation module in the acceleration unit utilizes exponential and logarithmic approximation results during approximation calculations, eliminating floating-point exponentiation and division operations, effectively reducing hardware resource consumption.
The application discloses a kind of FPGA acceleration card power consumption test method, device and electronic equipment, the method includes: based on CPLD receiving the first instruction sent by BMC, first instruction at least includes modification instruction and target voltage value;According to modification instruction and target voltage value, modify the on-board voltage value of FPGA acceleration card;Based on the second instruction sent by BMC, second instruction at least includes save effective instruction;According to second instruction, save current on-board voltage value and make it effective;Based on current on-board voltage value, upgrade the firmware of FPGA acceleration card and execute power consumption stress test;By BMC and the CPLD of FPGA acceleration card are interacted to realize the modification on-board voltage value of FPGA acceleration card, then high-power consumption version of FW corresponding voltage value is burned again to carry out high-power consumption test, and the test scene of FPGA acceleration card is enriched, and the stability of FPGA acceleration card is ensured.
The application provides a CPU and FPGA virtual-real combination simulationverification method and system, comprising: creating a CPU simulation model and loading a CPU-side binary executable file; deploying FPGA code through FPGA acceleration hardware based on the CPU simulation model; dynamically configuring the connection relationship and connection interface of the CPU simulation model and the FPGA hardware, and converting the calling interface of the FPGA code into a network interfacetransceiver; matching the running speed of the CPU simulation model and the FPGA acceleration hardware, and then running an external test device to obtain a test result. Through software and hardware collaborative adjustment and data caching technology, the application ensures the relative uniformity of the simulation timing, improves the reliability and correctness of the joint verification, and effectively supports the simulation deduction of the system.
The application discloses an MTLA-Transformer hardware accelerator based on FPGA and relates to the technical field of FPGA acceleration and deep learninginference optimization. The accelerator comprises a controller module, an attention calculation module and a feedforward network module. The controller module is used for completing data scheduling of on-chip cache and off-chip memory and KV cache update management. The attention calculation module comprises a reusable matrix multiplication and addition calculation array, a position coding submodule, a HyperNet calculation submodule, a KV cache management submodule, a Softmax submodule and a residual normalization submodule. Projection calculation, fractional calculation and output projection and other matrix multiplication and addition operation multiplexing are realized through a systolic array. Position coding rotation factor precalculation lookup table, HyperNet position correlation result offline precalculation and Softmax scaling coefficient fusion are adopted to reduce online calculation amount and memory access overhead. The feedforward network module completes two-layer linear transformation and activation operation and realizes residual connection normalization processing. The application can improve incremental inferencethroughput and reduce hardware resource overhead under the premise of ensuring inferencecorrectness and is suitable for edge end low-power real-time inference scenarios.
The application relates to the computer technical field and discloses a memory system, a control method and a server, the system comprising: a CPU and an FPGA acceleration card, the CPU and the FPGA acceleration card both supporting a CXL protocol; the CPU and the FPGA acceleration card being connected through a PCIe interface supporting the CXL protocol; the CPU being configured with a plurality of local memory strips through a memory strip slot; the FPGA acceleration card being configured with a plurality of extended memory strips through a memory connection component or a memory connector; and the FPGA acceleration card being used for receiving memory data sent by the CPU, and performing logical calculation on the memory data according to the attribute characteristics of the memory data based on the extended memory strips to obtain target memory data. The system provided by the above scheme realizes expansion of the memory capacity of the server by additionally arranging the FPGA acceleration card, reduces the difficulty of expansion of the memory capacity of the server, and improves the overall performance and response speed of the server.
The invention discloses an FPGA (Field Programmable Gate Array) accelerated intelligent matching technology based on dynamic weight of an AI (Artificial Intelligence) large model. The method comprises the following steps: receiving user demand data; calling an AI large model to carry out deep semantic understanding and feature extraction; carrying out weighted calculation on the feature vectors by adopting a dynamic weight ratio to obtain a matching degree; and outputting a matching result. Wherein the dynamic weight factor is generated in real time by the AI large model according to the context of the current scene. According to the invention, by introducing a dynamic weight mechanism and logic auditing, accurate, efficient and credible intelligent matching is realized, and the technical problems of information asymmetry, low matching precision and high fraud risk in traditional intermediary services are effectively solved.
The application relates to the field of artificial intelligencehardware acceleration and wireless communication technology, and particularly discloses a low-delayFPGA acceleration method for graph neural network inference, which comprises the following steps: initializing FPGA hardware resources, loading satellitecommunication channel data and weight and bias parameters of a graph neural network model from an off-chip memory to an on-chip memory, and initializing each calculation module in a calculation engine; performing matrix operation in the graph neural network by using a parallel calculation engine in the FPGA, wherein the parallel calculation engine comprises a plurality of contraction arrays, each contraction array is composed of a plurality of processing elements, and is used for performing matrix multiplication calculation of a full connection layer in parallel; parallel operation is realized between data processing and data transmission by adopting a double-buffering technology, and a plurality of calculation layers are merged into a calculation group by using a layer fusion technology, so that the storage and transmission of intermediate data are reduced; and the calculated beamforming data is output to the off-chip memory, so that the acceleration task is completed.
The technical scheme of the invention discloses a cache folding grid FPGA accelerationsystem and a cache folding grid FPGA acceleration method. According to the method, under the condition that on-chip storage resources are limited, through a data organization and access mechanism in which on-chip storage and off-chip storage are coordinated, extensible and efficient solving of the problem of minimum cut and maximum flow of a large-scale grid diagram is achieved; the method is characterized in that a label mapping cache mechanism is introduced aiming at an existing implementation mode of loading grid graph data into an on-chip memory of an FPGA (Field Programmable Gate Array) at one time; according to the method, a boundary buffer area is arranged in a boundary processor and used for temporarily storing key states related to cross-block boundaries and updating information to be submitted;