The application relates to the field of
artificial intelligence hardware acceleration and
wireless communication technology, and particularly discloses a low-
delay FPGA acceleration method for graph neural network
inference, which comprises the following steps: initializing FPGA hardware resources, loading
satellite communication channel data and weight and bias parameters of a graph neural
network model from an off-
chip memory to an on-
chip memory, and initializing each calculation module in a calculation engine; performing matrix operation in the graph neural network by using a parallel calculation engine in the FPGA, wherein the parallel calculation engine comprises a plurality of contraction arrays, each contraction array is composed of a plurality of
processing elements, and is used for performing
matrix multiplication calculation of a full connection layer in parallel; parallel operation is realized between
data processing and
data transmission by adopting a double-buffering technology, and a plurality of calculation
layers are merged into a calculation group by using a layer fusion technology, so that the storage and transmission of intermediate data are reduced; and the calculated
beamforming data is output to the off-
chip memory, so that the acceleration task is completed.