FPGA-based digital holographic particle field real-time segmentation method and system

CN122820744APending Publication Date: 2026-09-25XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610986242.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0010]本发明的目的在于解决现有数字全息粒子场分割技术计算复杂度高、边缘设备无法实时处理、能效比低、硬件适配性差的技术问题,提供一种基于FPGA的数字全息粒子场实时分割方法及系统,通过轻量化网络设计与FPGA硬件架构、映射策略协同优化,在保证分割精度的前提下,大幅提升粒子场分割推理速度与硬件能效

Benefits of technology

[0034]1、本发明面向数字全息粒子场分割专属任务特性定制LAFNet-s轻量化网络,从算法源头削减浮点运算开销,解决通用分割网络参数冗余、噪声鲁棒性弱、粒子尺度适配性差的问题,通过深度可分离卷积、大卷积核拆分、分段线性激活、组合式上采样多层硬件友好改造,在保留粒子高精度特征提取能力的前提下,大幅降低网络推理计算开销,适配边缘FPGA资源约束。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820744A_ABST
    Figure CN122820744A_ABST
Patent Text Reader

Abstract

The application discloses a kind of FPGA-based digital holographic particle field real-time segmentation method and system, belongs to holographic imaging and AI hardware acceleration technical field.Digital holographic particle field data is acquired and preprocessed, and inference is input into lightweight segmentation network LAFNet-s;Network inference is executed using FPGA hardware accelerator.Accelerator uses pulsating array as core computing unit, and is configured with three levels of on-chip storage hierarchy and data post-processing module.For deep convolution, point-by-point convolution and standard three-dimensional convolution, a differentiated hardware mapping strategy is designed.Through algorithm-hardware collaborative design, the segmentation accuracy is maintained, and the inference speed and energy efficiency ratio are improved.The Dice coefficient retention rate of FPGA fixed-point inference and PyTorch floating-point inference is 96.21%, the inference speed can reach 33.6ms / frame, the power consumption is 3.5W, and the energy efficiency ratio is 8.51FPS / W.The problems of high computational complexity and difficulty in real-time processing on edge devices in traditional particle field segmentation are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of digital holographic imaging, embedded systems and artificial intelligence hardware acceleration technology, and specifically relates to a real-time segmentation method and system for digital holographic particle fields based on FPGA. Background Technology

[0002] Digital holographic particle field detection technology is an important research direction at the intersection of fluid mechanics and optical imaging, and is widely used in scientific and engineering fields such as environmental monitoring, spray combustion diagnosis, and industrial online detection. This technology records particle field holograms and reconstructs the three-dimensional distribution of particles using numerical reconstruction methods, offering unique advantages such as non-contact operation, high precision, and full-field three-dimensional measurement. However, with the increasing complexity of detection scenarios and the rising demands for real-time performance, traditional physical model-based numerical reconstruction methods face problems such as low computational efficiency and weak noise resistance, making it difficult to meet the real-time processing requirements of high frame rates and high-resolution particle fields.

[0003] In recent years, deep learning methods have driven significant progress in the field of image segmentation. Convolutional neural networks (CNNs), with their superior feature extraction and representation capabilities, can directly learn particle features from holograms, achieving end-to-end particle recognition and segmentation. Research shows that deep learning-based holographic processing methods outperform traditional methods in both reconstruction accuracy and processing efficiency. However, deep learning models designed for particle field recognition tasks typically have high computational complexity, with the number of parameters and computational cost increasing dramatically with network depth. In resource-constrained edge computing scenarios, how to efficiently deploy high-precision neural network models has become a key bottleneck restricting the practical application of this technology.

[0004] Field-Programmable Gate Arrays (FPGAs), with their reconfigurable characteristics, have become an ideal platform for edge deployment of neural networks. Compared with GPUs, FPGAs offer higher energy efficiency in power-constrained scenarios; compared with Application-Specific Integrated Circuits (ASICs), FPGAs offer greater flexibility and shorter development cycles. In recent years, scholars both domestically and internationally have conducted extensive research on the design of FPGA accelerators for neural networks. In terms of hardware architecture optimization, researchers have advanced the development of FPGA neural network accelerators from dimensions such as systolic arrays, multiple cache levels, and dynamic reconfiguration; in terms of algorithm mapping strategies, researchers have explored algorithm mapping methods such as Winograd transform, octave convolution, and sparse computation; and in terms of data flow optimization, researchers have proposed data flow organization methods such as fixed row flow, fixed weight flow, and fixed output flow.

[0005] However, existing FPGA accelerators are mostly geared towards 2D image classification or object detection tasks, with limited research on customized acceleration for 3D particle field segmentation networks. There is also a lack of systematic discussion on hardware mapping strategies for lightweight operators such as skip connections, upsampling operations, and depthwise separable convolutions. Specifically, existing research has the following shortcomings:

[0006] First, in terms of segmentation networks, existing deep learning methods mostly focus on general image segmentation tasks, with few network designs specifically designed for the characteristics of holographic particle field images, such as complex noise, sparse particles, and significant scale variations. The ability to extract multi-scale particle structure features needs to be improved.

[0007] Second, regarding network deployment, existing segmentation networks generally have a large number of parameters and computational loads, while the on-chip resources of edge devices such as FPGAs are limited, making it difficult to directly deploy complex networks onto hardware platforms for real-time inference. Therefore, it is necessary to perform lightweight modifications to the network to adapt it to the resource constraints of FPGAs while maintaining accuracy.

[0008] Third, in terms of hardware acceleration, existing FPGA accelerators are mostly geared towards 2D image classification or object detection tasks. There is a lack of research on customized acceleration for 3D particle field segmentation networks (especially encoder-decoder structures), and there is a lack of systematic discussion on hardware mapping strategies for lightweight operators such as skip connections, upsampling operations, and depth-separable convolution.

[0009] In summary, there is an urgent need for a lightweight network design and its FPGA hardware acceleration scheme for digital holographic particle field segmentation tasks, in order to solve the speed and power consumption bottlenecks in real-time particle field monitoring. Summary of the Invention

[0010] The purpose of this invention is to solve the technical problems of high computational complexity, inability of edge devices to process in real time, low energy efficiency, and poor hardware adaptability of existing digital holographic particle field segmentation technology. It provides a real-time digital holographic particle field segmentation method and system based on FPGA. Through lightweight network design and FPGA hardware architecture and mapping strategy co-optimization, the particle field segmentation inference speed and hardware energy efficiency are greatly improved while ensuring segmentation accuracy.

[0011] To achieve the above objectives, the present invention adopts the following technical solution:

[0012] A real-time segmentation method for digital holographic particle fields based on FPGA is divided into two main stages: offline model solidification and online real-time inference. The online inference stage includes data block preprocessing, lightweight network inference, FPGA hardware pipeline acceleration, and parallel computation steps of multi-convolutional differential mapping. In the offline stage, the lightweight segmentation network LAFNet-s is pre-performed with BN layer fusion and symmetric linear fixed-point quantization to solidify the model, and the solidified network is then used for online inference.

[0013] This invention designs a lightweight segmentation network, LAFNet-s, with an encoder-decoder architecture as its core. It features comprehensive lightweight modifications and hardware adaptation optimizations tailored to FPGA hardware resources and the computational characteristics of 3D particle fields. Simultaneously, it constructs an FPGA accelerator architecture that includes a pulsating array computing core, three-level reconfigurable on-chip storage, a double-buffered pipeline mechanism, a multi-mode post-processing unit, and state machine instruction scheduling. Dedicated hardware mapping and data flow organization methods are designed for standard 3D convolution, depthwise convolution, and pointwise convolution, respectively, solving the problems of low computational utilization, high memory access latency, and inability to adapt to 3D particle field segmentation networks caused by traditional single mapping strategies.

[0014] The lightweight segmentation network LAFNet-s design includes: replacing standard 3D convolution with depthwise separable convolution, splitting spatial convolution and channel convolution, significantly reducing the number of network parameters and multiplication / accumulation computation; effectively decomposing large receptive field, large-size convolutional kernels into a multi-level cascaded structure of small convolutional kernels, compressing parameter redundancy while maintaining the same receptive field and multi-scale feature extraction capabilities; replacing the exponential sigmoid activation of the SE attention module with piecewise linear Hardsigmoid activation, avoiding exponential and division operations, and implementing it in hardware only through comparators, adders, and multipliers; and replacing high-latency transposed convolution upsampling with a combined upsampling structure of nearest neighbor interpolation and convolutional smoothing, significantly reducing the computational overhead of the decoder upsampling stage.

[0015] The FPGA hardware accelerator uses a fixed-size systolic array as its core computing unit, coupled with a reconfigurable three-level on-chip memory architecture and a double-buffered memory access hiding mechanism, to achieve full pipelined parallelism for data loading, convolution computation, post-processing, and result write-back. To address the independent computational characteristics of the three types of convolution operators, differentiated mapping strategies are adopted, including channel isolation, column broadcasting, and subarray depth pipelines, respectively. This ensures that different types of convolution can maximize the utilization of array parallel resources, improving hardware computing power utilization and inference throughput.

[0016] The offline BN layer fusion processing unifies and merges all convolutional BN layer parameters into the convolutional layer weights and biases, eliminating the need for separate BN normalization operations during the inference phase; symmetric linear fixed-point quantization relies on the training calibration set to statistically analyze the numerical range of each layer, compressing 32-bit floating-point parameters and activation values ​​into low-bit-width fixed-point data, adapting to the hardware characteristics of FPGA fixed-point arithmetic.

[0017] Furthermore, the symmetric linear fixed-point quantization can be any bit width of 4bit, 8bit, 12bit, or 16bit. Different bit widths only adjust the quantization scaling factor and the value range, and the overall network structure and hardware mapping logic do not need to be changed. As a preferred implementation, 8bit symmetric linear quantization is used, and the quantization value range is [-128, 127].

[0018] Furthermore, the systolic array can be configured with various specifications such as 8×8, 16×16, 24×24, and 32×32. The array size only changes the number of parallel computing channels and spatial positions in a single operation, while the three types of convolutional differential mapping logic remain unchanged. As a preferred embodiment, the systolic array is selected with a 16×16 specification.

[0019] Furthermore, the finite state machine can be configured with a 6-state, 9-state, 12-state, or 16-state architecture, and the reduced instruction set bit width can be configured to 64-bit, 128-bit, or 256-bit. Only the state transition logic and instruction field length need to be adjusted, and the overall pipeline scheduling framework does not need to be reconstructed. As a preferred implementation, a 9-state finite state machine is used, equipped with a 128-bit reduced instruction set.

[0020] Furthermore, the weighted cache adopts a hierarchical partitioned storage architecture, consisting of several independent banks, which can be configured with 8 banks, 16 banks, 24 banks, or 64 banks; as a preferred embodiment, 16 independent banks are set up with a total capacity of 32KB.

[0021] Furthermore, the combined upsampling module supports custom three-dimensional upsampling ratios, which can be selected as (2,2,2), (1,2,2), (3,3,1), or (2,2,1). Only the address generation unit mapping parameters need to be modified, and the pixel copying and convolution smoothing hardware logic does not need to be reconstructed. As a preferred embodiment, the upsampling ratio is fixed at (2,2,1).

[0022] Furthermore, the activation unit reuses hardware computing resources and achieves multiple working mode switching through two control signals S1 and S0: when S1S0=00, LeakyReLU activation is performed and the negative half-axis slope coefficient is fixed at 0.01; when S1S0=01, Hardsigmoid activation is performed; when S1S0 is 10 or 11, there is no activation processing and the original feature data is directly output.

[0023] A real-time segmentation system for digital holographic particle fields based on FPGA is used to execute the aforementioned real-time segmentation method. Relying on an algorithm-hardware co-architecture, it realizes fully automatic real-time processing of holographic particle field data from acquisition and preprocessing, feature inference, hardware-accelerated computation to result output. The system includes a data preprocessing module, a lightweight inference module, an FPGA hardware acceleration module, and a result output module.

[0024] The data preprocessing module is used to acquire the original image data of the digital holographic particle field and complete the normalization and block normalization preprocessing.

[0025] The lightweight inference module is used to carry the trained and quantized LAFNet-s lightweight segmentation network to realize multi-scale feature extraction and segmentation inference of particle fields; and to send the convolution operation tasks and feature data of each layer to the FPGA hardware acceleration module.

[0026] The FPGA hardware acceleration module is equipped with a configurable systolic array, three-level reconfigurable on-chip memory, a double-buffered pipeline unit, a multi-mode post-processing unit, and a multi-state instruction scheduling and control unit. It relies on three types of convolutional differential mapping strategies to complete the hardware parallel acceleration of all convolutional operators in the network.

[0027] The result output module is used to receive the segmentation feature map output by the FPGA hardware acceleration module, complete data caching, feature splicing verification, output particle field segmentation results in real time, and complete data caching and off-chip write-back.

[0028] Furthermore, the data preprocessing module has built-in image normalization and size cropping subunits, which can adaptively adjust the size of the slices according to the resolution of the holographic acquisition device to adapt to the input of original holographic images of different specifications.

[0029] Furthermore, the lightweight inference module has a built-in network parameter cache that can simultaneously store multiple sets of LAFNet-s model parameters with different quantization bit widths, supporting rapid on-site switching between accuracy and power consumption levels.

[0030] Furthermore, the FPGA hardware acceleration module is configured with multiple independent weight cache partitions. The number of partitions can be flexibly increased or decreased according to the network depth. Parallel read and write operations between partitions do not block each other, further reducing weight read latency.

[0031] Furthermore, the result output module has an external storage interface and a communication interface with the host computer, which can store the segmentation results locally, or upload detection data such as particle segmentation contours and particle counts to the industrial control terminal in real time.

[0032] This invention can be used for online real-time inference of a lightweight segmentation network LAFNet-s that has been pre-fused with BN layers and solidified with symmetric linear fixed-point quantization.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] 1. This invention is a customized LAFNet-s lightweight network designed for the specific task characteristics of digital holographic particle field segmentation. It reduces floating-point operation overhead from the source of the algorithm and solves the problems of parameter redundancy, weak noise robustness, and poor particle scale adaptability of general segmentation networks. Through depthwise separable convolution, large convolution kernel splitting, piecewise linear activation, and combined upsampling multi-layer hardware-friendly modification, it significantly reduces the network inference computation overhead while retaining the ability to extract high-precision particle features, and adapts to edge FPGA resource constraints.

[0035] 2. This invention integrates BN layer fusion and general symmetric linear fixed-point quantization in the offline stage, eliminates redundant BN normalization operations in the inference stage, and constructs a dedicated FPGA accelerator architecture adapted to encoder-decoder segmentation networks. Through three-level reconfigurable storage, skip-connection feature caching, double-buffered pipeline mechanism, and instruction-based state machine scheduling, it achieves full-link hardware acceleration for multiple types of convolution, upsampling, pooling, activation, and feature concatenation, thus overcoming the technical deficiency of existing FPGA accelerators that are only adapted to two-dimensional classification and detection networks and cannot efficiently support three-dimensional particle field segmentation networks.

[0036] 3. This invention configures a pulsating array hardware architecture and designs differentiated hardware mapping and data flow organization strategies for three types of operators: depthwise convolution, pointwise convolution, and standard 3D convolution. This completely solves the problems of poor adaptability of a single mapping method to different convolution operators, low computing power utilization, and high inference latency, thereby maximizing the parallel computing capabilities of the pulsating array and significantly improving hardware inference throughput.

[0037] 4. This invention achieves lightweight model deployment through BN layer fusion and 8-bit fixed-point quantization, and combines the advantages of FPGA low-power parallel computing to achieve a balance between segmentation accuracy, inference speed, and overall system energy efficiency. In actual testing with optimized 8-bit quantization and a 16×16 systolic array configuration, the floating-point precision retention rate of the Dice coefficients reaches 96.21%. With a standard input size of 64×64×20 3D segmented images, the single-frame inference time is 33.6ms, power consumption is 3.5W, and the energy efficiency ratio is 8.51FPS / W. All indicators fully meet the real-time detection requirements of edge scenes. Attached Figure Description

[0038] Figure 1 This is a block diagram of the overall system architecture of the present invention.

[0039] Figure 2 This is a block diagram of the core architecture of the accelerator of this invention.

[0040] Figure 3This is a schematic diagram comparing the timing of the double buffering mechanism in this invention. (a) shows the timing of the traditional single-buffered (One Buffer) mechanism; (b) shows the timing of the Ping-pong double buffering mechanism in this invention.

[0041] Figure 4 This is a schematic diagram of the data flow of the depth convolution hardware mapping strategy in this invention.

[0042] Figure 5 This is a schematic diagram of the data flow for the pointwise convolution hardware mapping strategy in this invention.

[0043] Figure 6 This is a schematic diagram of the data flow of the standard 3D convolution hardware mapping strategy in this invention.

[0044] Figure 7 This is a schematic diagram of the spatial mapping relationship of the upsampling units in this invention.

[0045] Figure 8 This is a circuit diagram of the multi-mode activation function unit in this invention. Detailed Implementation

[0046] The embodiments of the present invention will now be described in detail and completely with reference to the accompanying drawings. These embodiments are implemented based on the technical solution of the present invention, providing detailed implementation methods and specific operation processes. However, the scope of protection of the present invention is not limited to the following embodiments. Hardware and model parameters such as quantization bit width, systolic array specifications, number of finite state machine states, reduced instruction bit width, number of weighted cache banks, and 3D upsampling rate can be flexibly replaced according to FPGA resources, real-time requirements, and accuracy requirements.

[0047] I. Design of the lightweight segmentation network LAFNet-s

[0048] To address the challenges of complex noise, sparse particles, and significant scale variations in holographic particle field images, this invention designs a lightweight segmentation network, LAFNet-s, with large receptive field feature extraction and channel attention enhancement as its core. This network uses U-Net as its backbone and employs an encoder-decoder architecture.

[0049] To meet the deployment requirements of FPGA hardware, the network structure is optimized as follows:

[0050] (a) Depthwise separable convolution replacement

[0051] The standard 3D convolution is decomposed into depthwise convolution and pointwise convolution to reduce the number of network parameters and computational cost. A convolution kernel size is... The number of input channels is The number of output channels is Taking a standard 3D convolutional layer as an example (where , , These are the width, height, and depth of the output feature map of this layer, respectively, and their parameter values ​​are... computational load However, by replacing it with depthwise separable convolution, the number of parameters is reduced. computational load The ratio of the number of parameters to the computational cost of both, relative to standard convolution, is [missing information]. ,for Convolution, with approximately 1 / 3 the number of parameters of standard convolution. Compared to standard 3D convolutions of the same size, depthwise separable convolutions can significantly reduce the number of network parameters and the computational cost of multiplication and accumulation, effectively reducing the complexity of model inference and adapting to the fixed-point parallel computing characteristics of FPGAs.

[0052] (ii) Equivalent decomposition of large kernel convolution

[0053] To balance the feature receptive field with hardware computational overhead, the original The convolution kernel is equivalent to three cascaded kernels. Small convolutional kernel. Single. The number of convolution parameters is Three The number of convolution parameters is approximately The number of parameters is reduced to approximately [amount] of large kernel convolution. While retaining the global receptive field of the original large kernel convolution and ensuring the ability to extract features of small particles at multiple scales, the number of parameters and computation of convolutional layers are greatly reduced, avoiding the problems of high latency and large resource consumption in the hardware implementation of large kernel convolution.

[0054] (III) Hardware-friendly modification of activation function

[0055] Replace the standard Sigmoid activation function in the SE module with the Hardsigmoid function based on a piecewise linear approximation, whose expression is: This function avoids exponential and division operations, and can be implemented in an FPGA using only adders, comparators, and multipliers, significantly reducing hardware implementation complexity.

[0056] (iv) Upsampling method optimization

[0057] Replace the transposed convolution with nearest neighbor upsampling and Combinations of convolutions. Nearest neighbor upsampling achieves scaling solely through spatial copying, without multiplication operations, resulting in a simple and fully pipelining memory access pattern; subsequently... Convolution smooths and compensates for details of the magnified features, ensuring the accuracy of feature map resolution restoration while controlling hardware overhead, and balancing inference speed and segmentation accuracy.

[0058] (V) BN Layer Fusion

[0059] During training, each convolutional layer is followed by a batch normalization (BN) layer to normalize the feature distribution. During inference, the BN layer incurs additional computational overhead. This invention integrates all parameters of the BN layer into the weights and biases of the preceding convolutional layers during the offline phase, merging the computational logic of the two layers through unified scaling and offset conversion. After fusion, the inference process no longer requires separate BN normalization operations, eliminating the independent calculation steps of the BN layer, reducing intermediate feature data read / write operations, and lowering the computational power consumption of FPGA real-time inference. A minimal stability constant is set during the calculation to avoid the denominator being zero when taking the square root, ensuring stable and error-free fusion computation.

[0060] (vi) Symmetric linear fixed-point quantization

[0061] Symmetric linear quantization is employed, and representative images from the training set are selected as the calibration set to statistically analyze the weights and activation value ranges of each layer. The quantization scaling factor and zero points are calculated, mapping 32-bit floating-point data to low-bit-width fixed-point integers. 8-bit quantization is preferred, with a value range of [-128, 127]. Besides 8-bit, 4-bit, 12-bit, and 16-bit symmetric linear quantization can also be used: 4-bit quantization minimizes storage and computational overhead, suitable for low-power, minimalist edge devices; 12-bit quantization balances accuracy and hardware resources, suitable for high-precision particle measurement scenarios; and 16-bit quantization suffers from extremely low accuracy loss, suitable for ultra-high resolution and weak, small particle recognition scenarios. Different bit widths only change the quantization scaling factor and value range; the BN fusion and convolutional differential hardware mapping architectures do not require modification and can all be adapted to the FPGA accelerator of this invention. After solidification, the quantization weights and biases are stored in the FPGA's off-chip DDR for online inference loading.

[0062] II. FPGA Hardware Accelerator Architecture Design

[0063] This invention addresses the structural characteristics of the LAFNet-s network, including its multi-type convolutions, multi-level skip connections, and multi-mode post-processing. It designs a dedicated FPGA accelerator architecture for the LAFNet-s network. This architecture consists of a convolution computation module, a differential mapping scheduling module, a three-level reconfigurable on-chip storage module, a multi-mode data post-processing module, and a global control module. These modules work together to achieve hardware acceleration of all network layer operators and pipelined parallel inference, as detailed below:

[0064] (a) Convolution Calculation Module

[0065] The convolution calculation module is the core of the accelerator of this invention, and it uses a configurable fixed-size pulsating array as the core matrix calculation unit.

[0066] The choice of systolic array size involves a trade-off between computational parallelism, resource overhead, and timing convergence. While a 32×32 array offers higher parallel computing capabilities, large-scale arrays significantly increase the risk of wiring congestion, and the increased data broadcasting and partial transfer paths between PE units can easily lead to timing violations. An 8×8 array, while having lower resource consumption, suffers from insufficient parallelism, requiring more computation cycles when processing LAFNet-s networks, which is detrimental to inference latency optimization. This invention preferably uses a 16×16 size, achieving a good balance between parallelism and resource overhead. Besides 16×16, systolic arrays can be configured with 8×8, 24×24, 32×32, and other specifications. Adjusting the array size only changes the number of channels and spatial locations for single-run parallel computations; the logic of the three differentiated mapping strategies—channel isolation, column broadcasting, and subarray depth pipeline—remains unchanged, and all can adapt to parallel acceleration of depthwise convolution, pointwise convolution, and standard 3D convolution.

[0067] (ii) Differentiated hardware mapping strategy

[0068] For different convolution types, this invention employs differentiated hardware mapping strategies.

[0069] 1. Depth Convolution Mapping Strategy

[0070] Depthwise convolution has the characteristic of independent channel computation, where each input channel corresponds to an independent convolution kernel for processing, and the output channels correspond one-to-one with the input channels. This invention adopts a channel isolation and spatial parallel mapping strategy, dedicating each row of the systolic array to process an independent input channel, while the PEs within the row process the spatial position of that channel in parallel.

[0071] Specifically, each batch is processed once. There are input channels, and the input features of each channel are divided into segments of size . The feature blocks, where The convolution kernel spatial size, This refers to the number of spatial locations processed in a single batch. The convolution stride is [stagger value]. Weights are statically loaded into the row registers before computation begins and remain unchanged throughout the channel computation. Feature data is streamed in a sliding window manner, with data multiplexing achieved through in-row systolic propagation. In each clock cycle, the PE (Programmer, Executor, and Processor) within each row performs operations synchronously, after [processing steps]. After one cycle of pipelined computation, each PE generates a complete output pixel result. After a single batch of computation is completed, each row of PEs generates... The results are collected simultaneously and written back to the output buffer.

[0072] 2. Pointwise convolution mapping strategy

[0073] The mathematical essence of pointwise convolution is a fully connected linear transformation across channels, where the dimension is [dimensionality missing] at each spatial location. The input feature vector and dimension are Multiply the weight matrices.

[0074] This invention employs an input feature column broadcasting and weight pulsating loading mechanism. The input feature vector is loaded from the left side of the array, and each feature element is vertically distributed to all PEs in the corresponding column via an intra-column broadcast bus. The weight matrix is ​​divided into row blocks, input from the top of the array, and pulsatingly propagated vertically. After each PE completes multiplication and addition operations, the partial sums are accumulated within the row. For the number of input channels... If the number of columns is less than the number of physical columns in the PE array, the system shuts down the column clocks for the corresponding redundant input channels and logically isolates them from the broadcast network.

[0075] 3. Standard 3D Convolution Mapping Strategy

[0076] Standard 3D convolution computation involves dense multiplication and accumulation operations of the input feature map in 3D space and channel dimensions. This invention statically divides the systolic array into several independent 3×3 subarrays, each subarray responsible for the complete convolution computation of one output channel at a spatial location. In a preferred embodiment, a 16×16 systolic array is divided into 16 independent 3×3 subarrays.

[0077] The computation process unfolds sequentially according to depth slices, with feature data supplied in the form of a three-level depth slice buffer. Within each subarray, the nine PEs perform nine parallel multiply-accumulate operations per cycle. The computation is pipelined according to depth slices 0, 1, and 2, with intermediate results accumulated across time steps. After completing the depth pipeline for one input channel, its partial sum is written back to the buffer, and accumulation continues when switching to the next input channel. The final output is obtained after the loop.

[0078] (iii) On-chip storage module

[0079] The on-chip storage module adopts a three-level storage hierarchy and a multi-cache collaborative organization method to form a three-level data path from external storage, on-chip cache to computing unit.

[0080] 1. Input feature caching

[0081] A double-buffered structure is adopted, and a reconfigurable data organization pattern is designed for three convolution modes. In the standard convolution mode, the buffer is divided into 16 independent feature window storage volumes, each storing a complete feature window. The feature window arranges data according to a depth-first principle. In depthwise convolution mode, a channel-isolated storage architecture is used, allocating an independent storage block for each input channel. Each channel region block stores the feature values ​​of that channel at all spatial locations within the entire feature map or the current processing block. In pointwise convolution mode, a column-based storage architecture is used, dividing the storage volume according to the input channel dimension. The reconstructed feature vectors are distributed to the computation array through an intra-column broadcast network.

[0082] 2. Weight caching

[0083] A hierarchical grouping architecture is adopted, dividing the system into standard convolutional regions, depthwise convolutional regions, and pointwise convolutional regions based on the convolution type. Physically, it consists of several independent banks. In the preferred embodiment, 16 banks are set up, with a total capacity of 32KB and a capacity of 2KB per bank, all of which can be independently addressed and accessed. The number of banks and the total capacity can be adjusted according to the total network weights, such as 8 banks / 16KB, 24 banks / 48KB, and 64 banks / 128KB, while maintaining the hierarchical grouping and partitioning storage architecture, still enabling independent high-speed retrieval of the three types of convolutional weight partitions.

[0084] 3. Output Feature Buffer

[0085] The output feature cache employs a design combining cache line organization and a dual-buffering architecture. Each cache line has a fixed byte length and is used to store a complete, multi-channel output feature block, the size of which is aligned with the system bus width and off-chip transmission granularity. Two independent hardware paths are constructed on the output side: a feedforward path pushes skip features generated in the encoder path to the feature concatenation cache; a write-back path writes output features that need to be stored long-term or used for subsequent calculations back to off-chip DDR memory via a DMA controller.

[0086] 4. Feature splicing caching

[0087] To address the need for skip connections in networks, this invention employs two dedicated double buffers, L1 and L2, to store the output features of the encoder's L1 and L2 layers, respectively. When the decoder issues a request from the corresponding layer, the feature concatenation buffer concatenates the stored skip features with the features of the current decoding path. This separate design of the double buffers avoids mutual interference between features from different layers and supports parallel reading of multiple feature blocks within a single cycle, improving the throughput of the concatenation operation.

[0088] (iv) Data post-processing module

[0089] The data post-processing module integrates a pooling unit, an upsampling unit, an activation function unit, and an accumulation unit. The input of the data post-processing module is connected to the output of the systolic array, and it receives partial and / or intermediate calculation results from the systolic array. The pooling unit, upsampling unit, activation function unit, and accumulation unit are arranged sequentially according to a preset data pipeline order. Data paths are selected or bypassed between units via multiplexers, and the activated units are flexibly selected based on the current network layer configuration, ultimately outputting feature maps or intermediate results.

[0090] 1. Pooling Unit

[0091] A pipelined design based on a comparator array is employed to achieve 3D global max pooling. This unit consists of three core modules: an input interface and address generation unit, a comparator array, and an output formatting unit. The address generation unit uses three nested counters (depth, height, and width) to generate a complete traversal address sequence in 3D space, with the traversal order being depth-first, height-second, and width-last, ensuring coverage of the entire feature map space. Each channel in the comparator array is equipped with an independent comparison path, including a current input value register and a maximum value register; all channels operate in parallel without interference.

[0092] 2. Upsampling unit

[0093] A combined upsampling method using nearest neighbor interpolation and convolutional smoothing is employed. In the preferred embodiment, the upsampling factor is fixed at (2, 2, 1), meaning that only the height and width directions are magnified by a factor of 2, while the depth dimension remains unchanged. (Position in the input feature map) The pixel value corresponds to a 2×2×1 pixel block in the output feature map, where the values ​​at all four positions are equal to the original pixel value at the input position. The upsampling unit calculates the storage address in the output buffer for each copied output pixel. A channel-parallel copying architecture is used, with all channels executing the same copying logic simultaneously, continuing in a pipeline manner until all pixels in the input feature map have been traversed.

[0094] 3. Activation function unit

[0095] The design employs a shared arithmetic unit and a multiplexer to reconstruct the data flow. Three operating modes are flexibly switched using two control signals S1 and S0: when S1S0=00, LeakyReLU is activated, and the negative half-axis slope coefficient is fixed at 0.01; when S1S0=01, Hardsigmoid is activated; when S1S0 is 10 or 11, it enters the inactive mode and directly outputs the original input.

[0096] 4. Accumulation Unit

[0097] It is responsible for integrating the partial sums of the systolic array outputs into a complete output feature map. Efficient management of intermediate results is achieved through the Partial Sum Buffer (PSB), which contains 16 independent storage units, each corresponding to a spatial location, supporting parallel read, modify, and write-back operations within a single cycle. The accumulation unit supports two operating modes: standard channel accumulation mode for accumulation across input channels, and Shortcut element-wise addition mode for residual connection scenarios.

[0098] (v) Control Module

[0099] This invention employs a multi-state finite state machine for unified scheduling, coupled with a configurable bit-width reduced instruction set. A preferred embodiment uses a 9-state finite state machine with a 128-bit reduced instruction set. Instruction fields include opcode, layer number, number of input / output channels, feature map size, convolution kernel size, stride, padding method, and off-chip memory base address. The control module coordinates and controls each functional module through instruction prefetching and scheduling mechanisms. State transitions are jointly controlled by external start signals and completion flags of each functional module, achieving fully automatic timing scheduling across all modules.

[0100] Example 1: FPGA Accelerator Hardware Architecture

[0101] like Figure 1 As shown, the FPGA accelerator system of this invention adopts a heterogeneous computing platform based on the AXI4 bus under the top-level system. The FPGA, as a dedicated acceleration unit, is connected to the processor subsystem through a high-performance AXI4-HP interface, receives control commands and transmits feature maps and weight data at high speed, and accesses configuration registers through a general-purpose AXI4-GP interface to achieve flexible system control.

[0102] In the diagram, CLK (system clock signal) and RESET (global reset signal) are the basic control signals of the system. The system hardware carrier is a programmable logic chip. The FPGA chip has an internal AXI bus interconnect matrix to realize the routing and interaction of various AXI bus signals.

[0103] FPGA accelerators include five types of AXI standard bus interfaces: AXI4-HP (AXI4 high-speed, high-performance data bus), AXI4-GP (AXI4 general-purpose configuration bus), AXI4-Lite (AXI4 lightweight register configuration bus), AXI4-Full (AXI4 complete memory mapping bus), and AXI4-Stream (AXI4 streaming high-speed data bus).

[0104] The FPGA chip integrates an AXI4 Direct Memory Access (DMA) module, which is configured with MM-Slave (memory-mapped slave port) and MM-Master (memory-mapped master port), and is equipped with two AXI4-Stream data stream channels: S2MM (peripheral-to-memory data stream channel) and MM2S (memory-to-peripheral data stream channel). The neural network acceleration core module is located on the right side of the FPGA chip, which is also configured with MM-Slave (memory-mapped slave port) and MM-Master (memory-mapped master port) for bus interaction.

[0105] The DMA module and the NN ACCELERATOR module share the AXI4 streaming data bit width conversion unit. This unit is configured with an MM-Master (memory-mapped master port) and an MM-Slave (memory-mapped slave port) on either side to match the data stream bit widths of the upstream and downstream modules. The FPGA, as a dedicated acceleration unit, interconnects with the processor subsystem via a high-performance AXI4-HP interface. It receives control commands from the processor and performs high-speed transmission of feature maps and weight data. Simultaneously, it accesses and reads internal configuration registers via a general-purpose AXI4-GP interface, enabling flexible configuration of accelerator hardware parameters and real-time reading of operating status.

[0106] The system workflow is as follows:

[0107] Step 1: The processor issues configuration instructions via AXI4-GP and the AXI4 bus interconnect matrix, configures the accelerator control register through the AXI4-Lite bus, sets the current network layer parameters, and starts DMA transfer.

[0108] Step 2: The DMA module relies on the MM-Master (memory-mapped master port), AXI4 bus interconnect matrix, and AXI4-Full bus to read quantized weights and feature data from external DDR (double data rate dynamic memory). The data is converted into AXI4-Stream streaming data through the DMA's internal MM2S (memory-to-peripheral) channel and transmitted to the AXI4 streaming data bit width conversion unit. After bit width matching and normalization, it is sent to the MM-Slave (memory-mapped slave port) of the neural network acceleration core through the MM-Master (memory-mapped master port) and AXI4-Stream bus of this unit, and finally sent to the convolution calculation module to complete the multiplication and accumulation operation.

[0109] Step 3: After the convolutional computation module inside the neural network acceleration core completes the parallel multiplication and accumulation calculation, the result is processed by the data post-processing module, such as pooling, upsampling, activation and other operations; the processed data is output through the MM-Master (memory-mapped master port) of the NN ACCELERATOR and sent to the AXI4 streaming data bit width conversion unit through the AXI4-Stream bus to complete the bit width adaptation.

[0110] Step 4: The final data after bit-width conversion is processed in two categories: If it is intermediate network feature, it is temporarily stored in the on-chip double-buffered BRAM cache for direct use by the next layer, reducing the memory access overhead caused by repeated access to the external DDR; if it is the final segmented output feature data of the network, it is sent into the DMA through the bit-width conversion unit MM-Slave (memory-mapped slave port) and the DMA module S2MM (peripheral to memory) channel, and then written back to the external DDR memory for storage by relying on the DMA module MM-Slave (memory-mapped slave port), AXI4-Full bus, and AXI4 bus interconnect matrix.

[0111] Throughout the process, the double buffering mechanism ensures that while one set of data is being processed in the computing unit, the next set of data has been preloaded into the on-chip cache, thus allowing data transmission and computation to overlap in time and effectively hiding off-chip memory access latency.

[0112] like Figure 2 As shown, the neural network acceleration core of this invention includes: a convolution calculation module, an on-chip storage module, a data post-processing module, and a control module.

[0113] In terms of data flow, the on-chip feature storage controller and the on-chip weight storage controller receive external data through a memory-mapped AXI4 bus interface and cache it in the on-chip feature cache and on-chip weight cache, respectively. The data from the on-chip feature cache and on-chip weight cache is input to the matrix multiplication unit (i.e., the systolic array convolution calculation module) for parallel multiplication and addition operations. The output of the matrix multiplication unit is input to the data post-processing module, which integrates a pooling unit, an upsampling unit, an activation function unit, a skip connection unit, and an accumulation unit to perform various functional operations after convolution. The post-processed result is written to the on-chip output cache through the on-chip output storage controller and finally written back to the external DDR through the memory-mapped AXI4 bus interface.

[0114] In terms of control flow, the state machine serves as the control module, connecting to the aforementioned storage controllers, matrix multiplication units, and data post-processing modules to coordinate the timing and data flow between these modules. The convolution calculation module employs a systolic array, the on-chip storage utilizes a three-level storage hierarchy and multiple caches in collaboration, and the control module employs a 9-state finite state machine for unified scheduling.

[0115] Example 2: Double Buffering Mechanism

[0116] like Figure 3 As shown, this invention employs a double-buffering mechanism to achieve pipelined parallelism in data transmission and computation. The double-buffering mechanism consists of two independent BRAM buffers (Buffer0 and Buffer1), and a state machine is used to complete the pipelined scheduling of data transmission and computation tasks.

[0117] Figure 3 In this context, Cycle represents the hardware clock cycle; Load ifm is the input feature loading operation; and Compute ifm is the convolution calculation operation. Figure 3 In diagram (a), the timing of the traditional single-buffered architecture is shown. Data loading and convolution calculation are executed serially: Cycle0 only performs input feature loading, Cycle1 must wait for loading to complete before starting convolution calculation; subsequently, Cycle2 and Cycle3 continuously alternate between loading and calculation tasks. The loading and calculation processes occupy hardware resources in a time-sharing manner, the PE calculation unit has a large number of idle cycles, and the overall throughput efficiency of the pipeline is low.

[0118] Figure 3 (b) in the figure shows the Ping-pong double buffering timing used in this invention. Ifm0 and ifm1 in the figure correspond to the two independent BRAM buffers Buffer0 and Buffer1 in the hardware configuration, which can realize the timing overlap and parallelism of data loading and convolution calculation.

[0119] The specific workflow of the dual-buffered core of this invention is as follows:

[0120] When the PE reads data from Buffer0 for computation, the DMA controller synchronously writes the next batch of data from DDR to Buffer1 via a burst transfer mechanism. Once the current batch of data is computed, the system does not need to wait for the data transfer to finish; it directly switches to Buffer1 to start the next round of computation. Simultaneously, the DMA writes the next batch of data to the released Buffer0. This parallel mode allows the high latency of DDR to be covered by the computation process, fundamentally avoiding idle periods caused by the PE waiting for data, ensuring the continuous and efficient operation of the computing unit. Compared to the traditional single-buffered structure, the dual-buffered mechanism of this invention eliminates the serial waiting between the data loading stage and the computation stage, significantly improving pipeline throughput efficiency.

[0121] Example 3: Hardware Mapping of Depthwise Convolution

[0122] like Figure 4 As shown, the batch parallel model of depthwise convolution processes each batch. There are 1 input channel. The input features of each channel are divided into feature blocks, and the size of the feature blocks is 1. ,in The spatial size of the convolution kernel (in this embodiment) ), The number of spatial locations processed in a single batch (in this embodiment) ), The convolution stride (in this embodiment) ).

[0123] During the calculation process, the front of the array Row processing units (PEs) are activated, with each row PE dedicated to processing an independent input channel. The 16 processing units within each row PE are used for parallel computation of that channel. Output positions.

[0124] The data flow is as follows: feature blocks for each channel are streamed into their respective PE rows via feature sliding, with data pulsed from left to right within the row. The feature windows received by adjacent PEs are spatially contiguous, and due to sliding overlap, a large amount of feature data can be directly reused between PEs without repeatedly reading from the on-chip cache. The weights corresponding to each channel are statically stored in the shared cache of each row of PEs before the computation begins, and remain unchanged during the computation of the entire channel by broadcasting weights continuously to all PEs within the row.

[0125] During each clock cycle, the PEs within each row synchronously perform multiply-accumulate operations. Each PE retrieves an index from its local register. The feature vectors are obtained from the broadcast bus, and the corresponding weight vectors are retrieved, then multiply-accumulate operations are performed. After 27 cycles of pipelined computation, each PE generates a complete output pixel result. After a single batch of computation is completed, each row of PEs generates... The results are collected simultaneously and written to the output block, and finally stored in the output buffer.

[0126] Example 4: Hardware Mapping of Pointwise Convolution

[0127] like Figure 5 As shown, the pointwise convolution batch processing model decomposes the complete computation into regularized batch tasks. Among them... The number of input channels for a single batch processing. This represents the number of output channels for a single batch. At each spatial location, a complete matrix multiplication is performed using a double loop.

[0128] The systolic array is configured in fully connected matrix multiplication mode, and the data flow mechanism is as follows: the input feature vector is broadcast into the column, and the weight matrix is ​​vertically systolic loaded in the row.

[0129] Specifically, the current spatial location The dimensional input feature vector is loaded from the left side of the array. Each feature element... Vertically distributed to the first via the in-column broadcast bus In a column, all PEs from different output channels within the same column receive the exact same input feature values ​​at the same time. The weight matrix is ​​divided into rows into blocks, each block... Line, corresponding The nth output channel. The current weight block's nth... Rows are input from the top of the array and propagate pulsatingly along the vertical direction. As a weight row flows through each row PE, that row PE captures the weight element in the weight row corresponding to its column index. .

[0130] Each PE performs a multiply-accumulate operation after receiving the feature value and weight value. After the calculation is completed, the sum of all parts of each row of PE is collected and accumulated within one cycle to obtain the output. These results are written to the output buffer, the feature vector of the next spatial location is broadcast into the array, and the weight block can be reloaded or reused.

[0131] Example 5: Hardware Mapping of Standard 3D Convolution

[0132] like Figure 6 As shown, The processing unit PE array is statically divided into 16 independent Subarrays. Each subarray is responsible for the complete convolution computation of one output channel at a spatial location.

[0133] The feature data stream is provided by independent sliding window buffers, which temporally feed each subarray according to its depth slice. The input feature window oscillates horizontally between PEs within the subarray to complete the convolution window computation. The weight data stream loads all 27 weights of the current channel at once from the top weight memory and broadcasts them vertically to all 16 subarrays. Each PE locks its corresponding weight value based on its spatial location and the depth slice being processed.

[0134] The computation and accumulation process exhibits multi-level parallelism. Spatially, the nine PEs within each subarray perform nine multiply-accumulate operations in parallel per cycle; in terms of depth, the computation is pipelined according to slices 0, 1, and 2, with intermediate results accumulated across time steps; in terms of channels, after completing the depth pipeline of an input channel, its partial sum is temporarily stored in a partial sum buffer after partial sum accumulation, and accumulation continues when switching to the next input channel. The final output is obtained after the loop.

[0135] Example 6: Specific Implementation of the Upsampling Unit

[0136] In this invention, the upsampling unit is located in the decoder path and is responsible for restoring the low-resolution feature map to a high resolution. A nearest-neighbor upsampling method is employed, expanding the spatial resolution through data replication. The upsampling factor is fixed at (2, 2, 1), meaning that only the height and width directions are magnified by a factor of 2, while the depth dimension remains unchanged.

[0137] like Figure 7 As shown, the mapping relationship of nearest neighbor upsampling can be described as: position in the input feature map The pixel value at that location corresponds to a pixel in the output feature map. The pixel block. Four locations within this block. , , , The values ​​are all equal to the input position. The original pixel value at that location.

[0138] The upsampling unit calculates the storage address in the output buffer for each copied output pixel. The output buffer uses row-major storage, and its output position... The formula for calculating the linear address at a given location is:

[0139]

[0140] For each input pixel The corresponding four output positions generate the following addresses:

[0141]

[0142] Subsequently, the upsampling unit writes the pixel value at the current input pixel into the four corresponding output positions in parallel in the output buffer according to the generated address.

[0143] Example 7: Specific Implementation of the Activation Function Unit

[0144] Based on the LAFNet-s network architecture, the use cases for activation functions are mainly divided into two categories: LeakyReLU activation in the backbone network and Hardsigmoid activation in the SE module. The negative half-axis slope coefficient of LeakyReLU is fixed at 0.01. Furthermore, some convolutional layers do not require activation functions. Therefore, the activation function unit settings of this invention support three switchable operating modes.

[0145] like Figure 8 As shown, the activation function unit of this invention employs a design that reconstructs the data flow using a shared arithmetic unit and a multiplexer. By reusing core computational resources such as adders and multipliers, hardware overhead is effectively reduced. This unit uses two control signals, S1 and S0, to flexibly switch between three operating modes. The operating logic of each mode is as follows:

[0146] 1. When S1S0=00, the unit operates in LeakyReLU mode. The input data X enters the comparator for sign determination: if X≥0, X is output directly; if X<0, a linear transformation with a slope α=0.01 in the negative interval is achieved through the lower branch multiplier, and 0.01X is output.

[0147] 2. When S1S0=01, the unit operates in Hardsigmoid mode. The input data X is first added to the constant 3 by the upper branch adder to obtain X+3. Then, the ReLU6 clamping operation is implemented through two-stage multiplexers MUX1 and MUX2 to restrict the result to the interval [0,6]. Finally, the scaling operation of ×1 / 6 is implemented through the lower branch multiplier, and the output is normalized to the Hardsigmoid activation value in the range [0,1].

[0148] 3. When S1S0 is 10 or 11, the unit enters the inactive mode, all arithmetic units are bypassed, and the input data X is directly output.

[0149] Through the above design, the present invention enables the reuse of three activation functions on a single hardware unit, effectively reducing hardware resource consumption while maintaining low latency characteristics in each working mode.

[0150] exist Figure 8 In the diagram, ! represents logical NOT operation and & represents logical AND operation; S1 and S0 are two-bit mode control signals, and their combined values ​​determine the unit's operating mode; X is the unit's input raw feature data; X1 is the intermediate operation value after adding input X and the constant 3; MUX0, MUX1, MUX2, MUX3, and MUX4 are five-way two-to-one multiplexers. For adders; This is a multiplier; the logical expression next to the multiplexer is the selection condition for that path; the value 0.1667 is the decimal approximation of the fraction 1 / 6.

[0151] Example 8: Detailed workflow of the control module state machine

[0152] The control module of this invention employs a 9-state finite state machine (FSM) to uniformly schedule the accelerator's data handling, computation execution, and result write-back processes. The definitions and transition logic of each state are as follows:

[0153] IDLE (Idle State): The system enters this state after a reset, and all control signals are cleared. When an external start signal start=1 is detected, the state transitions to LOAD_CFG; otherwise, it remains idle.

[0154] LOAD_CFG (Configuration Loading State): Reads the computation parameters of the current layer from the instruction memory. Parameters include the number of input / output channels, feature map size, convolution kernel size, stride, padding method, and off-chip memory base address. After configuration loading is complete, the signal cfg_done=1, and the state transitions to LOAD_IFM.

[0155] LOAD_IFM (Input Feature Loading State): The input feature map is moved from off-chip DDR to on-chip input feature buffer via the DMA controller. For the decoder layer, this state also initiates a read operation of the Feature Concatenation Buffer (FCB), loading the skip connection features into the concatenation engine. After the transfer is complete, the signal dma_rd_done=1, and the state transitions to LOAD_WEIGHT.

[0156] LOAD_WEIGHT (Weight Loading Status): Starts the weight loading module, loading the convolutional kernel weights of the current layer from off-chip DDR to the weight cache. After the weight loading is complete, the signal weight_done=1, and the status transitions to EXECUTE.

[0157] EXECUTE (Execution State): Based on the current layer type (standard convolution, depthwise convolution, pointwise convolution, pooling, upsampling, etc.), the corresponding computation unit is started. Then, it unconditionally transitions to the WAIT_EXECUTE state.

[0158] WAIT_EXECUTE (Waiting for execution): Waiting for the computation unit to complete the computation of the current data block. When execute_done=1, the state transitions to STORE_OFM.

[0159] STORE_OFM (Output Feature Write-Back State): Initiates the DMA write channel to write the calculation results in the on-chip output feature buffer back to the off-chip DDR. For the encoder layer, this state also initiates the FCB write operation, storing the output features into the FCB for use by the subsequent decoder layer. After the write-back is complete, the block counter is used to determine the state: if there are still unprocessed data blocks in the current layer (block_last=0), the state returns to LOAD_IFM; if all data blocks in the current layer have been processed (block_last=1), the state transitions to NEXT_LAYER.

[0160] NEXT_LAYER (Next Layer State): Updates the layer counter and re-fetches the parameters for the next layer from the instruction memory. If the current layer is the last layer (layer_last=1), it transitions to the DONE state; otherwise, it returns to the LOAD_CFG state.

[0161] DONE (Completion Status): The done signal is pulled high to send an interrupt to the processor, notifying that the inference task is complete. It returns to the IDLE state after the external start signal is cleared.

[0162] Through the above finite state machine design, this invention achieves precise scheduling of multi-network layer and multi-data block pipelines, ensuring efficient operation of the accelerator.

[0163] Example 9: Detailed Implementation Steps for Model Quantization and Batch Normalization Layer Fusion

[0164] To further reduce hardware resource consumption and improve inference speed, this invention performs model lightweighting processing before deploying the trained LAFNet-s network to the FPGA. Specifically, this includes Batch Normalization (BN) layer fusion and 8-bit fixed-point quantization. BN layer fusion merges each convolutional layer (including depthwise convolutions and pointwise convolutions) with its subsequent batch normalization layer. 8-bit fixed-point quantization uniformly quantizes the weight parameters and activation values ​​of each layer.

[0165] 1. BN layer fusion

[0166] During the training phase, batch normalization (BN) layers effectively accelerate model convergence by stabilizing the inter-layer input distribution. However, during the inference phase, the presence of BN layers introduces additional computational overhead. Therefore, this invention employs a BN layer fusion technique, merging the trained BN layer parameters with the corresponding convolutional layer parameters to form a single fused convolutional layer.

[0167] Specifically, let the weights of the convolutional layer be... , bias is The scaling parameters for the BN layer are: The offset parameter is The statistical mean of the training set is The variance is The new weights after fusion and new bias Calculate using the following formula:

[0168]

[0169]

[0170] in, This is a numerically stable term, taking the value of After fusion, the original network layer sequence is simplified from convolutional layer → BN layer to a single fused convolutional layer, eliminating the independent computational overhead of BN layer during inference, avoiding off-chip transfer of intermediate results, and eliminating the need to configure separate hardware modules for BN layer.

[0171] 2. 8-bit symmetric linear quantization

[0172] This invention employs an 8-bit symmetric linear uniform quantization scheme to map the weight parameters and activation values ​​in the network from 32-bit floating-point numbers to 8-bit fixed-point integers. For a given layer, its floating-point value... With quantized integers The mapping relationship between them is as follows:

[0173]

[0174]

[0175] in, This indicates rounding operations. This is the quantization scaling factor. This is the zero-point offset. The scaling factor and zero-point offset are determined based on the maximum and minimum values ​​of the layer's weights:

[0176]

[0177]

[0178] For 8-bit quantization, the integer range is set to... To ensure quantization accuracy, this invention selects 100 representative images from the training set as a calibration set. The range of activation values ​​for each layer is statistically analyzed through forward propagation, and the quantization parameters for each layer are calculated accordingly. The quantized weights and activation values ​​are stored in off-chip DDR in 8-bit fixed-point format and are read and calculated by the FPGA accelerator during inference.

[0179] Example 10: Performance Verification

[0180] The accelerator of this invention was deployed on a Xilinx Zynq UltraScale+ ZCU102 development board. The target operating frequency was set to 200MHz. The test data came from holograms acquired by a digital holographic recording system, and digital holographic particle field image data containing three-dimensional particle distribution information was obtained after numerical reconstruction using the angular spectrum method.

[0181] The particle detection rates of the LAFNet-s network on the PyTorch floating-point model and the FPGA fixed-point model were compared, and the results are shown in Table 1. Table 1 shows that the average particle detection rate of the FPGA fixed-point inference is 92.29%, which is only 2.95 percentage points lower than that of the PyTorch floating-point model, indicating that the accelerator can stably and effectively identify particle targets.

[0182] Table 1 Comparison of particle detection rates under LAFNet-s PyTorch floating-point model and FPGA fixed-point model

[0183]

[0184] The Dice coefficient was used to quantitatively evaluate the hardware and software inference results, and the results are shown in Table 2. Table 2 shows that the average Dice coefficient for FPGA fixed-point inference is 0.8195, and compared with the PyTorch floating-point model, the accuracy retention rate reaches 96.21%, indicating that the accuracy loss introduced by quantization and hardware mapping is within an acceptable range.

[0185] Table 2 Comparison of Dice coefficients of LAFNet-s on PyTorch and FPGA platforms under different particle fields.

[0186]

[0187] Synthesis and implementation were completed on the Vivado 2022.2 platform, with the target operating frequency set to 200MHz. No severe routing congestion occurred after placement and routing, and timing convergence was good. The hardware resource consumption of this invention is shown in Table 3. As can be seen from Table 3, the resource utilization of this invention is low, demonstrating good deployment feasibility on the target FPGA platform, and providing sufficient margin for subsequent network expansion and algorithm iteration.

[0188] Table 3 Accelerator Hardware Resource Consumption Statistics

[0189]

[0190] As shown in Table 4, the inference speed of this invention is 33.6 ms / frame, the power consumption is 3.5W, and the energy efficiency ratio reaches 8.51 FPS / W. Under the same test conditions, compared with the CPU platform (Intel i7-13700H), the inference speed of this invention is improved by approximately 7.1 times; compared with the GPU platform (NVIDIA RTX 3090), the energy efficiency ratio is improved by approximately 9.2 times. The above experimental results show that this invention achieves low power consumption and high energy efficiency real-time inference while ensuring segmentation accuracy, and can meet the real-time processing requirements of edge computing scenarios.

[0191] Table 4 Comparison of Inference Performance and Energy Efficiency of CPU / GPU / FPGA

[0192]

[0193] In summary, this invention effectively solves the technical bottlenecks of traditional digital holographic particle field segmentation algorithms, such as high computational load, poor real-time performance, and low energy efficiency in edge deployment, through algorithm-hardware co-optimization. While ensuring high-precision segmentation results, it achieves low-power, high-real-time edge inference, which can be widely applied to engineering scenarios such as fluid dynamics particle dynamic measurement, industrial spray diagnostics, and environmental particle monitoring.

[0194] It should be noted that adjusting parameters such as quantization bit width, array size, storage capacity, and upsampling rate will only change inference accuracy, single-frame latency, hardware power consumption, and on-chip resource usage. The core of this invention, the lightweight LAFNet-s network structure, the BN fusion offline solidification process, the three-level reconfigurable on-chip storage, the ping-pong double-buffered memory access hiding mechanism, and the three types of convolutional differentiated hardware mapping schemes, remain unchanged and can all achieve real-time segmentation at the edge of the digital holographic particle field.

[0195] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A real-time segmentation method for digital holographic particle fields based on FPGA, characterized in that, Includes the following steps: Step S1: Acquire digital holographic particle field image data and perform block preprocessing on the image data; Step S2: Input the preprocessed image data into the pre-trained lightweight segmentation network LAFNet-s for feature inference to obtain the particle segmentation result map. The LAFNet-s network adopts an encoder-decoder architecture and includes depthwise separable convolutional layers, large kernel convolution equivalent decomposition modules, piecewise linear activation functions, and nearest neighbor upsampling modules. Step S3: Utilize the FPGA hardware accelerator to perform the depthwise convolution, pointwise convolution, and standard 3D convolution operations in step S2. The FPGA hardware accelerator uses a fixed-size pulsating array as the core computing unit, and is equipped with three-level reconfigurable on-chip memory, a double-buffered pipeline mechanism, a multi-mode data post-processing module, and a multi-state finite state machine scheduling architecture to achieve pipelined parallel computing. Step S4: For the three core convolution operators of standard 3D convolution, depthwise convolution, and pointwise convolution, a differentiated hardware mapping strategy is used to drive the pulsating array to perform parallel computation and complete the real-time segmentation of the digital holographic particle field.

2. The method according to claim 1, characterized in that, The construction of the lightweight segmentation network LAFNet-s includes the following optimization strategies: (1) Replace the standard 3D convolutional layers in the original network LAFNet with depth-separable convolutional layers, which consist of depth convolution and pointwise convolution; (2) The large-size convolutional kernels in the original network are equivalently decomposed into three cascaded small-size convolutional kernels; (3) Replace the Sigmoid activation function in the SE module with the Hardsigmoid function based on piecewise linear approximation; (4) Replace the upsampling operation from transposed convolution with a combination of nearest neighbor upsampling and convolution.

3. The method according to claim 1, characterized in that, The FPGA hardware accelerator includes a configurable fixed-size systolic array, a three-level on-chip memory hierarchy, a double-buffered pipeline mechanism, a data post-processing module, and a control module. (1) A configurable fixed-size systolic array, wherein each processing unit PE of the systolic array includes a multiplier, an adder and a partial sum register; the input end of the systolic array is connected to the input feature buffer and the weight buffer respectively, and the output end of the systolic array is connected to the data post-processing module; the double buffer pipeline mechanism is configured in the input feature buffer and the weight buffer respectively; (2) A three-level on-chip storage hierarchy, including input feature cache, weight cache and output feature cache, forms a high-speed on-chip data interaction path, which respectively adapts to the independent functional requirements of feature data loading, weight data storage and inference result caching; (3) The dual-buffered pipeline mechanism consists of two independent BRAM buffers forming a ping-pong alternating working architecture. By switching between data loading and computing tasks, the latency of off-chip DDR memory access is hidden, and the data transmission and hardware computing pipelines are parallelized. (4) Data post-processing module, which integrates pooling unit, upsampling unit, activation function unit and accumulation unit, to complete the post-processing of convolutional features in pipeline sequence; (5) Control module, which is connected to the pulse array, the three-level on-chip storage hierarchy and the data post-processing module respectively, is used to schedule the timing and data flow of each module. The control module adopts a multi-state finite state machine to realize the unified timing scheduling of the entire module, and is equipped with a configurable bit-width simplified instruction set to realize fully automatic scheduling of network layer parameter configuration, data transportation, calculation execution and result writing back.

4. The method according to claim 1, characterized in that, The differentiated hardware mapping strategy includes: (1) The depthwise convolution adopts a channel isolation mapping strategy, which maps each row of processing units (PE) of the systolic array to the input channel one by one. A single row of PE independently processes the spatial convolution operation of a single input channel. The weights are statically loaded and broadcast within the row, and the feature data is input in a sliding window manner. (2) Pointwise convolution adopts column broadcasting and weight broadcasting mapping strategy. The input feature vector is input from the left side of the systolic array and broadcast by column. The rows of the weight matrix flow in from the top of the array and are systolic. (3) The standard 3D convolution adopts a depth pipeline and subarray partitioning mapping strategy, which statically divides the pulsating array into several independent subarrays. Each subarray is responsible for the convolution calculation of one output channel at one spatial location. The calculation process is executed in a time-series pipelined manner according to the depth slice.

5. The method according to claim 3, characterized in that, The input feature cache has a reconfigurable data organization mode for different convolution modes: in the standard convolution mode, the cache area is divided into several independent feature window storage volumes, and the data arrangement follows the depth-first principle; in the depthwise convolution mode, a channel-isolated storage architecture is adopted, and an independent storage block is allocated to each input channel; in the pointwise convolution mode, a column-based storage architecture based on the input channel dimension is adopted, the storage volume is divided according to the input channel, and the reconstructed feature vector is distributed to the systolic array through the in-column broadcast network.

6. The method according to claim 3, characterized in that, The weight cache adopts a hierarchical grouping architecture, which is divided into standard convolution area, depthwise convolution area and pointwise convolution area according to the convolution type. Physically, it consists of several independent banks, and each bank can be accessed independently.

7. The method according to claim 3, characterized in that, The upsampling unit uses the nearest neighbor interpolation algorithm combined with convolutional smoothing to restore feature resolution. It supports configurable three-dimensional upsampling rate, maps a single pixel in the input feature map to the corresponding pixel block, and executes the same pixel copying logic in parallel across all feature channels.

8. The method according to claim 3, characterized in that, The activation function unit adopts a computation unit multiplexing architecture and realizes the switching of multiple activation working modes and non-activation pass-through mode through control signals; the multiple activation working modes include at least two types: LeakyReLU and Hardsigmoid.

9. The method according to claim 1, characterized in that, The network deployment phase involves lightweight model solidification, specifically including Batch Normalization (BN) layer fusion and symmetric linear fixed-point quantization. The BN layer fusion is achieved by fusing the scaling, offset, mean, and variance parameters of the BN layer during the training phase to the corresponding convolutional layer weights and biases, thereby eliminating the independent computational overhead of BN during the inference phase. The symmetric linear fixed-point quantization maps 32-bit floating-point weights and activation values ​​to low-bit-width fixed-point data. By statistically analyzing quantization parameters through a calibration set, storage overhead and computation bit width are reduced under the premise of controllable accuracy loss, thus adapting to the FPGA fixed-point parallel computing architecture.

10. A real-time segmentation system for digital holographic particle fields based on FPGA, characterized in that, The system is used to execute the real-time segmentation method according to any one of claims 1-9, and the system is an algorithm-hardware co-acceleration architecture, including a data preprocessing module, a lightweight inference module, an FPGA hardware acceleration module and a result output module. The data preprocessing module is used to collect the original image data of the digital holographic particle field and complete the normalization and block normalization preprocessing to generate network compliant input data; The lightweight inference module is used to load the LAFNet-s lightweight segmentation network that has been trained, BN fusion and 8-bit quantization completed, to realize multi-scale noise robust feature extraction and end-to-end segmentation inference of holographic particle field. The FPGA hardware acceleration module has a built-in configurable systolic array computing core, a three-level reconfigurable on-chip storage unit, a multi-mode post-processing unit and an instruction scheduling and control unit. It achieves hardware acceleration of the full-layer operators of the LAFNet-s network and pipelined parallel inference through three types of convolutional differential mapping strategies. The result output module is used to receive the segmentation feature map after hardware-accelerated inference, complete data caching, splicing verification and off-chip DDR write-back, and output the particle field segmentation result in real time.