YOLOv10 in-orbit remote sensing image target detection network optimization method and FPGA acceleration system
By improving the YOLOv10 network model and designing the FPGA acceleration system, the complexity, accuracy and timeliness of the YOLOv10 on-orbit remote sensing image object detection network are solved when deploying on FPGA, and more efficient resource utilization and detection performance are achieved.
Patent Information
- Application Number
- CN202510308034.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
When deploying on FPGAs, YOLOv10 on-orbit remote sensing image object detection network faces problems such as complex network structure, low detection accuracy for small objects, and difficult to guarantee the timeliness of on-orbit detection.
Improvements to the YOLOv10 network model include using the FasterNet network module to replace the BottleNeck module, replacing the activation function, adopting fast normalized fusion technology, layer fusion and square weight fixed-point quantization, and designing FPGA acceleration systems, including AXI Bus, Merger, YOLO Subsystem and Fetcher modules.
It reduces network complexity and computing volume, improves small object detection accuracy, improves FPGA resource utilization and on-orbit detection timeliness, and reduces power consumption and hardware resource requirements.
Smart Images

Figure BSA0000300024680000021 
Figure BSA0000300024680000022 
Figure BSA0000300024680000041
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural network accelerators, and in particular to a method for implementing a YOLOv10 target detection network optimization on an FPGA and a system for performing FPGA acceleration on the optimized network. Background Art
[0002] In recent years, many researchers have conducted extensive research in the field of ship target detection in remote sensing images and achieved remarkable results. In particular, deep learning target detection algorithms based on convolutional neural networks have greatly improved the performance of remote sensing ship target detection compared to traditional algorithms. The YOLO (You Only Look Once) series of target detection algorithms can achieve end-to-end real-time detection. Its latest version, YOLOv10, has improved accuracy, but at the same time introduced more parameters and greater computational complexity. In the on-orbit remote sensing image target detection scenario, GPUs have an advantage in computing efficiency, but their power consumption often cannot meet the needs of embedded and low-power devices. FPGAs have the characteristics of low power consumption, high parallelism, and programmability, making them suitable for deploying deep learning models.
[0003] However, deploying YOLOv10 directly on FPGA still faces the following challenges.
[0004] (1) Complex network structure: Compared with previous YOLO versions, YOLOv10 has more network model parameters, is larger in size, has a more complex model structure, and takes longer to run.
[0005] (2) Low detection accuracy for small targets: The targets detected in on-orbit remote sensing images are relatively small, and the YOLOv10 network model has poor detection results for small targets.
[0006] (3) It is difficult to ensure the timeliness of on-orbit detection: As a mobile and time-sensitive target, ships must be detected with high timeliness. How to achieve hardware acceleration on a satellite-borne hardware platform with limited resources and power consumption to reduce system computing delay is a difficulty in on-orbit ship target detection.
[0007] In view of the above problems, there is an urgent need for a method to optimize the characteristics of the YOLOv10 network and an FPGA acceleration system to effectively reduce network complexity, improve small target detection accuracy, increase target detection speed, and reduce power consumption and hardware resource requirements. Summary of the invention
[0008] The object of the present invention is to propose an optimization method for the on-orbit remote sensing image target detection network of YOLOv10 and an FPGA acceleration system in view of the above challenges. The present invention optimizes the on-orbit remote sensing image target detection network of YOLOv10 to reduce the network complexity and computational amount, improve the detection accuracy of small targets, and increase the utilization rate of FPGA resources; at the same time, by designing an FPGA acceleration system to accelerate the optimized network, the on-orbit detection timeliness is improved.
[0009] To achieve the above object, the technical solution adopted by the present invention is: an optimization method for the ship target detection network of YOLOv10 in on-orbit remote sensing images, characterized by comprising the following steps:
[0010] (1) Improve the on-orbit remote sensing image target detection network model of YOLOv10 to reduce the model complexity and improve the accuracy of small target detection:
[0011] (1.1) Use the FasterNet network module to replace the Bottleneck module in the original C2f network to reduce the network computational amount and the number of parameters and accelerate the calculation process;
[0012] (1.2) Use the LeakyReLU activation function to replace the original SiLU activation function;
[0013] (1.3) Improve the Concat network using Fast Normalized Fusion to obtain the Concat_FN and Concat_FN3 networks; Fast Normalized Fusion performs fast normalized fusion on i input feature maps, and the formula is as follows:
[0014]
[0015] where O is the fused output feature map, w i is the learnable weight of the i-th input feature map, I i is the i-th input feature map, ε is a very small constant, and ∑ j w j is the sum of all feature map weights;
[0016] The number of input feature maps of the Concat_FN network is i = 2, and the number of input feature maps of the Concat_FN3 network is i = 3;
[0017] (1.4) Replace the Concat network in the 15th and 18th layers of the original YOLOv10 with the Concat_FN network described in (1.3); replace the Concat network in the 24th layer of the original YOLOv10 with the Concat_FN3 described in (1.3), and extract feature maps from the networks in the 6th, 13th, and 24th layers for weighted splicing; to improve the detection accuracy of small targets;
[0018] (2) Use the DOTA dataset to train the improved YOLOv10 network model described in (1.4) to obtain the best model weight data;
[0019] (3) Perform layer fusion and weight squared fixed-point quantization on the trained YOLOv10 network model described in (2) to improve the hardware implementation efficiency and accelerate the inference speed:
[0020] (3.1) Fuse the convolutional layer and the BN layer in the trained YOLOv10 network model described in (2) to reduce the number of intermediate result caches and optimize the calculation process; the outputs of the convolutional and BN layers after layer fusion are as follows:
[0021]
[0022] Where: Y is the output after convolutional BN fusion, x is the feature input, w is the convolutional weight, b is the convolutional bias, is the average value of x, σ 2 is the variance of x, γ and β are learnable parameters, and ε is a very small value to prevent the denominator from being zero;
[0023] (3.2) Perform squared fixed-point quantization on the fused weights described in (3.1), convert the multiplication operation into a shift operation to facilitate implementation using the lookup table of the FPGA, thereby improving the hardware implementation efficiency and accelerating the inference speed; export the weights as a binary BIN file for the initialization of the neural network accelerator.
[0024] In addition, the present invention also proposes a FPGA acceleration system for the YOLOv10 on-orbit remote sensing image target detection network, which is characterized by including modules such as AXI Bus, Merger, YOLO Subsystem, and Fetcher implemented on a single FPGA, and an external DDR3 memory;
[0025] The DDR3 memory is used to store the remote sensing images to be recognized and the YOLOv10 network weight data;
[0026] The MIG module is used to access the DDR3 memory through the AXI bus;
[0027] The AXI Bus is used to connect the Merger, Fetcher, and MIG modules;
[0028] The YOLO Subsystem module is used to implement the core computing function of the FPGA acceleration system for the YOLOv1 0 on-orbit remote sensing image target detection network;
[0029] The Fetcher module is used to allocate weight data streams for the YOLO Subsystem module according to its operating status;
[0030] The Merger module decodes the detection boxes of the output results of the YOLO Subsystem to obtain the final target positions, target categories, and confidence levels.
[0031] Furthermore, the specific implementation of the YOLO Subsystem is as follows:
[0032] (1) Sub-network modular IP design: Use the Vitis HLS tool to design the IP of each sub-network module using the space-for-time strategy, including: Conv, C2f_Faster, SCDown, C2fCIB, SPPF, PSA, Upsample, Concat_FN, Concat_FN3, BiFPN_Add3, and V10Detect, etc.;
[0033] (2) YOLO Subsystem instantiation: Add the IP of the sub-network modules described in (1) to the Backbone function and the Head function in the order of the Backbone and Head structures, and further instantiate these two functions into the YOLO Subsystem module;
[0034] (3) Data stream allocation: Provide data streams for each sub-network module through the data stream allocation function, including input feature map data streams, weight data streams, normalized weight data streams, etc.;
[0035] (4) Convolution calculation optimization: Optimize the convolution calculation process, expand the convolution process, and perform pipelining on it to improve the calculation performance;
[0036] (5) FPGA resource optimization: Utilize the LUT, DSP, and BRAM of the FPGA for dynamic resource allocation, optimize the FPGA resource usage efficiency through adaptive scheduling and parallel computing, reduce power consumption, and improve system performance. Description of the Drawings
[0037] Figure 1 It is the flowchart of the specific implementation of the present invention
[0038] Figure 2 It is the network structure diagram of the improved YOLOv10 of the present invention
[0039] Figure 3 Structure diagram of the hardware acceleration system described in the present invention
[0040] Figure 4 Operation flowchart of the hardware acceleration system described in the present invention Specific implementation manners
[0041] The embodiments of the present invention will be described in detail below. These embodiments are exemplary and are only used to explain the present invention, and should not be construed as a limitation to the present invention. The method for optimizing the YOLOv10 on-orbit remote sensing image target detection network and the FPGA acceleration system of the present invention will be described in detail below with reference to the accompanying drawings of the specification.
[0042] The flowchart of the specific implementation manner of the present invention is as Figure 1 shown, and includes the following steps.
[0043] S1: Improve the YOLOv10 on-orbit remote sensing image target detection network model to reduce the model complexity and improve the accuracy of small target detection.
[0044] (1.1) Use the FasterNet network module to replace the Bottleneck module in the original C2f network to reduce the network calculation amount and the number of parameters and accelerate the calculation process;
[0045] (1.2) Use the LeakyReLU activation function to replace the original SiLU activation function;
[0046] The formula of LeakyRelu is as follows:
[0047]
[0048] where β is a scaling factor when x is negative, and the default value is 0.01.
[0049] (1.3) Use Fast Normalized Fusion to improve the Concat network to obtain the Concat_FN and Concat_FN3 networks; Fast Normalized Fusion is to perform fast normalized fusion on i input feature maps, and the formula is as follows:
[0050]
[0051] where O is the output feature map after fusion, w i is the learnable weight of the i-th input feature map, I i is the i-th input feature map, ∈ is a very small constant, and ∑ j w j is the sum of all feature map weights;
[0052] The number of input feature maps of the Concat_FN network is i = 2, and the number of input feature maps of the Concat_FN3 network is i = 3.
[0053] (1.4) Replace the Concat network in the 15th and 18th layers of the original YOLOv10 with the Concat_FN network described in (1.3); replace the Concat network in the 24th layer of the original YOLOv10 with the Concat_FN3 described in (1.3), and extract feature maps from the networks in the 6th, 13th, and 24th layers for weighted splicing; to improve the detection accuracy of small targets.
[0054] S2: Use the DOTA dataset to train the improved YOLOv10 network model described in (1.4) to obtain the best model weight data.
[0055] The DOTA dataset is a large image dataset for object detection in aerial images, annotated with 15 common object categories, including: airplane, ship, storage tank, baseball field, tennis court, basketball court, ground runway, port, bridge, large vehicle, small vehicle, helicopter, roundabout, football field, and basketball court.
[0056] S3: Perform layer fusion and weight square fixed-point quantization on the trained improved YOLOv10 network model described in (S2) to improve the hardware implementation efficiency and accelerate the inference speed.
[0057] (3.1) Fuse the convolutional layer and the BN layer in the trained improved YOLOv10 network model described in (S2) to reduce the number of intermediate result caches and optimize the calculation process; the outputs of the convolutional and BN layers after layer fusion are as follows:
[0058]
[0059] Among them: Y is the output after convolutional BN fusion, x is the feature input, w is the convolutional weight, b is the convolutional bias, is the average value of x, a 2 is the variance of x, γ and β are learnable parameters, and ε is a very small value to prevent the denominator from being zero; after layer fusion, one convolutional operation can represent one convolutional operation plus one BN operation, reducing the number of caches of the intermediate structure.
[0060] (3.2) Perform square fixed-point quantization on the weights after fusion described in (3.1), convert the multiplication operation into a shift operation, to facilitate implementation using the lookup table of the FPGA, thereby improving the resource utilization rate and accelerating the inference speed.
[0061] The specific implementation of square quantization is as follows: for a weight ω with a bit width of m, the quantization value scaling factor α multiplies the quantization series Map the original weight value to the m-th power of 2;
[0062] The quantization function is as follows:
[0063]
[0064] where the function h(·) maps the value range of the weight value from [-1, +1] to [0, 1];
[0065] The weight value 2 after square quantization b is equivalent to shifting a by b times when multiplied by the activation function value α, and the formula is as follows:
[0066]
[0067] S4: Design the core operation module YOLO Subsystem of the FPGA acceleration system.
[0068] (4.1) Sub-network modular IP design: Use the Vitis HLS tool to design the IP of each sub-network module using the strategy of trading space for time, including: Conv, C2f_Faster, SCDown, C2fCIB, SPPF, PSA, Upsample, Coneat_FN, Concat_FN3, BiFPN_Add3, and V10Detect, etc.
[0069] Each sub-network IP uses parametric design. Taking the Conv function module as an example, its parameters include in_sizes (input width), out_channels (number of output channels), in_channels (number of input channels), kernel_size (convolution kernel size), stride (convolution stride), padding (whether to pad), and use_activation (whether to use the activation function) as inputs, and control the convolution calculation process according to the parameters; the function data stream interfaces include input, output, weight_stream, bias_stream, bn_weights_stream, bn_mean_stream, and bn_var_stream, which are the input feature map data stream, output feature map data stream, weight data stream, bn weight data stream, bn average data stream, and bn normalized variance data stream respectively. The shape of the output feature map can be calculated according to the parameters. Assume the width and height of the input image are H in ×H in , the padding of the convolution kernel is P, the convolution kernel size is K, and the convolution stride is S, then the shape of the output feature map H out ×H out is:
[0070]
[0071] (4.2) YOLO Subsystem instantiation: Add the sub-network module IPs described in (4.1) to the Backbone function and the Head function in the order of the improved YOLOv10 network structure shown in Figure 2 , and further instantiate these two functions into the YOLO Subsystem module.
[0072] (4.3) Data flow allocation: Provide data flows for each sub-network module through the data flow allocation function, including input feature map data flow, weight data flow, normalized weight data flow, etc.
[0073] (4.4) Convolution calculation optimization: Optimize the convolution calculation process, expand the convolution process, and pipeline it to improve the calculation performance.
[0074] For the Conv convolution calculation function with the largest amount of computation, use the UNROLL command to partially expand the for loop of the convolution process, and use the #pragma PIPELINE pipeline to pipeline the for loop in the convolution process to improve the calculation performance. At the same time, use stream data flow type variables for weight data, output data, and feature map data, and the stream type uses the FIFO mode for caching.
[0075] (4.5) FPGA resource optimization: Utilize the LUT, DSP, and BRAM of the FPGA for dynamic resource allocation, optimize the FPGA resource usage efficiency through adaptive scheduling and parallel computing, reduce power consumption, and improve the system performance.
[0076] After the resource optimization is completed, package the YOLO Subsystem into an IP core and perform synthesis implementation.
[0077] S5: Design the FPGA acceleration part of the YOLOv10 on-orbit remote sensing image target detection network.
[0078] As Figure 3 shown, a YOLOv10 on-orbit remote sensing image target detection network FPGA acceleration system of the present invention includes modules such as AXI Bus, Merger, YOLO Subsystem, and Fetcher implemented on a single FPGA, and an external DDR3 memory;
[0079] The DDR3 memory is used to store the remote sensing images to be recognized and the YOLOv10 network weight data;
[0080] The MIG module is used to access the DDR3 memory through the AXI bus;
[0081] The AXI Bus is used to connect the Merger, Fetcher, and MIG modules;
[0082] The YOLO Subsystem module is used to implement the core computing function of the FPGA acceleration system for the YOLOv10 on-orbit remote sensing image target detection network;
[0083] The Fetcher module is used to allocate weight data streams for the YOLO Subsystem module according to its running state;
[0084] The Merger module decodes the detection boxes of the output results of the YOLO Subsystem to obtain the final target position, target category, and confidence level.
[0085] Use Vivado to add IPs of modules such as AXI Bus, Merger, YOLO Subsystem, and Fetcher, and make connections to design the FPGA acceleration system for the YOLOv10 on-orbit remote sensing image target detection network, and perform synthesis and implementation to generate the system configuration bitstream file.
[0086] S6: Build the FPGA acceleration system for the YOLOv10 on-orbit remote sensing image target detection network.
[0087] On the FPGA development board containing the DDR3 memory, download the configuration bitstream generated in (S5) to the FPGA chip, and detect the images to be detected stored in the DDR3 memory according to the Figure 4 shown process. The main steps are as follows:
[0088] (6.1) The Fetcher module, according to the running state of the YOLO Subsystem module, controls the MIG to read weight data and image feature map data from the off-chip memory DDR in advance, and transmits the data to the YOLO Subsystem module;
[0089] (6.2) The YOLO Subsystem module performs accelerated calculation on the YOLOv10 network;
[0090] (6.3) The Merger module decodes the detection boxes of the accelerated calculation results from the YOLO Subsystem to obtain the final target position, target category, and confidence level;
[0091] (6.4) The Merger module transmits the detection result information in (6.3) to the main control computer.
Claims
1. A YOLOv10 on-orbit remote sensing image ship target detection network optimization method, characterized in that: The following steps are involved: (1) Improve the YOLOv10 on-orbit remote sensing image target detection network model to reduce model complexity and improve the accuracy of small target detection: (1.1) Use FasterNet network module to replace BottleNeck module in the original C2f network to reduce the amount of network calculation and parameters and speed up the calculation process; (1.2) Use the LeakyReLU activation function to replace the original SiLU activation function; (1.3) Fast Normalized Fusion is used to improve the Concat network to obtain the Concat_FN and Concat_FN3 networks; Fast Normalized Fusion is to perform fast normalized fusion on i types of input feature maps, and the formula is as follows: Where O is the output feature map after fusion, w i is the learnable weight of the i-th input feature map, I i is the i-th input feature map, ε is a very small constant, ∑ j w j is the sum of all feature map weights; The number of input feature maps of the Concat_FN network is i=2, and the number of input feature maps of the Concat_FN3 network is i=3; (1.4) Use the Concat_FN network described in (1.3) to replace the Concat network of the 15th and 18th layers of the original YOLOv10; use the Concat_FN3 described in (1.3) to replace the Concat network of the 24th layer of the original YOLOv10, and extract feature maps from the networks of the 6th, 13th and 24th layers for weighted concatenation; to improve the detection accuracy of small targets; (2) Use the DOTA dataset to train the improved YOLOv10 network model described in (1.4) to obtain the optimal model weight data; (3) Perform layer fusion and weight square fixed-point quantization on the trained YOLOv10 network model described in (2) to improve hardware implementation efficiency and accelerate reasoning speed: (3.1) The convolution layer and the BN layer in the trained YOLOv10 network model described in (2) are fused to reduce the number of intermediate result caches and optimize the calculation process; the output of the convolution and BN layers after layer fusion is as follows: Among them: Y is the output after convolution BN fusion, x is the feature input, w is the convolution weight, b is the convolution bias, is the mean value of x, σ 2 is the variance of x, γ and β are learnable parameters, and ε is the minimum value that prevents the denominator from being zero; (3.2) The fused weights described in (3.1) are squared fixed-point quantized, and the multiplication operation is converted into a shift operation to facilitate the implementation using the FPGA lookup table, thereby improving the hardware implementation efficiency and accelerating the inference speed.
2. A YOLOv10 on-orbit remote sensing image target detection network FPGA acceleration system, characterized in that: Includes modules such as AXI Bus, Merger, YOLO Subsystem and Fetcher implemented on a single FPGA, as well as external DDR3 memory; The DDR3 memory is used to store the remote sensing image to be identified and the YOLOv10 network weight data; The MIG module is used to access the DDR3 memory through the AXI bus; The AXI Bus is used to connect the Merger, Fetcher and MIG modules; The YOLO Subsystem module is used to implement the core computing functions of the YOLOv10 on-orbit remote sensing image target detection network FPGA acceleration system; The Fetcher module is used to allocate a weight data stream to the YOLO Subsystem module according to its operating status; The Merger module decodes the detection frame of the output result of the YOLO Subsystem to obtain the final target position, target category and confidence.
3. A YOLOv10 on-orbit remote sensing image target detection network FPGA acceleration system, characterized in that: The specific implementation of the YOLOSubsystem is as follows: (1) Sub-network modular IP design: Use the VitisHLS tool to adopt the space-for-time strategy to design each sub-network module IP, including: Conv, C2f_Faster, SCDown, C2fCIB, SPPF, PSA, Upsample, Concat_FN, Concat_FN3, BiFPN_Add3 and V10Detect; (2) YOLO Subsystem instantiation: Add the subnetwork module IP described in (1) to the Backbone function and the Head function in the order of the Backbone and Head structures, and further instantiate these two functions into the YOLO Subsystem module; (3) Data flow allocation: The data flow allocation function is used to provide data flow for each sub-network module, including input feature map data flow, weight data flow, normalized weight data flow, etc. (4) Convolution calculation optimization: Optimize the convolution calculation process, expand the convolution process, and pipeline it to improve computing performance; (5) FPGA resource optimization: Utilize the FPGA's LUT, DSP, and BRAM to dynamically allocate resources, optimize FPGA resource utilization efficiency through adaptive scheduling and parallel computing, reduce power consumption, and improve system performance.