Lightweight neural network design and transplantation deployment method for remote sensing ship target detection

By adopting the R-YOLOv5 model based on the rotating target box and neural network search technology in remote sensing ship target detection, the model is optimized lightly, solving the problem of low operation efficiency of deep convolutional neural networks on embedded devices, and achieving efficient remote sensing ship target detection.

CN120164008APending Publication Date: 2025-06-17CHINA ACADEMY OF SPACE TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510137821.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing deep convolutional neural network cannot be effectively transplanted into devices with limited embedded computing resources in remote sensing ship target detection due to the large amount of parameters and calculations in the satellite target detection, resulting in the inability to realize real-time remote sensing services.

Method used

The R-YOLOv5 model based on the rotary target box is adopted, and the model backbone network is automatically and efficiently searched and lightweighted by the Cell space-based search method and gradient descent-based search strategy, reducing the amount of model parameters and calculations.

Benefits of technology

While maintaining the performance of the neural network model, it reduces the interference of human factors, obtains the best model parameters, realizes the lightweight design and transplant deployment of the model, and improves the speed and accuracy of object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164008A_ABST
    Figure CN120164008A_ABST
Patent Text Reader

Abstract

The invention relates to a lightweight neural network design and transplantation deployment method for remote sensing ship target detection. The method comprises the following steps: S1, designing an R-YOLOv5 remote sensing ship target detection model based on a rotating target frame; s2, performing automatic and efficient search on a backbone network of the R-YOLOv5 remote sensing ship target detection model based on a micro neural network search method of a Cell space and a search strategy based on gradient descent, and performing lightweight optimization on the model; and S3, carrying out operator adaptation and transplantation optimization deployment on the searched lightweight neural network model on an intelligent chip computing platform. The remote sensing ship target detection lightweight model design method is provided based on the deep neural network search technology, the method can be suitable for remote sensing ship rotating target detection, and interference of human factors is reduced and optimal model parameters are obtained while the neural network model performance is better maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent remote sensing applications, and in particular to a lightweight neural network design and transplantation deployment method for remote sensing ship target detection. Background Art

[0002] Remote sensing ship target detection technology can provide timely monitoring, early warning and intelligence support by monitoring ship activities in the ocean in real time, which is of great significance to the maintenance of China's marine rights and interests. It plays an irreplaceable role especially in aspects such as maritime security monitoring, prevention of marine accidents, maintenance of marine boundaries, search and rescue, and protection of marine resources.

[0003] Traditional remote sensing satellite data processing adopts the working mode of collecting data by satellite and then transmitting it to the ground for intelligent processing. This application mode has low timeliness and cannot meet the increasingly concerned real-time remote sensing service requirements of users. In recent years, China's space industry has developed vigorously, and the number of remote sensing satellites has been increasing continuously. On-orbit intelligent processing of remote sensing satellites has become an inevitable choice for the development of remote sensing satellites, and it has important strategic significance for enhancing China's global rapid perception and response capabilities.

[0004] Deep learning neural network technology has achieved great success in the task of remote sensing ship detection. In order to obtain better performance, the existing deep convolutional neural network has an increasingly complex neural network model and an increasingly large number of parameters. With the complexity, depth and size of the convolutional neural network model increasing exponentially, the number of parameters and the amount of calculation of typical network models are very large, resulting in these huge network models can only be used on ground cluster computing platforms and cannot be well transplanted to devices with limited on-board embedded computing resources for operation. Model lightweight and acceleration aim to reduce the number of parameters and the amount of calculation of the model to improve the inference speed of the model with minimal impact on the model performance. The lightweight design and transplantation deployment of deep neural network models have become one of the urgent problems to be solved in on-board remote sensing image intelligent processing.

[0005] The invention patent "Lightweight Method, System and Target Detection Method for Deep Convolutional Neural Network", application number CN113420651A. This invention proposes an effective lightweight scheme applicable to the FasterRCNN target detection framework. Aiming at the characteristics of the FasterRCNN network architecture, a deep convolutional neural network lightweight technology combining deep sparse low-rank and tensor TT decomposition theory is proposed. The lightweight method of deep sparse low-rank separable convolution is used to perform "layer-by-layer channel pruning, layer-by-layer retraining, and layer-by-layer tuning" on the feature extraction backbone network part of the FasterRCNN network. It realizes a high compression ratio of the target detection model and improves the speed and accuracy of target detection.

[0006] Patent for Invention "Training Method, System, Device and Medium for Lightweight Deep Neural Network", Application No. CN116187420A. The invention provides a training method, system, device and medium for a lightweight deep neural network. The method performs low-bit quantization on the floating-point network, deletes the layers and branches of the floating-point network corresponding to the N combined layers, and selects M combined layers among the N combined layers to delete the layers and branches of the floating-point network corresponding thereto, thereby realizing structural pruning and channel pruning of the floating-point network, and achieving multi-level quantization pruning processing based on the floating-point network, which is beneficial to improving the training efficiency.

[0007] Existing methods mainly adopt model compression methods such as multi-level quantization pruning processing based on floating-point networks and reducing memory usage by low-bit quantized weights. Using these methods not only requires algorithm designers to have rich professional domain knowledge, but also requires a large number of repeated test experiments to explore a large parameter space, weigh the model size, operation efficiency and model performance to obtain the best model parameters. Otherwise, the model accuracy will drop significantly and may not achieve the best effect in actual deployment, resulting in extremely high research and development costs and time costs for lightweight model development.

[0008] In addition, some existing model lightweighting methods are only effective for a certain type of object detection framework algorithm. For example, the lightweighting scheme for the Faster RCNN object detection framework is only applicable to the lightweight design of the Faster RCNN object detection framework model and does not support the lightweighting of the rotating box YOLO series remote sensing ship target detection model. Summary of the Invention

[0009] To solve the above technical problems existing in the prior art, the object of the present invention is to provide a method for designing and transplanting and deploying a lightweight neural network for remote sensing ship target detection. Based on the deep neural network search technology, a method for designing a lightweight model for remote sensing ship target detection is proposed, which can be applied to the detection of rotating targets of remote sensing ships, and while better maintaining the performance of the neural network model, reduces the interference of human factors and obtains the best model parameters.

[0010] To achieve the above object of the invention, the present invention provides a method for designing and transplanting and deploying a lightweight neural network for remote sensing ship target detection, including the following steps:

[0011] Step S1, design an R-YOLOv5 remote sensing ship target detection model based on a rotating target box;

[0012] Step S2, based on the differentiable neural network search method in the Cell space and the search strategy based on gradient descent, automatically and efficiently search the backbone network of the R-YOLOv5 remote sensing ship target detection model and perform lightweight optimization on the model;

[0013] Step S3: Perform operator adaptation, transplantation, and optimization deployment of the obtained lightweight neural network model on the intelligent chip computing platform.

[0014] According to a technical solution of the present invention, in the R-YOLOv5 remote sensing ship target detection model based on a rotated target box, a method based on five-parameter regression is adopted to achieve target box detection in any direction by adding an angle parameter θ. The angle parameter θ represents the acute angle between the x-axis and the adjacent side w, and the position of the rectangular rotation box is defined as (x, y, w, h, θ).

[0015] According to a technical solution of the present invention, the number of channels of the output structure of the Head of the R-YOLOv5 remote sensing ship target detection model is 3*(an + 5 + 10), where 3 represents 3 Anchors, an represents the encoding of the angle, and the loss function consists of the bbox regression loss L CIoU , the object confidence loss L obj , the class loss L cls , the angle loss L csl and is composed of:

[0016]

[0017] where K, S2, and B are the number of output feature maps, cells, and anchors on each cell respectively; α* is the weight of the corresponding term; represents whether the k-th output feature map, the i-th cell, and the j-th anchor box are positive samples. If it is a positive sample, it is 1, otherwise it is 0.

[0018] According to a technical solution of the present invention, in the step S2, it specifically includes:

[0019] Step S21: Determine the search space based on the Cell space;

[0020] Step S22: Implement a search strategy based on gradient descent;

[0021] Step S23: Screen the model based on the evaluation strategy to complete model lightweighting.

[0022] According to a technical solution of the present invention, in the step S22, implementing a search strategy based on gradient descent specifically includes:

[0023] Step S221: Mix the discrete candidate operations defined in the Cell and use Softmax to complete the continuous processing of the discrete search space;

[0024] Step S222: Perform double-layer optimization on the architecture parameters and weight parameters, and alternately update the architecture parameters and weight parameters using approximate iterative optimization steps;

[0025] Step S223: Discretize the searched better architecture, and replace continuous operations with the most likely discrete operations;

[0026] Step S223: Retrain the selected discrete architecture using the original dataset.

[0027] According to a technical solution of the present invention, in the step S3, it specifically includes:

[0028] Step S31: Convert the network model generated by the deep learning network search tool into a network model file in the standard pytorch deep learning framework;

[0029] Step S32: Develop and adapt custom operators on the intelligent chip computing platform, develop, package, and register for the rotation box operator, and integrate the custom operator with the framework;

[0030] Step S33: Quantize the model using the symmetric quantization algorithm, and perform online layer-by-layer inference or online fusion inference using the quantized pth weight file;

[0031] Step S34: After the online inference is correctly executed, use the operator fusion method to compile the network into a complete static graph and compile and export to generate an offline model file;

[0032] Step S35: Develop and deploy the offline inference program using optimization means such as multi-thread optimization, multi-core scheduling, and multi-card parallelism, and load and run the offline model.

[0033] According to a technical solution of the present invention, for the symmetric quantization algorithm, determine the initial scaling range according to the quantization bit width, and then calculate and adjust the scaling range according to the input data to obtain the quantization factor Q, and convert the input floating-point type data x into a fixed-point number V'x. The calculation formula is expressed as: V x ′ = Q × (V x - V min );

[0034] After the fixed-point calculation is completed, it can be converted back to a floating-point number through V x = Q -1 × V x ′ + V min

[0035] According to a technical solution of the present invention, the multi-thread optimization includes: using three threads to separately complete the pre-processing of input data, the inference on the intelligent chip computing platform side, and the post-processing task. The pre-processing thread puts the processed data into the pre-processing cache. The inference monitoring thread fetches data from the pre-processing cache, and after the inference is completed, puts the result into the post-processing cache. The post-processing thread fetches data from the post-processing cache for processing.

[0036] According to a technical solution of the present invention, the multi-card parallel computing includes: cutting the remote sensing image into multiple sub-images with sizes much smaller than the remote sensing image, and running the model on multiple intelligent acceleration cards through a distributed running program to process the parallel inference of a large amount of data.

[0037] According to an aspect of the present invention, there is provided a remote sensing ship target detection system, including:

[0038] A lightweight neural network model obtained based on the method for designing, transplanting, and deploying the lightweight neural network for remote sensing ship target detection as described in any one of the above technical solutions;

[0039] An intelligent chip computing platform for deploying and executing model inference.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] The present invention proposes a method for designing, transplanting, and deploying the lightweight neural network for remote sensing ship target detection. Based on the deep neural network search technology, a method for designing the lightweight model for remote sensing ship target detection is proposed, which can be applied to the detection of rotating targets of remote sensing ships. While better maintaining the performance of the neural network model, it reduces the interference of human factors. The present invention designs a complete set of methods for transplanting, optimizing, and deploying the model, which can transplant and deploy the lightweight neural network model generated by the deep learning network search tool on the domestic intelligent computing platform, effectively reducing the use and implementation costs of the neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 A flowchart schematically showing the method for designing, transplanting, and deploying the lightweight neural network for remote sensing ship target detection according to an embodiment of the present invention;

[0044] Figure 2Schematically showing a schematic diagram of the five parameters (x, y, w, h, θ) of the optical rectangular rotation target box according to an embodiment of the present invention;

[0045] Figure 3 Schematically showing a block diagram of a lightweight neural network search architecture according to an embodiment of the present invention;

[0046] Figure 4 Schematically showing a schematic diagram of the internal operation architecture of a Cell according to an embodiment of the present invention;

[0047] Figure 5 Schematically showing a schematic diagram of the candidate operation structure of a Cell according to an embodiment of the present invention;

[0048] Figure 6 Schematically showing a flow chart of model transplantation and deployment according to an embodiment of the present invention;

[0049] Figure 7 Schematically showing a flow chart of lightweight model transplantation and deployment according to an embodiment of the present invention. Detailed implementation manners

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] Such as Figure 1As shown in the figure, a lightweight neural network design and transplantation deployment method for remote sensing ship target detection according to the present invention aims to design a lightweight remote sensing ship target detection neural network model, transplant and deploy it to a domestic intelligent chip computing platform for operation, and form a complete technical solution for the overall design process of the domestic intelligent chip lightweight target detection model. The overall design solution of the present invention includes three parts. First, a R-YOLOv5 remote sensing ship target detection deep neural network structure based on a rotated bounding box is designed according to the characteristics of remote sensing ships with rotated close neighbors. Second, considering that the backbone network of the model has a stacking feature and accounts for the largest proportion of the parameters in the model, in order to realize the transplantation and deployment of the remote sensing ship target detection model on the edge platform, a differentiable neural network search method based on the Cell space and a search strategy based on gradient descent are designed to automatically and efficiently search the backbone network of the model based on the R-YOLOv5 remote sensing ship target detection model, optimize the lightweight of the model, improve the search efficiency, reduce the search cost, and realize the lightweight processing of the model. Finally, the obtained lightweight neural network model needs to be operator-adapted, transplanted, optimized and deployed on a domestic intelligent chip computing platform (such as the Cambrian neural network chip computing platform).

[0052] I. Remote Sensing Ship Rotated Target Detection and Recognition Algorithm

[0053] Remote sensing ship targets are often closely adjacent in many cases, especially near some ports along the shore, where two adjacent ships are usually parked together. For the target of remote sensing ships with the characteristic of arbitrary direction rotation, the original YOLOv5 series models use horizontal bounding boxes for target detection and output four information (x, y, w, h) for the positioning of the prediction bounding box, that is, the center coordinates and the length and width of the prediction bounding box. The horizontal bounding box detection will introduce a large amount of redundant background features and information of the same type of targets, and the boundary expression of the ship target is fuzzy, which is not only not conducive to the convergence during network training, but also prone to false alarms and affects the detection accuracy of the network. Therefore, in order to solve the problem of fine-grained remote sensing ship target detection and recognition, an intelligent detection and recognition algorithm for rotated close neighbor remote sensing ships needs to be designed. The present invention designs a R-YOLOv5 architecture on the basis of the original YOLOv5 architecture and introduces a rotated bounding box to perform more accurate characterization and learning of the target.

[0054] The rotation direction information of the target is introduced into the detection process, the positioning method of the anchor box is redesigned, and the structure of the R-YOLOv5 deep neural network based on the rotated bounding box (that is, the R-YOLOv5 remote sensing ship target detection model based on the rotated bounding box) is designed by transforming the detection head of the deep neural network. In the present invention, a method based on five-parameter regression is adopted to realize the detection of the bounding box in any direction by adding an additional angle parameter θ, where θ represents the acute angle between the x-axis and the adjacent side w. As Figure 2 shown in the figure, the position of the rectangular rotated bounding box is defined as (x, y, w, H, θ).

[0055] The output structure of the Head of the original YOLOv5 trained on the COCO dataset consists of three feature maps: 255*H*W, 255*2H*2W, and 255*4H*4W. The 255*H*W with the smallest size is responsible for detecting large targets, 255*2H*2W is responsible for detecting medium targets, and the 255*4H*4W with the largest size is responsible for detecting small targets. The number of channels 255 = 3*(5 + 80), where 3 represents 3 Anchors, 5 represents four position information (x, y, w, h) and a confidence (indicating the probability that there may be an object in the current grid), and 80 represents the probabilities of 80 categories in the COCO dataset. Considering that the number of categories in this project is 10 and rotation target detection is to be performed, after introducing angle information on the basis of the original YOLO output head, the number of channels of the Head output structure is 3*(an + 5 + 10), where 3 represents 3 Anchors and an represents the encoding of the angle.

[0056] The loss function consists of 4 parts, namely the bbox regression loss L CIoU 、the object confidence loss L obj 、the class loss L cls 、and the angle loss L csl ,so there is:

[0057]

[0058] Among them, K, S2, and B are the number of output feature maps, cells, and anchors on each cell respectively; α* is the weight of the corresponding item; represents whether the k-th output feature map, the i-th cell, and the j-th anchor box are positive samples. If it is a positive sample, it is 1, otherwise it is 0.

[0059] For the bbox regression loss, the traditional IOU_Loss mainly considers the overlapping area between the detection box and the target box. On the basis of IOU_Loss, GIOU_Loss solves the problem when the bounding boxes do not overlap. On the basis of IOU_Loss and GIOU_Loss, DIOU_Loss considers the information of the distance between the center points of the bounding boxes. And CIOU_Loss further considers the scale information of the aspect ratio of the predicted box and the target box on the basis of DIOU_Loss, and can obtain better optimization performance. Therefore, CIOU_Loss is adopted in the bbox regression loss.

[0060] II. Remote Sensing Ship Target Detection Lightweight Model Search Strategy

[0061] The present invention automatically searches for the backbone network of the remote sensing ship rotation target detector through neural network search technology. By using the differentiable architecture search technology based on cells, that is, selecting the cell-based search space, the gradient descent-based search strategy, and the low-fidelity evaluation strategy to complete the search design of the backbone network, the search efficiency can be improved, the search cost can be reduced, and the lightweight design of the remote sensing ship target detection model can be realized.

[0062] The present invention selects the differentiable neural network search technology based on cells to complete the lightweight design of the backbone network of the remote sensing ship rotation target detection model (remote sensing ship target detection model), and uses the Neck and Head structures of the R-YOLOv5 model as the backend part of the backbone network. The macro architecture of the backbone network is based on the R-YOLOv5 backbone network and is designed as a stack of multiple cells. The lightweight neural network architecture search scheme involves three aspects: search space, search strategy, and performance evaluation strategy. The search scheme architecture is as Figure 3 shown.

[0063] (1) Search space design

[0064] The present invention uses neural network search technology to search for the search space of the remote sensing ship rotation target detection model. The search space adopts the cell-based search space technology. The inside of the cell is regarded as a super network, and each search samples a sub-network from the super network to form a complete candidate architecture for performance evaluation. The search space directly affects the difficulty of network structure optimization. It is necessary to construct an efficient target detection search space. The specific process of determining the search space based on the cell space includes: constructing a suitable search space according to the characteristics of the data set and the information of the existing target detection backbone network features; selecting suitable candidate operations according to the characteristics of the data set. For example, different types of convolutional layers such as ordinary convolutions or depthwise separable convolutions of different sizes can be included in the search space; according to the analysis of the existing target detection backbone network features, select suitable network structure candidates, including network depth, number of kernels, etc., to control the network size and the number of parameters.

[0065] The Cell design of the search space of the present invention draws on the design idea of R-YOLOv5, which can ensure lightweight while maintaining accuracy, and reduce the computational bottleneck and memory cost. In the Cell design of the present invention, the gradient flow is divided into two parts, and at the same time, a connection method similar to the direct connection in Resnet is adopted, so that the gradient information of different layers is effectively fused. Through operations such as switching connections and conversions, the diversity of gradient information at different levels in the network is made richer. One of the gradient flows contains a residual structure stacked N times, and the residual structure can solve the problem of network degradation caused by the increase in the number of network layers, making the network construction deeper. The stacking times of the residual blocks are jointly determined by the search strategy and the predefined template. The schematic diagram of the internal operation architecture of the Cell is as Figure 4 shown.

[0066] Convert the schematic diagram of the internal operation architecture of the Cell into a standard schematic diagram of the internal structure of the Cell with edges as candidate operations and nodes as feature maps, as Figure 5 shown. In the Cell designed in the present invention, there is a total of 1 input node, 1 output node and 6 internal nodes.

[0067] In the present invention, the edges that need to search for the operation structure are (0,1), (0,2), (2,3), (3,4), (6,7). Among them, node 5 represents the feature map obtained by performing an Add operation on the feature maps of node 2 and node 4, and node 6 represents the feature map obtained by performing a Concat operation on the feature map of node 5 and the feature map of node 1. The candidate operation set in the lightweight neural architecture search space for remote sensing ship rotation target detection includes conventional convolutions of different sizes, depthwise separable convolutions, CBS modules, CBL modules, CBR modules, etc. Among them, the CBS module is a serial structure of convolution Conv, batch normalization BatchNorm and activation function SiLU, the CBL module is a serial structure of convolution Conv, batch normalization BarchNorm, activation function LeakyReLU, and the CBR module is a serial structure of convolution Conv, batch normalization BarchNorm, activation function ReLU. Considering the lightweight design, the Cell can search for the stacking number of the residual structure. In R-YOLOv5, the stacking number of the residual structure and the number of channels are respectively controlled by two parameters, the network depth coefficient and the network width coefficient. The search range of the network depth coefficient is set to (0.10, 1.00), and the search range of the network width coefficient is set to (0.10, 1.00), so as to improve the performance and efficiency of the searched neural architecture. The candidate parameter table of the search space is shown in Table 1 below.

[0068] Table 1 Candidate Parameter Table of the Search Space

[0069]

[0070] To achieve lightweight search of the architecture, a 1×1 Conv structure and a lightweight operator, spatial separable convolution, are added to the search space. Spatial Separable Convolutions decomposes the convolution kernel into several small convolution kernels, that is, an n×n convolution is calculated in two steps of 1×n and n×1, thereby saving computational costs. Similar to spatial separable convolution is depthwise separable convolution. The core idea is to decompose a complete convolution operation into two steps, namely depthwise convolution and pointwise convolution. The calculation of depthwise convolution is very simple. It uses a convolution kernel for each channel of the input feature map, and then stitches together the outputs of all convolution kernels to obtain its final output. Since the number of output channels of the convolution operation is equal to the number of convolution kernels, and only one convolution kernel is used for each channel in depthwise convolution, the number of output channels of a single channel after the convolution operation is also 1. Then, if the number of channels of the input feature map is N, after using a convolution kernel for each of the N channels separately, N feature maps with 1 channel are obtained. These N feature maps are then stitched together in sequence to obtain an output feature map with N channels. The operation of pointwise convolution is very similar to that of a conventional convolution operation. The size of its convolution kernel is 1×1×M, where M is the number of channels of the previous layer. Therefore, the convolution operation here will perform a weighted combination of the previous map in the depth direction to generate a new feature map. There are as many output feature maps as there are convolution kernels.

[0071] The various convolutions in the search space designed by the present invention include conventional convolutions, depthwise separable convolutions, and spatial separable convolutions of different sizes. Pooling includes max pooling, etc. The search for the Neck or Head structure is reflected in the network depth coefficient and the network width coefficient. The candidate operations in the search space designed by the present invention can not only effectively extract data features, but also minimize computational costs and accelerate model training and inference.

[0072] (2) Search Strategy Design

[0073] The search strategy defines how to find the optimal network structure in the search space, which is essentially a hyperparameter iterative optimization problem. The search strategy of the present invention adopts a search strategy based on gradient descent. The specific process is search space continuousization, gradient descent optimization, architecture discretization, and retraining.

[0074] 1) Search space continuity: Mix the discrete candidate operations defined in the Cell and use Softmax to complete the continuity processing of the discrete search space. The operation o at a specific position (i, j) is jointly represented by the candidate operation set and the architecture parameters, as shown in the following formula. The continuity of the search space transforms the original discrete search space into a continuous search space, providing a basic condition for using gradient descent to accelerate the optimization of the objective function.

[0075]

[0076] Among them, is a set of candidate operation sets (such as convolution, max pooling), is the architecture weight in the mixed edge, and the edge (i, j) is the set of candidate operations o of the conversion node (i,j) .

[0077] 2) Double-layer gradient optimization: Perform double-layer optimization on the architecture parameters and weight parameters. Regard the entire internal structure of the Cell as a super network, perform gradient descent optimization on all operations together, and all operations share the weight parameters to accelerate the training process. Gradient descent optimization uses the training loss function on the training set to optimize the architecture parameter weights, and optimizes the architecture encoding through the loss function on the validation set. The optimization objective function is shown in the following formula.

[0078]

[0079] Among them, represents the training loss, and represents the validation loss: These two losses are jointly determined by the architecture encoding α and the weight w; the goal of architecture search is to find α that minimizes the validation loss ; the architecture weight * is obtained by minimizing the training loss .

[0080] The above double-layer gradient optimization has a strict order. In order to enable both to achieve the optimization strategy simultaneously, the design scheme of the present invention is: adopt an approximate iterative optimization step to alternately update the two parameters. First, fix the architecture parameter α in the training set to optimize the weight parameter w, and then fix w in the validation set to update the network structure parameter α, and repeat until both values are ideal. When using the differentiable neural architecture search technology to solve the double-layer optimization problem, alternately optimize the architecture parameters and network weights, with a complexity of Ο(|ω|*|α|), and the computational cost is extremely high. The present invention uses a one-step approximate optimization method commonly used in meta-learning to approximately optimize the objective function, and obtains the approximate objective function as shown in the following formula:

[0081] ​

[0082] This approximation method adjusts ω by using a single step of training, so that the result is close to ω * In the differentiable neural architecture search, the algorithm trains the supernet for 50 epochs. In each epoch, ω is first fixed, and the architecture parameter α is updated through a complete forward and backward propagation. This architecture parameter is used as the approximate optimization value, and α is fixed on this basis to optimize w.

[0083] 3) Architecture discretization: Discretize the optimal architecture after the search, that is, replace the continuous operation with the most likely discrete operation. After the gradient descent optimization reaches the optimal point, the differentiable architecture search technology completes the training of all sub-networks in the supernet and the architecture performance evaluation. At this time, the operations corresponding to the optimal sub-network can be discretized according to the architecture weight α to obtain a discrete architecture. That is, replace the continuous operation with the most likely discrete operation. As shown below:

[0084]

[0085] 4) Retraining: The discretized architecture obtained by screening needs to be retrained. After obtaining the optimal discretized architecture, the optimal architecture needs to be retrained using the original data set.

[0086] (3) Evaluation strategy design

[0087] Since Cell is designed as a supernet with weight sharing features, the one-time training of weights is completed by optimizing the entire supernet. Each search samples a subnet with weights from the supernet. The performance evaluation strategy only needs to perform a low-fidelity evaluation on the candidate network to obtain an estimate of the performance of the target detector composed of the subnet. Low-fidelity evaluation cannot obtain the final performance of the architecture to be tested, but its performance ranking can support the screening of neural architectures that meet the requirements.

[0088] The one-time training of the Cell supernet will be conducted on the dataset selected for this topic. Before the search begins, the dataset is split into a training set and a validation set at a ratio of 7:3, and the Cell supernet training is conducted on this basis. The dataset is a remote sensing ship target detection dataset, and its characteristics basically meet the target dataset. The dataset has 10 categories. The dataset is annotated with a directed target bounding box. The annotation information includes the source of the data image, the ground sampling distance, the four-point coordinates of the bounding box, the category to which the target belongs, etc.

[0089] The evaluation of the search algorithm requires constraint optimization of the speed and accuracy of the candidate model to achieve the lightweight optimization goal. In multi-objective optimization techniques, the balance factor between accuracy and speed is an important parameter, which determines the weight allocation for the two objectives, and this balance factor can be adjusted according to specific requirements and application scenarios. This solution introduces a weight factor, regards accuracy and speed as two independent objectives, and then controls the degree of emphasis of the optimization algorithm on the two objectives by adjusting these weight factors. In this way, the trade-off between accuracy and speed can be balanced according to actual needs. The present invention combines the detection accuracy and speed of the remote sensing ship rotation target detection model to jointly guide the lightweight architecture search. The inference speed of the model is T c , the inference speed of the candidate architecture is T h , and regard their ratio as θ = T h / T c . When the inference speed of the candidate architecture is better than that of the baseline model, in the optimization goal is rewarded, that is, makes the gradient descent faster, so as to accelerate the inference of the searched model. Optimization objective function:

[0090]

[0091] The low-fidelity evaluation of the present invention uses an incomplete network and a short training time to complete the performance evaluation of the candidate architecture. The searched Cell structure is stacked four times according to the configuration of the lowest network width and depth to form the candidate backbone network Backbone. After adding the model backend Neck and Head, a candidate object detector is formed, and a low-fidelity estimation of the candidate architecture is carried out. After the Cell supernetwork training is completed, the search strategy completes the architecture discretization, so as to obtain the optimal architecture and retrain it.

[0092] III. Method for transplanting and deploying lightweight model for remote sensing ship target detection

[0093] In order to be able to transplant and deploy the deep neural network model generated by the deep learning network search tool on the domestic intelligent computing platform, the present invention designs a complete model transplanting and deploying method based on the Cambrian neural network chip computing platform. The specific model transplanting and deploying flow chart is as Figure 6 shown, including:

[0094] Step 1: First, the network model generated by the deep learning network search tool needs to be converted into a network model file in the standard pytorch deep learning framework;

[0095] Step 2: Then, custom operator development and adaptation are carried out on the domestic computing platform. In the present invention, for the development, encapsulation, and registration of the rotation box operator, the custom operator is integrated with the framework;

[0096] Step 3: Quantize the model using the symmetric quantization algorithm, and perform online layer-by-layer inference or online fusion inference using the quantized pth weight file;

[0097] Step 4: After the online inference is correctly executed, use the operator fusion method to compile the network into a complete static graph and compile and export to generate an offline model file;

[0098] Step 5: Finally, use optimization means such as multi-thread optimization, multi-core scheduling, and multi-card parallelism to develop and deploy the offline inference program, and load and run the offline model.

[0099] 1) Design of the quantization method for the domestic intelligent chip computing platform: Deep learning networks generally contain many dense floating-point operations, such as convolution and fully connected. To avoid the operations of costly floating-point arithmetic units, the common approach is to adopt quantization technology and use fixed-point arithmetic units to complete this function, while reducing the calculation time and power consumption. When performing model quantization, the quantization bit width needs to be determined first, and the quantization factor size is determined by this bit width and the input data. The algorithm for the entire model quantization process is as Figure 7 shown.

[0100] The initialization scaling range refers to determining an initial mapping range according to the quantization bit width. INT8 corresponds to (-127, 128), and INT16 corresponds to (-32768, 32767). Then, adjust this scaling range according to the input data to make the mapping more accurate. Finally, calculate the quantization factor Q = R / S according to the formula of the symmetric quantization algorithm, where R is the scaling range and S is 2 to the power of x, and x is the quantization bit width. After obtaining the quantization factor Q, the input floating-point type data x can be converted into a fixed-point number V x ′ , and the calculation process can be defined by the following formula:

[0101] V x ′ = Q × (V x - V min )

[0102] After the fixed-point calculation is completed, if it is necessary to convert V x ′ back to a floating-point number, only the corresponding inverse operation needs to be performed to restore it. The calculation process is as follows:

[0103] V x = Q -1 × V x ′ + V min

[0104] As can be seen from the flow chart, there will be certain errors after the floating-point numbers are quantized and restored. However, such errors will not have a significant impact on the final result. For operations that support multiple different input data types, without significantly affecting the algorithm accuracy, by combining different quantization bit widths and reducing the accuracy of the input data, the overall throughput of the algorithm can be improved.

[0105] 2) Multi-thread optimization: On the domestic intelligent chip computing platform, the lightweight algorithm inference process for remote sensing ship rotating target detection includes three processes, namely pre-processing of input data, inference on the domestic intelligent chip computing platform side, and post-processing of results. Using multi-thread technology, three threads are respectively used to complete these three types of tasks. After the pre-processing thread completes the task, it puts the processed data into the pre-processing cache. The thread responsible for monitoring the inference on the intelligent computing platform first takes the pre-processed data from the pre-processing cache. After finding that the inference is completed, it puts the inference result into the post-processing cache. After the post-processing thread obtains the data in the post-processing cache, it completes the post-processing of the results. For the target detection algorithm, the pre-processing and post-processing take relatively short time, while the inference time on the intelligent computing platform has an order of magnitude difference from the former. Ideally, these three types of tasks will form a multi-thread pipeline. When the number of executed tasks is large enough, the time consumed by the pre-processing and post-processing of the target detection algorithm will be masked by the inference time, that is, the time consumed for a complete execution of the target detection process of the target detection algorithm is the inference time on the domestic intelligent chip computing platform.

[0106] 3) Multi-card parallel computing: Since the size of the remote sensing image is large, usually the method of cutting a large image of the original size into a series of sub-images for processing is adopted. A large image of 5000*5000 can be cut into hundreds of sub-images of 800*800. A single domestic intelligent acceleration card cannot bear the parallel inference of such a large amount of data, so it is necessary to use the multi-card parallel method to run the model on multiple domestic intelligent acceleration cards. By running the program distributively, the advantages of multiple cards can be fully utilized and the performance of each intelligent acceleration card can be fully exerted.

[0107] According to one aspect of the present invention, there is provided a remote sensing ship target detection system, including:

[0108] A lightweight neural network model obtained based on the lightweight neural network design and transplantation deployment method for remote sensing ship target detection as described in any one of the above technical solutions;

[0109] An intelligent chip computing platform, such as the Cambrian neural network chip computing platform, for deploying and executing model inference.

[0110] It should be noted that in the present invention, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.

[0111] It should also be noted that the above is the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those skilled in the art of this technology, once the basic creative concept of the present invention is known, without departing from the principle described in the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as within the protection scope of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

Claims

1. A lightweight neural network design and transplantation deployment method for remote sensing ship target detection, characterized in that: The following steps are involved: Step S1, designing an R-YOLOv5 remote sensing ship target detection model based on a rotating target frame; Step S2: Based on the Cell space differentiable neural network search method and the gradient descent-based search strategy, the backbone network of the R-YOLOv5 remote sensing ship target detection model is automatically and efficiently searched, and the model is lightweight optimized; Step S3: Operator adaptation and transplantation optimization deployment of the lightweight neural network model obtained by the search are performed on the intelligent chip computing platform.

2. The lightweight neural network design and transplantation deployment method for remote sensing ship target detection according to claim 1 is characterized in that: In the R-YOLOv5 remote sensing ship target detection model based on rotated target frame, a five-parameter regression-based method is adopted to realize target frame detection in any direction by adding the angle parameter θ. The angle parameter θ represents the acute angle between the x-axis and the neighboring edge w. The position of the rectangular rotated frame is defined as (x, y, w, h, θ).

3. The lightweight neural network design and transplantation deployment method for remote sensing ship target detection according to claim 2 is characterized in that: The number of channels of the Head output structure of the R-YOLOv5 remote sensing ship target detection model is 3*(an+5+10), where 3 represents 3 anchors, an represents the encoding of the angle, and the loss function is composed of the bbox regression loss L CIoU , target confidence loss L obj , category loss L cls , Angle loss L csl Composition, then: Among them, K, S2, B are the number of output feature maps, cells and anchors on each cell respectively; α* is the weight of the corresponding item; Indicates whether the k-th output feature map, the i-th cell, and the j-th anchorbox are positive samples. If they are positive samples, it is 1, otherwise it is 0.

4. The lightweight neural network design and transplantation deployment method for remote sensing ship target detection according to claim 1 is characterized in that: The step S2 specifically includes: Step S21, determining a search space based on the Cell space; Step S22: implementing a search strategy based on gradient descent; Step S23: Filter the model based on the evaluation strategy to complete the model lightweighting.

5. The lightweight neural network design and transplantation deployment method for remote sensing ship target detection according to claim 4 is characterized in that: In step S22, a search strategy based on gradient descent is implemented, which specifically includes: Step S221, mix the discrete candidate operations defined in Cell, and use Softmax to complete the continuous processing of the discrete search space; Step S222, performing double-layer optimization on the architecture parameters and the weight parameters, and using an approximate iterative optimization step to alternately update the architecture parameters and the weight parameters; Step S223, discretize the searched optimal architecture, and use the most likely discrete operation to replace the continuous operation; Step S223: retrain the screened discrete architecture using the original data set.

6. The lightweight neural network design and transplantation deployment method for remote sensing ship target detection according to claim 1 is characterized in that: The step S3 specifically includes: Step S31, converting the network model generated by the deep learning network search tool into a standard pytorch deep learning framework network model file; Step S32: Develop and adapt custom operators on the smart chip computing platform, develop, package, and register the rotating frame operator, and integrate the custom operator with the framework; Step S33: quantize the model using a symmetric quantization algorithm, and use the quantized pth weight file to perform online layer-by-layer reasoning or online fusion reasoning; Step S34: After the online reasoning is correctly executed, the network is compiled into a complete static graph using operator fusion and compiled and exported to generate an offline model file; Step S35: Develop and deploy an offline reasoning program using multi-threaded optimization, multi-core scheduling, and multi-card parallel optimization methods, and load the offline model to run.

7. The lightweight neural network design and transplantation deployment method for remote sensing ship target detection according to claim 6 is characterized in that: The symmetric quantization algorithm determines the initialization scaling range according to the quantization bit width, and then calculates and adjusts the scaling range according to the input data to obtain the quantization factor Q, converts the input floating-point data x into a fixed-point number V'x, and the calculation formula is expressed as: V x ′ =Q×(V x -V min ); After the fixed-point calculation is completed, the V x =Q -1 ×V x ′ +V min Convert it back to floating point.

8. The lightweight neural network design and transplantation deployment method for remote sensing ship target detection according to claim 6 is characterized in that: The multi-threaded optimization includes: using three threads to complete input data pre-processing, reasoning on the smart chip computing platform, and result post-processing tasks respectively, the pre-processing thread puts the processed data into the pre-processing cache, the reasoning monitoring thread obtains data from the pre-processing cache, and after the reasoning is completed, the result is placed in the post-processing cache, and the post-processing thread obtains data from the post-processing cache for processing.

9. The lightweight neural network design and transplantation deployment method for remote sensing ship target detection according to claim 6 is characterized in that: The multi-card parallel computing includes: cutting the remote sensing image into multiple sub-images with sizes much smaller than the remote sensing image, running the model on multiple smart acceleration cards through a distributed running program to process parallel reasoning of large amounts of data.

10. A remote sensing ship target detection system, characterized in that: include: A lightweight neural network model obtained based on the lightweight neural network design and transplantation deployment method for remote sensing ship target detection according to any one of claims 1 to 9; Intelligent chip computing platform for deploying and executing model reasoning.

Citation Information

Patent Citations

  • Lightweight deep neural network training method, system, device and medium

    CN116187420A