Lightweight remote sensing image multi-task target sensing method
By designing a lightweight remote sensing image multi-task target perception method on the domestic computing power platform, using improved neural network architecture and multi-core inference acceleration program, the problem of poor perception of small and medium-sized targets and dense targets in drone image perception is solved, and efficient and accurate multi-task perception capabilities are achieved.
Patent Information
- Application Number
- CN202510022299.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-09
AI Technical Summary
It is difficult for the existing technology to achieve high sensitivity perception of small and dense targets on domestic computing power platforms, and the drone image perception faces problems such as high altitude image resolution, large changes in target scales, challenges in the natural environment, complex target shapes, diverse target postures, and serious background interference.
A lightweight remote sensing image multi-task target perception method is designed, using the improved Shufflenetv2 backbone network, inverted pyramid neck structure and lightweight detection head, combined with multi-core inference acceleration program and edge collaborative inference deployment tool, to realize the efficient deployment of multi-task perception model on the Rockchip Micro 3588 chip.
It has achieved high sensitivity detection of key targets such as people and cars on the domestic computing power platform and high-precision segmentation of buildings and bridges, improved the drone target detection capabilities and reduced the hardware computing power demand.
Smart Images

Figure CN119964050A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and computer vision, and in particular to a real-time image perception method for a lightweight multi-task unmanned aerial vehicle. Background Art
[0002] At present, it has become the norm for drones to carry sensor equipment and intelligent computing processing boards. They can run intelligent algorithms to provide effective support for maritime traffic, fishery management, fire supervision, danger perception, etc. However, the intelligent algorithms used for different tasks basically run independently, and most of them run on foreign computing platforms. In addition, compared with natural scene image perception, drone image perception has some special challenges. UAV high-altitude images have high resolution, large changes in target scale, natural environment challenges, complex target shapes, diverse target postures, and serious background interference.
[0003] Most current perception algorithms have poor perception effects on small and densely packed targets, are sensitive to target scale and shape, and are not optimized for deployed equipment. For example, when performing target detection on key targets such as people and vehicles, and image segmentation tasks on buildings and bridges, most existing technologies require the use of multiple models to achieve this.
[0004] Therefore, how to achieve high-sensitivity perception of small and dense targets on the domestic computing power platform (with Rockchip 3588 as the computing power chip) has become a technical problem that needs to be urgently solved by the builders of domestic computing power platforms. Summary of the invention
[0005] In response to the above technical problems, the present invention provides a lightweight remote sensing image multi-task target perception method, which can detect key targets such as people and vehicles, and can perform image segmentation tasks on buildings and bridges. It also optimizes the deployment of drone image data and Rockchip 3588 airborne terminal equipment, and innovatively designs a lightweight multi-task drone real-time perception algorithm, which reduces the hardware computing power requirements and improves the target detection capability of high-altitude and long-flight drones. Through the edge collaborative reasoning deployment tool, the terminal algorithm model can be quickly deployed and iterated.
[0006] The technical solution of the present invention is: a lightweight remote sensing image multi-task target perception method, including a board of a Rockchip 3588 chip deployed with an inference program, including the following steps:
[0007] S1. Collect remote sensing images, perform detection and segmentation annotation,
[0008] Using drones to collect remote sensing image data, performing detection and segmentation annotation on the remote sensing images, and obtaining training sets, test sets, and validation sets;
[0009] S2, training multi-task perception model,
[0010] The multi-task perception model is trained using a constructed neural network, wherein the constructed neural network includes a backbone network, a neck structure, a head structure, and a segmentation structure.
[0011] Then, the multi-task perception model is trained using the training set obtained in step S1 and appropriate training parameters to obtain an optimal perception model in pt format, and the optimal perception model in pt format is converted and optimized to obtain an onnx model;
[0012] S3. Use the edge collaborative reasoning tool to convert the onnx model obtained by training conversion into the Rockchip 3588 model to obtain the RKNN model, and deploy the RKNN model and reasoning program to the board of the Rockchip 3588 chip;
[0013] S4. When the inference program is running on the drone side, the data input uses the mpp module to decode the visible light camera input data, and calls the multi-core NPU to infer the input video to obtain the inference result, and then post-processes the inference result;
[0014] S5, detection and segmentation results are displayed in superposition.
[0015] The detection and segmentation results in S45 are marked on the current frame using opencv to automatically overlay the result data. Specifically, the detection results are marked with a rectangular box and target category, and the segmentation results are marked with different colors of the region.
[0016] S6, push the stream to output the superimposed result video,
[0017] The automatic superposition result data obtained in step S5 is used to create a pipeline using the Gstream library, and the automatic superposition result data is cyclically read to superimpose the image to create an RTSP stream to obtain a video stream;
[0018] Further, in step S1, a data set is prepared for the remote sensing image, and data enhancement processing is performed thereon. Specifically, a data set is prepared for the remote sensing image, and data enhancement processing is performed thereon, including the following steps:
[0019] S11, the remote sensing image includes a public data set, actual flight collection data and map data, and high-altitude image data with four types of targets, namely, people, vehicles, bridges and roads, are obtained from the remote sensing image at different angles and at an altitude of 100m-800m;
[0020] S12, using annotation software to mark the people and vehicle targets in the high-altitude image data with rectangular frames, and at the same time mark the category of each target to achieve the detection annotation, and segment and mark the roads and bridges to achieve the segmentation annotation to obtain training data;
[0021] S13, enhancing the training data by rotation, cropping and data augmentation to obtain a training data set;
[0022] S14. Divide the training data set into a training set, a test set and a validation set in a ratio of 8:1:1, wherein the training set is used for training the model, and the test set is used for evaluating the model.
[0023] Furthermore, in step S13, the enhancement method of scaling and splicing is introduced to further detect small targets.
[0024] Furthermore, the backbone network: uses an improved Shufflenetv2 as the backbone network, specifically, uses a convolution network with a large 6×6 convolution kernel and a step size of 2 for downsampling at the beginning of the input downsampling, which has a small improvement in detection accuracy, changes the activation function in the entire network to ReLU6, adds an spp module after the output stage5 module to form a new feature extraction network, crops the 1×1 convolution structure in the downsampling in Shufflenetv2 and the ordinary convolution block structure, uses a reduced hidden layer channel and modifies the SPPF module of spatial pyramid pooling;
[0025] The neck structure: The neck structure adopts an inverted pyramid structure, and features of different scales are spliced and fused from bottom to top. Finally, the feature map generated by the spp module is dimensionally upgraded and connected to the CBR module of convolution + BN + ReLU after the stage3 and stage5 feature maps for dimension reduction and splicing with the final feature map. The backbone and neck enhance the model representation capability while improving the model reasoning speed. In order to accelerate the reasoning speed, different convolution modes are used in the blocks during training and reasoning.
[0026] The head structure: a lighter detection head is designed, the structure includes three detection heads, and a 3×3 convolution is used to calculate the target box loss, category loss and confidence loss;
[0027] The segmentation structure: in the segmentation module, after obtaining the features output by FPN, a segmentation head is constructed, a 3×3 convolution layer is used, and a 1×1 convolution layer is used in the last layer, and a mask coefficient is added.
[0028] Furthermore, the multi-task perception model in step S2 is trained, including the following steps:
[0029] S21. Train the target detection network. The training framework is Pytorch, the stochastic gradient descent method is used, and the optimizer is Adam.
[0030] S22, the training batch size is 32, the learning rate is 0.0001, and the cosine annealing function is used to dynamically update the learning rate;
[0031] S23. Use Dropout in the backbone network architecture, specifically, except for the first convolution layer, the rest use 10%;
[0032] S24. In terms of detection loss function, classification loss and regression loss use Varifocal Loss and GIou Loss. Adding DFL Loss enables the network to quickly focus on the value near the marked position and improve the probability. Among them, VarifocalLoss:
[0033]
[0034] p is the predicted IACS score, q is the target IoU score, for the positive samples in training, q is set to the IoU between the generated bbox and gt box, and for the negative samples in training, the training target q of all categories is 0, α is a variable parameter, and α is taken as 0.75 in this project.
[0035] DFL Loss:
[0036]
[0037] Among them, y is the regression label;
[0038] In the discrete range [y0,y n ] within {y0, y1,...,y n}, estimated value for
[0039]
[0040] Where P is the probability distribution,
[0041] Then the DFL loss can be expressed as follows
[0042] DFL(P(y i ),P(y i+1 ))=-((y i+1 -y)log(P(y i ))+(yy i )log(P(y i+1 )))
[0043] i is any integer,
[0044] The total loss is the weighted sum of the three losses as follows:
[0045]
[0046] α is 0.001, β is 0.0005, γ is 0.001, t, N~j, e, is the normalization of the target score,
[0047] In the segmentation task, the loss function uses the cross entropy loss function as shown below
[0048]
[0049] Where x is the activation value, N is the feature dimension, label is the label corresponding to the segmentation task, ranging from (0 to C-1), where C is the number of categories;
[0050] S25. Convert the optimal pt model file trained by the Pytorch framework into an onnx file;
[0051] S26, trim the post-processing of detection and segmentation tasks in onnx and use onnxsim to merge and optimize the trimmed onnx model;
[0052] S27. Use onnx-simplifier to perform non-identity operator deletion on the obtained onnx model.
[0053] Furthermore, in step S3, the edge collaborative reasoning tool is used to convert the optimal detection onnx model into the Rockchip 3588 model, and the specific steps of deploying the converted model and reasoning program to the terminal device are as follows:
[0054] S31. The edge collaborative reasoning tool consists of two parts: one is the end-side intelligent agent, and the other is the client.
[0055] The end-side intelligent agent includes the inference deployment package receiving and decompression service and the inference program start and stop control service components, which are mainly deployed in the RK3588 embedded end-side device;
[0056] The client includes model quantization and conversion, model publishing and model updating, inference service start and stop, and model evaluation functions, as well as model warehouse, inference program warehouse, and data warehouse, and is mainly deployed on PCs or servers;
[0057] S32. Use the model quantization and conversion module to convert the optimal detection onnx model into a model in rknn format. During the conversion process, turn on the quantization mode, and finally convert the onnx model file into an INT8 rknn file. The model quantization and conversion encapsulates the export_rknn interface of the RKNN-Toolkit2 tool. RKNN-Toolkit2 is a development kit for model conversion, reasoning, and performance evaluation provided to users.
[0058] S33, storing the converted rknn model file in the model warehouse;
[0059] S34, storing the inference program in the inference program warehouse, where the inference program must include a service startup script;
[0060] S35: When publishing or updating a model, select the rknn model file and inference program to be deployed. The model publishing and model updating modules package the rknn model file and the inference program, and send a model publishing request to the inference deployment service. After receiving the deployment package, the inference deployment service decompresses it and sends an inference program start request to the inference program start and stop control service.
[0061] S36: After receiving the request, the inference service start-stop control service starts the model inference program and generates a new model inference service;
[0062] S37. In the inference service startup module, the start and stop of the deployed model inference service can be controlled;
[0063] S38. In the model evaluation module, evaluation image data and the real labels corresponding to the images can be selected from the data warehouse and sent to the model inference service for reasoning to obtain annotation information. The model evaluation module can obtain the accuracy evaluation index of the model by comparing the real labels and the annotation information generated by model reasoning, and the performance evaluation index of the model can be obtained according to the reasoning time returned by the model reasoning service.
[0064] Furthermore, in step S4, when the model is running on the drone side, the mpp module is used to decode the visible light camera input data, and the multi-core NPU is called to infer the input video to obtain the inference result, and then the inference result is post-processed, including the following steps:
[0065] S41, read the video stream from the camera, and decode it using the hard decoding mpp module of the embedded terminal device;
[0066] S42, scaling, normalizing, and channel conversion preprocessing are performed on the mpp decoded data, and the data is stored in a post-data preprocessing queue;
[0067] S43. Create three inference threads according to the hardware characteristics of RK3588. Schedule one NPU core in each thread. Each inference thread obtains data from the preprocessing thread in turn for model inference.
[0068] S44, the data post-processing module obtains the inference result of each inference thread, and restores it to a segmented area and annotated box corresponding to the size of the original image;
[0069] S45, the detection and segmentation results include object category, location and confidence information, which are superimposed on the video image.
[0070] The present invention mainly provides solutions for the high-precision and high-efficiency deployment of multi-task perception tasks of unmanned aerial vehicles on the airborne Rockchip 3588 chip. Due to the limitation of computing resources, unlike the perception method of multiple models performing multiple tasks, the present invention optimizes the three-layer architecture of the backbone network, neck structure, and head structure in the target detection algorithm based on the characteristics of the Rockchip 3588 chip, and realizes a lightweight multi-task unmanned aerial vehicle real-time image perception network, which can simultaneously perform target detection, building and bridge segmentation tasks. Using the shared backbone network features, a multi-task feature connection module is designed for segmentation tasks to speed up segmentation processing; a new neck network is designed, the upsample in the RepBi-PAN structure is replaced by convolution, and the activation function is changed to ReLU, which is suitable for detection and segmentation tasks; based on the 3-core NPU characteristics of Rockchip 3588, a multi-core inference acceleration program is designed, the input video frame is input into the queue and then distributed to each core for parallel processing, which greatly improves resource utilization and processing speed, and finally pushes the stream output in sequence. Real-time and accurate perception of the environment during the flight of the unmanned aerial vehicle is achieved. We also designed an edge collaborative reasoning deployment tool, which formed functional modules for model quantization and conversion, model deployment, and model evaluation, and realized the end-side model deployment process management as well as the management of model and reasoning program results. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0072] Figure 1 is a flow chart of the present invention,
[0073] Figure 2 is a network structure diagram of the present invention,
[0074] Figure 3 is a workflow diagram of multi-threaded reasoning acceleration in the present invention,
[0075] Figure 4 It is a working principle diagram of the edge collaborative reasoning deployment tool of the present invention;
[0076] Figure 2 The “*” that appears in the number indicates the number of repetitions, and “⊕” indicates splicing. DETAILED DESCRIPTION
[0077] The following is combined with Figure 1-4 The technical solution of the present invention is further illustrated by specific implementation methods.
[0078] A lightweight remote sensing image multi-task target perception method includes a board of a Rockchip 3588 chip deployed with an inference program, including the following steps:
[0079] S1. Collect remote sensing images, perform detection and segmentation annotation,
[0080] Use drones to collect remote sensing image data, perform detection and segmentation annotation on the remote sensing images, and obtain training sets, test sets, and validation sets. The two tasks of detection and segmentation are explained as follows: Detection and annotation: identify targets (such as people and vehicles) in remote sensing images and annotate them with rectangular frames; segmentation and annotation: annotate bridges and roads with lines.
[0081] In step S1, a data set is prepared for the remote sensing image, and data enhancement processing is performed on it. Specifically, a data set is prepared for the remote sensing image, and data enhancement processing is performed on it. A data set is prepared for the remote sensing image, and data enhancement processing is performed on it, including the following steps:
[0082] S11. Remote sensing images include public data sets, actual flight data and map data. High-altitude image data with four types of targets, namely people, vehicles, bridges and roads, are obtained from remote sensing images at different angles and at altitudes of 100m-800m.
[0083] S12. Use labeling software (labelme software) to mark people and vehicle targets in the aerial image data with rectangular boxes, and mark the category of each target to achieve detection and labeling, and segment and label roads and bridges to achieve segmentation and labeling to obtain training data;
[0084] S13. Enhance the training data by rotation, cropping and data augmentation to obtain a training data set.
[0085] In addition, the enhancement methods of scaling and splicing have better detection effects on small objects;
[0086] S14. Divide the training data set into a training set, a test set, and a validation set in a ratio of 8:1:1. The training set is used to train the model, the test set is used to evaluate the model, and the validation set is used for validation. It should be noted that the above division method retains a certain proportion for subsequent testing and validation, and a large amount of image materials are used for training.
[0087] S2, training multi-task perception model,
[0088] The multi-task perception model is trained using the constructed neural network, which includes a backbone network (Backbone), a neck structure (Neck), a head structure (Head) and a segmentation structure.
[0089] Then use the training set obtained in step S1 and appropriate training parameters to train the multi-task perception model to obtain the optimal perception model in pt format, and convert and optimize the optimal perception model in pt format to obtain the onnx model;
[0090] The aforementioned backbone network: the improved Shufflenetv2 is used as the backbone network. Specifically, a convolution network with a large 6×6 convolution kernel and a step size of 2 is used for downsampling at the beginning of the input, which has a small improvement in detection accuracy. The activation function in the entire network is changed to ReLU6, and the spp module is added after the output stage5 module to form a new feature extraction network. The downsampling in Shufflenetv2 and the 1×1 convolution structure in the ordinary convolution block structure are cropped. Use a reduced hidden layer channel and modify the SPPF module of spatial pyramid pooling.
[0091] Neck structure: The neck structure adopts an inverted pyramid structure, which splices and fuses features of different scales from bottom to top. Finally, the feature map generated by the spp module is upgraded and connected to the convolution + BN + ReLU (CBR) module for dimensionality reduction and splicing with the final feature map. The backbone and neck enhance the model's representation ability while improving the model's reasoning speed. In order to speed up the reasoning speed, different convolution modes are used in the blocks during training and reasoning.
[0092] Head structure: Design a lighter detection head. The structure contains three detection heads and uses 3×3 convolution to calculate the target box loss, category loss, and confidence loss.
[0093] Segmentation structure: In the segmentation module, after obtaining the features output by FPN, the segmentation head is constructed, using a 3×3 convolution layer, and a 1×1 convolution layer in the last layer, and a mask coefficient is added.
[0094] In this step, the multi-task perception model is trained, including the following steps:
[0095] S21. Train the target detection network. The training framework is Pytorch, the stochastic gradient descent method is used, and the optimizer is Adam.
[0096] S22, training batch size is 32, learning rate is 0.0001, and the cosine annealing function is used to dynamically update the learning rate.
[0097] S23. Use Dropout in the backbone network architecture. Specifically, except for the first convolution layer, the rest use 10%.
[0098] S24. In terms of detection loss function, classification loss and regression loss use Varifocal Loss (VFL), GIouLoss, and adding DFL Loss enables the network to quickly focus on the value near the marked position and improve the probability.
[0099] Among them, Varifocal Loss:
[0100]
[0101] p is the predicted IACS score, q is the target IoU score, for the positive samples in training, q is set to the IoU between the generated bbox and gt box (gt IoU), and for the negative samples in training, the training target q of all categories is 0, α is a variable parameter, and α is taken as 0.75 in this project.
[0102] DFL Loss:
[0103]
[0104] Among them, y is the regression label;
[0105] In the discrete range [y0,y n ] within {y0, y1,...,y n}, estimated value for
[0106]
[0107] Where P is the probability distribution,
[0108] Then the DFL loss can be expressed as follows:
[0109] DFL(P(y i ),P(y i+1 ))=-((y i+1 -y)log(P(y i ))+(yy i )log(P(y i+1 )))
[0110] i is any integer,
[0111] The total loss is the weighted sum of the three losses as follows:
[0112]
[0113] α is 0.001, β is 0.0005, γ is 0.001, t, N~j, e For the normalization of the target score, the loss function in the segmentation task uses the cross entropy loss function as shown below
[0114]
[0115] Where x is the activation value, N is the feature dimension, label is the label corresponding to the segmentation task, ranging from (0 to C-1), where C is the number of categories;
[0116] S25. Convert the optimal pt model file trained by the Pytorch framework into an onnx file.
[0117] S26, trim the detection and segmentation tasks in onnx and use onnxsim to merge and optimize the trimmed onnx model.
[0118] S27. Use onnx-simplifier to perform non-identity operator deletion on the obtained onnx model.
[0119] S3. Use the edge collaborative reasoning tool to convert the onnx model obtained through training into the Rockchip 3588 model to obtain the RKNN model, and deploy the RKNN model and reasoning program to the board (end-side device) of the Rockchip 3588 chip.
[0120] Furthermore, in step S3, the edge collaborative reasoning tool is used to convert the optimal detection onnx model into the Rockchip 3588 model, and the specific steps of deploying the converted model and reasoning program to the terminal device are as follows:
[0121] S31,Edge collaborative reasoning tools consist of two parts, one is the end-side intelligent agent, and the other is the client, such as Figure 4 ,
[0122] The end-side intelligent agent includes the inference deployment package receiving and decompression service and the inference program start and stop control service components, which are mainly deployed in the RK3588 embedded end-side device;
[0123] The client includes model quantization and conversion, model publishing and model updating, inference service start and stop, and model evaluation functions, as well as model warehouse, inference program warehouse, and data warehouse, and is mainly deployed on PCs or servers;
[0124] S32. Use the model quantization and conversion module to convert the optimal detection onnx model into a model in rknn format. During the conversion process, turn on the quantization mode, and finally convert the onnx model file into an INT8 rknn file. The model quantization and conversion encapsulates the export_rknn interface of the RKNN-Toolkit2 tool. RKNN-Toolkit2 is a development kit for model conversion, reasoning, and performance evaluation provided to users.
[0125] S33, storing the converted rknn model file in the model warehouse;
[0126] S34, storing the inference program in the inference program warehouse, where the inference program must include a service startup script;
[0127] S35: When publishing or updating a model, select the rknn model file and inference program to be deployed. The model publishing and model updating modules package the rknn model file and the inference program, and send a model publishing request to the inference deployment service. After receiving the deployment package, the inference deployment service decompresses it and sends an inference program start request to the inference program start and stop control service.
[0128] S36: After receiving the request, the inference service start-stop control service starts the model inference program and generates a new model inference service;
[0129] S37. In the inference service startup module, the start and stop of the deployed model inference service can be controlled;
[0130] S38. In the model evaluation module, evaluation image data and the real labels corresponding to the images can be selected from the data warehouse and sent to the model inference service for reasoning to obtain annotation information. The model evaluation module can obtain the accuracy evaluation index of the model by comparing the real labels and the annotation information generated by model reasoning, and the performance evaluation index of the model can be obtained according to the reasoning time returned by the model reasoning service.
[0131] S4, when the reasoning program is running on the drone side, Figure 3 As shown, the data input uses the mpp module to decode the visible light camera input data, and calls the multi-core NPU (neural network processor) to infer the input video to obtain the inference result, and then post-processes the inference result;
[0132] Furthermore, in step S4, when the model is running on the drone side, the mpp module is used to decode the visible light camera input data, and the multi-core NPU is called to infer the input video to obtain the inference result, and then the inference result is post-processed, including the following steps:
[0133] S41, read the video stream from the camera, and decode it using the hard decoding mpp module of the embedded terminal device;
[0134] S42, scaling, normalizing, and channel conversion preprocessing are performed on the mpp decoded data, and the data is stored in a post-data preprocessing queue;
[0135] S43. According to the hardware characteristics of RK3588, three inference threads are created, one NPU core is scheduled in each thread, and each inference thread obtains data from the preprocessing thread in turn for model inference; the present invention adopts multi-core scheduling measures to achieve real-time algorithm operation under the limited computing power resources of domestic chips.
[0136] S44, the data post-processing module obtains the inference result of each inference thread, and restores it to a segmented area and annotated box corresponding to the size of the original image;
[0137] S45, the detection and segmentation results include object category, location and confidence information, which are superimposed on the video image.
[0138] S5, detection and segmentation results are displayed in superposition.
[0139] The detection and segmentation results in S45 are marked on the current frame using opencv to automatically overlay the result data. Specifically, the detection results are marked with a rectangular box and target category, and the segmentation results are marked with different colors of the region.
[0140] S6, push the stream to output the superimposed result video;
[0141] The automatic superposition result data obtained in step S5 is used to create a pipeline using the Gstream library, and the automatic superposition result data is cyclically read to superimpose the image to create an RTSP stream to obtain a video stream.
[0142] The present invention is aimed at the lightweight multi-task drone real-time image perception network of Rockchip 3588 device. The backbone network adopts the pruned shufflenetv2 network. According to the operation mode of the embedded device operator, the network layer that consumes more time is pruned, and the improved SPP module is added to make up for the accuracy loss. A lightweight detection module is designed, the inverted pyramid mode is adopted, the reverse convolution is used as the upsampling module, and the convolution + BN + ReLU is used as the dimension reduction module for feature fusion. An integrated segmentation head is designed. After obtaining the features after FPN output, a 3*3 convolution layer is used, and a 1*1 convolution layer is used in the last layer, and the mask coefficient is added to obtain the segmentation result. An edge collaborative reasoning deployment tool is designed, forming the model quantization and conversion, model deployment, and model evaluation function modules, realizing the end-side model deployment process management and model and reasoning program achievement management. On Rockchip 3588, the overall efficiency can reach 31fps, the detection accuracy is 91.3%, and the segmentation accuracy reaches 71%.
[0143] Finally, the method of the present invention can realize data transmission to the control end, so that the acquisition equipment (such as unmanned aerial vehicles, vehicles, etc.) can run, shoot, and intelligently perceive during the movement, and at the same time form a data stream to efficiently transmit data to the user end.
[0144] It should be noted that the above specific implementations are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art should understand that various modifications, equivalent substitutions, changes, etc. can be made to the present invention based on the technical content disclosed in this application document. However, as long as these changes do not deviate from the spirit of the present invention, they should be within the scope of protection of the present invention. In addition, some terms used in the specification and claims of this application are not restrictive, but are only for the convenience of description.
Claims
1. A lightweight remote sensing image multi-task target perception method, including a board with a Rockchip 3588 chip deployed with an inference program, It is characterized in that The following steps are involved: S1. Collect remote sensing images, perform detection and segmentation annotation, Using drones to collect remote sensing image data, performing detection and segmentation annotation on the remote sensing images, and obtaining training sets, test sets, and validation sets; S2, training multi-task perception model, The multi-task perception model is trained by using a constructed neural network, wherein the constructed neural network includes a backbone network, a neck structure, a head structure and a segmentation structure. Then, the multi-task perception model is trained using the training set obtained in step S1 and appropriate training parameters to obtain an optimal perception model in pt format, and the optimal perception model in pt format is converted and optimized to obtain an onnx model; S3. Use the edge collaborative reasoning tool to convert the onnx model obtained by training conversion into the Rockchip 3588 model to obtain the RKNN model, and deploy the RKNN model and reasoning program to the board of the Rockchip 3588 chip; S4. When the inference program is running on the drone, the data input uses the mpp module to decode the visible light camera input data, and calls the multi-core NPU to infer the input video to obtain the inference result, and then post-processes the inference result; S5, detection and segmentation results are displayed in superposition. The detection and segmentation results in S45 are marked on the current frame using opencv to automatically overlay the result data. Specifically, the detection results are marked with a rectangular box and target category, and the segmentation results are marked with different colors of the region. S6, push the stream to output the superimposed result video, The automatic superposition result data obtained in step S5 is used to create a pipeline using the Gstream library, and the automatic superposition result data is cyclically read to superimpose the image to create an RTSP stream to obtain a video stream.
2. According to claim 1, a lightweight remote sensing image multi-task target perception method is characterized in that: In step S1, a data set is prepared for the remote sensing image, and data enhancement processing is performed on it. Specifically, a data set is prepared for the remote sensing image, and data enhancement processing is performed on it, including the following steps: S11, the remote sensing image includes a public data set, actual flight collection data and map data, and high-altitude image data with four types of targets, namely, people, vehicles, bridges and roads, are obtained from the remote sensing image at different angles and at an altitude of 100m-800m; S12, using annotation software to mark the people and vehicle targets in the high-altitude image data with rectangular frames, and at the same time indicate the category of each target, to achieve the detection annotation, and to segment and mark the roads and bridges, to achieve the segmentation annotation, and obtain training data; S13, enhancing the training data by rotation, cropping and data augmentation to obtain a training data set; S14. Divide the training data set into a training set, a test set and a validation set in a ratio of 8:1:1, wherein the training set is used for training the model, and the test set is used for evaluating the model.
3. The lightweight remote sensing image multi-task target perception method according to claim 2 is characterized in that: In step S13, the enhancement method of scaling and splicing is introduced to further detect small targets.
4. The lightweight remote sensing image multi-task target perception method according to claim 1, characterized in that: The backbone network: uses the improved Shufflenetv2 as the backbone network, specifically, uses a 6×6 large convolution kernel and a step size of 2 convolution network for downsampling at the beginning of the input, which has a small improvement in detection accuracy, changes the activation function in the entire network to ReLU6, adds an spp module after the output stage5 module to form a new feature extraction network, cuts the 1×1 convolution structure in the downsampling in Shufflenetv2 and the ordinary convolution block structure, uses a reduced hidden layer channel and modifies the SPPF module of spatial pyramid pooling; The neck structure: The neck structure adopts an inverted pyramid structure, and features of different scales are spliced and fused from bottom to top. Finally, the feature map generated by the spp module is dimensionally upgraded and connected to the CBR module of convolution + BN + ReLU after the stage3 and stage5 feature maps for dimension reduction and splicing with the final feature map. The backbone and neck enhance the model representation capability while improving the model reasoning speed. In order to accelerate the reasoning speed, different convolution modes are used in the blocks during training and reasoning. The head structure: a lighter detection head is designed, the structure includes three detection heads, and a 3×3 convolution is used to calculate the target box loss, category loss and confidence loss; The segmentation structure: in the segmentation module, after obtaining the features output by FPN, a segmentation head is constructed, a 3×3 convolution layer is used, and a 1×1 convolution layer is used in the last layer, and a mask coefficient is added.
5. The lightweight remote sensing image multi-task target perception method according to claim 1, characterized in that: The multi-task perception model in step S2 is trained, including the following steps: S21. Train the target detection network. The training framework is Pytorch, the stochastic gradient descent method is used, and the optimizer is Adam. S22, the training batch size is 32, the learning rate is 0.0001, and the cosine annealing function is used to dynamically update the learning rate; S23. Use Dropout in the backbone network architecture, specifically, except for the first convolution layer, the rest use 10%; S24. In terms of detection loss function, classification loss and regression loss use Varifocal Loss and GIou Loss. Adding DFL Loss enables the network to quickly focus on the value near the marked position and improve the probability. Varifocal Loss: p is the predicted IACS score, q is the target IoU score, for the positive samples in training, q is set to the IoU between the generated bbox and gt box, and for the negative samples in training, the training target q of all categories is 0, α is a variable parameter, and α is taken as 0.75 in this project. DFL Loss: Among them, y is the regression label; In the discrete range [y0,y n ] within {y0, y1,...,y n }, estimated value for Where P is the probability distribution, Then DFLloss can be expressed as follows DFL(P(and i ),P(and i+1 ))=-((and i+1 -y)log(P(y i ))+(yy i )log(P(and i+1 ))) i is any integer, The total loss is the weighted sum of the three losses as follows: α is 0.001, β is 0.0005, γ is 0.001, t, N~j, e, is the normalization of the target score, In the segmentation task, the loss function uses the cross entropy loss function as shown below Where x is the activation value, N is the feature dimension, label is the label corresponding to the segmentation task, ranging from (0 to C-1), where C is the number of categories; S25. Convert the optimal pt model file trained by the Pytorch framework into an onnx file; S26, trim the post-processing of detection and segmentation tasks in onnx and use onnxsim to merge and optimize the trimmed onnx model; S27. Use onnx-simplifier to perform non-identity operator deletion on the obtained onnx model.
6. The lightweight remote sensing image multi-task target perception method according to claim 1, characterized in that: In step S3, the edge collaborative reasoning tool is used to convert the optimal detection onnx model into the Rockchip 3588 model, and the specific steps of deploying the converted model and reasoning program to the terminal device are as follows: S31. The edge collaborative reasoning tool consists of two parts: one is the end-side intelligent agent, and the other is the client. The end-side intelligent agent includes the inference deployment package receiving and decompression service and the inference program start and stop control service components, which are mainly deployed in the RK3588 embedded end-side device; The client includes model quantization and conversion, model publishing and model updating, inference service start and stop, and model evaluation functions, as well as model warehouse, inference program warehouse, and data warehouse, and is mainly deployed on PCs or servers; S32. Use the model quantization and conversion module to convert the optimal detection onnx model into a model in rknn format. During the conversion process, turn on the quantization mode, and finally convert the onnx model file into an INT8 rknn file. The model quantization and conversion encapsulates the export_rknn interface of the RKNN-Toolkit2 tool. RKNN-Toolkit2 is a development kit for users to provide model conversion, reasoning and performance evaluation; S33, storing the converted rknn model file in the model warehouse; S34, storing the inference program in the inference program warehouse, where the inference program must include a service startup script; S35: When publishing or updating a model, select the rknn model file and inference program to be deployed. The model publishing and model updating modules package the rknn model file and the inference program, and send a model publishing request to the inference deployment service. After receiving the deployment package, the inference deployment service decompresses it and sends an inference program start request to the inference program start and stop control service. S36: After receiving the request, the inference service start-stop control service starts the model inference program and generates a new model inference service; S37. In the inference service startup module, the start and stop of the deployed model inference service can be controlled; S38. In the model evaluation module, evaluation image data and the real labels corresponding to the images can be selected from the data warehouse and sent to the model inference service for reasoning to obtain annotation information. The model evaluation module can obtain the accuracy evaluation index of the model by comparing the real labels and the annotation information generated by model reasoning, and the performance evaluation index of the model can be obtained according to the reasoning time returned by the model reasoning service.
7. The lightweight remote sensing image multi-task target perception method according to claim 1, characterized in that: In step S4, when the model is running on the drone side, the mpp module is used to decode the visible light camera input data, and the multi-core NPU is called to infer the input video to obtain the inference result, and then the inference result is post-processed, including the following steps: S41, read the video stream from the camera, and decode it using the hard decoding mpp module of the embedded terminal device; S42, scaling, normalizing, and channel conversion preprocessing are performed on the mpp decoded data, and the data is stored in a post-data preprocessing queue; S43. Create three inference threads according to the hardware characteristics of RK3588. Schedule one NPU core in each thread. Each inference thread obtains data from the preprocessing thread in turn for model inference. S44, the data post-processing module obtains the inference result of each inference thread, and restores it to a segmented area and annotated box corresponding to the size of the original image; S45, the detection and segmentation results include object category, location and confidence information, which are superimposed on the video image.
Citation Information
Patent Citations
Lightweight multi-task video stream real-time reasoning method and system
CN115661712A
SAR (Synthetic Aperture Radar) ship detection method and system based on space-ground synchronous neural network
CN118053080A
Remote sensing image target detection method and system based on improved YOLOv8
CN118691967A
Semantic segmentation and height estimation method for remote sensing scene
CN118865130A
Multi-task joint perception network model and detection method for traffic road surface information
US20240420487A1
Cited By
Method for generating label containing moving target tracking image
CN121074891A
End side target detection method, system and equipment based on reinforcement learning
CN121280977A
Deep learning model training and deployment optimization method based on NPU
CN121766383A
A method for training and deploying an NPU-based deep learning model
CN121766383B