Training method of remote sensing image target detection model and related equipment
By improving the YOLOv8 model and GWD loss function, the problems of slow calculation rate and inaccurate scale in remote sensing image object detection are solved, and the detection accuracy and speed are improved, especially in multi-scale object detection.
Patent Information
- Application Number
- CN202510494942.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-05
AI Technical Summary
The existing remote sensing image object detection algorithm has problems such as slow calculation rate, inaccurate RP scale, and manual design and no generalization, which affects the detection accuracy.
The improved YOLOv8 model is adopted, and the multi-scale detection head is added. Combined with the improved GWD loss function, the real box and prediction box information are replaced by Wasserstein distance and Gaussian distribution, and the object detection is optimized using the penalty term of the CIoU loss function to improve the detection performance of the model on different scales.
The recognition accuracy and speed of remote sensing image object detection is improved, the detection ability of multi-scale targets is enhanced, and missed detection and missed detection are reduced.
Smart Images

Figure CN120431313A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a training method for a remote sensing image target detection model and related equipment. Background Art
[0002] Research on object detection in remote sensing images involves automatically identifying and locating objects within them. This technology has garnered widespread attention. As a hot topic in remote sensing image processing, object detection in remote sensing images is a key application of artificial intelligence in remote sensing. This technology can enhance the practical application of remote sensing technology and significantly contribute to related research and applications.
[0003] Traditional object detection algorithms suffer from drawbacks such as the need for manual feature design and a lack of generalization, prompting interest in deep learning. While deep learning has significantly improved object detection accuracy, it still faces challenges such as slow computational speed and inaccurate RP scaling. Summary of the Invention
[0004] The present invention provides a method for training a remote sensing image target detection model and related equipment, which can improve the recognition accuracy of the model for target detection in remote sensing images.
[0005] A first aspect of the present invention provides a method for training a remote sensing image target detection model, comprising:
[0006] Obtain remote sensing image datasets;
[0007] Preprocessing each remote sensing image in the remote sensing image dataset to obtain a training sample dataset;
[0008] Constructing an initial training model, wherein the initial training model is a training model constructed based on an improved YOLOv8 model, and the initial training model includes a multi-scale detection head, wherein the multi-scale detection head is used to detect targets of different scales in a downsampled feature map of the target resolution;
[0009] Constructing a target loss function corresponding to the initial training model;
[0010] Initialization parameters are set for the initial training model, and the initial training model after the initialization parameters are set is iteratively trained based on the training sample set and the target loss function to generate a remote sensing image target detection model.
[0011] A second aspect of the present invention provides a training device for a remote sensing image target detection model, comprising:
[0012] Acquisition module, used to obtain remote sensing image datasets;
[0013] A preprocessing module, configured to preprocess each remote sensing image in the remote sensing image dataset to obtain a training sample dataset;
[0014] A first construction module is used to construct an initial training model, where the initial training model is a training model constructed based on an improved YOLOv8 model, and the initial training model includes a multi-scale detection head, where the multi-scale detection head is used to detect targets of different scales in a downsampled feature map of target resolution;
[0015] A second construction module is used to construct a target loss function corresponding to the initial training model;
[0016] The training module is used to set initial parameters for the initial training model, and iteratively train the initial training model after the initialization parameters are set based on the training sample set and the target loss function to generate a remote sensing image target detection model.
[0017] In one possible design, the first building block is specifically configured to:
[0018] Add a 4x downsampling detection head based on the YOLOv8s model;
[0019] The 32-fold downsampled detection head of the YOLOv8s model is removed to obtain the initial training model. The initial training model includes a backbone network structure, a neck network structure, and a head network structure. The backbone network structure adopts the CSPDarknet-53 architecture, including multiple convolutional layers and residual blocks. The neck network structure includes a PAN-FPN structure, which removes the upsampled convolution on the basis of the PAN structure and adds PAN to the FPN. The head network structure uses two independent branches for target classification and prediction regression.
[0020] In one possible design, the second building block is specifically configured to:
[0021] The objective loss function is determined by the following formula:
[0022]
[0023] Among them, L GWD is the target loss function, IoU is the intersection-over-union ratio between the predicted box and the real box, (b,b gt) is the center point of the predicted box and the true box, ρ is the Euclidean distance between the predicted box and the true box, c is the minimum diagonal length of the bounding box between the predicted box and the true box, α is a parameter for balancing the ratio, and v is a parameter for describing the consistency of the aspect ratio between the predicted box and the true box;
[0024]
[0025]
[0026] m1, m2 are the coordinate information of the predicted box and the real box, is the diagonal vector of the length and width of the predicted box and the real box, is the Frobenius norm, is the 2-norm, x, y, w, and h represent the center coordinate x, center coordinate y, width, and height of the representative box respectively; (xmin1, xmin2) and (ymin1, ymin2) are the coordinate points of the upper left corner of the predicted box and the real box; (xmax1, xmax2) and (ymax1, ymax2) are the coordinate points of the lower right corner of the predicted box and the real box.
[0027] In one possible design, the training module is specifically used to:
[0028] Determining evaluation indicators of the remote sensing image target detection model;
[0029] Dividing the training sample set into a training set, a validation set, and a test set according to a preset ratio;
[0030] Iteratively training the initial training model after initialization parameter setting based on the training set and the target loss function to obtain a first detection model;
[0031] The first detection model is verified and tested based on the evaluation index, the verification set and the test set to obtain the remote sensing image target detection model.
[0032] In one possible design, the training module is further used to:
[0033] receiving initial remote sensing images;
[0034] Preprocessing the initial remote sensing image to obtain a target feature map;
[0035] The target feature map is input into the remote sensing image target detection model to obtain a detection result.
[0036] In one possible design, the training module inputs the target feature map into the remote sensing image target detection model to obtain a detection result, including:
[0037] Performing a convolution operation on the target feature map, and upsampling the target feature map after the convolution operation to obtain an initial upsampled feature map;
[0038] Fusing the initial up-sampled feature map with the first preset feature map to obtain an initial fused feature map;
[0039] Upsampling the initial fusion feature map to obtain a target upsampled feature map;
[0040] Fusing the target upsampled feature map with the second preset feature map to obtain a target feature map;
[0041] Performing small target detection based on the target feature map to obtain a small target detection result;
[0042] Downsampling the target feature map, fusing it with the first preset feature map at the same scale, and performing small target detection on the fused feature map to obtain a small target detection result;
[0043] Downsampling the target feature map, fusing it with the third preset feature map and the fourth preset feature map respectively, and performing medium target detection and large target detection on the fused feature maps respectively to obtain medium target detection results and large target detection results;
[0044] Among them, the tiny target detection result, the small target detection result, the medium target detection result and the large target detection result are all the detection results.
[0045] In summary, it can be seen that in the embodiment provided by the present invention, the improved YOLOv8 model adopts a multi-scale detection head, which not only modifies the input anchor box size, but also adopts 4x, 8x, and 16x down-sampling feature maps, which can predict targets in remote sensing images at different scales, greatly improving the multi-scale target detection performance of the algorithm; in addition, the target loss function adopts an improved GWD loss function, which introduces the Wasserstein distance and replaces the original true box and the predicted box with a Gaussian distribution to obtain the gap between the target information and the predicted information, and modifies the width and height value regression to the upper left corner position information and the lower right corner position information on the basis of the Gaussian distribution, and replaces the parameter τ with IoU, which can adaptively adjust the change of the parameter value, and effectively reduce the overlap of the center of the true box and the predicted box in combination with the penalty term of the CIoU loss function, so that the focus is faster on the target samples that are difficult to identify, and the recognition accuracy of the model for the detected target can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A flowchart of a method for training a remote sensing image target detection model provided by an embodiment of the present invention;
[0047] Figure 2 A schematic diagram of the network structure of the YOLOv8 algorithm provided in an embodiment of the present invention;
[0048] Figure 3 A schematic diagram of the feature fusion structure of the YOLOv8s-P2 model provided in an embodiment of the present invention;
[0049] Figure 4 A schematic diagram of the virtual structure of a training device for a remote sensing image target detection model provided by an embodiment of the present invention;
[0050] Figure 5 A schematic diagram of the hardware structure of a training device for a remote sensing image target detection model provided by an embodiment of the present invention;
[0051] Figure 6 A schematic diagram of an electronic device according to an embodiment of the present invention;
[0052] Figure 7 A schematic diagram of an embodiment of a computer-readable storage medium provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0054] In the following description, the specific embodiments of the present invention will be described with reference to steps and symbols performed by one or more computers, unless otherwise stated. Therefore, these steps and operations will be mentioned several times as being performed by a computer, and the computer execution referred to herein includes the operation of a computer processing unit by electronic signals representing data in a structured form. This operation converts the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise change the operation of the computer in a manner familiar to testers in the field. The data structure in which the data is maintained is a physical location in the memory, which has specific characteristics defined by the data format. However, the principles of the present invention are described in the above text, which does not represent a limitation, and testers in the field will understand that the various steps and operations described below can also be implemented in hardware.
[0055] The principles of the present invention may be implemented and operated using many other general-purpose or special-purpose computing and communication environments or configurations. Examples of well-known computing systems, environments, and configurations suitable for use with the present invention include, but are not limited to, handheld phones, personal computers, servers, multiprocessor systems, microcomputer-based systems, mainframe computers, and distributed computing environments, including any of the aforementioned systems or devices.
[0056] The terms "first", "second" and "third" in the present invention are used to distinguish different objects rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions.
[0057] The goal of object detection is to detect objects of interest in an image or video. This includes identifying the location of the object in the image, determining the location of the object, and determining the category of the object. Depending on the detection method, object detection algorithms can be divided into one-stage algorithms and two-stage algorithms.
[0058] A one-stage algorithm directly extracts features from an image and then uses a classifier or regressor to predict the location and category of the object. Common one-stage algorithms include YOLO and SSD.
[0059] Second-stage algorithms divide the object detection problem into two subtasks: object proposal and object classification. First, all regions that may contain an object must be extracted, known as object proposals. Then, each proposed region is classified to determine whether it contains an object and its class. Common two-stage algorithms include R-CNN, Fast R-CNN, and Faster R-CNN.
[0060] In comparison: the one-stage algorithm usually has faster detection speed and lower computational complexity, but slightly lower detection accuracy; while the two-stage algorithm has higher detection accuracy, but slower speed and higher computational complexity.
[0061] The present invention adopts the YOLOv8 algorithm based on the improved loss function to identify targets in remote sensing images, which has faster network convergence speed and higher detection accuracy. By modifying the size of the target detection layer embedded in the network, it is used to increase the location and detail information of the target, reduce missed detections and false detections, and improve detection accuracy.
[0062] In addition, the present invention adopts an improved GWD loss function, which not only uses Gaussian distribution to replace the gap between the true box information and the predicted box information in the original YOLOv8 loss function, but also uses the upper left and lower right corner coordinate information of the true box and the upper left and lower right corner coordinate information of the predicted box as the regression function, and uses the exponential function to convert it into the interval [0,1]. The modified function replaces IoU as the bounding box regression of target detection, and optimizes the deviation between the true box and the predicted box through the penalty term, so that it can focus more quickly on target samples that are difficult to identify, thereby accelerating the network convergence speed.
[0063] The following describes the training method of the remote sensing image target detection model provided by an embodiment of the present invention from the perspective of the training device of the remote sensing image target detection model. The training device of the remote sensing image target detection model can be a server or a service unit in the server. For the sake of simplicity, the following describes the training device of the remote sensing image target detection model as a server as an example.
[0064] See also Figure 1 , Figure 1 A flowchart of a method for training a remote sensing image target detection model provided by an embodiment of the present invention includes:
[0065] 101. Obtain remote sensing image datasets.
[0066] In this embodiment, the server can obtain a remote sensing image dataset, which can be a collection of remote sensing images queried on Google Earth, or a collection of remote sensing images obtained from other data sources. There is no specific limitation. The remote sensing image dataset includes, for example, various types of buildings, vehicles and other targets. In the present invention, a remote sensing image dataset including 852 pictures involving 4 different categories is used as an example for explanation.
[0067] It should be noted that the annotation information of each remote sensing image in the remote sensing image dataset is initially stored in XML format. After the server obtains the annotation information in XML format, it can convert the annotation information in XML format into TXT format to facilitate subsequent model training.
[0068] 102. Preprocess each remote sensing image in the remote sensing image dataset to obtain a training sample set.
[0069] In this embodiment, after obtaining the remote sensing image dataset, the server can preprocess each remote sensing image in the remote sensing image dataset to obtain a training sample set, where the preprocessing includes but is not limited to data cleaning, normalization, missing value interpolation and other operations.
[0070] 103. Build the initial training model.
[0071] In this embodiment, the server can construct an initial training model, which is a training model constructed by an improved YOLOv8 model. The initial training model includes a multi-scale detection head, which is used to detect targets of different scales in the downsampled feature map of the target resolution. In order to facilitate understanding, the following Figure 2 and Figure 3 The network structure of the initial training model provided by the embodiment of the present invention is described in detail:
[0072] See also Figure 2 , Figure 2 This is a schematic diagram of the network structure of the YOLOv8 algorithm provided in an embodiment of the present invention. The network structure of the YOLOv8 algorithm mainly consists of a backbone, a neck, and a head. The backbone network structure, the neck network structure, and the head network structure are described in detail below:
[0073] 1. Backbone network structure:
[0074] The backbone network of YOLOv8 is as follows Figure 2 As shown in the Backbone module in the figure, the CSPDarknet-53 architecture is adopted, which consists of 53 convolutions, including some residual blocks. These residual blocks can help the YOLOv8 algorithm better capture the details and contextual information in remote sensing images. The C2f module is used to fuse feature maps of different scales to extract rich feature information and have rich gradient flows.
[0075] 2. Neck network structure:
[0076] The neck network is composed of PAN-FPN structure as follows Figure 2 As shown in the Neck module in the figure, the YOLOv8 algorithm removes the convolution after upsampling on the basis of the PAN structure, achieving a lightweight effect without changing the original performance. The traditional FPN adopts a top-down approach to transmit deep semantic information, which will lead to the loss of some target positioning information. The PAN-FPN in this application adds PAN to FPN, and realizes the enhancement of path information through the complementarity of shallow information and deep information.
[0077] 3. Head network structure:
[0078] The YOLOv8 detection head adopts a decoupled head structure, which uses two independent branches for target classification and prediction regression. Figure 2 The Head module is shown in the figure, and then a 1×1 convolution layer is used to complete the classification and positioning tasks. At the same time, YOLOv8 uses anchor-free detection. This structure can make the model converge faster and improve accuracy. Figure 2The Conv structure in the YOLOv8 model, the Bottleneck structure in the C2f module, and the SPPF structure in the pooling layer are also shown.
[0079] It should be noted that in the YOLOv8 algorithm model, the feature map used for target detection can perceive the minimum range of the original image is 8×8, which makes the model easily miss small targets with width and height less than 8 pixels. To improve this problem, the present invention adds a 4x downsampling detection head on the basis of the YOLOv8s model, named YOLOv8s-P2. The minimum target resolution that YOLOv8s-P2 can detect is 4×4. On this basis, the 32x downsampling detection head is removed to strengthen the recognition of small target samples, so that the model can predict targets of different scales on 4x, 8x, and 16x downsampled feature maps, greatly improving the algorithm's multi-scale target detection performance.
[0080] Since the improved model adds a small-scale detection head, the feature fusion method at the Neck end has also changed accordingly, but the overall structure still follows the FPN+PAN structure, such as Figure 3 As shown, Figure 3 Schematic diagram of the feature fusion structure of the YOLOv8s-P2 model provided in an embodiment of the present invention, Figure 3 C2, C3, and C4 correspond to the 4x, 8x, and 16x downsampled feature maps extracted by the backbone network, respectively. F3 and F4 are feature fusion layers, and P2, P3, and P4 are detection layers.
[0081] 104. Construct the target loss function corresponding to the initial training model.
[0082] In this embodiment, the CIoU loss function has defects in handling border size differences and sample imbalance, and is optimized and adjusted by introducing Wasserstein distance and penalty terms, as follows:
[0083] First, the CIoU loss function cannot distinguish between bounding boxes with the same center and the same aspect ratio but different sizes. The Wasserstein distance is introduced to replace the original true box and the predicted box with Gaussian distribution to obtain the gap between the target information and the predicted information. Based on the Gaussian distribution, the width and height value regression is modified to the upper left corner position information and the lower right corner position information. Secondly, the Gaussian Wasserstein distance regression loss (GWD) is normalized using the exponential formula and the GWD is mapped to A function similar to the IoU loss is obtained, namely the CIoU loss function. The calculation formula of the CIoU loss function is as follows:
[0084]
[0085] Where IoU is the intersection-over-union ratio between the predicted box and the real box, (b,b gt ) is the center point of the predicted box and the ground-truth box, ρ is the Euclidean distance between them, c is the minimum diagonal length of the bounding box between the predicted box and the ground-truth box, α is a parameter for balancing the ratios, and v is a parameter that describes the consistency of the aspect ratio between the predicted box and the ground-truth box. The CIoU loss function is relatively robust, can better handle objects of different shapes and sizes, and can also reduce the risk of overfitting.
[0086] The calculation formula of the GWD loss function is as follows:
[0087]
[0088] Among them, m1 and m2 are the coordinate information of the predicted box and the real box. is the diagonal vector of the length and width of the predicted box and the real box, is the Frobenius norm, It is the 2-norm, and x, y, w, and h represent the center point coordinate x, center point coordinate y, width, and height of the representative box respectively.
[0089] The calculation formula of the improved loss function is as follows:
[0090]
[0091]
[0092] Where IoU is the intersection-over-union ratio between the predicted box and the real box, (b,b gt ) is the center point of the predicted box and the real box, ρ is the Euclidean distance between the two, c is the minimum diagonal length of the external box between the predicted box and the real box, α is a parameter used to balance the ratio, v is a parameter used to describe the consistency of the aspect ratio of the predicted box and the real box; m1, m2 are the coordinate information of the predicted box and the real box, is the diagonal vector of the length and width of the predicted box and the real box, is the Frobenius norm, It is a 2-norm, x, y, w, and h represent the center coordinate x, center coordinate y, width, and height of the representative box respectively; (xmin1, xmin2) and (ymin1, ymin2) are the coordinate points of the upper left corner of the predicted box and the real box; (xmax1, xmax2) and (ymax1, ymax2) are the coordinate points of the lower right corner of the predicted box and the real box.
[0093] 105. Initialize the parameters of the initial training model, and iteratively train the initial training model after the initialization parameters are set based on the training sample set and the target loss function to generate a remote sensing image target detection model.
[0094] In this embodiment, the server can set the initialization parameters of the initial training model, such as the anchor box size, optimizer, and number of iterations, and then iteratively train the initial training model after the initialization parameters are set based on the training sample set and the target loss function to obtain a remote sensing image target detection model.
[0095] It should be noted that before iterative training, the server can also determine the evaluation indicators of the remote sensing image target detection model; and divide the training sample set into a training set, a validation set, and a test set according to a preset ratio. Here, the training sample set is divided in an 8:1:1 ratio, of which 689 images are used to construct the training set, 77 images are used for the validation set, and the remaining 86 images are used for the test set.
[0096] Finally, the initial training model after initialization parameter setting is iteratively trained based on the training set and the target loss function to obtain the first detection model; the first detection model is verified and tested based on the evaluation index, validation set and test set to obtain the remote sensing image target detection model.
[0097] It should also be noted that Mosaic data augmentation was used during the training of the remote sensing image object detection model to enrich the training samples and enhance the model's generalization performance. The experimental evaluation criteria used precision, recall, F1, and mean average precision (mAP) as objective evaluation criteria. The formulas for calculating precision and recall are as follows.
[0098]
[0099] TP: The label value is True and the model prediction is Positive;
[0100] FN: The label value is False and the model prediction is Negative;
[0101] FP: The label value is False and the model prediction is Positive;
[0102] TN: The label value is True, and the model prediction is Negative.
[0103] AP is calculated based on the true label and predicted probability of each category. The area formed by the corresponding Precision and Recall values is called Average Precision (AP); mAP is the average AP value of all categories, and its value range is [0,1]. The higher the values of the above two indicators, the better the detection performance of the model.
[0104] In one embodiment, after training the remote sensing image target detection model, the server further performs the following operations:
[0105] receiving initial remote sensing images;
[0106] Preprocess the initial remote sensing image to obtain the target feature map;
[0107] The target feature map is input into the remote sensing image target detection model to obtain a detection result.
[0108] In this embodiment, after training the remote sensing image target monitoring model, the server can receive the actual initial remote sensing image, preprocess the initial remote sensing image, obtain the target feature map, and finally input the target feature map into the trained remote sensing image target detection model to obtain the target detection result in the initial remote sensing image.
[0109] It should be noted that when performing target detection on the target feature map, the server performs a convolution operation on the target feature map, and upsamples the target feature map after the convolution operation to obtain an initial upsampled feature map; the initial upsampled feature map is fused with the first preset feature map to obtain an initial fused feature map; the initial fused feature map is upsampled to obtain a target upsampled feature map; the target upsampled feature map is fused with the second preset feature map to obtain a target feature map; small target detection is performed based on the target feature map to obtain a small target detection result; the target feature map is downsampled and fused with the first preset feature map at the same scale, and small target detection is performed on the fused feature map to obtain a small target detection result; the target feature map is downsampled and fused with the third preset feature map and the fourth preset feature map respectively, and medium target and large target detection are performed on the fused feature maps respectively to obtain medium target detection results and large target detection results.
[0110] For ease of understanding, the following is combined Figure 3 The target detection in the remote sensing image target detection model is described in detail. Figure 3 It can be seen that the Neck network first performs upsampling fusion on the input feature map, then downsampling fusion, and then predicts the fused feature map separately. The specific process is as follows:
[0111] 1. Upsampling fusion:
[0112] To achieve fine-grained feature detection, the 32*32*512 feature map input from the Backbone network is first subjected to convolution and other operations, then upsampled and fused with a 64*64*256 feature map to obtain a finer-grained 64*64*768 feature map. The 64*64*768 feature map is further upsampled to obtain a 128*128*256 feature map, which is then fused with the 128*128*128 feature map. Small object detection is performed on this fused large-scale feature map, and the 128*128*128 detection layer finally outputs the detection results.
[0113] 2. Downsampling Fusion:
[0114] The 128*128*128 detection layer's feature map is first downsampled and fused with the 64*64*256 feature map at the same scale. The 64*64*256 detection layer then outputs the detection results for small objects. The network continues downsampling and is then fused with the 32*32*512 and 32*32*256 feature maps, respectively. Medium and large objects are detected on the resulting 16x downsampled feature maps, and the 32*32*512 detection layer finally outputs the prediction results.
[0115] In summary, it can be seen that in the embodiment provided by the present invention, the improved YOLOv8 model adopts a multi-scale detection head, which not only modifies the input anchor box size, but also adopts 4x, 8x, and 16x down-sampling feature maps, which can predict targets in remote sensing images at different scales, greatly improving the multi-scale target detection performance of the algorithm; in addition, the target loss function adopts an improved GWD loss function, which introduces the Wasserstein distance and replaces the original true box and the predicted box with a Gaussian distribution to obtain the gap between the target information and the predicted information, and modifies the width and height value regression to the upper left corner position information and the lower right corner position information on the basis of the Gaussian distribution, and replaces the parameter τ with IoU, which can adaptively adjust the change of the parameter value, and effectively reduce the overlap of the center of the true box and the predicted box in combination with the penalty term of the CIoU loss function, so that the focus is faster on the target samples that are difficult to identify, and the recognition accuracy of the model for the detected target can be improved.
[0116] The above describes the embodiment of the present invention from the perspective of a training method for a remote sensing image target detection model. The following describes the embodiment of the present invention from the perspective of a training device for a remote sensing image target detection model.
[0117] See also Figure 4 , Figure 4 Schematic diagram of the virtual structure of a computing device for excitation of an array antenna intermediate frequency unit in an embodiment of the present invention. The training device 400 of the remote sensing image target detection model includes:
[0118] Acquisition module 401, used to acquire remote sensing image dataset;
[0119] A preprocessing module 402 is used to preprocess each remote sensing image in the remote sensing image dataset to obtain a training sample dataset;
[0120] A first construction module 403 is configured to construct an initial training model, wherein the initial training model is a training model constructed based on an improved YOLOv8 model, and the initial training model includes a multi-scale detection head, wherein the multi-scale detection head is configured to detect targets of different scales in a downsampled feature map of target resolution;
[0121] The second construction module 404 is used to construct the target loss function corresponding to the initial training model;
[0122] The training module 405 is used to set initial parameters for the initial training model, and iteratively train the initial training model after the initialization parameters are set based on the training sample set and the target loss function to generate a remote sensing image target detection model.
[0123] In one possible design, the first building module 403 is specifically configured to:
[0124] Add a 4x downsampling detection head based on the YOLOv8s model;
[0125] The 32-fold downsampled detection head of the YOLOv8s model is removed to obtain the initial training model. The initial training model includes a backbone network structure, a neck network structure, and a head network structure. The backbone network structure adopts the CSPDarknet-53 architecture, including multiple convolutional layers and residual blocks. The neck network structure includes a PAN-FPN structure, which removes the upsampled convolution on the basis of the PAN structure and adds PAN to the FPN. The head network structure uses two independent branches for target classification and prediction regression.
[0126] In one possible design, the second building module 404 is specifically configured to:
[0127] The objective loss function is determined by the following formula:
[0128]
[0129] Among them, L GWD is the target loss function, IoU is the intersection-over-union ratio between the predicted box and the real box, (b,b gt) is the center point of the predicted box and the true box, ρ is the Euclidean distance between the predicted box and the true box, c is the minimum diagonal length of the bounding box between the predicted box and the true box, α is a parameter for balancing the ratio, and v is a parameter for describing the consistency of the aspect ratio between the predicted box and the true box;
[0130]
[0131] m1, m2 are the coordinate information of the predicted box and the real box, is the diagonal vector of the length and width of the predicted box and the real box, is the Frobenius norm, is the 2-norm, x, y, w, and h represent the center coordinate x, center coordinate y, width, and height of the representative box respectively; (xmin1, xmin2) and (ymin1, ymin2) are the coordinate points of the upper left corner of the predicted box and the real box; (xmax1, xmax2) and (ymax1, ymax2) are the coordinate points of the lower right corner of the predicted box and the real box.
[0132] In one possible design, the training module 405 is specifically configured to:
[0133] Determining evaluation indicators of the remote sensing image target detection model;
[0134] Dividing the training sample set into a training set, a validation set, and a test set according to a preset ratio;
[0135] Iteratively training the initial training model after initialization parameter setting based on the training set and the target loss function to obtain a first detection model;
[0136] The first detection model is verified and tested based on the evaluation index, the verification set and the test set to obtain the remote sensing image target detection model.
[0137] In one possible design, the training module 405 is further configured to:
[0138] receiving initial remote sensing images;
[0139] Preprocessing the initial remote sensing image to obtain a target feature map;
[0140] The target feature map is input into the remote sensing image target detection model to obtain a detection result.
[0141] In one possible design, the training module 405 inputs the target feature map into the remote sensing image target detection model to obtain a detection result including:
[0142] Performing a convolution operation on the target feature map, and upsampling the target feature map after the convolution operation to obtain an initial upsampled feature map;
[0143] Fusing the initial up-sampled feature map with the first preset feature map to obtain an initial fused feature map;
[0144] Upsampling the initial fusion feature map to obtain a target upsampled feature map;
[0145] Fusing the target upsampled feature map with the second preset feature map to obtain a target feature map;
[0146] Performing small target detection based on the target feature map to obtain a small target detection result;
[0147] Downsampling the target feature map, fusing it with the first preset feature map at the same scale, and performing small target detection on the fused feature map to obtain a small target detection result;
[0148] Downsampling the target feature map, fusing it with the third preset feature map and the fourth preset feature map respectively, and performing medium target detection and large target detection on the fused feature maps respectively to obtain medium target detection results and large target detection results;
[0149] Among them, the tiny target detection result, the small target detection result, the medium target detection result and the large target detection result are all the detection results.
[0150] above Figure 4 The training device for a remote sensing image target detection model in an embodiment of the present invention has been described from the perspective of modular functional entities. The training device for a remote sensing image target detection model in an embodiment of the present invention will be described in detail from the perspective of hardware processing. Refer to FIG. 500 , which is a schematic diagram of an embodiment of a training device 500 for a remote sensing image target detection model in an embodiment of the present invention. The training device 500 for a remote sensing image target detection model includes:
[0151] Input device 501, output device 502, processor 503 and memory 504 (wherein the number of processor 503 can be one or more, Figure 5 In some embodiments of the present invention, the input device 501, the output device 502, the processor 503 and the memory 504 may be connected via a communication bus or other means, wherein: Figure 5 The communication bus connection is taken as an example.
[0152] By calling the operation instructions stored in the memory 504, the processor 503 is configured to execute the following steps:
[0153] Obtain remote sensing image datasets;
[0154] Preprocessing each remote sensing image in the remote sensing image dataset to obtain a training sample dataset;
[0155] Constructing an initial training model, wherein the initial training model is a training model constructed based on an improved YOLOv8 model, and the initial training model includes a multi-scale detection head, wherein the multi-scale detection head is used to detect targets of different scales in a downsampled feature map of the target resolution;
[0156] Constructing a target loss function corresponding to the initial training model;
[0157] Initialization parameters are set for the initial training model, and the initial training model after the initialization parameters are set is iteratively trained based on the training sample set and the target loss function to generate a remote sensing image target detection model.
[0158] By calling the operation instructions stored in the memory 504, the processor 503 is also used to execute Figure 1 Any method in the corresponding embodiment.
[0159] See also Figure 6 , Figure 6 A schematic diagram of an electronic device according to an embodiment of the present invention.
[0160] like Figure 6 As shown, an embodiment of the present invention provides an electronic device, including a memory 610, a processor 620, and a computer program 611 stored in the memory 610 and executable on the processor 620. When the processor 620 executes the computer program 611, the following steps are implemented:
[0161] Obtain remote sensing image datasets;
[0162] Preprocessing each remote sensing image in the remote sensing image dataset to obtain a training sample dataset;
[0163] Constructing an initial training model, wherein the initial training model is a training model constructed based on an improved YOLOv8 model, and the initial training model includes a multi-scale detection head, wherein the multi-scale detection head is used to detect targets of different scales in a downsampled feature map of the target resolution;
[0164] Constructing a target loss function corresponding to the initial training model;
[0165] Initialization parameters are set for the initial training model, and the initial training model after the initialization parameters are set is iteratively trained based on the training sample set and the target loss function to generate a remote sensing image target detection model.
[0166] In a specific implementation process, when the processor 620 executes the computer program 611, it can achieve Figure 1 Any implementation manner in the corresponding embodiments.
[0167] Since the electronic device introduced in this embodiment is a device used to implement a computing device for excitation of a mid-frequency point unit of an array antenna in an embodiment of the present invention, based on the method introduced in the embodiment of the present invention, technical personnel in this field can understand the specific implementation of the electronic device of this embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of the present invention will not be introduced in detail here. As long as the equipment used by technical personnel in this field to implement the method in the embodiment of the present invention falls within the scope of protection of the present invention.
[0168] Please refer to Figure 700, which is a schematic diagram of an embodiment of a computer-readable storage medium provided by an embodiment of the present invention.
[0169] As shown in FIG. 700 , an embodiment of the present invention further provides a computer-readable storage medium 700 on which a computer program 711 is stored. When the computer program 711 is executed by a processor, the following steps are implemented:
[0170] Obtain remote sensing image datasets;
[0171] Preprocessing each remote sensing image in the remote sensing image dataset to obtain a training sample dataset;
[0172] Constructing an initial training model, wherein the initial training model is a training model constructed based on an improved YOLOv8 model, and the initial training model includes a multi-scale detection head, wherein the multi-scale detection head is used to detect targets of different scales in a downsampled feature map of the target resolution;
[0173] Constructing a target loss function corresponding to the initial training model;
[0174] Initialization parameters are set for the initial training model, and the initial training model after the initialization parameters are set is iteratively trained based on the training sample set and the target loss function to generate a remote sensing image target detection model.
[0175] In the specific implementation process, the computer program 711 is executed by the processor to implement Figure 1 Any implementation manner in the corresponding embodiments.
[0176] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0177] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0178] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0179] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0180] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0181] The embodiment of the present invention also provides a computer program product, which includes computer software instructions. When the computer software instructions are executed on a processing device, the processing device executes the following Figure 1 The process in the corresponding embodiment.
[0182] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0183] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for a remote sensing image target detection model, characterized in that: include: Obtain remote sensing image datasets; Preprocessing each remote sensing image in the remote sensing image dataset to obtain a training sample dataset; Constructing an initial training model, where the initial training model is a training model constructed based on an improved YOLOv8 model, and the initial training model includes a multi-scale detection head, where the multi-scale detection head is used to detect targets of different scales in a downsampled feature map of target resolution; Constructing a target loss function corresponding to the initial training model; Initialization parameters are set for the initial training model, and the initial training model after the initialization parameters are set is iteratively trained based on the training sample set and the target loss function to generate a remote sensing image target detection model.
2. The method according to claim 1, characterized in that The constructing of the initial training model includes: Add a 4x downsampling detection head based on the YOLOv8s model; The 32-fold downsampled detection head of the YOLOv8s model is removed to obtain the initial training model. The initial training model includes a backbone network structure, a neck network structure, and a head network structure. The backbone network structure adopts the CSPDarknet-53 architecture, including multiple convolutional layers and residual blocks. The neck network structure includes a PAN-FPN structure, which removes the upsampled convolution on the basis of the PAN structure and adds PAN to the FPN. The head network structure uses two independent branches for target classification and prediction regression.
3. The method according to claim 1, characterized in that The objective loss function corresponding to the initial training model is constructed as follows: The objective loss function is determined by the following formula: Among them, L CGWD is the target loss function, IoU is the intersection-over-union ratio between the predicted box and the real box, (b,b gt ) is the center point of the predicted box and the true box, ρ is the Euclidean distance between the predicted box and the true box, c is the minimum diagonal length of the bounding box between the predicted box and the true box, α is a parameter for balancing the ratio, and v is a parameter for describing the consistency of the aspect ratio between the predicted box and the true box; m1, m2 are the coordinate information of the predicted box and the real box, is the diagonal vector of the length and width of the predicted box and the real box, is the Frobenius norm, is the 2-norm, x, y, w, and h represent the center coordinate x, center coordinate y, width, and height of the representative box respectively; (xmin1, xmin2) and (ymin1, ymin2) are the coordinate points of the upper left corner of the predicted box and the real box; (xmax1, xmax2) and (ymax1, ymax2) are the coordinate points of the lower right corner of the predicted box and the real box.
4. The method according to claim 1, wherein The iterative training of the initial training model after the initialization parameters are set based on the training sample set and the target loss function to generate a remote sensing image target detection model includes: Determining evaluation indicators of the remote sensing image target detection model; Dividing the training sample set into a training set, a validation set, and a test set according to a preset ratio; Iteratively training the initial training model after initialization parameter setting based on the training set and the target loss function to obtain a first detection model; The first detection model is verified and tested based on the evaluation index, the verification set and the test set to obtain the remote sensing image target detection model.
5. The method according to any one of claims 1 to 3, characterized in that The method further comprises: receiving initial remote sensing images; Preprocessing the initial remote sensing image to obtain a target feature map; The target feature map is input into the remote sensing image target detection model to obtain a detection result.
6. The method according to claim 5, characterized in that Inputting the target feature map into the remote sensing image target detection model to obtain a detection result includes: Performing a convolution operation on the target feature map, and upsampling the target feature map after the convolution operation to obtain an initial upsampled feature map; Fusing the initial up-sampled feature map with the first preset feature map to obtain an initial fused feature map; Upsampling the initial fusion feature map to obtain a target upsampled feature map; Fusing the target upsampled feature map with the second preset feature map to obtain a target feature map; Performing small target detection based on the target feature map to obtain a small target detection result; Downsampling the target feature map, fusing it with the first preset feature map at the same scale, and performing small target detection on the fused feature map to obtain a small target detection result; Downsampling the target feature map, fusing it with the third preset feature map and the fourth preset feature map respectively, and performing medium target detection and large target detection on the fused feature maps respectively to obtain medium target detection results and large target detection results; Among them, the tiny target detection result, the small target detection result, the medium target detection result and the large target detection result are all the detection results.
7. A training device for a remote sensing image target detection model, characterized in that: include: Acquisition module, used to obtain remote sensing image datasets; A preprocessing module, configured to preprocess each remote sensing image in the remote sensing image dataset to obtain a training sample dataset; A first construction module is used to construct an initial training model, where the initial training model is a training model constructed based on an improved YOLOv8 model, and the initial training model includes a multi-scale detection head, where the multi-scale detection head is used to detect targets of different scales in a downsampled feature map of target resolution; A second construction module is used to construct a target loss function corresponding to the initial training model; The training module is used to set initial parameters for the initial training model, and iteratively train the initial training model after the initialization parameters are set based on the training sample set and the target loss function to generate a remote sensing image target detection model.
8. The device according to claim 7, characterized in that The first building block is specifically used for: Add a 4x downsampling detection head based on the YOLOv8s model; The 32-fold downsampled detection head of the YOLOv8s model is removed to obtain the initial training model. The initial training model includes a backbone network structure, a neck network structure, and a head network structure. The backbone network structure adopts the CSPDarknet-53 architecture, including multiple convolutional layers and residual blocks. The neck network structure includes a PAN-FPN structure, which removes the upsampled convolution on the basis of the PAN structure and adds PAN to the FPN. The head network structure uses two independent branches for target classification and prediction regression.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is configured to implement the training method for the remote sensing image target detection model as described in any one of claims 1 to 6 when executing a computer management program stored in the memory.
10. A computer-readable storage medium storing a computer management program, characterized in that: When the computer management program is executed by the processor, the steps of the training method of the remote sensing image target detection model described in any one of claims 1 to 6 are implemented.