An optical remote sensing ship target detection method
By adding a deep learning method assisted by frequency feature in optical remote sensing images, a ship target detection model assisted by frequency information is constructed, which solves the problems of severe changes in the ship target scale and dense arrangement of small targets, and achieves high-precision ship target detection.
Patent Information
- Application Number
- CN202211470360.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-11-23
AI Technical Summary
In the prior art, in optical remote sensing ship images, the ship target scale changes violently and the small targets are arranged intensively, making it difficult to extract effective features and the ship target detection accuracy difficult to meet the requirements.
Using a deep learning method assisted by adding frequency feature, an optical remote sensing ship target detection model with added frequency information is constructed, frequency information is extracted using Laplacian convolution, and ship target detection is performed in combination with the YOLOv6 model, and target prediction is performed using the rotary box YOLOhead.
High-precision detection of ship targets in optical remote sensing images is achieved, and the detection accuracy is improved.
Smart Images

Figure CN115761500B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and particularly relates to an optical remote sensing ship target detection method. Background Art
[0002] Ships are important tools for maritime trade transportation. In order to ensure the success of tasks such as maritime situation awareness, maritime rescue, and safeguarding maritime rights and interests, it is of great significance to use optical remote sensing satellites for ship detection.
[0003] Most existing target detection methods use deep learning models. First, a target dataset close to the real application scenario is constructed, then an appropriate model is designed for training according to the characteristics of the targets in the dataset, and finally a detection model that meets the requirements is obtained. The deep learning model automatically extracts the features of the targets by training. In the remote sensing images of the maritime scene, the human eye distinguishes ship targets from the surrounding environment mainly relying on texture, contour, and intensity information. The present invention adopts a deep learning model assisted by adding frequency features, which can achieve high-precision detection of ship targets in optical remote sensing images. Summary of the Invention
[0004] The object of the present invention is to address the problem that in optical remote sensing ship images, the scales of ship targets change drastically, small targets are densely arranged, resulting in difficult extraction of effective features and further difficult to meet the requirements of ship target detection accuracy. An optical remote sensing ship target detection method is proposed. The method adopts a deep learning method assisted by adding frequency features, which can achieve high-precision detection of ship targets in optical remote sensing images.
[0005] The present invention is realized through the following technical solutions. The present invention proposes an optical remote sensing ship target detection method, which specifically includes:
[0006] Step 1: Obtain remote sensing images containing ships of a unified size, and perform rotated bounding box annotation on the ship targets in each remote sensing image to form a remote sensing image dataset;
[0007] Step 2: Divide the dataset into three parts: a training set, a validation set, and a test set according to a certain ratio;
[0008] Step 3: Construct an optical remote sensing ship target detection model assisted by adding frequency information: extract the frequency information of the remote sensing image, and then fuse the feature map containing the frequency information with the original feature map and input it into the target detection model;
[0009] Step 4: Use the training set and the validation set to train the target detection model, where the training set is used to adjust the model parameters, the validation set is used to adjust the model hyperparameters, and then use the test set to confirm the final effect of the trained model; finally, a target detection model that meets the requirements is obtained;
[0010] Step 5: Send the optical satellite remote sensing image to be detected into the trained object detection model for object prediction, so as to obtain the object detection result.
[0011] Further, the specific content of the said Step 3 is as follows:
[0012] First, downsample the optical remote sensing image by 1 / 2 and 1 / 4. Then, perform Laplacian convolution operations on the original image and the 1 / 2 and 1 / 4 downsampled images to obtain frequency information at three scales. Further, upsample the feature maps at the 1 / 2 and 1 / 4 scales of the obtained frequency information, fuse the original image and the image corresponding to the original image containing frequency information, and finally input the fused feature map into the object detection model.
[0013] Further, the Laplacian convolution uses a second-order convolution operator, and the parameters of the convolution kernel can be fine-tuned by the network, and gradient clipping is added.
[0014] Further, in Step 3, the feature map after fusing frequency information is input into the object detection model based on YOLOv6; the model is divided into three parts: Backbone, Neck, and Head; among them, the Backbone is used for feature extraction of the input image and outputs feature maps at three scales; the Neck is used for feature fusion of the feature maps at different scales output by the Backbone and outputs feature maps at three scales with enhanced features; the Head is used for judging the feature maps output by the Neck and predicting ship objects from three scales.
[0015] Further, for the Head, a rotated box YOLOhead is adopted, and the rotated box YOLOhead uses a decoupled method to obtain detection information and obtains classification results and spatial results respectively.
[0016] Further, the spatial result is the size and position information of the object, expressed as (x, y, w, h, θ), where the center point coordinates of the object are (x, y), the width and height of the object are (w, h), and the rotation angle is (θ); starting from the lowest vertex of the object box, the extension line along the positive x-axis is used as the reference line, and moving counterclockwise, the first side of the object box is defined as the wide side w of the object box, and the adjacent side is defined as the long side h; the angle between the wide side w of the object box and the reference line is the object offset angle θ, and the range is [-90, 0). For rotated box regression, when the true object box θ → 90°, and the predicted object box θ' → 0°, the angle loss will approach 90°; however, at this time, the predicted object box approaches the true object box, and only needs to rotate a small angle clockwise to be regressed to the true object box. Therefore, the angle loss L θ is:
[0017]
[0018] Among them, t θ is the predicted target box angle, and t θ' is the true target box angle; then the overall target box regression loss L reg is:
[0019]
[0020] Among them, t w , t h , t x , t y are respectively the width, height and center point coordinates of the predicted target box, and t w' , t h' , t x' , t y' are respectively the width, height and center point coordinates of the true target box.
[0021] Furthermore, in step four, the model is pre-trained on the open-source dataset DOTA to ensure that the model can converge effectively. Then, the pre-trained weights are imported into the object detection model, and the training set is input into the above-mentioned constructed object detection model for training. The model loss function will be continuously iterated until convergence. The validation set is used to verify the effect of the trained model every fixed number of rounds to obtain the optimal weights and hyperparameters of the network; finally, the test set is used to test and confirm the final result of the model.
[0022] Furthermore, in step five, first, the remote sensing image is segmented: assuming the image size is M×N, where M and N are the width and length of the image block respectively; setting the width of the cropping window as m and the length as n, and making the cropping window slide along the length direction of the image with a step size of k, the remote sensing image is cropped into several image blocks; then, the segmented image blocks are respectively sent into the model for target prediction; finally, the predicted image blocks are stitched together, and the predicted boxes in the overlapping area are fused using NMS.
[0023] The present invention provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the optical remote sensing ship target detection method are implemented.
[0024] The present invention provides a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, the steps of the optical remote sensing ship target detection method are implemented.
[0025] The present invention has the following beneficial effects:
[0026] The optical remote sensing ship target detection method described in the present invention adopts a deep learning method with frequency feature assistance, and can achieve high-precision detection of ship targets in optical remote sensing images. Description of the Drawings
[0027] Figure 1 It is a schematic diagram for extracting Laplacian convolution frequency features;
[0028] Figure 2 It is a schematic diagram of an optical remote sensing ship target detection model based on YOLOv6;
[0029] Figure 3 It is a schematic diagram of the representation method of the rotated bounding box. Specific Embodiments
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] Combined with Figures 1-3 , the present invention proposes an optical remote sensing ship target detection method, and the method specifically includes:
[0032] Step 1: Obtain remote sensing images containing ships of a unified size, and perform rotated bounding box annotation on the ship targets in each remote sensing image to form a remote sensing image dataset;
[0033] Step 2: Divide the dataset into three parts: a training set, a validation set, and a test set according to a certain ratio;
[0034] Step 3: Build an optical remote sensing ship target detection model with frequency information assistance: extract frequency information from the remote sensing images, and then fuse the feature maps containing frequency information with the original feature maps and input them into the target detection model;
[0035] The specific content of Step 3 is as follows:
[0036] First, downsample the optical remote sensing image by 1 / 2 and 1 / 4. Then, perform Laplacian convolution operations on the original image and the 1 / 2 and 1 / 4 downsampled images to obtain frequency information at three scales. Further, upsample the feature maps at the 1 / 2 and 1 / 4 scales with the obtained frequency information, fuse the original image and the image corresponding to the original image containing frequency information. Finally, input the fused feature map into the object detection model to obtain the ship target to be detected. The Laplacian convolution uses a second-order convolution operator, and the parameters of the convolution kernel can be fine-tuned by the network, and gradient clipping is added to ensure that the values after fine-tuning are not much different from the original convolution kernel, thus ensuring the effective extraction of frequency information.
[0037] Step 4: Use the training set and the validation set to train the object detection model. The training set is used to adjust the model parameters, and the validation set is used to adjust the model hyperparameters. Then, use the test set to confirm the final effect of the trained model; finally, obtain an object detection model that meets the requirements.
[0038] Step 5: Send the optical satellite remote sensing image to be detected into the trained object detection model for object prediction, so as to obtain the object detection result.
[0039] Dataset construction:
[0040] Use the Jilin-1 optical remote sensing satellite to obtain high-resolution remote sensing images containing ships, annotate the images with rotated bounding boxes, and save the annotation information as xml format files. The collected remote sensing images and their corresponding annotation information constitute the optical remote sensing ship dataset. The dataset is divided into a training set, a validation set, and a test set. The training set is used for training the model parameters, the validation set is used for adjusting the model hyperparameters during the model training process, and the test set is used for testing the final effect of the trained model.
[0041] Construction of an optical remote sensing ship object detection model based on YOLOv6:
[0042] The present invention first extracts frequency information from the original remote sensing image, and then fuses the feature map integrated with frequency information and the original image and inputs them into the object detection model. In this embodiment, the feature map after fusing frequency information is input into the object detection model based on YOLOv6. As Figure 2 shown, the model is divided into three parts: Backbone, Neck, and Head. Among them, the Backbone is used for feature extraction of the input image and outputs feature maps at three scales. The Neck is used for feature fusion of the feature maps at different scales output by the Backbone and outputs feature maps at three scales with enhanced features. The Head is used for judging the feature maps output by the Neck and predicting ship targets from three scales.
[0043] The backbone is mainly composed of RepVGGBlock and RepBlock. RepBlock is stacked by RepVGGBlock. The backbone can be divided into five stages: Steam, ERBlock_2, ERBlock_3, ERBlock_4, ERBlock_5, and SimSPFF. Among them, the three scale feature maps are the outputs of ERBlock_3, ERBlock_4, and SimSPFF respectively.
[0044] Neck first upsamples the SimSPFF feature map, further fuses the features with the ERBlock_4 feature map, then uses the upsampled fused feature map to fuse the features with the ERBlock_3 feature map, and finally outputs three scale feature maps with enhanced features by downsampling twice and fusing the intermediate feature maps in the upsampling process.
[0045] For the Head, when predicting the target with a rectangular box in YOLOv6, redundant features will be noticed. Especially for ship targets with a large aspect ratio, too many redundant features will directly lead to a decline in the detection performance of the algorithm. Therefore, a rotated box YOLOhead is proposed.
[0046] The rotated box YOLOhead obtains detection information in a decoupled manner, and obtains classification results and spatial results respectively. The spatial result is the size and position information of the target, expressed as (x, y, w, h, θ), as Figure 3 shown, where the center point coordinates of the target are (x, y), the width and height of the target are (w, h), and the rotation angle is (θ); specifically, starting from the bottom vertex of the target box, its extension line along the positive x-axis is used as the reference line, and moving counterclockwise, the first side of the target box is defined as the wide side w of the target box, and the adjacent side is defined as the long side h; the angle between the width w of the target box and the reference line is the target offset angle θ, and the range is [-90, 0). For the rotated box regression, when the true target box θ → 90°, and the predicted target box θ' → 0°, the angle loss will approach 90°; however, at this time, the predicted target box approaches the true target box, and only needs to rotate a small angle clockwise to return to the true target box. Therefore, the angle loss L θ is:
[0047]
[0048] where, t θ is the predicted target box angle, and t θ' is the true target box angle; then the entire target box regression loss L reg is:
[0049]
[0050] where tw , t h , t x , t y are the width, height, and the coordinates of the center point of the predicted target bounding box, respectively. t w' , t h' , t x' , t y' are the width, height, and the coordinates of the center point of the ground truth target bounding box, respectively.
[0051] Model training, validation, and testing:
[0052] Due to the high time and cost for collecting and annotating optical remote sensing datasets and the small dataset size, to ensure effective convergence during training, a method of pre-training on relevant large datasets and then training on local small samples is used to obtain the model. First, the model is pre-trained in the open-source dataset DOTA to ensure effective convergence. Then, the pre-trained weights are imported into the object detection model. The training set is input into the above-mentioned constructed object detection model for training. The model loss function will continuously iterate until convergence. The validation set is used to verify the effectiveness of the trained model every fixed number of rounds to obtain the optimal weights and hyperparameters of the network. Finally, the test set is used to test and confirm the final results of the model.
[0053] Using the model in the optical remote sensing scenario:
[0054] Since the original size of the optical satellite remote sensing image is too large, the remote sensing image is first segmented: Assume the image size is M×N, where M and N are the width and length of the image block, respectively. Set the width of the cropping window as m and the length as n, and let the cropping window slide along the length direction of the image with a step size of k to crop the remote sensing image into several image blocks. Then, the segmented image blocks are respectively fed into the model for target prediction. Finally, the predicted image blocks are stitched together, and the predicted bounding boxes in the overlapping regions are fused using NMS.
[0055] The present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the optical remote sensing ship target detection method are implemented.
[0056] The present invention provides a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, the steps of the optical remote sensing ship target detection method are implemented.
[0057] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the method described in the present invention is intended to include but not limited to these and any other suitable types of memory.
[0058] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a high-density digital video disc (DVD)), or a semiconductor medium (such as a solid state disc (SSD)), etc.
[0059] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0060] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0061] The above has introduced in detail an optical remote sensing ship target detection method proposed by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An optical remote sensing ship target detection method, characterized in that, The method specifically includes: Step 1: Obtain remote sensing images of ships with a unified size, and perform rotated bounding box annotation on ship targets in each remote sensing image to form a remote sensing image dataset; Step 2: Divide the dataset into three parts: a training set, a validation set, and a test set according to a certain ratio; Step 3: Construct an optical remote sensing ship target detection model with frequency information assistance: extract frequency information from the remote sensing image, and then fuse the feature map containing frequency information with the original feature map and input it into the target detection model; Step 4: Use the training set and the validation set to train the target detection model. The training set is used to adjust the model parameters, and the validation set is used to adjust the model hyperparameters. Then, use the test set to confirm the final effect of the trained model; finally, obtain a target detection model that meets the requirements; Step 5: Send the optical satellite remote sensing image to be detected into the trained target detection model for target prediction to obtain the target detection result; The specific content of Step 3 is: First, downsample the optical remote sensing image by 1 / 2 and 1 / 4, and then perform Laplacian convolution operations on the original image and the 1 / 2 and 1 / 4 downsampled images to obtain frequency information at three scales. Further, upsample the 1 / 2 and 1 / 4 scale feature maps containing the obtained frequency information, fuse the original image and the image corresponding to the original image containing frequency information, and finally input the fused feature map into the target detection model; In Step 3, the feature map after fusing frequency information is input into the target detection model based on YOLOv6; the model is divided into three parts: Backbone, Neck, and Head. Among them, Backbone is used for feature extraction of the input image and outputs feature maps at three scales; Neck is used for feature fusion of the feature maps at different scales output by Backbone and outputs three feature maps with enhanced features; Head is used for judging the feature maps output by Neck and predicting ship targets from three scales.
2. The method according to claim 1, wherein The Laplacian convolution uses a second-order convolution operator, and the parameters of the convolution kernel can be fine-tuned by the network, and gradient clipping restrictions are added.
3. The method according to claim 1, characterized in that For Head, a rotated bounding box YOLOhead is adopted, and the rotated bounding box YOLOhead obtains detection information in a decoupled manner and obtains classification results and spatial results respectively.
4. The method according to claim 3, wherein The spatial result is the size and position information of the target, expressed as (x, y, w, h, θ), where (x, y) are the coordinates of the center point of the target, (w, h) are the width and height of the target, and θ is the rotation angle; starting from the bottom vertex of the target box, with the extension line in the positive x-axis direction as the reference line, moving counterclockwise, the first side of the target box is defined as the wide side w of the target box, and the adjacent side is defined as the long side h; the angle between the wide side w of the target box and the reference line is the target offset angle θ, with a range of [-90, 0). For the rotation box regression, when the true target box θ → 90 ° When, the predicted target box θ ' → 0, the angle loss will approach 90°; however, at this time the predicted target box approaches the true target box, and only by rotating a small angle clockwise can it be regressed to the true target box. Therefore, the angle loss L θ is set as: where t θ is the predicted target box angle, and t θ' is the true target box angle; then the overall target box regression loss L reg is as follows: where t w 、t h 、t x 、t y are the width, height, and center point coordinates of the predicted target bounding box, and t w' 、t h' 、t x' 、t y' are the width, height, and center point coordinates of the ground truth target bounding box respectively.
5. The method according to claim 4, characterized in that, In Step 4, pre-train the model on the open-source dataset DOTA to ensure that the model can converge effectively, then import the pre-trained weights into the target detection model, input the training set into the above-constructed target detection model for training, and the model loss function will continuously iterate until convergence. Use the validation set to verify the effect of the trained model every fixed number of rounds to obtain the optimal weights and hyperparameters of the network; finally, use the test set to test and confirm the final result of the model.
6. The method according to claim 5, wherein In Step 5, first divide the remote sensing image into blocks: assume that the image size is M×N, where M and N are the width and length of the image block respectively; Set the width of the cropping window to m and the length to n, and make the cropping window slide along the length direction of the image with a step size of k to crop the remote sensing image into several image patches; then send the cropped image patches into the model for target prediction respectively; finally, splice the predicted image patches and use NMS to fuse the prediction boxes in the overlapping areas.
7. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-6.
8. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, it implements the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Sea surface remote sensing image ship detection method based on a feature pyramid
CN109800716A
Ship target automatic detection and identification method and system in natural scene
CN112464883A