Natural scene ship target detection method and system

Through the three-stage deep learning detection algorithm, including YOLOv5x, DINO and YOLOv5nano network models, the ships in natural scenes are phased and angle corrections are performed, which solves the problems of low ship recognition accuracy and low detection efficiency in the prior art, and achieves high accuracy and high efficiency ship detection.

CN120032102AActive Publication Date: 2025-05-23COSCO SHIPPING TECH CO LTD

Patent Information

Application Number
CN202510067527.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-23
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The existing ship detection methods are low in recognition accuracy and low detection efficiency due to interference factors in natural scenarios, and fail to effectively identify small ships in distant places, which is easy to miss reports.

Method used

A three-stage deep learning detection algorithm is adopted, including the YOLOv5x network model, the DINO target detection model and the YOLOv5nano network model, and the ship's angle is corrected through contour extraction and affine transformation. Finally, the third stage detection is performed in the YOLOv5nano network model to verify and correct the detection results.

Benefits of technology

The ship detection rate and recognition accuracy have been greatly improved, especially in identifying long-distance and fuzzy ships, which significantly improves the accuracy, solving the problems of low recognition accuracy and low detection efficiency in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032102A_ABST
    Figure CN120032102A_ABST
Patent Text Reader

Abstract

The invention provides a natural scene ship target detection method and system, and the method comprises the steps: firstly obtaining a ship detection data set, a COCO data set and a plurality of to-be-detected natural scene pictures, and inputting the plurality of to-be-detected natural scene pictures into a YOLOv5x network model trained based on the ship detection data set, inputting the plurality of natural scene pictures containing the ships into a DINO target detection model trained based on a COCO data set to obtain a first target detection result of all the ships in each natural scene picture, and cutting out a cutting area containing the ships according to the first target detection result of each ship; calculating an included angle between a connecting line of the two contour points and the horizontal direction by adopting a contour searching function, a rotation matrix function and an affine transformation function, and rotating according to the calculated included angle to obtain a cutting area after ship direction correction; and finally, inputting the cutting area after ship direction correction into the trained YOLOv5nano network model for verification so as to complete ship target detection of a natural scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent shipping technology, and in particular to a natural scene ship target detection method and system. Background Art

[0002] Ships are important targets on the sea, and ship identification research has always been a hot topic. Ship identification and detection technology has broad application prospects. The existing ship detection algorithm directly passes natural scene images into a single-stage detection algorithm, but due to the many interference factors of natural scenes, the accuracy of the single-stage detection algorithm cannot meet the requirements of high-accuracy ship identification.

[0003] In addition, there is a detection method that uses the two-stage Faster-RCNN algorithm to detect and identify ships. However, since ship images are taken in natural scenes and ships at a distance are affected by many interference factors such as size, light, distance, and angle, the recognition accuracy is low. In addition, the deep learning method based on GoogleNet is used to detect ships, which has problems such as low detection efficiency, high power consumption, and failure to identify small boats in the distance, which is prone to missed reports.

[0004] In view of the above problems, there is an urgent need for a ship detection method with high detection efficiency, strong practicality and the ability to accurately identify small ships in the distance to meet the challenges of ship tracking and ship segmentation in current maritime images. Summary of the invention

[0005] In order to solve the problems of low recognition accuracy, low detection efficiency, and failure to identify small boats in the distance, which are easy to miss, etc. in the existing ship detection process, the present invention provides a natural scene ship target detection method, which uses a three-stage deep learning detection algorithm (YOLOv5x network model, DINO target detection model and YOLOv5nano network model) to perform phased recognition on the ships in the natural scene pictures to be tested, thereby greatly improving the ship detection rate. In addition, for difficult-to-detect ships, the recognition accuracy and accuracy of long-distance and blurred ships can be greatly improved. The present invention also relates to a natural scene ship target detection system.

[0006] The technical solution of the present invention is as follows:

[0007] A method for detecting ship targets in natural scenes, characterized by comprising the following steps:

[0008] Image and data set acquisition steps: obtain the ship detection data set, COCO data set, and the natural scene video image to be tested, and perform frame processing on the natural scene video image to be tested to obtain multiple natural scene pictures to be tested;

[0009] The first stage of ship image screening step: using the ship detection data set as a training set sample to train the YOLOv5x network model and the YOLOv5nano network model respectively, to obtain the trained YOLOv5x network model and the YOLOv5nano network model respectively, and inputting a plurality of natural scene images to be tested into the trained YOLOv5x network model for the first stage screening test, to obtain a plurality of natural scene images containing ships;

[0010] The second stage of ship fine detection and cropping area acquisition steps: using the COCO dataset to train the DINO target detection model to obtain a trained DINO target detection model, and inputting multiple natural scene pictures containing ships obtained by the first stage screening detection into the trained DINO target detection model for the second stage screening detection to obtain the first target detection results of all ships in each natural scene picture, and cropping the cropping area containing the ship according to the first target detection result of each ship;

[0011] Ship direction correction step: Use the contour search function to extract the ship contour from the cropped area containing the ship, and extract the points corresponding to the minimum value of the horizontal coordinate and the maximum value of the horizontal coordinate from the ship contour according to the first target detection result, as the first contour point and the second contour point, calculate the angle between the connecting line of the two contour points and the horizontal direction according to the first contour point and the second contour point, and then use the rotation matrix function to create a rotation matrix, based on the rotation matrix and using the affine transformation function to rotate the ship contour in the cropped area according to the calculated angle, so as to obtain the cropped area after the ship direction correction;

[0012] The third stage detection and detection result verification steps are as follows: the cropped area after ship direction correction is input into the lightweight trained YOLOv5nano network model for the third stage detection, and the second target detection result of the ship in the cropped area after ship direction correction is obtained, and the first target detection result output by the trained YOLOv5nano network model is used to verify whether there is a false detection or missed detection in the first target detection result output by the second stage screening detection of the trained DINO target detection model. If not, the second target detection result is used as the final ship target detection result; if so, the false detection or missed detection is corrected in combination with the first target detection result and the second target detection result to complete the ship target detection in the natural scene.

[0013] Preferably, in the first stage of the ship image screening step, a plurality of natural scene images to be tested are respectively input into the trained YOLOv5x network model for first stage screening detection, and the plurality of natural scene images containing ships are obtained, specifically including:

[0014] Multiple natural scene images to be tested are respectively input into the trained YOLOv5x network model, and all ship detection results in each natural scene image are output. The ship detection results include bounding boxes, category labels and confidence scores. The confidence scores of all ship detection results in a natural scene image are compared with the preset confidence threshold. If there is at least one ship detection result with a confidence score greater than or equal to the confidence threshold, the natural scene image is considered to contain ships and is saved and marked, and finally multiple natural scene images containing ships are obtained.

[0015] Preferably, in the third stage detection and detection result verification step, the correction of false detection or missed detection in combination with the first target detection result and the second target detection result includes: if a false detection exists, then according to the first target detection result output by the DINO target detection model, the false detection box that does not belong to the ship identified in the second target detection result output by the YOLOv5nano network model is removed from the final detection result; if a missed detection exists, the missed detection target identified by the YOLOv5nano network model is added to the first target detection result output by the DINO model, and the position of the target is re-marked.

[0016] Preferably, in the second stage of ship fine detection and cropping area acquisition step, before the multiple natural scene pictures containing ships obtained by the first stage screening and detection are respectively input into the trained DINO target detection model, the multiple natural scene pictures containing ships screened by YOLOv5x are first resized and standardized to meet the input requirements of the trained DINO target detection model.

[0017] Preferably, in the image and dataset acquisition step, the acquired ship detection dataset includes multiple different types of ship images and annotations, and the COCO dataset includes multiple different object categories, and the multiple different object categories include several combinations of people, bicycles, cars, motorcycles, airplanes, buses, trains, trucks, ships, traffic lights, fire hydrants, stop signs, parking meters, and benches.

[0018] A natural scene ship target detection system, characterized by comprising an image and data set acquisition module, a first-stage ship image screening module, a second-stage ship fine detection and cropping area acquisition module, a ship direction correction module, and a third-stage detection and detection result verification module connected in sequence.

[0019] The image and data set acquisition module acquires the ship detection data set, the COCO data set, and the video image of the natural scene to be tested, and performs frame processing on the video image of the natural scene to be tested to obtain multiple pictures of the natural scene to be tested;

[0020] The first stage includes a ship image screening module, which uses the ship detection data set as a training set sample to train the YOLOv5x network model and the YOLOv5nano network model respectively, and obtains the trained YOLOv5x network model and the YOLOv5nano network model respectively, and inputs a plurality of natural scene images to be tested into the trained YOLOv5x network model for the first stage screening test, and obtains a plurality of natural scene images containing ships;

[0021] The second-stage ship fine detection and cropping region acquisition module uses the COCO dataset to train the DINO target detection model to obtain a trained DINO target detection model, and inputs multiple natural scene pictures containing ships obtained by the first-stage screening detection into the trained DINO target detection model for the second-stage screening detection to obtain the first target detection results of all ships in each natural scene picture, and crops out the cropping region containing the ship according to the first target detection result of each ship;

[0022] The ship direction correction module uses a contour search function to extract a ship contour from a clipping area containing the ship, and extracts points corresponding to a minimum abscissa value and a maximum abscissa value from the ship contour according to the first target detection result, as first contour points and second contour points, calculates an angle between a line connecting the two contour points and a horizontal direction according to the first contour points and the second contour points, and then uses a rotation matrix function to create a rotation matrix, and based on the rotation matrix and using an affine transformation function, rotates the ship contour in the clipping area according to the calculated angle to obtain a clipping area after the ship direction is corrected;

[0023] The third-stage detection and detection result verification module inputs the cropped area after the ship direction correction into the lightweight trained YOLOv5nano network model for the third-stage detection, obtains the second target detection result of the ship in the cropped area after the ship direction correction, and verifies whether the first target detection result output by the second-stage screening detection of the trained DINO target detection model has false detection or missed detection according to the second target detection result output by the trained YOLOv5nano network model. If not, the second target detection result is used as the final ship target detection result; if so, the false detection or missed detection is corrected in combination with the first target detection result and the second target detection result to complete the ship target detection in the natural scene.

[0024] Preferably, in the first-stage ship image screening module, a plurality of natural scene images to be tested are respectively input into the trained YOLOv5x network model for first-stage screening and detection, and a plurality of natural scene images containing ships are obtained, specifically including:

[0025] Multiple natural scene images to be tested are respectively input into the trained YOLOv5x network model, and all ship detection results in each natural scene image are output. The ship detection results include bounding boxes, category labels and confidence scores. The confidence scores of all ship detection results in a natural scene image are compared with the preset confidence threshold. If there is at least one ship detection result with a confidence score greater than or equal to the confidence threshold, the natural scene image is considered to contain ships and is saved and marked, and finally multiple natural scene images containing ships are obtained.

[0026] Preferably, in the third-stage detection and detection result verification module, the correction of false detection or missed detection in combination with the first target detection result and the second target detection result includes: if a false detection exists, then according to the first target detection result output by the DINO target detection model, the false detection box that does not belong to the ship identified in the second target detection result output by the YOLOv5nano network model is removed from the final detection result; if a missed detection exists, the missed detection target identified by the YOLOv5nano network model is added to the first target detection result output by the DINO model, and the position of the target is re-marked.

[0027] Preferably, in the second-stage ship fine detection and cropping area acquisition module, before the multiple natural scene pictures containing ships obtained by the first-stage screening and detection are respectively input into the trained DINO target detection model, the multiple natural scene pictures containing ships screened by YOLOv5x are first resized and standardized to meet the input requirements of the trained DINO target detection model.

[0028] Preferably, the vessel detection dataset includes multiple different types of ship images and annotations, and the COCO dataset includes multiple different object categories, and the multiple different object categories include several combinations of people, bicycles, cars, motorcycles, airplanes, buses, trains, trucks, ships, traffic lights, fire hydrants, stop signs, parking meters, and benches.

[0029] The beneficial effects of the present invention are:

[0030] The present invention provides a method for detecting ship targets in natural scenes. The method comprises the following steps: firstly obtaining a ship detection dataset, a COCO dataset, and a video image of a natural scene to be tested, performing frame processing on the video image of the natural scene to be tested to obtain a plurality of pictures of the natural scene to be tested, and respectively training a YOLOv5x network model, a DINO target detection model, and a YOLOv5nano network model corresponding to a constructed three-stage deep learning detection algorithm for different collected data sets, and respectively training a YOLOv5x network model and a YOLOv5nano network model of different scales in the same series using the ship detection dataset. The two models have similar architectures but different depths and widths, and the trained YOLOv5x network model and the YOLOv5nano network model are obtained. YOLOv5x network model and YOLOv5nano network model are used to train the YOLOv5x network model. Multiple natural scene images to be tested are input into the trained YOLOv5x network model for the first stage of screening and detection to obtain multiple natural scene images containing ships. YOLOv5x is one of the largest models in the YOLOv5 series. It has the characteristics of both speed and accuracy. It is suitable for large target or complex scene detection tasks on high-performance hardware. In this stage, images with ships are quickly screened from massive images. The YOLOv5x network model is fine-tuned by using an annotated ship detection dataset containing ship images to ensure that the fine-tuned model can efficiently and accurately detect ships and can better adapt to specific application scenarios. The COCO dataset is then used to train the DINO target detection model to obtain a trained DINO target detection model, and multiple natural scene pictures containing ships obtained from the first stage of screening and detection are respectively input into the trained DINO target detection model for the second stage of screening and detection to obtain the first target detection results of all ships in each natural scene picture, and the cropped area containing the ship is cropped according to the first target detection result of each ship. The YOLOv5x network model is fine-tuned using the COCO dataset containing various ship images to improve the model's recognition accuracy for small ships. This stage can efficiently detect small boats that are very difficult to detect, and accurately detect all ships to achieve the purpose of inspecting all ships that should be inspected.Then, the contour search function is used to extract the ship contour from the cropped area containing the ship, and according to the first target detection result, the points corresponding to the minimum and maximum values ​​of the horizontal coordinates are respectively extracted from the ship contour as the first contour point and the second contour point. The angle between the line connecting the two contour points and the horizontal direction is calculated according to the first contour point and the second contour point. Then, the rotation matrix function is used to create a rotation matrix. Based on the rotation matrix and the affine transformation function, the ship contour in the cropped area is rotated according to the calculated angle to obtain the cropped area after the ship direction is corrected. By adjusting the ship to an angle close to parallel to the horizontal line, it helps to reduce the error of the model in detection and classification. The ship angle is corrected through contour extraction and affine transformation. By unifying the ship angle, the fluctuation of the detection results caused by different ship directions can be reduced, making the performance of the model more consistent at various angles, thereby improving the stability and effect of the overall system. Finally, the cropped area after ship direction correction is input into the trained YOLOv5nano network model for the third stage detection, and the second target detection result of the ship in the cropped area after ship direction correction is obtained. YOLOv5nano is a lightweight version of the model, which aims to provide fast target detection capabilities for resource-constrained environments. The model reduces the number of parameters and the amount of calculation by reducing the depth and width of the network, significantly reducing the delay time and the required computing resources. The first target detection result output by the YOLOv5nano network model is used to verify whether there is a false detection or missed detection in the first target detection result output by the DINO target detection model. If so, the false detection or missed detection is corrected in combination with the first target detection result and the second target detection result to complete the ship target detection in the natural scene. In this way, the YOLOv5nano network model can verify and improve the output of the DINO target detection model to improve the overall accuracy of ship detection.

[0031] The present invention uses a three-stage deep learning detection algorithm to identify ships, improve the ship detection rate, and at the same time, also improve the accuracy of ship detection. For difficult-to-detect ships, it can greatly improve the recognition precision and accuracy of long-distance and blurred ships. The present invention uses the trained yolov5x network model to screen multiple natural scene pictures to be tested, detects almost all ships at sea, and then uses the trained DINO target detection model to further perform fine detection, greatly improving the detection rate. In addition, after the picture is detected by the trained DINO target detection model, the picture of the target area is cropped, and after affine transformation, the target is turned positive, and then it is input into the trained yolov5nano network model for further detection, which greatly improves the accuracy.

[0032] The present invention also relates to a natural scene ship target detection system, which corresponds to the above-mentioned natural scene ship target detection method, and can be understood as a system for realizing the above-mentioned natural scene ship target detection method, including an image and data set acquisition module, a first-stage ship image screening module, a second-stage ship fine detection and cropping area acquisition module, a ship direction correction module, and a third-stage detection and detection result verification module connected in sequence, each module works in coordination with each other, and uses the yolov5x network model to perform the first-stage screening and detection of multiple natural scene images to be tested, detects almost all ships at sea, and then uses the trained DINO target detection model to further perform the second-stage fine detection, greatly improving the detection rate. In addition, after the image is detected by the trained DINO target detection model, the image of the target area is cropped, and the target is converted to the positive target after affine transformation, and then input into the yolov5nano network model for further third-stage detection, which greatly improves the accuracy. By utilizing a three-stage deep learning detection algorithm (YOLOv5x network model, DINO target detection model, and YOLOv5nano network model) to perform phased recognition of ships in the natural scene images to be tested, the ship detection rate is greatly improved. In addition, for difficult-to-detect ships, the recognition accuracy and precision of distant and blurred ships can be greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a flow chart of the natural scene ship target detection method of the present invention.

[0034] Figure 2 It is a schematic diagram of the first stage of screening images with ships by the trained YOLOv5x network model of the present invention.

[0035] Figure 3 It is a schematic diagram of the second stage of fine screening of all ships by the trained DINO target detection model of the present invention.

[0036] Figure 4 It is a schematic diagram of the present invention for correcting the direction of a ship.

[0037] Figure 5 It is a schematic diagram of further third-stage detection of ships in the cropped area by the trained YOLOv5nano network model of the present invention. DETAILED DESCRIPTION

[0038] The present invention will be described below in conjunction with the accompanying drawings.

[0039] The present invention relates to a method for detecting ship targets in natural scenes. The flowchart of the method is as follows: Figure 1As shown, picture → YOLOv5x → DINO → affine transformation → YOLOv5nano, for multiple natural scene pictures to be tested, a three-stage deep learning detection algorithm is used, such as Figure 1 The YOLOv5x (i.e., YOLOv5x network model), DINO (i.e., DINO target detection model), and YOLOv5nano (i.e., YOLOv5nano network model) shown perform phased recognition of ships, and between the second-stage detection of DINO and the third-stage detection of YOLOv5nano, the angle of the ship is corrected through contour extraction and affine transformation. This solution is targeted at difficult-to-detect ships, and can greatly improve the recognition of distant and blurred ships, improve the ship detection rate, and at the same time, improve the accuracy of ships. High detection rate: The present invention first uses the yolov5x network model (or the yolov5x model algorithm) to perform the first-stage screening and detection of the image, and sets a low threshold to screen out the images with ships, and then uses the DINO network model to perform the second-stage fine detection to screen out almost all ships at sea, greatly improving the detection rate. High accuracy: After the image is detected by the DINO network model (or the DINO model algorithm), the image of the target area is cropped, and the target is transformed through affine transformation, and then input into the YOLOv5nano network model (or the yolov5nano algorithm) for further third-stage detection, which greatly improves the accuracy. Specifically, the natural scene ship target detection method includes the following steps in sequence:

[0040] 1. Image and data set acquisition steps: Obtain the ship detection data set, COCO data set, and the video images of the natural scene to be tested, and perform frame processing on the video images of the natural scene to be tested to obtain multiple pictures of the natural scene to be tested.

[0041] Specifically, first obtain a ship detection dataset and a COCO dataset, wherein the ship detection dataset includes a variety of different types of ship images and annotations, and the COCO dataset includes a plurality of different object categories, such as a combination of people, bicycles, cars, motorcycles, airplanes, buses, trains, trucks, ships, traffic lights, fire hydrants, stop signs, parking meters, benches, etc. Then install a video acquisition camera near a navigable port or on a ship, and take pictures of nearby sailing ships to obtain natural scene video images in different scenes as the natural scene video images to be tested, and perform frame processing on the natural scene video images to be tested to obtain a plurality of natural scene pictures to be tested, that is, obtain a natural scene picture of each frame.

[0042] 2. The first stage of screening steps for images with ships: Use the ship detection dataset as the training set samples to train the constructed YOLOv5x network model and YOLOv5nano network model respectively, and obtain the trained YOLOv5x network model and YOLOv5nano network model respectively, and input multiple natural scene images to be tested into the trained YOLOv5x network model for the first stage of screening and detection, and obtain multiple natural scene images containing ships.

[0043] Specifically, Figure 2 As shown in the figure, the YOLOv5x network model is first trained using the ship detection data set as a training set sample to obtain a trained YOLOv5x network model, and then multiple natural scene images to be tested are preprocessed, such as resizing, normalization, etc., to meet the input requirements of the trained YOLOv5x network model, and then the preprocessed multiple natural scene images to be tested are respectively input into the trained YOLOv5x network model for the first stage of screening and detection, and the ship detection results of each natural scene image are output. Using the characteristics of yolov5x that both speed and accuracy are taken into account, images with ships can be quickly screened out from a large number of images. Ocean-going ships sail on the sea, and most of the time there is no ship on the sea, but there may be small ships in the distance, so a very fast and high-precision model is needed to filter out most of the images without ships, retaining images with ships, or even images with only small boats in the distance.

[0044] The ship detection result includes a bounding box, a category label, and a confidence score. Then, the confidence scores of all ship detection results in a natural scene image are compared with a preset confidence threshold (such as 0.5). If there is at least one ship detection result with a confidence score greater than or equal to the confidence threshold (such as 0.5), the detected ship result is considered valid, and the natural scene image is marked as a picture containing a ship, that is, the picture with a ship detected by YOLOv5x shown in b). Otherwise, the natural scene image is not marked, that is, the picture without a ship shown in a), and finally multiple natural scene images containing ships are obtained.

[0045] 3. Second-stage ship fine detection and cropping area acquisition steps: Use the COCO dataset to train the DINO target detection model to obtain a trained DINO target detection model, and input multiple natural scene pictures containing ships obtained in the first-stage screening and detection into the trained DINO target detection model for the second-stage screening and detection to obtain the first target detection results of all ships in each natural scene picture, and crop the cropping area containing the ship according to the first target detection result of each ship.

[0046] This step finely detects all ships and cuts out specific areas, such as Figure 3 As shown in the figure, first use the COCO dataset to train the DINO object detection model (preferably the 5scale-swin-L large model of the DINO object detection model) to obtain a trained DINO object detection model. After fine-tuning the trained DINO object detection model on the COCO dataset, the mAP 50 can reach 63. Among them, mAP (mean Average Precision) is an index in object detection, which is used to measure the detection accuracy of the model. mAP 50 represents the average precision when the IoU (Intersection over Union) is greater than 0.5 in the detections of different categories. The mAP 50 reaching 63 indicates that various scales and small targets can be detected, achieving the purpose of detecting all that should be detected. Then, resize and standardize multiple natural scene images containing ships selected by the YOLOv5x network model to meet the input requirements of the trained DINO model. Then, input multiple natural scene images containing ships into the trained DINO object detection model respectively to obtain the first object detection results of all ships in each natural scene image, and crop out the cropping area containing the ship according to the first object detection result of each ship.

[0047] IV. Ship direction correction steps: Use the contour search function to extract the ship contour from the cropping area containing the ship, and extract the points corresponding to the minimum abscissa and the maximum abscissa from the ship contour respectively as the first contour point and the second contour point. Calculate the angle between the line connecting the two contour points and the horizontal direction according to the first contour point and the second contour point. Then use the rotation matrix function to create a rotation matrix, and based on the rotation matrix, use the affine transformation function to rotate the ship contour in the cropping area according to the calculated angle to obtain the cropping area after ship direction correction.

[0048] Specifically, as Figure 4 shown in the figure, correct the angle of the target ship through contour extraction and affine transformation. First, use the contour search function cv2.findContours to extract the ship contour from the cropping area, and extract the points corresponding to the minimum abscissa and the maximum abscissa from the ship contour respectively as the first contour point (xmin, y1) and the second contour point (xmax, y2). Calculate the angle ɑ between the line connecting the two contour points and the horizontal direction according to the first contour point and the second contour point, and calculate according to the following formula:

[0049]

[0050] In the above formula, when y2 > y1, k = 0; when y2 < y1, k = 1.

[0051] Then, the rotation matrix function cv2.getRotationMatrix2D is used to create a rotation matrix. Based on the rotation matrix and the affine transformation function cv2.warpAffine, the ship in the cropping area is rotated according to the calculated angle ɑ to correct the direction of the ship, that is, rotate the ship counterclockwise by an angle ɑ and correct it to the horizontal direction to obtain the cropping area after the ship direction is corrected.

[0052] 5. Third-stage detection and verification steps of detection results: Input the cropped area after ship direction correction into the lightweight trained YOLOv5nano network model for third-stage detection, and obtain the second target detection result of the ship in the cropped area after ship direction correction, and verify whether the first target detection result output by the second-stage screening detection of the trained DINO target detection model has false detection or missed detection according to the second target detection result output by the trained YOLOv5nano network model. If so, the false detection or missed detection is corrected in combination with the first target detection result and the second target detection result to complete the ship target detection in the natural scene.

[0053] Specifically, this step uses the lightweight yolov5nano network model to further detect the cropped area box containing the target object output by the DINO network model to ensure the accuracy of ship detection. Figure 5 As shown, each cropped area is first resized and normalized to meet the input requirements of the YOLOv5nano model; then the cropped area after the ship direction correction is input into the YOLOv5nano network model to obtain the second target detection result of the ship in the cropped area after the ship direction correction, and the first target detection result output by the DINO target detection model is verified based on the second target detection result output by the YOLOv5nano network model to see if there is any false detection or missed detection. If not, the second target detection result is used as the final ship target detection result; if so, the false detection or missed detection is corrected in combination with the first target detection result and the second target detection result to complete the ship target detection in the natural scene. For example,

[0054] 1) False detection processing

[0055] Confirming false positives: Identify false positive boxes that do not belong to ships from the results of YOLOv5nano.

[0056] Eliminate false positives: Based on the detection results of the DINO model, remove these false positive boxes from the final detection results.

[0057] 2) Missed detection processing

[0058] Identify missed detections: Find the ship targets that were not detected by the DINO model but were detected by YOLOv5nano from the results of YOLOv5nano.

[0059] Supplement missed detections: Add missed detection targets identified by YOLOv5nano to the detection results of the DINO model and re-label the locations of these targets.

[0060] 3) Comprehensive results:

[0061] Merge: Combine the first target detection result output by the DINO target detection model and the second target detection result output by the YOLOv5nano network model, update and correct the target detection box, and ensure that the final result is as accurate as possible. By using the YOLOv5nano network model to improve the detection results of the DINO target detection model, the accuracy of the overall detection system can be effectively enhanced.

[0062] The present invention also relates to a natural scene ship target detection system, which corresponds to the above-mentioned natural scene ship target detection method and can be understood as a system for implementing the above-mentioned method. The system includes an image and data set acquisition module, a first-stage ship image screening module, a second-stage ship fine detection and cropping area acquisition module, a ship direction correction module, and a third-stage detection and detection result verification module connected in sequence. Specifically,

[0063] The image and data set acquisition module acquires the ship detection data set, the COCO data set, and the video image of the natural scene to be tested, and performs frame processing on the video image of the natural scene to be tested to obtain multiple pictures of the natural scene to be tested;

[0064] The first stage includes a ship image screening module, which uses the ship detection data set as a training set sample to train the YOLOv5x network model and the YOLOv5nano network model respectively, and obtains the trained YOLOv5x network model and the YOLOv5nano network model respectively, and inputs a plurality of natural scene images to be tested into the trained YOLOv5x network model for the first stage screening test, and obtains a plurality of natural scene images containing ships;

[0065] The second-stage ship fine detection and cropping region acquisition module uses the COCO dataset to train the DINO target detection model to obtain a trained DINO target detection model, and inputs multiple natural scene pictures containing ships obtained by the first-stage screening detection into the trained DINO target detection model for the second-stage screening detection to obtain the first target detection results of all ships in each natural scene picture, and crops out the cropping region containing the ship according to the first target detection result of each ship;

[0066] The ship direction correction module uses a contour search function to extract a ship contour from a clipping area containing the ship, and extracts points corresponding to a minimum abscissa value and a maximum abscissa value from the ship contour according to the first target detection result, as first contour points and second contour points, calculates an angle between a line connecting the two contour points and a horizontal direction according to the first contour points and the second contour points, and then uses a rotation matrix function to create a rotation matrix, and based on the rotation matrix and using an affine transformation function, rotates the ship contour in the clipping area according to the calculated angle to obtain a clipping area after the ship direction is corrected;

[0067] The third-stage detection and detection result verification module inputs the cropped area after the ship direction correction into the lightweight trained YOLOv5nano network model for the third-stage detection, obtains the second target detection result of the ship in the cropped area after the ship direction correction, and verifies whether the first target detection result output by the second-stage screening detection of the trained DINO target detection model has false detection or missed detection according to the second target detection result output by the trained YOLOv5nano network model. If not, the second target detection result is used as the final ship target detection result; if so, the false detection or missed detection is corrected in combination with the first target detection result and the second target detection result to complete the ship target detection in the natural scene.

[0068] Preferably, in the first-stage ship image screening module, multiple natural scene images to be tested are respectively input into the trained YOLOv5x network model for first-stage screening and detection, and multiple natural scene images containing ships are obtained, specifically including:

[0069] Multiple natural scene images to be tested are respectively input into the trained YOLOv5x network model, and all ship detection results in each natural scene image are output. The ship detection results include bounding boxes, category labels and confidence scores. The confidence scores of all ship detection results in a natural scene image are compared with the preset confidence threshold. If there is at least one ship detection result with a confidence score greater than or equal to the confidence threshold, the natural scene image is considered to contain ships and is saved and marked, and finally multiple natural scene images containing ships are obtained.

[0070] Preferably, in the third stage detection and detection result verification module, the correction of false detection or missed detection in combination with the first target detection result and the second target detection result includes: if there is a false detection, then according to the first target detection result output by the DINO target detection model, the false detection box that does not belong to the ship identified in the second target detection result output by the YOLOv5nano network model is removed from the final detection result; if there is a missed detection, the missed detection target identified by the YOLOv5nano network model is added to the first target detection result output by the DINO model, and the position of the target is re-marked.

[0071] Preferably, in the second-stage ship fine detection and cropping area acquisition module, before the multiple natural scene pictures containing ships obtained by the first-stage screening and detection are respectively input into the trained DINO target detection model, the multiple natural scene pictures containing ships screened by YOLOv5x are first resized and standardized to meet the input requirements of the trained DINO target detection model.

[0072] Preferably, the ship detection dataset includes multiple different types of ship images and annotations, and the COCO dataset includes multiple different object categories, and the multiple different object categories include several combinations of people, bicycles, cars, motorcycles, airplanes, buses, trains, trucks, ships, traffic lights, fire hydrants, stop signs, parking meters, and benches.

[0073] The present invention provides an objective and scientific natural scene ship target detection method and system. By utilizing a three-stage deep learning detection algorithm (YOLOv5x network model, DINO target detection model and YOLOv5nano network model), the ships in the natural scene pictures to be tested are sequentially identified in stages, thereby greatly improving the ship detection rate. In addition, for difficult-to-detect ships, the recognition precision and accuracy of long-distance and blurred ships can be greatly improved.

[0074] It should be noted that the above-described specific implementations can enable those skilled in the art to more fully understand the invention, but do not limit the invention in any way. Therefore, although this specification has described the invention in detail with reference to the drawings and embodiments, those skilled in the art should understand that the invention can still be modified or replaced by equivalents. In short, all technical solutions and improvements that do not deviate from the spirit and scope of the invention should be included in the protection scope of the patent for the invention.

Claims

1. A method for detecting ship targets in natural scenes, characterized in that: The following steps are involved: Image and data set acquisition steps: obtain the ship detection data set, COCO data set, and the natural scene video image to be tested, and perform frame processing on the natural scene video image to be tested to obtain multiple natural scene pictures to be tested; The first stage of ship image screening step: using the ship detection data set as a training set sample to train the YOLOv5x network model and the YOLOv5nano network model respectively, to obtain the trained YOLOv5x network model and the YOLOv5nano network model respectively, and inputting a plurality of natural scene images to be tested into the trained YOLOv5x network model for the first stage screening test, to obtain a plurality of natural scene images containing ships; The second stage of ship fine detection and cropping area acquisition steps: using the COCO dataset to train the DINO target detection model to obtain a trained DINO target detection model, and inputting multiple natural scene pictures containing ships obtained by the first stage screening detection into the trained DINO target detection model for the second stage screening detection to obtain the first target detection results of all ships in each natural scene picture, and cropping the cropping area containing the ship according to the first target detection result of each ship; Ship direction correction step: Use the contour search function to extract the ship contour from the cropped area containing the ship, and extract the points corresponding to the minimum value of the horizontal coordinate and the maximum value of the horizontal coordinate from the ship contour according to the first target detection result, as the first contour point and the second contour point, calculate the angle between the connecting line of the two contour points and the horizontal direction according to the first contour point and the second contour point, and then use the rotation matrix function to create a rotation matrix, based on the rotation matrix and using the affine transformation function to rotate the ship contour in the cropped area according to the calculated angle, so as to obtain the cropped area after the ship direction correction; The third stage detection and detection result verification steps are as follows: the cropped area after ship direction correction is input into the lightweight trained YOLOv5nano network model for the third stage detection, and the second target detection result of the ship in the cropped area after ship direction correction is obtained, and the first target detection result output by the trained YOLOv5nano network model is used to verify whether there is a false detection or missed detection in the first target detection result output by the second stage screening detection of the trained DINO target detection model. If not, the second target detection result is used as the final ship target detection result; if so, the false detection or missed detection is corrected in combination with the first target detection result and the second target detection result to complete the ship target detection in the natural scene.

2. The natural scene ship target detection method according to claim 1 is characterized in that: In the first stage of the ship image screening step, multiple natural scene images to be tested are respectively input into the trained YOLOv5x network model for the first stage screening test, and multiple natural scene images containing ships are obtained, specifically including: Multiple natural scene images to be tested are respectively input into the trained YOLOv5x network model, and all ship detection results in each natural scene image are output. The ship detection results include bounding boxes, category labels and confidence scores. The confidence scores of all ship detection results in a natural scene image are compared with the preset confidence threshold. If there is at least one ship detection result with a confidence score greater than or equal to the confidence threshold, the natural scene image is considered to contain ships and is saved and marked, and finally multiple natural scene images containing ships are obtained.

3. The natural scene ship target detection method according to claim 1, characterized in that: In the third stage detection and detection result verification step, the false detection or missed detection is corrected in combination with the first target detection result and the second target detection result, including: if a false detection exists, the false detection box that does not belong to the ship identified in the second target detection result output by the YOLOv5nano network model is removed from the final detection result according to the first target detection result output by the DINO target detection model; if a missed detection exists, the missed detection target identified by the YOLOv5nano network model is added to the first target detection result output by the DINO model, and the position of the target is re-marked.

4. The natural scene ship target detection method according to any one of claims 1 to 3, characterized in that: In the second stage of ship fine detection and cropping area acquisition steps, before the multiple natural scene pictures containing ships obtained by the first stage screening and detection are respectively input into the trained DINO target detection model, the multiple natural scene pictures containing ships screened by YOLOv5x are first resized and standardized to meet the input requirements of the trained DINO target detection model.

5. The natural scene ship target detection method according to claim 1, characterized in that: In the image and dataset acquisition step, the acquired ship detection dataset includes multiple different types of ship images and annotations, and the COCO dataset includes multiple different object categories, and the multiple different object categories include several combinations of people, bicycles, cars, motorcycles, airplanes, buses, trains, trucks, ships, traffic lights, fire hydrants, stop signs, parking meters, and benches.

6. A natural scene ship target detection system, characterized in that: It includes the sequentially connected image and data set acquisition module, the first-stage ship image screening module, the second-stage ship fine detection and cropping area acquisition module, the ship direction correction module and the third-stage detection and detection result verification module. The image and data set acquisition module acquires the ship detection data set, the COCO data set, and the video image of the natural scene to be tested, and performs frame processing on the video image of the natural scene to be tested to obtain multiple pictures of the natural scene to be tested; The first stage includes a ship image screening module, which uses the ship detection data set as a training set sample to train the YOLOv5x network model and the YOLOv5nano network model respectively, and obtains the trained YOLOv5x network model and the YOLOv5nano network model respectively, and inputs a plurality of natural scene images to be tested into the trained YOLOv5x network model for the first stage screening test, and obtains a plurality of natural scene images containing ships; The second-stage ship fine detection and cropping region acquisition module uses the COCO dataset to train the DINO target detection model to obtain a trained DINO target detection model, and inputs multiple natural scene pictures containing ships obtained by the first-stage screening detection into the trained DINO target detection model for the second-stage screening detection to obtain the first target detection results of all ships in each natural scene picture, and crops out the cropping region containing the ship according to the first target detection result of each ship; The ship direction correction module uses a contour search function to extract a ship contour from a clipping area containing the ship, and extracts points corresponding to a minimum abscissa value and a maximum abscissa value from the ship contour according to the first target detection result, as first contour points and second contour points, calculates an angle between a line connecting the two contour points and a horizontal direction according to the first contour points and the second contour points, and then uses a rotation matrix function to create a rotation matrix, and based on the rotation matrix and using an affine transformation function, rotates the ship contour in the clipping area according to the calculated angle to obtain a clipping area after the ship direction is corrected; The third-stage detection and detection result verification module inputs the cropped area after the ship direction correction into the lightweight trained YOLOv5nano network model for the third-stage detection, obtains the second target detection result of the ship in the cropped area after the ship direction correction, and verifies whether the first target detection result output by the second-stage screening detection of the trained DINO target detection model has false detection or missed detection according to the second target detection result output by the trained YOLOv5nano network model. If not, the second target detection result is used as the final ship target detection result; if so, the false detection or missed detection is corrected in combination with the first target detection result and the second target detection result to complete the ship target detection in the natural scene.

7. The natural scene ship target detection system according to claim 6, characterized in that: In the first stage of the ship image screening module, multiple natural scene images to be tested are respectively input into the trained YOLOv5x network model for the first stage screening test, and multiple natural scene images containing ships are obtained, including: Multiple natural scene images to be tested are respectively input into the trained YOLOv5x network model, and all ship detection results in each natural scene image are output. The ship detection results include bounding boxes, category labels and confidence scores. The confidence scores of all ship detection results in a natural scene image are compared with the preset confidence threshold. If there is at least one ship detection result with a confidence score greater than or equal to the confidence threshold, the natural scene image is considered to contain ships and is saved and marked, and finally multiple natural scene images containing ships are obtained.

8. The natural scene ship target detection system according to claim 6, characterized in that: In the third-stage detection and detection result verification module, the correction of false detection or missed detection is performed in combination with the first target detection result and the second target detection result, including: if a false detection exists, the false detection box that does not belong to the ship identified in the second target detection result output by the YOLOv5nano network model is removed from the final detection result according to the first target detection result output by the DINO target detection model; if a missed detection exists, the missed detection target identified by the YOLOv5nano network model is added to the first target detection result output by the DINO model, and the position of the target is re-marked.

9. The natural scene ship target detection system according to any one of claims 6 to 8, characterized in that: In the second-stage ship fine detection and cropping area acquisition module, before the multiple natural scene pictures containing ships obtained by the first-stage screening and detection are respectively input into the trained DINO target detection model, the multiple natural scene pictures containing ships screened by YOLOv5x are first resized and standardized to meet the input requirements of the trained DINO target detection model.

10. The natural scene ship target detection system according to claim 6, characterized in that: The ship detection dataset includes multiple different types of ship images and annotations, and the COCO dataset includes multiple different object categories, which include several combinations of people, bicycles, cars, motorcycles, airplanes, buses, trains, trucks, ships, traffic lights, fire hydrants, stop signs, parking meters, and benches.

Citation Information

Patent Citations

  • Remote sensing image ship detection method and device based on attention model

    CN114677596A

  • Small ship detection method and system

    CN117274925A

  • Target detection method based on OTS-DETR lightweight model

    CN118262087A

  • Infrared ship detection method based on improved RT-DETR algorithm

    CN119169453A

  • A system of detecting abnormal action

    KR102628689B1

Cited By

  • Bridge area water area ship target detection method, equipment and medium

    CN121214301A

  • A bridge water area ship target detection method, device and medium

    CN121214301B