A water column detection method for impact point based on improved YOLOv4 algorithm

By improving the YOLOv4 algorithm and combining it with Huffman line detection, DBSCAN and K-means clustering as well as a hybrid attention mechanism, the problem of low accuracy in water column detection and positioning is solved, and high-precision water column detection and positioning is achieved, which is suitable for the automatic picking system of naval artillery training.

CN115909072BActive Publication Date: 2025-09-30NAVAL UNIV OF ENG PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211504834.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-09-30
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

The existing technology has problems with low water column detection and positioning accuracy and large positioning deviation, which affects the accuracy of shooting evaluation especially in naval artillery training.

Method used

An improved YOLOv4 algorithm is used in combination with the Huffman line detection method to constrain sensitive areas. The DBSCAN clustering algorithm and K-means clustering analysis are used to optimize the prior bounding box. A hybrid attention mechanism is added to the path aggregation network to improve feature extraction capabilities.

Benefits of technology

The detection accuracy and positioning accuracy of the water column are significantly improved, and it has a high-precision and robust detection effect, which is suitable for the automatic picking system of the water column at the impact point.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115909072B_ABST
    Figure CN115909072B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting water columns at impact points based on an improved YOLOv4 algorithm. The method comprises the following steps: first, detecting a sea-sky line by using a Huffman line detection method, and constraining sensitive areas in a current detection image; second, performing cluster analysis on a water column marker frame by using a density-based DBSCAN clustering algorithm and a distance-based K-means clustering algorithm, obtaining a better prior anchor frame, and inputting the frame into a YOLOv4 network; and third, adding a hybrid attention mechanism to a path aggregation network to improve the detection accuracy and positioning accuracy of weakly characteristic water column targets, and extracting fuzzy water columns by using an improved YOLOv4 PANet structure. By improving YOLOv4 and using the improved YOLOv4 algorithm to detect and locate characteristic water columns, the method can effectively improve the detection accuracy and positioning accuracy of water columns, and has the characteristics of high detection and positioning accuracy and good detection effect robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water column detection, and in particular to a method for detecting water columns at impact points based on an improved YOLOv4 algorithm. Background Art

[0002] For a long time, the Navy has generally used manual methods to detect the offset distance between the impact point and the target during naval artillery training. This method suffers from low accuracy, subjectivity, and low efficiency. It is urgent to establish a new high-precision automatic target picking system to guide troops in scientific and effective training. In the automatic target selection system, the accuracy of water column detection and positioning will directly affect the shooting evaluation results.

[0003] Water column detection belongs to mobile target detection, and the application of traditional algorithms in mobile target detection has been widely studied. In "Computer Engineering and Applications", the author sorted out the development and current status of target detection algorithms, and proposed research prospects; traditional algorithms include scale-invariant feature transform (SIFT) algorithm, PCA-SIFT algorithm, speeded up robust feature (SURF) algorithm, etc.; and pointed out that PCA-SIFT algorithm and SURF algorithm simplify the matching process of SIFT respectively, so the speed is relatively fast, but the detection accuracy and positioning accuracy based on feature point matching will also decrease; in "Infrared Physics and Technology", the author proposed a target detection method based on a new cascade classifier, which consists of a multi-scale block-local binary pattern (MB-LBP) feature classifier, a SIFT feature classifier and a SURF The feature classifier is composed of high accuracy. In "Ma Qin, Zhang Xingzhong, Li Haifang, Deng Hongxia. Research on motion target detection based on spectral residual and clustering method [J]. Computer Engineering and Science, 2018, 40(10): 1867-1873. [7]", the authors proposed a moving target detection algorithm based on spectral residual algorithm and K-means clustering algorithm. The method first extracts the accelerator features every two frames and registers the images. Secondly, the spectral residual algorithm is used to extract visual salient features from the frame difference results to remove noise and pseudo-moving targets caused by inaccurate matching. Finally, after morphological processing, the improved K-means clustering algorithm is introduced to cluster the discontinuous contours to form a complete target.

[0004] In order to overcome the limitations of traditional algorithms, people have developed target detection based on deep learning to detect mobile targets; for example, in "Fang Luping, He Jiang, Zhou Guomin. A review of target detection algorithms; "Computer Engineering and Applications", 2018, pp. 54: 11-18", the authors summarized the target detection algorithms based on deep learning, including the Faster Regional Convolutional Network (Faster R-CNN) algorithm, SSD algorithm, YOLO, YOLOv2, YOLOv3, YOLOv4, YOLOv5 algorithm, etc., and pointed out that the YOLO series and the single-shot multi-box detector (SSD) algorithm both follow the R-CNN series algorithm. The method is used to pre-train the classification of large data sets, and then find adjustments on small data sets; in "Li Mengjiao, Wang Hao, Wan Zhibo, Steel Bar Surface Defect Detection Based on Improved YOLOv4, "Computer and Electrical Engineering", 2022, 102: 45-53", the authors proposed a complex pedestrian detection model based on the improved YOLOv4 algorithm to solve the problems of pedestrian posture, scale diversity and pedestrian occlusion; the model uses an improved K-means clustering algorithm to analyze the real frame size of the pedestrian data set, and then uses the path aggregation network (PANet) to perform multi-scale feature fusion to improve the detection effect; in " Li Tan, Lv Xinyue, Lian Xiaofeng, Wang Ge, YOLOv4_Drone: Target Detection in UAV Images Based on Improved YOLOv4 Algorithm, "Computers and Electrical Engineering", 2021, 93:45:52", the authors proposed a steel strip surface defect detection method based on the improved "Look Only Once" version 4 (YOLOv4) algorithm, which embeds the attention mechanism into the backbone network structure, modifies the path aggregation network into a customized receptive field block structure, and strengthens the feature extraction function of the network model; in "Bochkovskiy, A; Wang CY; Liao, HYMYOLOv4: Target Detection in UAV Images Based on Improved YOLOv4 Algorithm", ... In the paper "Optimal Speed ​​and Accuracy of Water Column Detection [J], arXiv, 2020, 204, 109-124", the authors used the hollow convolution technique to resample the feature image to improve feature extraction and target detection performance; in addition, the ultra-lightweight subspace attention mechanism (ULSAM) was used to generate different attention feature maps for each subspace of the feature map for multi-scale feature representation. Finally, soft non-maximum suppression (Soft-NMS) was introduced to reduce the occurrence of missed targets due to occlusion. Although the above methods have their own advantages in water column detection, they still have problems such as low water column detection and positioning accuracy and large positioning deviation.

[0005] Therefore, it is urgent to design a new water column detection algorithm to detect and locate the water column to solve the problems existing in the above-mentioned existing technologies. Summary of the Invention

[0006] In response to the above-mentioned problems, the present invention aims to provide a method for detecting water columns at impact points based on an improved YOLOv4 algorithm. By improving YOLOv4 and using the improved YOLOv4 algorithm to detect and locate characteristic water columns, the method can effectively improve the detection accuracy and positioning accuracy of water columns, and has the characteristics of high detection and positioning accuracy and good detection robustness.

[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] A water column detection method for impact point based on improved YOLOv4 algorithm, including

[0009] Step 1: Use the Huffman line detection method to detect the sea-sky line and constrain the sensitive area of ​​water column detection in the current detection image;

[0010] Step 2: Use the density-based DBSCAN clustering algorithm + the distance-based K-means clustering algorithm to perform cluster analysis on the water column marker frame to obtain a better prior anchor frame and input it into the YOLOv4 network;

[0011] Step 3: Add a hybrid attention mechanism to the path aggregation network, improve the YOLOv4 network structure, and use the improved YOLOv4 PANet structure to extract the fuzzy water column.

[0012] Preferably, the process of detecting the sea-sky line by using the Hoffman line detection method in step 1 and constraining the sensitive area of ​​the water column detection in the current detection image includes:

[0013] Step 1.1. Detect sea-sky line using Huffman line detection method

[0014] (1) Using the Huffman line detection method to detect the duality of points and lines, the number of curve intersections in the Huffman space is detected, and a "voter" matrix is ​​constructed to approximate the Huffman space;

[0015] (2) add 1 to the vote representing the linear position of the valid point in the image among the voters;

[0016] (3) By sorting the votes, the point with the most votes is selected, and then the point is converted into a line in the rectangular coordinate system to complete the sea-sky line detection;

[0017] In step 1.2, the area above and below the sea-sky line, which is 1 / 3 of the image height, is set as the sensitive area, thereby constraining the water column detection area in the current image.

[0018] Preferably, the process of constraining the sensitive area in the current sea-sky-line detection image described in step 1.2 includes:

[0019] Step 1.21. First, reduce the original sea-sky-line detection image to 480*640;

[0020] Step 1.22. Next, use the Canny algorithm to calculate the feature value of each pixel in the image using a 3x3 horizontal and vertical operator template. Then, use binarization to convert the image into a binary image, showing the texture characteristics of the image edge image.

[0021] Step 1.23. Finally, use the Hoffman line detection algorithm to detect the longest line segment in the map;

[0022] Step 1.24. If the number of target pixels in the longest line segment in the map detected in step 1.23 is greater than 2 / 3 of the line segment length, it is considered to be sea level.

[0023] Preferably, the process of detecting the longest line segment in the map using the Hoffman line detection algorithm described in step 1.23 includes:

[0024] (1) First, for the target point P1, draw a straight line L1 through the point with α = 0 degrees (counterclockwise is positive);

[0025] (2) Calculate the distance ρ from the origin to the line L1;

[0026] (3) Statistical angle θ between the perpendicular line from the origin to the line L1 and the x-axis (positive in counterclockwise direction);

[0027] (4) Loop through steps (1) to (3), taking α as the variable and a step size of 10 degrees for point P1, and taking the angles from 0 to 180 degrees to count the distances ρ and angles θ of all points;

[0028] (5) Loop through step (4) and take all target points P i , count the distance ρ and angle θ of all points;

[0029] (6) In the Huffman space, the point with the largest number of identical values ​​of (ρ, θ) is converted into a rectangular coordinate system, which is the line to be extracted.

[0030] Preferably, the process of clustering the water column marker frame using the density-based DBSCAN clustering algorithm + the distance-based K-means clustering algorithm described in step 2 includes:

[0031] Step 2.1. Set the neighborhood radius of the DBSCAN algorithm to 0.1 and the number of nearest neighbor samples to 25. Use the DBSCAN algorithm to perform density partitioning on clusters with the water column marker box size.

[0032] Step 2.2. Use the K-means clustering algorithm to divide the clusters with the highest density in DBSCAN according to the K value;

[0033] Step 2.3. Input the target annotation box size dataset as the initial dataset into the clustering algorithm, and obtain a set of prior anchor boxes according to steps 2.1 and 2.2.

[0034] Preferably, the K value in step 2.2 is 5, and the process of dividing the high-density clusters according to the K value using the K-means clustering algorithm includes:

[0035] Step 2.21. Randomly generate five cluster centers in the data. Calculate the distance from each sample to these five cluster centers and assign the corresponding sample to the cluster corresponding to the cluster center with the smallest distance.

[0036] Step 2.22. After all samples are classified, recalculate the cluster center of each cluster for each of the five clusters, that is, the centroid of all samples in each cluster, and repeat the above operation until the cluster center does not change.

[0037] Preferably, the process of adding a hybrid attention mechanism to the path aggregation network in step 3 and extracting the fuzzy water column using the improved YOLOv4 PANet structure includes:

[0038] Step 3.1. Improve YOLOv4 by adding a hybrid attention mechanism module to each of the three branches at the end of the YOLOv4 feature fusion network.

[0039] Step 3.2 uses the improved YOLOv4PANet structure to extract the fuzzy water column.

[0040] Preferably, the improved YOLOv4 algorithm network structure described in step 3.1 includes a Backbone network, a Neck network, and a Head network.

[0041] The backbone network is a CSPParknet53 network, which consists of five CSPNet residual blocks, and the activation function of the CSPParknet53 network is a Mish activation function;

[0042] The Neck network consists of an SPP module and a PANet, wherein the SPP module performs maximum pooling operations of different scales: 13×13, 9×9, 5×5, and 1×4; the PANet is improved based on the feature pyramid network;

[0043] The final output of the Head network is 13×13, 26×26, and 52×5252 detection heads, which detect large, medium, and small targets respectively.

[0044] Preferably, the process of extracting the fuzzy water column using the improved YOLOv4PANet structure described in step 3.2 includes:

[0045] Step 3.21. Assume that the size of the input image is 608*608*3, and the size of the input feature map to the first CBAM module is 152*152*256;

[0046] Step 3.22. The input feature map is passed through two parallel MaxPool layers and Avgpool layers, changing the feature map from 152*152*252*1*256 to 1*1*256;

[0047] Step 3.23. The feature map passes through the Share MLP module, first compressing the number of channels to 1*1*16, and then expanding it to 1*1*256;

[0048] Step 3.24. Add the two results obtained from the feature map of the ReLU activation function and obtain the output of the channel attention mechanism through the feature map of the sigmoid activation function.

[0049] Step 3.25. Multiply the output with the initial input feature map to obtain a feature map enhanced by the channel attention mechanism.

[0050] Step 3.26. Convert the attention-enhanced feature map to a 152*152*1 feature map through max pooling and average pooling. Concatenate the two feature maps and then convert them to a 1-channel feature map through a 7*7 convolution. Obtain a normalized weight map through the sigmoid function. Multiply the weight map by the channel-attention-enhanced feature map to obtain a spatial attention feature map of size 152*152*256.

[0051] Step 3.27. Using the feature map generated by the spatial attention mechanism, YOLO Head analyzes the position loss, classification loss, and confidence loss to complete the water column detection of the impact point.

[0052] The beneficial effects of the present invention are as follows: the present invention discloses a method for detecting water columns at impact points based on an improved YOLOv4 algorithm. Compared with the prior art, the improvements of the present invention are:

[0053] The present invention proposes a method for detecting water columns at impact points based on an improved YOLOv4 algorithm. When used, firstly, the method detects the sea-sky line through the Huffman line detection method, constrains the sensitive areas in the current detection image, and improves the detection accuracy of water column detection; secondly, the DBSCAN+K-means clustering analysis algorithm is used to improve the size selection of the prior bounding box of YOLOv4, making the prior bounding box more typical; finally, CBAM is added to PANet to improve the detection accuracy and positioning accuracy of weak feature water column targets; experimental results show that the above algorithm can effectively improve the detection accuracy and positioning accuracy of water columns at impact points, and has the advantages of high detection and positioning accuracy and good detection robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is an algorithm flow chart of the water column detection method of the impact point based on the improved YOLOv4 algorithm of the present invention.

[0055] Figure 2 This is the algorithm structure diagram of the YOLOV4 algorithm of the present invention.

[0056] Figure 3 This is the Huffman linear detection framework diagram of the present invention.

[0057] Figure 4 This is the marine antenna detection diagram based on Huffman space of the present invention.

[0058] Figure 5 Schematic diagram of the clustering algorithm of the present invention.

[0059] Figure 6 This is the clustering result diagram of the DBSCAN+K-means clustering algorithm run by the present invention.

[0060] Figure 7 This is the prior anchor graph generated by the present invention based on the DBSCAN+K-means clustering algorithm.

[0061] Figure 8 This is a schematic diagram of the structure of the Neck module of the YOLOV4 algorithm of the present invention.

[0062] Figure 9 This is a comparison chart of the small water column detection results of the present invention.

[0063] Figure 10 This is the ablation experiment effect diagram of the improved YOLOv4 algorithm of the present invention.

[0064] Figure 11 This is a diagram showing the comparative operation results of the algorithm in Example 2 of the present invention.

[0065] Figure 12 This is a pixel deviation curve diagram of locating water column targets using the target detection algorithm of the present invention.

[0066] Among them: Figure 3 In FIG. 1 , FIG. (a) is a schematic diagram of four scattered points in the rectangular coordinate system of the present invention, and FIG. (b) is a schematic diagram of the conversion of the rectangular coordinate system of the present invention into the Hoffman space; Figure 4 In the figure, Figures (a), (b), and (c) are sea-sky-line diagrams of different inclinations of the present invention, and Figures (d), (e), and (f) are Huffman line detection result diagrams of the present invention;

[0067] exist Figure 5 In the figure, Figure (a) is the K-means clustering result diagram of the present invention, and Figure (b) is the DBSCAN clustering result diagram of the present invention;

[0068] exist Figure 7 In the figure, (a) is the water column annotation box diagram of the present invention, and (b) is the prior box diagram determined by the present invention based on K-means. It can be seen that the detection box of the YOLOv4 algorithm is significantly larger than the annotation box. (c) is the prior bounding box diagram determined by the DBSCAN+K-means algorithm of the present invention. It can be seen that the detection box of the YOLOv4 algorithm is basically consistent with the annotation box.

[0069] exist Figure 9 In the figure, (a) is the detection result of the YOLOv4 algorithm of the present invention. It can be seen that the small water column is not detected. Figure (b) is the detection result of the improved YOLOv4 algorithm of the present invention. It can be seen that all water columns are detected.

[0070] exist Figure 10 In the figure, Figure (a) is a small water column detection effect diagram of the present invention, and Figure (b) is a fuzzy water column detection effect diagram of the present invention;

[0071] exist Figure 11 Among them, Figure (a) is a clear water column operation result diagram of the present invention, Figure (b) is a fuzzy water column operation result diagram of the present invention, Figure (c) is a small wave operation result diagram of the present invention, and Figure (d) is a large water column operation result diagram of the present invention. DETAILED DESCRIPTION

[0072] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0073] Example 1: Refer to the attached Figure 1-12 The water column detection method of the impact point based on the improved YOLOv4 algorithm shown in FIG.

[0074] Step 1: Detect the sea-sky line using the Huffman line detection method to constrain the sensitive area of ​​water column detection in the current detection image, thereby improving the detection accuracy of water column detection.

[0075] The Huffman line detection process requires a combination of region growing theory, Huffman space transformation, and statistical principles. The basic principle of the region growing algorithm is to merge pixels with similar attributes in the image to achieve the segmentation goal between the object and the background. Since the pixels that make up the same object have a high degree of similarity, the object and background can be segmented through region growing.

[0076] Step 1.1 Use Hoffman line detection method to detect sea-sky line

[0077] The basic principle of Huffman line detection is to use the duality of points and lines. In the Cartesian coordinate system, a straight line can be represented by its slope and intercept on the y-axis. A straight line can be represented by its perpendicular length to the origin and the angle θ between the perpendicular axis and the x-axis in the Huffman space. That is, the coordinates of any point on the straight line in the Huffman space are (ρ, θ)

[0078] ρ=xcosθ+ysinθ (1)

[0079] In this embodiment, xoy represents a rectangular coordinate system, (ρ, θ) represents a Huffman space coordinate, θoρ represents a Huffman space coordinate system, and (x0, y0) represents the coordinates of any point in the rectangular coordinate system;

[0080] A straight line passing through a point on a plane is represented as a curve in Hough space, expressed as

[0081] ρ=x0 cosθ+y0 sinθ (2)

[0082] Therefore, the possibility of the existence of a linear target can be converted into the number of curve intersections obtained by formula (1); Figure 3 As shown in the figure, the straight line 2 passing through the three points is represented as the intersection of the three curves in the Huffman space. The simple explanation of the Huffman line detection idea is to rotate the straight line through any point in the rectangular coordinate system (a certain step size, rotate to 180 degrees), and then count which line has the most points. The line with the most points can be directly extracted. The function of Huffman line detection is to extract the straight line in the binary image. Since the sea-sky line has straight line characteristics, it can be extracted.

[0083] In the actual calculation process, it is necessary to discretize the independent variables and construct a "voter" matrix to approximate the Hough space. Then, 1 is added to the votes representing the linear positions of valid image points in the voters. Finally, the row positions can be filtered by sorting the votes to complete the sea-sky line detection.

[0084] Step 1.2 Use the Huffman line detection method to constrain the sensitive areas in the current detection image (sea-sky line)

[0085] Step 1.21. First, reduce the original detection image to 480*640;

[0086] Step 1.22. Next, use the Canny algorithm to calculate the feature value of each pixel in the image using a 3x3 horizontal and vertical operator template. Then, use binarization to convert the image into a binary image, showing the texture characteristics of the image edge image.

[0087] Step 1.23. Finally, use the Hoffman line detection algorithm to detect the longest line segment in the map;

[0088] (1) First, for the target point P1, draw a straight line L1 through the point with α = 0 degrees (counterclockwise is positive);

[0089] (2) Calculate the distance ρ from the origin to the line L1;

[0090] (3) Statistical angle θ between the perpendicular line from the origin to the line L1 and the x-axis (positive in counterclockwise direction);

[0091] (4) Loop through steps (1) to (3), taking α as the variable and a step size of 10 degrees for point P1, and taking the angles from 0 to 180 degrees to count the distances ρ and angles θ of all points;

[0092] (5) Loop through step (4) and take all target points P i , count the distance ρ and angle θ of all points;

[0093] (6) In the Huffman space, the point with the largest number of identical values ​​of (ρ, θ) is converted into a rectangular coordinate system, which is the line to be extracted;

[0094] Step 1.24. If the number of target pixels in the longest line segment detected by the Hoffman line detection algorithm is greater than 2 / 3 of the line segment length, it is considered to be sea level. Considering that sea-sky-line detection is sensitive to blue, the blue channel of the image is used as input. The detection effect is as follows: Figure 4 As shown;

[0095] from Figure 4 It can be seen that the sea-sky-line detection algorithm based on Huffman space can clearly display the texture of the sea-sky-line and has good detection ability for the sea-sky-line with large inclination;

[0096] Step 2: Use the density-based DBSCAN clustering algorithm + the distance-based K-means clustering algorithm (K-means) to cluster the water column marker frame to obtain a better prior anchor frame, which is input into the YOLOv4 network to improve the water column positioning accuracy.

[0097] When using the YOLOv4 network, it is necessary to pre-set a priority bounding box to determine the aspect ratio of the detection target and generate a specified number of prior anchor boxes in each grid;

[0098] Currently, the YOLOv4 algorithm uses the K-means clustering algorithm to determine the size of the prior anchor box. The larger the number of cluster centers, the better the fit with the dataset. However, when the K value is greater than the threshold, the clustering effect is not significantly improved. The idea of ​​the K-means clustering algorithm is to randomly select K samples from the initial data as the initial cluster centers, and then calculate the distance d from each sample to the cluster center. The distance d is usually calculated based on the center point of a two-dimensional box.

[0099] In actual situations, water columns have a process of growth and dissipation, and stay for a long time in the initial growth stage and the final dissipation stage. Therefore, the number density of small and medium-sized water columns is relatively large; the target selection accuracy is the error between the water column and the target, which is reflected in the positioning accuracy in the initial stage of water column growth; since the K-means clustering algorithm clusters the Euclidean distance and ignores the density, a density-based clustering algorithm should be adopted; currently commonly used density-based clustering algorithms include the mean shift algorithm and the DBSCAN algorithm; since the anchor frame of the water column in this embodiment only needs to roughly divide the water column into two categories: the growth process (high density) and the dissipation process (low density), it is not necessary to locate the position with the highest density. The mean shift algorithm is not easy to divide the water column according to the sample density by setting the bandwidth; therefore, the DBSCAN algorithm is selected in this embodiment;

[0100] Before using the DBSCAN algorithm for cluster analysis, two parameters need to be determined: one is the radius of the nearest neighbor region, which represents the range of the circular field with a fixed point as the center point; the other is the minimum number of points contained in the neighboring region; because the DBSCAN algorithm infers the number of clusters based on the data, it does not need to determine the number of clusters in advance, and it can generate clusters of any shape; the size of the label box is clustered using the K-means clustering algorithm and the DBSCAN algorithm, and the results are shown below. Figure 5 As shown;

[0101] exist Figure 5 In the K-means algorithm, the K value is 4, the radius of the adjacent region of the DBSCAN algorithm is 0.1, and the number of samples is at least 25. According to the classification results, K-means divides the points evenly based on the Euclidean distance, ignoring the density information. DBSCAN divides by density, but it is not easy to adjust the parameters to further divide the points with high density.

[0102] In order to ensure high positioning accuracy in the early stage of water column growth, the number of clusters in the early stage should be as large as possible. Based on the above reasons, this embodiment proposes the DBSCAN+K-means clustering algorithm to determine the appropriate prior bounding box;

[0103] Step 2.1. First, use the DBSCAN algorithm to perform density division on clusters with water column marker box size according to clustering

[0104] Step 2.11. Set the neighborhood radius of the DBSCAN algorithm to 0.1 and the number of nearest neighbors to 25 to partition the samples. The DBSCAN algorithm works as follows: Select a point from the sample, given a radius epsilon and a minimum number of nearest neighbors min_points. If the point has at least min_points neighbors within its neighborhood circle of radius epsilon, then move the center of the circle to the next sample point. If a sample point does not meet these conditions, reselect a new sample point. Iterate clustering according to the set radius epsilon and min_points.

[0105] Step 2.2. Then use the K-means clustering algorithm to divide the clusters with high density according to the K value. The division results are as follows: Figure 5 As shown; in this embodiment, the K value is 5 to divide the data; the principle of K-means clustering algorithm:

[0106] Step 2.21. First, randomly generate five cluster centers in the data. Then, calculate the distance from each sample in the data to these five cluster centers and assign the corresponding sample to the cluster corresponding to the cluster center with the smallest distance.

[0107] Step 2.22. After all samples are classified, recalculate the cluster center of each of the five clusters, that is, the centroid of all samples in each cluster. Repeat the above steps until the cluster center does not change.

[0108] Step 2.3. In this embodiment, the target annotation box size (length and width of the target annotation box) marked in the target annotation box dataset is used as the initial dataset to input into the clustering algorithm. According to steps 2.1 and 2.2, a set of better prior anchor boxes is finally obtained. Figure 6 As shown;

[0109] Step 3: Add CBAM to PANet to improve the detection accuracy and positioning accuracy of weak feature water column targets, and use the improved YOLOv4 PANet structure to extract fuzzy water columns

[0110] Step 3.1 Add hybrid attention mechanism (CBAM) to the path aggregation network (PANet) to improve the YOLOv4 network structure

[0111] This embodiment adds a CBAM module to each of the three branches at the end of the YOLOv4 feature fusion network; based on the characteristics of the small water column integrated with the CBAM module, the weight of the water column is increased, while the weight of the non-target area is suppressed, thereby improving the overall detection accuracy of the water column; the improvements of the PANet module are as follows: Figure 7 As shown;

[0112] Step 3.2: Use the improved YOLOv4PANet structure to extract the fuzzy water column

[0113] Step 3.21. Combination Figure 2 and Figure 7 , assuming that the size of the input image is 608*608*3, the size of the input feature map input to the first CBAM module is 152*152*256;

[0114] Step 3.22. The input feature map passes through two parallel MaxPool layers and Avgpool layers, changing the feature map from 152*152*252*1*256 to 1*1*256;

[0115] Step 3.23. The feature map then passes through the Share MLP module, first compressing the number of channels to 1*1*16 and then expanding it to 1*1*256;

[0116] Step 3.24. Obtain two results based on the feature map of the ReLU activation function, add these two results, and then obtain the output of the channel attention result through the feature map of the sigmoid activation function;

[0117] Step 3.25. Multiply the output result with the initial input feature map to obtain the feature map enhanced by the channel attraction mechanism;

[0118] Step 3.26. Convert the feature map enhanced by the attention mechanism to a 152*152*1 feature map through max pooling and average pooling. Then concatenate the two feature maps and convert them into a 1-channel feature map through a 7*7 convolution. Then, use the sigmoid function to obtain the feature map of intelligent attention. Finally, multiply the output by the feature map enhanced by the attention mechanism to 152*152*256. At this time, the feature map has enhanced the information of the water column.

[0119] Step 3.27. Using the YOLO head to analyze the position loss, classification loss, and confidence loss can speed up the convergence speed and improve the accuracy. Figure 7 The improvement of the small water column test results is as follows Figure 8 As shown;

[0120] Preferably, the network structure of the improved YOLOv4 algorithm described in this embodiment mainly consists of three parts: backbone network, neck network, and head network; the network structure of the algorithm is as follows Figure 1 As shown:

[0121] The backbone network is the CSPParknet53 network, which is mainly composed of 5 CSPNet residual blocks. The introduction of CSPNet residual not only reduces network calculations but also enhances feature extraction. In terms of activation function, the CSPParknet53 network uses the Mish activation function to further enhance the propagation of deep network information.

[0122] The Neck network consists of an SPP (Spatial Pyramid) module and a PANet. The SPP module performs maximum pooling operations of different scales, 13×13, 9×9, 5×5, and 1×4, effectively increasing the network's receptive field and extracting significant contextual features. The PANet is mainly based on the Feature Pyramid Network (FPN) and adds a bottom-up path on the basis of FPN to enhance the fusion between different feature maps.

[0123] The final output of the Head network is 13×13, 26×26, and 52×5252 detection heads to detect large, medium, and small targets respectively.

[0124] Finally, the position and size of the predicted box are obtained using the pre-set prior box size and relative offset.

[0125] Example 2: Step 4: In order to verify the accuracy and feasibility of the water column detection method of the impact point based on the improved YOLOv4 algorithm as described in Example 1, this example is designed for verification.

[0126] 1. Test platform introduction

[0127] The experimental environment in this example uses an i7-10875H CPU, an RTX2060 GPU, and 16GB of memory. Python is used in the Windows 11 operating system. The deep learning framework is based on TensorFlow 2, and uses Cuda and Cudnn to accelerate the network model.

[0128] 2. Evaluation indicators

[0129] In the experiment, precision (P), recall (R), and average precision (AP) were used as evaluation indicators to measure the quality of the model. The IOU threshold of 0.6 was used as the detection threshold, that is, the overlap area between the detection box and the ground truth frame exceeded 60%, which was considered a positive detection example. The calculation formula is as follows:

[0130]

[0131]

[0132]

[0133] Among them, TP represents true examples, that is, the number of positive sample water columns actually detected is also the number of positive sample water columns, FN represents false negative examples, that is, samples detected as negative but actually positive, and FP represents false positive samples, that is, the detected water column target is a positive sample but actually is a negative sample.

[0134] The detection bias is defined as the distance between the center point of the detection box and the center point of the size box when the IOU is greater than 0.6; the specific definition is as follows:

[0135]

[0136] δ accuracy (x 1c ,y 1c )(x 2c ,y 2c ) where represents the detection deviation, and represents the center coordinates of the detection box and the marking box, respectively. When the target is not detected or the deviation is greater than the width of the marking box, the detection accuracy is recorded as the width of the marking box.

[0137] All algorithms (Faster R-CNN, SSD, YOLOv4, YOLOv5, and Improved YOLOv4) were tested in Python;

[0138] Among them: The experimental parameters of each algorithm are selected as follows:

[0139] (1) The parameters for training Faster R-CNN are as follows: the input image size is 608*608, Resnet50 is used in the backbone; the ratio of training set, validation set and test set is 6:2:2; 600 images are trained, batch_size is 4, epoch is 500, and the Faster R-CNN takes about 3.4G memory and takes about 8.6 hours to train, with a final loss function value of 0.3542;

[0140] (2) The parameters for training SSD are as follows: the input image size is 608*608, the backbone is VGG; the ratio of training set, validation set and test set is 6:2:2; 600 images are trained, batch_size is 4, epoch is 500, training SSD takes about 3.4G memory, training time is about 7.8 hours, and the final loss function value is 0.5437;

[0141] (3) The parameters for training YOLOv5 are set as follows: the input image size is 608*608, the backbone consists of CBL, bottleneck ekCSP / C3 and SPP / SPPF, etc.; the ratio of training set, validation set and test set is 6:2:2; training is 600 images, the batch size is 4, the epoch is 500, the memory consumption is about 3.3G, the training time is about 8-9 hours, and the final loss function value is 0.2663;

[0142] (4) The parameters for training YOLOv4 and the improved YOLOv4 are set as follows: the input image size is 608*608, the backbone is composed of DarkNet 53+ResX, the ratio of training set, validation set and test set is 6:2:2; 600 images are trained, the batch size is 4, and the epoch is 500; the training cost of the YOLOv4 algorithm is about 3.5G memory, the training time is about 8.2 hours, and the final loss function value is 0.3124; the training cost of the improved YOLOv4 algorithm is about 3.8G memory, the training time is about 9.1 hours, and the final loss function value is 0.2324;

[0143] 3. Experimental Results

[0144] 3.1 This example sets up an ablation experiment and a level comparison test. In the ablation experiment, YOLOv4 is compared with the following improvements:

[0145] (1) DBSCAN+K-means,

[0146] (2) CBAM-BasedYOLOv4PANet structure improvement;

[0147] The results are as follows Figure 10 As shown;

[0148] A horizontal comparison test was used: the faster R-CNN, SSD, YOLOv4, and YOLOv5 algorithms were used as comparison algorithms to detect water columns. A total of 200 images were detected, and the detection results are as follows: Figure 11 As shown, the statistical data Figure 12 and as shown in Table 1;

[0149] Table 1: Object detection algorithm operation data table

[0150]

[0151] As can be seen from Table 1, the detection speed of YOLOv4, YOLOv5, and the improved YOLOv4 algorithm meets the real-time requirement (above 30); the mAP of Faster R-CNN, SSD algorithm, YOLOv4, YOLOv5, and the improved YOLOv4 algorithm are all above 90%, but the improved YOLOv4 algorithm has the best performance; in addition, the improved YOLOv4 algorithm has the best positioning accuracy;

[0152] from Figure 12 It can be seen that the faster R-CNN, SSD, YOLOv4, YOLOv5, and improved YOLOv4 algorithms in this embodiment have higher positioning accuracy, which is about 10 pixels overall. However, the improved YOLOv4 has the best positioning accuracy. When the number of image frames is 100-130 frames, all algorithms have large positioning deviations due to image blur.

[0153] 5. Conclusion

[0154] Example 1 of the present invention proposes an improved YOLOv4 water column detection algorithm for the specific scenario of water column detection at the impact point. First, this method detects the sea-sky line using the Huffman line detection method, constraining the sensitive areas in the current detection image, thereby improving the detection accuracy of water column detection. Second, the DBSCAN+K-means clustering analysis algorithm is used to improve the size selection of the prior bounding box of YOLOv4, making the prior bounding box more typical. Finally, CBAM is added to PANet to improve the detection accuracy and positioning accuracy of weak-feature water column targets. Compared with the Faster R-CNN, SSD, YOLOv4, and YOLOv5 algorithms, this algorithm can effectively improve the detection accuracy and positioning accuracy of water columns, providing an effective water column detection and positioning method for USV target acquisition.

[0155] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for detecting water column at impact point based on an improved YOLOv4 algorithm, characterized by: include Step 1: Use the Huffman line detection method to detect the sea-sky line and constrain the sensitive area of ​​water column detection in the current detection image; Step 2: Use the density-based DBSCAN clustering algorithm and the distance-based K-means clustering algorithm to perform cluster analysis on the water column marker frame to obtain a better prior anchor frame and input it into the YOLOv4 network; Step 3: Add a hybrid attention mechanism to the path aggregation network, improve the YOLOv4 network structure, and use the improved YOLOv4 PANet structure to extract the fuzzy water column; The process of adding a hybrid attention mechanism to the path aggregation network described in step 3 and extracting fuzzy water columns using the improved YOLOv4 PANet structure includes: Step 3.

1. Improve YOLOv4 by adding a hybrid attention mechanism module to each of the three branches at the end of the YOLOv4 feature fusion network. The improved YOLOv4 algorithm network structure described in step 3.1 includes a backbone network, a neck network, and a head network. The backbone network is a CSPParknet53 network. The CSPParknet53 network is composed of five CSPNet residual blocks, and the activation function of the CSPParknet53 network is a Mish activation function. The Neck network consists of an SPP module and a PANet, wherein the SPP module performs maximum pooling operations of different scales: 13×13, 9×9, 5×5, and 1×4; the PANet is improved based on the feature pyramid network; The final output of the Head network is a 13×13, 26×26, and 52×5252 detection head, which detects large, medium, and small targets respectively; Step 3.2 uses the improved YOLOv4 PANet structure to extract the fuzzy water column; The process of extracting fuzzy water column using the improved YOLOv4 PANet structure described in step 3.2 includes: Step 3.

21. Assume the size of the input image is , the size of the input feature map input to the first CBAM module is ; Step 3.

22. Input feature map passes through two parallel MaxPool layers and Avgpool layers, and transforms the feature map from Change to ; Step 3.

23. The feature map passes through the Share MLP module and first compresses the number of channels to , and then expands to ; Step 3.

24. Add the two results obtained from the feature map of the ReLU activation function and obtain the output of the channel attention mechanism through the feature map of the sigmoid activation function. Step 3.

25. Multiply the output with the initial input feature map to obtain a feature map enhanced by the channel attention mechanism. Step 3.

26. Convert the feature map enhanced by the attention mechanism into Feature map, concatenate the two feature maps, and then pass The convolution is converted into a 1-channel feature map, and the normalized weight map is obtained by the sigmoid function. The weight map is multiplied by the feature map enhanced by the channel attention mechanism, and the size is obtained. The spatial attention mechanism feature map; Step 3.

27. Using the feature map generated by the spatial attention mechanism, YOLO Head analyzes the position loss, classification loss, and confidence loss to complete the water column detection of the impact point.

2. The method for detecting water column at impact point based on the improved YOLOv4 algorithm according to claim 1, characterized in that: The process of detecting the sea-sky line using the Hoffman line detection method in step 1 and constraining the sensitive area of ​​water column detection in the current detection image includes: Step 1.

1. Detect sea-sky line using Huffman line detection method (1) Using the Huffman line detection method to detect the duality of points and lines, the number of curve intersections in the Huffman space is detected, and a "voter" matrix is ​​constructed to approximate the Huffman space; (2) add 1 to the vote representing the linear position of the valid point in the image among the voters; (3) By sorting the votes, the point with the most votes is selected, and then the point is converted into a line in the rectangular coordinate system to complete the sea-sky line detection; In step 1.2, the area above and below the sea-sky line, which is 1 / 3 of the image height, is set as the sensitive area, thereby constraining the water column detection area in the current image.

3. The method for detecting water column at impact point based on the improved YOLOv4 algorithm according to claim 2, characterized in that: The process of constraining the water column detection area in the current image described in step 1.2 includes: Step 1.

21. First reduce the original sea-sky-line detection image to ; Step 1.

22. Then use the canny algorithm to pass The horizontal and vertical operator templates are used to calculate the feature values ​​of each pixel in the image, and then the image is converted into a binary image through binarization to show the texture characteristics of the image edge image; Step 1.

23. Finally, use the Hoffman line detection algorithm to detect the longest line segment in the map; Step 1.

24. If the number of target pixels in the longest line segment in the map detected in step 1.23 is greater than 2 / 3 of the line segment length, it is considered to be sea level.

4. The method for detecting water column at impact point based on the improved YOLOv4 algorithm according to claim 3, characterized in that: The process of detecting the longest line segment in the map using the Hoffman line detection algorithm described in step 1.23 includes: (1) First, for the target point P1, make the following Draw a straight line L1, counterclockwise is positive; (2) Statistical distance from the origin to line L1 ; (3) Statistical coordinate origin to the vertical line L1 and the angle between the x-axis , counterclockwise is positive; (4) Loop through steps (1) to (3), targeting point P1. is a variable with a step size of 10 degrees, and is taken from 0 to 180 to count the distances of all points. and angles ; (5) Loop through step (4) and take all target points P i , count the distances of all points and angles ; (6) In Hoffman space, The point with the largest number of identical values, converted into a rectangular coordinate system, is the line to be extracted.

5. The method for detecting water column at impact point based on the improved YOLOv4 algorithm according to claim 1, characterized in that: The process of clustering the water column marker frame using the density-based DBSCAN clustering algorithm and the distance-based K-means clustering algorithm described in step 2 includes: Step 2.

1. Set the neighborhood radius of the DBSCAN algorithm to 0.1 and the number of nearest neighbor samples to 25. Use the DBSCAN algorithm to perform density partitioning on clusters with the water column marker box size. Step 2.

2. Use the K-means clustering algorithm to divide the clusters with the highest density in DBSCAN according to the K value; Step 2.

3. Input the target annotation box size dataset as the initial dataset into the clustering algorithm, and obtain a set of prior anchor boxes according to steps 2.1 and 2.

2.

6. The method for detecting water column at impact point based on the improved YOLOv4 algorithm according to claim 5, characterized in that: The K value in step 2.2 is 5. The process of using the K-means clustering algorithm to divide the clusters with high density according to the K value includes: Step 2.

21. Randomly generate five cluster centers in the data. Calculate the distance from each sample to these five cluster centers and assign the corresponding sample to the cluster corresponding to the cluster center with the smallest distance. Step 2.

22. After all samples are classified, recalculate the cluster center of each cluster for each of the five clusters, that is, the centroid of all samples in each cluster, and repeat the above operation until the cluster center does not change.