Forward-looking sonar image recognition method based on improved mean filtering
By improving the combination of mean filtering and convolutional neural network, the forward-view sonar image is preprocessed and recognized, which solves the problem of noise in the image affecting target recognition and achieves higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202411987759.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-06-06
AI Technical Summary
There are a large number of noise and stray signals in the forward-view sonar image, which affects the accurate identification and positioning of targets in the water.
The forward-view sonar image is preprocessed using an improved mean filtering method to remove noise and preserve image edge details. The filtered image is then divided into small pieces and inputted to the convolutional neural network for recognition, and the recognition accuracy is further improved through non-maximum suppression and post-processing. The recognition results of multi-frame images are fused to output the final target information.
It effectively reduces noise and stray signals in front-view sonar images, improves image quality and target recognition accuracy, and enhances recognition accuracy and robustness.
Smart Images

Figure CN120107767A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of underwater robots, and in particular relates to a forward-looking sonar image recognition method based on improved mean filtering. Background Art
[0002] With the continuous deepening of ocean development and research, the application of underwater robots is becoming more and more extensive. Forward-looking sonar is one of the important sensors of underwater robots. It detects parameters such as the position, shape, and speed of objects in the water by emitting sound waves and receiving reflected echoes, thereby realizing functions such as underwater environment perception and obstacle avoidance.
[0003] However, due to the influence of noise and stray signals in seawater, there is often a lot of interference information in forward-looking sonar images, which affects the accurate recognition and positioning of underwater targets. Therefore, how to improve the accuracy of forward-looking sonar image recognition has become an urgent problem to be solved.
[0004] In recent years, deep learning has made great progress in sonar image detection. The YOLO series of convolutional neural networks have attracted widespread attention in the field of target recognition due to their fast operation, simple framework, small storage space, and low background false detection rate. It directly regresses the entire image and only needs to pass the network once to obtain the location and category information of the target, which greatly improves the recognition efficiency. Summary of the invention
[0005] In view of this, the present invention provides a forward-looking sonar image recognition method based on improved mean filtering, which can solve the problems existing in the existing forward-looking sonar image recognition method.
[0006] The technical solution for implementing the present invention is as follows:
[0007] The method for forward-looking sonar image recognition based on improved mean filtering includes the following steps:
[0008] Step 1, using improved mean filtering to process the forward-looking sonar image;
[0009] Step 2, the filtered forward-looking sonar image is cut into small blocks and input into the convolutional neural network;
[0010] Step 3, identifying the input image;
[0011] Step 4, screening the recognition results;
[0012] Step 5: The recognition results of multiple frames of sonar images are fused to output the final target information.
[0013] Furthermore, it also includes:
[0014] Acquire forward-looking sonar images;
[0015] Enhance the forward-looking sonar image.
[0016] Furthermore, the improved mean filtering process includes:
[0017] Step 11, taking each point of a frame of sonar image to be processed as the center, select a filter window of size w×w, remove the pixels with maximum and minimum values in the window, and the set of remaining k pixels is recorded as H;
[0018] Step 12, calculate the average value of the pixels in H, denoted as Mean(f(i, j)), (f(i, j) is the pixel value of the i-th row and j-th column;
[0019] Step 13: Calculate the absolute value of the difference between the grayscale value of each pixel in H and the average value, denoted as D k .
[0020] Step 14: Calculate all pixel points D in H k The average value of is denoted as T.
[0021] Step 15: Take T obtained in step 4 as a threshold. If the D corresponding to a certain pixel k is greater than the threshold, the weight is determined by D k Otherwise, it is determined by T. The expression is shown in Formula 4. Calculate the weight w corresponding to each pixel in H k (i, j) and normalized, and then the pixel value and weight are weighted according to Formula 5 as the output of the center point of the filter window.
[0022] H={f(i,j)|f(i,j)! =Max(f(i,j)) or f(i,j)! =Min(f(i,j))} (1)
[0023] D k =|H k -Mean(f(i, j))| (2)
[0024]
[0025]
[0026] Among them, Max(f(i, j)) is the maximum pixel value, Min(f(i, j)) is the minimum pixel value, D k is the absolute value of the pixel value of each point in the K set of points and the Mean(f(i, j)) calculated in the previous step; Max(D k , T) means taking D k , the maximum of the two values of T, H k(i, j) represents the pixel value corresponding to the point in the i-th row and j-th column in the set H, w k (i, j) represents the weight corresponding to the point in the i-th row and j-th column in the set H, f k (i, j) represents the weighted new pixel value of the point in the i-th row and j-th column in the set H.
[0027] Furthermore, the filtered forward-looking sonar image is cut into small blocks before being input into the convolutional neural network, specifically including: cutting a frame of image into multiple blocks as network input respectively, to avoid the target appearing exactly on the dividing line during cutting, leaving the overlapping parts between the slices, the size of the segmented image is 640×640, and the images are combined and restored after recognition.
[0028] Furthermore, the segmented images are identified, and the specific process is as follows:
[0029] Step 31, generating a convolutional neural network feature map and a region of interest according to the input image;
[0030] Step 32, obtaining a feature map using a region of interest pooling layer according to the convolutional neural network feature map and the region of interest;
[0031] Step 33, mapping the feature map into a one-dimensional feature vector;
[0032] Step 34, performing classification regression on the one-dimensional feature vector and outputting the result.
[0033] Furthermore, the screening of the recognition results includes:
[0034] Perform non-maximum suppression on the recognition results.
[0035] Furthermore, after screening the recognition results, the method further includes:
[0036] The information of multiple frames of targets is fused and the final target information is output, including:
[0037] Assume that the width of each frame is w 0 , height is h 0 , the image distance resolution ratio is obtained according to formula 6, where range is the working range of the sonar, and the position information of the detection frame is assumed to be (x, y, w, h), then the position of the suspected target is According to equations 7 and 8, it is converted into polar coordinates (ρ, θ);
[0038] According to equations 9 and 10, the distance dist and angle angle between the target and the sonar can be obtained. The longitude and latitude information of the sonar location is obtained according to the positioning device; the longitude and latitude of the target are obtained by combining the longitude and latitude of the sonar, dist, and angle. The recognition result of each frame is added to the set D. After that, each time a frame of sonar image is recognized, the longitude and latitude of the recognition result is combined with the longitude and latitude of each target in the set D to obtain the distance d between the two points. 0 , let d 1 is the distance threshold set, if d 0 Less than d 1 , then the two targets are considered to be the same target, and the recognition count of the target is increased and updated in set D, otherwise, it is added to set D. If sum is the set recognition count threshold, when the recognition count of a target in D is greater than sum, then the target is the final true target;
[0039] ratio=range / h 0 (6)
[0040]
[0041] ang le=θ (9)
[0042] dist=ρ×ratio (10)
[0043] Beneficial effects:
[0044] 1. The present invention adopts the above technical solution, which can effectively reduce the noise and stray signals in the forward-looking sonar image, improve the image quality, and enhance the accuracy of subsequent target recognition.
[0045] 2. The present invention utilizes an improved mean filtering algorithm to preserve image edge details and avoid losses caused by over-smoothing.
[0046] 3. The present invention can further reduce the influence of noise on target recognition by performing noise reduction processing on the filtered image.
[0047] 4. The present invention can further improve the accuracy and robustness of target recognition through means such as non-maximum suppression and post-processing.
[0048] 5. The present invention can further improve the accuracy of target recognition by fusing the recognition results of multiple frames of sonar images. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 A schematic flow chart of a forward-looking sonar image recognition method based on improved mean filtering provided by the present invention.
[0050] Figure 2It is a flow chart of the fusion of recognition results of multiple sonar images provided by the present invention. DETAILED DESCRIPTION
[0051] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0052] Embodiment 1:
[0053] In view of the problems existing in the related art, the embodiment of the present invention provides a forward-looking sonar image recognition method based on improved mean filtering, see Figure 1 The specific process shown is as follows:
[0054] S101. Process the forward-looking sonar image using an improved mean filter.
[0055] Specifically, firstly, a forward-looking sonar image is acquired. When acquiring the forward-looking sonar image, the forward-looking sonar image can be acquired by a sonar probe, and then the forward-looking sonar image is enhanced, that is, the contrast and clarity of the image are enhanced to improve the image quality. There are many ways of enhancement processing, such as histogram equalization, gamma correction, etc.
[0056] Next, the forward-looking sonar image is processed using an improved mean filter. The improved mean filter is an improvement on the mean filter, which can effectively remove Gaussian noise and other types of noise in the image while maintaining the sharpness of the image edge. Traditional mean filtering usually replaces the value of the central pixel with the average value of all pixels in a neighborhood of a fixed size. Although this method is simple and easy to use, it may cause over-smoothing or loss of important details for some complex scenes. To this end, we propose an improved mean filtering algorithm that can dynamically adjust the weight coefficient according to the different conditions of the pixel position, thereby more flexibly controlling the degree of smoothing.
[0057] In this embodiment, the improved mean filtering process includes the following steps:
[0058] The first step is to determine the filter window size. The filter window refers to an area used to calculate the weight coefficients of the pixels around each pixel. Generally speaking, the larger the filter window, the better the smoothing effect, but at the same time, more detail information may be lost. Therefore, we need to choose a suitable window size to ensure sufficient smoothing effect without excessively sacrificing detail information.
[0059] The second step is to remove the pixels with maximum and minimum values in the window, and the set of remaining k pixels is recorded as H;
[0060] The third step is to calculate the average value of the pixels in H;
[0061] The fourth step is to calculate the absolute value of the difference between the gray value of each pixel in H and the average value, recorded as D k ;
[0062] Step 5: Calculate the average of all absolute values obtained in step 4, denoted as T.
[0063] Step 6: Calculate the weight coefficient of the pixels around each pixel. The weight coefficient is an indicator that reflects the importance of the relative position of the pixel. Generally, the closer the pixel is to the center pixel, the higher the weight. In this embodiment, we use T obtained in step 5 as a threshold. If the D corresponding to a certain pixel is k is greater than the threshold, the weight is determined by D k Otherwise, it is determined by T to dynamically adjust the weight coefficient.
[0064] Step 7: Update the value of the current pixel based on the weight coefficient and the values of the pixels around each pixel. The update formula is as follows:
[0065]
[0066] Among them, w k (i, j) represents the weight corresponding to the pixel in the i-th row and j-th column, f k (i, j) is the weighted result of the pixel value in the i-th row and j-th column and the weight.
[0067] S102, dividing the filtered forward-looking sonar image into small blocks and inputting them into the convolutional neural network.
[0068] Specifically, take a forward-looking sonar image with an input image size of 2100*1280 as an example, and crop it into 6 small images of 640×640. The convolutional neural network here refers to YOLOv5s, which belongs to the YOLO series of convolutional neural networks. It has received widespread attention in the field of target recognition due to its fast operation, simple framework, small storage space, and low background false detection rate. It directly regresses the entire image and only needs to pass the network once to obtain the location information and category information of the target, which greatly improves the recognition efficiency.
[0069] In this example, we will cut the filtered forward-looking sonar image into multiple small blocks, each of which is an independent input sample, and input it into the same YOLOv5s model for classification prediction. The reason for cutting it into small blocks is that the entire image is too large, and inputting it all at once will cause problems such as memory overflow, and it is not conducive to model learning and convergence.
[0070] S103: Recognize the input image.
[0071] Specifically, after multiple iterations of training of the YOLOv5s model, a set of classifiers can be obtained, which can distinguish and identify different categories. At this point, we pass the input small blocks of the image into this model for prediction, and get the probability distribution corresponding to each small block, that is, the confidence score of each category.
[0072] S104: Screen the recognition results.
[0073] Specifically, the purpose of screening the recognition results is to eliminate false positives, that is, to filter out those misjudged results that are not real targets. Generally speaking, we use non-maximum suppression (NMS) to achieve this goal. The basic idea of NMS is: if the score in a target box is higher than a certain threshold, it is considered to be a real target, otherwise it is ignored. In other words, only those targets with high enough scores will be considered valid.
[0074] Furthermore, after screening the recognition results, it also includes:
[0075] Post-process the recognition results.
[0076] S105: Outputting final target information by fusing the recognition results of multiple frames of images.
[0077] Specifically, Figure 2 As shown, the following steps are included:
[0078] The first step is to assume that the width of each frame is w 0 , height is h 0 According to formula 6, the image distance resolution ratio is obtained, where range is the working range of the sonar, and the position information of the detection frame is assumed to be (x, y, w, h). Then the position of the suspected target is According to equations 7 and 8, it is converted into polar coordinates (ρ, θ);
[0079] In the second step, the distance dist and the angle angle between the target and the sonar can be obtained according to equations 9 and 10. The longitude and latitude information of the sonar location is obtained according to the positioning device; the longitude and latitude of the target are obtained by combining the longitude and latitude of the sonar, dist, and angle.
[0080] The third step is to add the recognition result of each frame to the set D. After that, each time a sonar image is recognized, the distance d between the two points is calculated by combining the longitude and latitude of the recognition result with the longitude and latitude of each target in the set D. 0 , let d 1 is the distance threshold set, if d 0 Less than d 1, then the two targets are considered to be the same target, and the recognition count of the target is increased and updated in set D, otherwise, it is added to set D. If sum is the set recognition count threshold, when the recognition count of a target in D is greater than sum, then the target is the final true target;
[0081] ratio=range / h 0 (6)
[0082]
[0083] ang le=θ_ (9)
[0084] dist=ρ×ratio (10)
[0085] The present invention adopts the above technical solution, which can effectively reduce the noise and stray signals in the forward-looking sonar image, improve the image quality, and enhance the accuracy of subsequent target recognition. At the same time, the improved mean filtering algorithm can retain the image edge details and avoid the loss caused by excessive smoothing. In addition, by performing noise reduction processing on the filtered image, the impact of noise on target recognition can be further reduced. Finally, the accuracy and robustness of target recognition can be further improved by means of non-maximum suppression and post-processing.
[0086] Embodiment 2:
[0087] Based on the same inventive concept, an embodiment of the present invention further provides a forward-looking sonar image recognition device based on improved mean filtering, comprising:
[0088] Improved mean filter module, used to process forward-looking sonar images using improved mean filter;
[0089] The image segmentation module is used to divide the filtered forward-looking sonar image into small blocks and input them into the convolutional neural network;
[0090] The target recognition module is used to recognize the input image;
[0091] A target screening module is used to screen the recognition results;
[0092] Target fusion module, used to filter out false targets;
[0093] The target display module is used to output the final target information.
[0094] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A forward-looking sonar image recognition method based on improved mean filtering, characterized in that: include: Use improved mean filtering on forward-looking sonar images; The filtered forward-looking sonar image is cut into small blocks and input into the convolutional neural network; Recognize the input image; filter the recognition results; output the final target information from the multi-frame sonar image recognition results.
2. The method according to claim 1, characterized in that Also includes: Acquire forward-looking sonar images; Enhance the forward-looking sonar image.
3. The method according to claim 1 or 2, characterized in that The improved mean filtering process includes: determining the size of the filtering window; calculating the weight coefficients of the pixels around each pixel; and updating the value of the current pixel according to the weight coefficients and the values of the pixels around each pixel.
4. The method according to claim 3, characterized in that The calculation of the weight coefficients of the pixels around each pixel includes: for any pixel to be filtered, the pixel to be filtered is set as the central pixel, and a rectangular window with a width of w and a length of w is taken with the central pixel as the center; the pixels with maximum and minimum values in the window are removed, and the set of k pixels left is recorded as H; the average value of the pixels in H is calculated, recorded as Mean(f(i,j)); the absolute value of the difference between the gray value of each pixel in H and the average value is calculated, recorded as D k ; Calculate all pixel points D in H k The average value is recorded as T; T is used as a threshold. If the D corresponding to a certain pixel k is greater than the threshold, the weight is determined by D k Determined by, otherwise, determined by T; the expression is shown in Formula 4, calculating the weight w corresponding to each pixel in H k (i, j) and normalized, and then the pixel value and weight are weighted as the output of the center point of the filter window according to Formula 5; H={f(i,j)|f(i,j)! =Max(f(i,j)) or f(i,j)! =Min(f(i,j))} (1) D k =|H k -Mean(f(i,j))| (2) Among them, Max(f(i,j)) is the maximum pixel value, Min(f(i,j)) is the minimum pixel value, D k is the absolute value of the pixel value of each point in the K set of points and the Mean(f(i,j)) calculated in the previous step; Max(D k ,T) means taking D k ,T is the maximum of the two values, H k (i,j) represents the pixel value corresponding to the point in the i-th row and j-th column in the set H, w k (i,j) represents the weight corresponding to the point in the i-th row and j-th column in the set H, f k (i, j) represents the weighted new pixel value of the point in the i-th row and j-th column in the set H.
5. The method according to claim 4, characterized in that The filtered forward-looking sonar image is cut into small blocks and input into the convolutional neural network. The specific process is as follows: A frame of image is divided into multiple blocks as the input of the network. In order to avoid the target appearing exactly on the dividing line during segmentation, the overlapping parts are retained between the slices. The size of the segmented image is 640×640. After recognition, the images are combined and restored.
6. The method according to claim 5, characterized in that The specific process of identifying the input image is as follows: Generate a convolutional neural network feature map and a region of interest according to an input image; obtain a feature map using a region of interest pooling layer according to the convolutional neural network feature map and the region of interest; map the feature map into a one-dimensional feature vector; Perform classification regression on the one-dimensional feature vector and output the result.
7. The method according to claim 6, characterized in that The screening of the recognition results includes: performing non-maximum suppression on the recognition results.
8. The method according to claim 7, characterized in that After screening the recognition results, the method further includes: The information of multiple frames of targets is fused and the final target information is output, including: Assuming that the width of each frame is w0 and the height is h0, the image distance resolution ratio is obtained according to formula 6, where range is the working range of the sonar and the position information of the detection frame is assumed to be (x, y, w, h). Then the position of the suspected target is According to equations 7 and 8, it is converted into polar coordinates (ρ, θ); According to formula 9 and formula 10, the distance dist and the angle angle between the target and the sonar can be obtained; the longitude and latitude information of the sonar location is obtained according to the positioning device; the longitude and latitude of the target are obtained by combining the longitude and latitude of the sonar, dist, and angle; the recognition result of each frame is added to the set D, and then each time a frame of sonar image is recognized, the longitude and latitude of the recognition result is combined with the longitude and latitude of each target in the set D to obtain the distance d0 between the two points, and d1 is set as the set distance threshold. If d0 is less than d1, the two targets are considered to be the same target, and the recognition times of the target are increased and updated in the set D, otherwise, it is added to the set D; If sum is the set recognition number threshold, when the recognition number of a target in D is greater than sum, then the target is the final true target; ratio=range / h0 (6) angle=θ (9) dist = ρ × ratio (10).