Forward-looking sonar image obstacle detection method based on encoder and reinforcement learning

By introducing an automatic encoder and reinforcement learning algorithm in front-view sonar image processing, the image processing strategy is dynamically adjusted, and the problem of insufficient detection accuracy and real-time in traditional methods is solved, and efficient, accurate and adaptive obstacle detection is achieved.

CN120107771APending Publication Date: 2025-06-06HEBEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256313.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, the detection of front-view sonar image obstacles has problems such as high noise, low contrast, and blurred details, resulting in insufficient detection accuracy and real-time. In addition, traditional image processing methods require manual adjustment of parameters, poor adaptability, high calculation complexity, and poor real-time performance.

Method used

Using an encoder-based and reinforcement learning method, dimensionality reduction and feature extraction are performed through automatic encoder, combined with reinforcement learning algorithm, the image processing actions are dynamically selected to achieve efficient, accurate, adaptive and intelligent obstacle detection of images.

Benefits of technology

It significantly improves the efficiency, accuracy and adaptability to complex environments, reduces manual intervention, and improves the reliability and real-timeness of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107771A_ABST
    Figure CN120107771A_ABST
Patent Text Reader

Abstract

The invention discloses a forward-looking sonar image obstacle detection method based on an encoder and reinforcement learning, and relates to the field of image processing and underwater detection, and the method comprises the following steps: obtaining a forward-looking sonar image, and carrying out the dimensionality reduction and feature extraction through the encoder, and obtaining low-dimensional features; on the basis of the low-dimensional features, related image quality indexes are calculated to obtain corresponding state values, and image processing actions are dynamically selected and executed by using a reinforcement learning algorithm; performing threshold segmentation and edge detection on the processed image, and extracting an obstacle contour; and outputting related information of the obstacle according to the obstacle contour. According to the method, efficient, accurate, self-adaptive and intelligent obstacle detection of the foresight sonar image is realized, the detection efficiency, accuracy and adaptability to a complex environment are remarkably improved, and technical support is provided for autonomous navigation and obstacle avoidance of an underwater robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing and underwater detection, and more specifically to a forward-looking sonar image obstacle detection method based on an encoder and reinforcement learning. Background Art

[0002] Forward-looking sonar is widely used in underwater robots, autonomous submersibles and ocean exploration equipment. It obtains real-time images of underwater environments through acoustic imaging technology. However, due to the complexity of the underwater environment, sonar images often have problems such as high noise, low contrast, and blurred details, which leads to severe challenges in the accuracy and real-time performance of obstacle detection.

[0003] In the prior art, obstacle detection in forward-looking sonar images mainly relies on traditional image processing methods, such as image filtering (Gaussian filtering, median filtering) and enhancement algorithms (histogram equalization) with manually set parameters. Such methods require pre-adjustment of parameters based on experience, such as filter kernel size, enhancement strength, etc., resulting in poor adaptability. In complex dynamic environments, fixed parameters are difficult to adapt to the dynamic changes in image characteristics and require frequent manual intervention, which is not only inefficient, but also prone to false detection or missed detection due to parameter mismatch. For novice users, it is particularly difficult to set parameters suitable for specific conditions, which further affects the popularization and application of the technology.

[0004] In addition, existing methods generally have the problems of high computational complexity and poor real-time performance. Traditional algorithms (such as full-image threshold segmentation and multi-scale edge detection) need to process high-resolution sonar images pixel by pixel, which is time-consuming and difficult to meet the real-time response requirements of underwater equipment. At the same time, the existing technology lacks the in-depth application of adaptive optimization technologies such as reinforcement learning, and cannot dynamically adjust the processing strategy according to the image quality, resulting in difficulty in balancing the contradiction between denoising and detail retention under noise interference, affecting the accuracy of obstacle contour extraction.

[0005] Therefore, how to design a forward-looking sonar image obstacle detection method based on encoder and reinforcement learning, which can achieve efficient and adaptive obstacle detection and precise positioning in complex underwater environments is a problem that technical personnel in this field urgently need to solve. Summary of the invention

[0006] In view of this, the present invention provides a forward-looking sonar image obstacle detection method based on encoder and reinforcement learning. Through automatic encoder dimensionality reduction and reinforcement learning optimization, efficient, accurate, adaptive and intelligent obstacle detection of forward-looking sonar images is achieved, which significantly improves the detection efficiency, accuracy and adaptability to complex environments.

[0007] In order to achieve the above object, the present invention adopts the following technical solution:

[0008] The present invention provides a forward-looking sonar image obstacle detection method based on an encoder and reinforcement learning, comprising the following steps:

[0009] S1, obtain the forward-looking sonar image, and perform dimensionality reduction and feature extraction through the encoder to obtain low-dimensional features;

[0010] S2. Based on the low-dimensional features, calculate relevant image quality indicators to obtain corresponding state values, and use a reinforcement learning algorithm to dynamically select and execute image processing actions;

[0011] S3, performing threshold segmentation and edge detection on the processed image to extract obstacle contours;

[0012] S4. Outputting relevant information of the obstacle according to the obstacle outline.

[0013] Furthermore, in S1, dimensionality reduction and feature extraction are performed by an encoder to obtain low-dimensional features, including:

[0014] S11, input the forward-looking sonar image to the input layer of the encoder;

[0015] S12. In at least two convolutional layers, extract local features of the input image by convolution operation, and each convolutional layer performs convolution operation on the image using a preset convolution kernel size and step size to generate a feature map;

[0016] S13, input the feature map output by the convolution layer to the maximum pooling layer, and reduce the spatial dimension of the feature map through the pooling operation;

[0017] S14. Input the feature map processed by the pooling layer to the output layer, and convert the feature map into a low-dimensional feature vector through the fully connected layer.

[0018] Furthermore, in S2, calculating the relevant image quality index to obtain the corresponding state value includes:

[0019] S211, respectively calculating the peak signal-to-noise ratio PSNR, the structural similarity index SSIM and the edge preservation index EPI of the current image;

[0020] S212, normalizing the peak signal-to-noise ratio PSNR, the structural similarity index SSIM, and the edge preservation index EPI; the normalization is expressed as:

[0021]

[0022] I represents the original value of the corresponding indicator, I max ,I min Respectively represent the preset maximum reference value and minimum reference value of the corresponding indicator;

[0023] S213, performing weighted summation on the normalized indicators to obtain a comprehensive state value S;

[0024] S=β 1 Norm (PSNR) + β 2 Norm(SSIM)+β 3 ·Norm(EPI)

[0025] Among them, β 1 , β 2 , β 3 represents the weight coefficient, and β 1 +β 2 +β 3 =1; Norm(PSNR), Norm(SSIM), and Norm(EPI) represent the normalized peak signal-to-noise ratio PSNR, structural similarity index SSIM, and edge preservation index EPI, respectively.

[0026] Furthermore, in S2, the image processing action is dynamically selected and executed using a reinforcement learning algorithm, including:

[0027] S221, construct action candidate set A = {a 1 ,a 2 …,a m}; where each action a i Corresponding to an image processing operation and parameter combination;

[0028] S222, combined with the current Q value vector, select the optimal action a * :

[0029]

[0030] Among them, s t Indicates the current status.

[0031] Furthermore, in S221, constructing an action candidate set includes:

[0032] Based on the historical data of image quality indicators, a clustering algorithm is used to classify the action effects, and the actions with the highest matching degree with the current state are selected to form a dynamic candidate set;

[0033] The dynamic candidate set is merged with the basic candidate set to form the action candidate set.

[0034] Furthermore, in S222, the current Q value vector is constructed, including:

[0035] Construct a state-action Q table, where each entry Q(s, a) represents the expected cumulative reward for selecting action a in state s.

[0036] Iteratively update the Q value:

[0037]

[0038] Among them, a t represents the currently selected action, α represents the learning rate, r t+1 Indicates execution of action a t The reward obtained after, γ represents the discount factor, Indicates the next state s t+1 All possible actions a i The maximum Q value.

[0039] Furthermore, the S222 further includes:

[0040] If the ε-greedy strategy is adopted, the action is randomly selected with the exploration rate ε; the dynamic adjustment of the exploration rate ε is expressed as:

[0041] ε=ε init ·exp(-k·t)

[0042] Among them, ε init represents the initial exploration rate, k represents the decay coefficient, and t represents the number of training steps.

[0043] Furthermore, S2 also includes evaluating the quality of the processed image through a reward function and updating the reinforcement learning strategy to optimize subsequent action selection;

[0044] The reward function is a multi-index weighted function, expressed as:

[0045] R=ω 1 ΔPSNR+ω 2 ΔSSIM+ω 3 ΔEPI

[0046] Among them, ω 1 ,ω 2 ,ω 3 represents the weight coefficient, ΔPSNR, ΔSSIM, and ΔEPI represent the changes of peak signal-to-noise ratio PSNR, structural similarity index SSIM, and edge preservation index EPI before and after image processing, respectively.

[0047] Furthermore, the S3 includes:

[0048] S31, performing threshold segmentation on the processed image by using Otsu threshold segmentation algorithm, and dividing the image into foreground and background;

[0049] S32, on the segmented foreground image, using the Canny algorithm to detect and extract the edge contour of the obstacle, and output a binary edge image;

[0050] S33: For the binary edge image, extract obstacle contours using an eight-neighborhood connected region labeling method, and filter regions whose areas are smaller than a preset pixel number threshold.

[0051] Furthermore, in S4, the obstacle-related information includes: obstacle area, centroid coordinates and angle.

[0052] It can be seen from the above technical solution that compared with the prior art, the technical solution of the present invention has the following advantages:

[0053] Beneficial effects:

[0054] 1. Through the convolutional layer and pooling layer structure of the encoder, the high-dimensional forward-looking sonar image is reduced to low-dimensional features, significantly reducing data redundancy and computational complexity. Combined with the dynamic action selection of reinforcement learning, the tedious process of manual parameter adjustment in traditional methods is avoided, thereby improving the overall processing efficiency.

[0055] 2. By calculating the normalized state values ​​of multiple indicators in real time and designing a multi-dimensional reward function, the image processing strategy can be dynamically adjusted according to changes in image quality. Its adaptability enables it to maintain high accuracy and robustness in the face of different noise interference or dynamic environments, solving the limitations of traditional fixed parameter methods.

[0056] 3. Adopt Otsu threshold segmentation and Canny edge detection, combined with normalized weighted evaluation of multiple indicators, to ensure the accuracy of obstacle contour extraction. By comprehensively optimizing image quality and edge preservation, the impact of noise on detection results is reduced, and the reliability of obstacle positioning and information output is improved, which is especially suitable for the precise detection needs of complex underwater scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0058] Figure 1 A flow chart of a forward-looking sonar image obstacle detection method based on an encoder and reinforcement learning provided in an embodiment of the present invention;

[0059] Figure 2 A schematic diagram of a process of obtaining low-dimensional features through an encoder provided in an embodiment of the present invention;

[0060] Figure 3 A schematic diagram of a process for calculating relevant image quality indicators to obtain status values ​​provided by an embodiment of the present invention;

[0061] Figure 4 A schematic diagram of a process of dynamically selecting and executing image processing actions using a reinforcement learning algorithm provided in an embodiment of the present invention;

[0062] Figure 5 A schematic diagram of the obstacle contour extraction process provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0063] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0064] like Figure 1 As shown, this embodiment provides a forward-looking sonar image obstacle detection method based on an encoder and reinforcement learning, comprising the following steps:

[0065] S1, obtain the forward-looking sonar image, and perform dimensionality reduction and feature extraction through the encoder to obtain low-dimensional features;

[0066] S2. Based on the low-dimensional features, calculate relevant image quality indicators to obtain corresponding state values, and use a reinforcement learning algorithm to dynamically select and execute image processing actions;

[0067] S3, performing threshold segmentation and edge detection on the processed image to extract obstacle contours;

[0068] S4. Outputting relevant information of the obstacle according to the obstacle outline.

[0069] This method uses an encoder to reduce the dimension of the image to reduce redundant calculations, combines it with reinforcement learning to dynamically optimize the filtering and enhancement strategies, and uses a multi-indicator normalized evaluation mechanism to provide real-time feedback on the processing effect, thereby achieving efficient and adaptive obstacle detection and precise positioning in complex underwater environments.

[0070] The following further describes each step in the above method in detail;

[0071] In this embodiment S1, a forward-looking sonar image is acquired, and dimension reduction and feature extraction are performed through an encoder to obtain low-dimensional features; the forward-looking sonar image refers to an image generated by a sonar device installed at the front end of an underwater vehicle or diving equipment, and is mainly used for environmental perception and obstacle detection in the front area.

[0072] like Figure 2 As shown, obtaining low-dimensional features includes:

[0073] S11, input the forward-looking sonar image to the input layer of the encoder;

[0074] S12. In at least two convolutional layers, extract local features of the input image by convolution operation, and each convolutional layer performs convolution operation on the image using a preset convolution kernel size and step size to generate a feature map;

[0075] S13, input the feature map output by the convolution layer to the maximum pooling layer, and reduce the spatial dimension of the feature map through the pooling operation;

[0076] S14. Input the feature map processed by the pooling layer to the output layer, and convert the feature map into a low-dimensional feature vector through the fully connected layer.

[0077] Specifically, the input layer receives the original forward-looking sonar image of size H×W×C; the first convolution layer uses a 3×3 convolution kernel, a step size of 1, a padding of 1, an activation function of ReLU, and an output feature map size of H×W×64; the first maximum pooling layer uses a 2×2 pooling window, a step size of 2, and an output feature map size of H / 2×W / 2×64; the second convolution layer uses a 3×3 convolution kernel, a step size of 1, a padding of 1, an activation function of ReLU, and an output feature map size of H / 2×W / 2×128; the second maximum pooling layer uses a 2×2 pooling window, a step size of 2, and an output feature map size of H / 4×W / 4×128; the fully connected compression layer flattens the feature map into a one-dimensional vector, compresses it to dimension D through the fully connected layer, the activation function is Sigmoid, and outputs a low-dimensional feature vector.

[0078] The encoder is used to reduce the dimension and extract features of the forward-looking sonar image, and the convolution layer and pooling layer are used to gradually extract the key features of the image, and the fully connected layer is used to compress the high-dimensional image into a low-dimensional feature vector. It only uses the encoder, which effectively reduces the computational complexity of image processing and improves the efficiency and accuracy of obstacle detection.

[0079] In this embodiment S2, based on the low-dimensional features, the relevant image quality index is calculated to obtain the corresponding state value, and the image processing action is dynamically selected and executed using the reinforcement learning algorithm;

[0080] like Figure 3 As shown, the relevant image quality indicators are calculated to obtain corresponding status values, including:

[0081] S211, respectively calculating the peak signal-to-noise ratio PSNR, the structural similarity index SSIM and the edge preservation index EPI of the current image;

[0082] Specifically, the peak signal-to-noise ratio (PSNR) is used to evaluate the similarity between the processed image and the original image. It reflects the degree of image distortion by calculating the mean square error of the difference in pixel values ​​between the two images. The higher the PSNR value, the better the image quality and the closer the processed image is to the original image. The structural similarity index (SSIM) evaluates the similarity between two images by considering the similarity in brightness, contrast, and structure. A high SSIM value indicates that the two images are very similar in structure. The edge preservation index (EPI) is obtained by calculating the gradients in the horizontal and vertical directions of the image. A high EPI value indicates that the processed image retains the edge information in the original image well, which helps to improve the accuracy of subsequent analysis steps.

[0083] S212, normalizing the peak signal-to-noise ratio PSNR, the structural similarity index SSIM, and the edge preservation index EPI; the normalization is expressed as:

[0084]

[0085] I represents the original value of the corresponding indicator, I max ,I min Respectively represent the preset maximum reference value and minimum reference value of the corresponding indicator;

[0086] S213, performing weighted summation on the normalized indicators to obtain a comprehensive state value S;

[0087] S=β 1 Norm (PSNR) + β 2 Norm(SSIM)+β 3 ·Norm(EPI)

[0088] Among them, β 1 , β 2 , β 3 represents the weight coefficient, and β 1 +β 2 +β 3 =1; Norm(PSNR), Norm(SSIM), and Norm(EPI) represent the normalized peak signal-to-noise ratio PSNR, structural similarity index SSIM, and edge preservation index EPI, respectively.

[0089] In the specific implementation, the normalized ranges of PSNR, SSIM and EPI are set to (20, 40), (0.6, 0.95) and (0.7, 1.2), respectively, and the weight coefficients β1 = 0.4, β2 = 0.4, and β3 = 0.2 to highlight the contribution of PSNR and SSIM to image quality. When the processed image PSNR = 32dB, SSIM = 0.85, and EPI = 0.9, the normalized comprehensive state value S = 0.4*(0.6) + 0.4*(0.714) + 0.2*(0.4) = 0.63.

[0090] Compared with conventional technical solutions, which often ignore edge preservation and cause reinforcement learning to over-smooth the image, the system achieves a balance between denoising (high PSNR) and detail preservation (high EPI) by weighting multiple indicators.

[0091] like Figure 4 As shown, the reinforcement learning algorithm is used to dynamically select and execute image processing actions, including:

[0092] S221, construct action candidate set A = {a 1 ,a 2 …,a m}; where each action a i Corresponding to an image processing operation and parameter combination; wherein, constructing an action candidate set includes: based on the historical data of the image quality index, using a clustering algorithm to classify the action effects, screening out the actions with the highest matching degree with the current state to form a dynamic candidate set; merging the dynamic candidate set with the basic candidate set to form an action candidate set.

[0093] Specifically, the action candidate set contains a variety of operation combinations, such as Gaussian filtering + histogram equalization, median filtering + gamma enhancement, etc., and further dynamically selects the top-3 actions that match the current state through K-means clustering to reduce invalid exploration. For example, when the image is blurred, the clustering model gives priority to sharpening operations.

[0094] S222, combined with the current Q value vector, select the optimal action a * :

[0095]

[0096] Among them, s t Indicates the current status.

[0097] Furthermore, the current Q value vector is constructed, including:

[0098] Construct a state-action Q table, where each entry Q(s, a) represents the expected cumulative reward for selecting action a in state s.

[0099] Iteratively update the Q value:

[0100]

[0101] Among them, a t represents the currently selected action, α represents the learning rate, r t+1 Indicates execution of action a t The reward obtained after, γ represents the discount factor, Indicates the next state s t+1 All possible actions a i The maximum Q value.

[0102] Furthermore, it also includes:

[0103] If the ε-greedy strategy is adopted, the action is randomly selected with the exploration rate ε; the dynamic adjustment of the exploration rate ε is expressed as:

[0104] ε=ε init ·exp(-k·t)

[0105] Among them, ε init represents the initial exploration rate, k represents the decay coefficient, and t represents the number of training steps.

[0106] In the specific implementation, the learning rate α = 0.1, the discount factor γ = 0.9, the initial exploration rate ε_init = 0.5, and the decay coefficient k = 0.001. The Q table is stored in a matrix, and the dimension of the state-action pair is 1000 × 10. For example, when the training step number t = 1000, ε = 0.5 × exp (-0.001 × 1000) = 0.184, and the random exploration ratio is gradually reduced.

[0107] Compared with traditional Q-learning (fixed ε = 0.1), the dynamic ε strategy fully explores in the early stage of training and tends to use the optimal strategy in the later stage.

[0108] Furthermore, S2 also includes evaluating the quality of the processed image through a reward function and updating the reinforcement learning strategy to optimize subsequent action selection;

[0109] The reward function is a multi-index weighted function, expressed as:

[0110] R=ω 1 ΔPSNR+ω 2 ΔSSIM+ω 3 ΔEPI

[0111] Among them, ω 1 ,ω 2 ,ω 3represents the weight coefficient, ΔPSNR, ΔSSIM, and ΔEPI represent the changes of peak signal-to-noise ratio PSNR, structural similarity index SSIM, and edge preservation index EPI before and after image processing, respectively.

[0112] In the specific implementation, the weight coefficients ω1=0.5, ω2=0.3, ω3=0.2, and the reward R=0.5×ΔPSNR+0.3×ΔSSIM+0.2×ΔEPI. If the PSNR is increased by 5dB, the SSIM is increased by 0.1, and the EPI is reduced by 0.05 after processing, then R=0.5×5+0.3×0.1-0.2×0.05=2.5+0.03-0.01=2.52.

[0113] In this embodiment S3, the processed image is subjected to threshold segmentation and edge detection to extract the obstacle contour; Figure 5 As shown, specifically including:

[0114] S31, perform threshold segmentation on the processed image by Otsu threshold segmentation algorithm, and divide the image into foreground and background; Otsu threshold segmentation algorithm is implemented by selecting a suitable threshold, so that the points with pixel values ​​greater than the threshold in the image are regarded as foreground, and the points with pixel values ​​less than the threshold are regarded as background. It determines the optimal segmentation threshold by maximizing the inter-class variance between the foreground and the background, thereby effectively separating the target object from the complex background.

[0115] S32, on the segmented foreground image, using the Canny algorithm to detect and extract the edge contour of the obstacle, and output a binary edge image;

[0116] Specifically, the Canny algorithm finds possible edge points by calculating the gradient of the image; then, it applies non-maximum suppression technology to remove insignificant edges and retain only points with the largest local gradient; finally, it uses double threshold detection to further refine the edges, ensuring that only continuous edges with sufficient strength are retained, generating a clear binary edge image.

[0117] S33, for the binary edge image, extract obstacle contours by eight-neighborhood connected region labeling method, and filter areas whose area is smaller than a preset pixel point number threshold, where the preset pixel point number is 100.

[0118] In this step, by combining the Otsu threshold segmentation algorithm and the Canny edge detection algorithm, the processed image can be accurately divided into foreground and background, and the edge contours of obstacles can be effectively extracted. In particular, the eight-neighborhood connected region labeling method is used to further accurately extract the obstacle contours and filter out small noise areas, which significantly improves the accuracy and robustness of obstacle contour extraction.

[0119] In S4 of this embodiment, the relevant information of the obstacle is output according to the obstacle outline; the relevant information of the obstacle includes: obstacle area, centroid coordinates and angle.

[0120] Specifically, based on the extracted obstacle contour, first, the obstacle area is calculated by counting the total number of pixels in the contour area to reflect its physical size. Then, the image geometric moment is used to calculate the centroid coordinates of the contour: by accumulating the position coordinates of each pixel point of the contour and normalizing them, the exact position of the centroid in the image coordinate system is obtained to identify the center point of the obstacle. The angle calculation is performed by fitting the minimum circumscribed rectangle of the contour, extracting the angle between its major axis and the horizontal axis, and representing the direction of the obstacle.

[0121] Furthermore, the above centroid coordinates are converted to actual spatial positions after coordinate transformation, which is convenient for underwater robot navigation. Principal component analysis is used to optimize direction estimation in angle calculation to avoid error accumulation when fitting the circumscribed rectangle. The final output information can be directly integrated into the relevant control system to support obstacle avoidance path planning, which can dynamically adjust the navigation direction according to the centroid distance and angle to ensure operational safety and real-time response in complex underwater environments.

[0122] The encoder and reinforcement learning-based forward-looking sonar image obstacle detection method in this implementation realizes accurate detection and positioning of obstacles in complex underwater environments through encoder dimensionality reduction, reinforcement learning-optimized image processing, multi-index evaluation and dynamic adjustment strategy, combined with efficient threshold segmentation and edge detection algorithms, and provides reliable obstacle avoidance support for underwater vehicles.

[0123] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0124] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A forward-looking sonar image obstacle detection method based on encoder and reinforcement learning, characterized in that: The following steps are involved: S1, obtain the forward-looking sonar image, and perform dimensionality reduction and feature extraction through the encoder to obtain low-dimensional features; S2. Based on the low-dimensional features, calculate relevant image quality indicators to obtain corresponding state values, and use a reinforcement learning algorithm to dynamically select and execute image processing actions; S3, performing threshold segmentation and edge detection on the processed image to extract obstacle contours; S4. Outputting relevant information of the obstacle according to the obstacle outline.

2. The method for detecting obstacles in forward-looking sonar images based on encoder and reinforcement learning according to claim 1, characterized in that: In S1, dimensionality reduction and feature extraction are performed by the encoder to obtain low-dimensional features, including: S11, input the forward-looking sonar image to the input layer of the encoder; S12. In at least two convolutional layers, extract local features of the input image by convolution operation, and each convolutional layer performs convolution operation on the image using a preset convolution kernel size and step size to generate a feature map; S13, input the feature map output by the convolution layer to the maximum pooling layer, and reduce the spatial dimension of the feature map through the pooling operation; S14. Input the feature map processed by the pooling layer to the output layer, and convert the feature map into a low-dimensional feature vector through the fully connected layer.

3. The method for detecting obstacles in forward-looking sonar images based on encoder and reinforcement learning according to claim 1, characterized in that: In S2, calculating the relevant image quality index to obtain the corresponding state value includes: S211, respectively calculating the peak signal-to-noise ratio PSNR, the structural similarity index SSIM and the edge preservation index EPI of the current image; S212, normalizing the peak signal-to-noise ratio PSNR, the structural similarity index SSIM, and the edge preservation index EPI; the normalization is expressed as: I represents the original value of the corresponding indicator, I max ,I min Respectively represent the preset maximum reference value and minimum reference value of the corresponding indicator; S213, performing weighted summation on the normalized indicators to obtain a comprehensive state value S; S=β1·Norm(PSNR)+β2·Norm(SSIM)+β3·Norm(EPI) Among them, β1, β2, and β3 represent weight coefficients, and β1+β2+β3=1; Norm(PSNR), Norm(SSIM), and Norm(EPI) represent the normalized peak signal-to-noise ratio PSNR, structural similarity index SSIM, and edge preservation index EPI, respectively.

4. The method for detecting obstacles in forward-looking sonar images based on encoder and reinforcement learning according to claim 1, characterized in that: In S2, a reinforcement learning algorithm is used to dynamically select and execute image processing actions, including: S221, construct action candidate set A = {a1, a2..., a m }; where each action a i Corresponding to an image processing operation and parameter combination; S222, combined with the current Q value vector, select the optimal action a * : Among them, s t Indicates the current status.

5. The method for detecting obstacles in forward-looking sonar images based on encoder and reinforcement learning according to claim 4, characterized in that: In S221, constructing an action candidate set includes: Based on the historical data of image quality indicators, a clustering algorithm is used to classify the action effects, and the actions with the highest matching degree with the current state are selected to form a dynamic candidate set; The dynamic candidate set is merged with the basic candidate set to form the action candidate set.

6. The method for detecting obstacles in forward-looking sonar images based on encoder and reinforcement learning according to claim 4, characterized in that: In S222, the current Q value vector is constructed, including: Construct a state-action Q table, where each entry Q(s, a) represents the expected cumulative reward for selecting action a in state s. Iteratively update the Q value: Among them, a t represents the currently selected action, α represents the learning rate, r t+1 Indicates execution of action a t The reward obtained after, γ represents the discount factor, Indicates the next state s t+1 All possible actions a i The maximum Q value.

7. The method for detecting obstacles in forward-looking sonar images based on encoder and reinforcement learning according to claim 4, characterized in that: The S222 further includes: If the ε-greedy strategy is adopted, the action is randomly selected with the exploration rate ε; the dynamic adjustment of the exploration rate ε is expressed as: e=e init ·exp(-k·t) Among them, ε init represents the initial exploration rate, k represents the decay coefficient, and t represents the number of training steps.

8. The method for detecting obstacles in forward-looking sonar images based on encoder and reinforcement learning according to claim 1, characterized in that: Said S2 also includes evaluating the quality of the processed image through a reward function and updating the reinforcement learning strategy to optimize the subsequent action selection; The reward function is a multi-index weighted function, expressed as: R=ω1·ΔPSNR+ω2·ΔSSIM+ω3·ΔEPI Among them, ω1, ω2, and ω3 represent weight coefficients, ΔPSNR, ΔSSIM, and ΔEPI represent the changes of peak signal-to-noise ratio PSNR, structural similarity index SSIM, and edge preservation index EPI before and after image processing, respectively.

9. The method for detecting obstacles in forward-looking sonar images based on encoder and reinforcement learning according to claim 1, characterized in that: The S3 includes: S31, performing threshold segmentation on the processed image by using Otsu threshold segmentation algorithm, and dividing the image into foreground and background; S32, on the segmented foreground image, using the Canny algorithm to detect and extract the edge contour of the obstacle, and output a binary edge image; S33: For the binary edge image, extract obstacle contours using an eight-neighborhood connected region labeling method, and filter regions whose areas are smaller than a preset pixel number threshold.

10. The method for detecting obstacles in forward-looking sonar images based on encoder and reinforcement learning according to claim 1, characterized in that: In S4, the obstacle-related information includes: obstacle area, centroid coordinates and angle.