Submersible path planning method and system based on visual identification
By constructing a three-dimensional map of the underwater environment and updating the target object information in real time, combining visual recognition technology and path planning algorithms, the problems of large amount of calculation and low efficiency of submersible path planning are solved, real-time path planning and efficient navigation of submersibles are realized.
Patent Information
- Application Number
- CN202510332015.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has a large amount of calculation and low processing efficiency when planning submersible paths in complex underwater environments, and cannot be adjusted in real time, which poses safety risks.
By constructing a three-dimensional map of the underwater environment, the target object information is updated in real time, and dynamic path planning is generated by combining visual recognition technology and path planning algorithms.
Real-time and efficient submersible path planning are achieved, the underwater environment perception ability is improved, and the safe navigation of the submersible in complex environments is ensured.
Smart Images

Figure CN120333434A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of underwater vehicles, and particularly to a path planning method and system for submersibles based on visual recognition. Background Art
[0002] With the increasing demands for underwater exploration and work, the need for autonomous navigation and intelligent control of submersibles has become increasingly urgent. When a submersible works in a complex underwater environment, it is necessary to effectively identify the distribution of objects in the underwater environment and plan the travel path of the submersible based on a preset task target point. However, path planning needs to comprehensively consider various factors, and when the environment changes, the planned path also needs to change with the environment.
[0003] The prior art generally uses heuristic algorithms to perform path planning for submersibles. However, when heuristic algorithms are applied to underwater environments with frequent changes, the computational amount of submersible path planning will increase sharply. And when heuristic algorithms face a large amount of data, the processing efficiency is very low, and it is impossible to modify the travel path of the submersible, presenting potential safety hazards. Summary of the Invention
[0004] In view of this, the present invention proposes a path planning method and system for submersibles based on visual recognition. By constructing a three-dimensional map of the underwater environment and real-time updating the information related to objects on the three-dimensional map of the underwater environment, a three-dimensional dynamic map of the underwater environment is obtained, and path planning is carried out on the three-dimensional dynamic map of the underwater environment, reducing the computational pressure of submersible path planning, improving the processing efficiency of submersible path planning, and realizing the real-time path planning of submersibles.
[0005] The technical solution of the present invention is implemented as follows: In the first aspect, the present invention provides a path planning method for submersibles based on visual recognition, including the following steps:
[0006] S1, collecting the original underwater image data of the water area where the submersible is located, and preprocessing the original underwater image data to obtain underwater image data;
[0007] S2, identifying the underwater image data based on a three-dimensional position recognition model to obtain the three-dimensional positions of target objects in the water area where the submersible is located, and identifying the underwater image data based on an underwater target recognition model to obtain the classification information of target objects in the water area where the submersible is located;
[0008] S3, constructing a three-dimensional map of the underwater environment through a simultaneous localization and mapping model, real-time updating the three-dimensional positions of the target objects and the classification information of the target objects, and integrating them into the three-dimensional map of the underwater environment to obtain a three-dimensional dynamic map of the underwater environment;
[0009] S4. Set the starting point and target point of the submersible according to the three-dimensional dynamic map of the underwater environment, and perform path planning on the submersible through a path planning algorithm to generate the travel path of the submersible.
[0010] Based on the above technical solution, preferably, step S1 includes:
[0011] Collect the original underwater image data of the water area where the submersible is located using a high-resolution underwater camera, perform noise reduction processing on the original underwater image data through Gaussian filtering to obtain the first underwater image data, perform image enhancement on the first underwater image data to obtain the second underwater image data, perform color correction on the second underwater image data to obtain the third underwater image data, perform image dehazing processing on the third underwater image data to obtain the fourth underwater image data, and perform geometric correction on the fourth underwater image data to obtain the underwater image data.
[0012] Based on the above technical solution, preferably, step S2 includes:
[0013] The construction process of the three-dimensional position recognition model includes:
[0014] Construct an initial three-dimensional position recognition model based on a convolutional neural network and a deep learning network. The initial three-dimensional position recognition model includes a first input layer, a 2D feature extraction layer, a depth estimation layer, a feature fusion layer, a 3D position regression layer, and a first output layer. The 2D feature extraction layer is constructed based on ResNet-50, the depth estimation layer is constructed based on MonoDepth2, the first input layer is respectively connected to the 2D feature extraction layer and the depth estimation layer, both the 2D feature extraction layer and the depth estimation layer are connected to the feature fusion layer, the feature fusion layer is connected to the 3D position regression layer, the 3D position regression layer is connected to the first output layer, and the first output layer outputs three-dimensional position coordinates;
[0015] Obtain the first historical underwater image dataset, and divide the first historical underwater image dataset into a first training set and a first validation set;
[0016] Train and validate the initial three-dimensional position recognition model according to the first training set and the first validation set, optimize the hyperparameters of the initial three-dimensional position recognition model according to the results of the training and validation, and obtain the optimized three-dimensional position recognition model after the training and validation are completed.
[0017] Based on the above technical solution, preferably, step S2 further includes:
[0018] The construction process of the underwater target recognition model includes:
[0019] Construct an initial underwater target recognition model, where the initial underwater target recognition model includes a second input layer, a Backbone layer, a Neck layer, a Head layer, and a second output layer. The second input layer is connected to the Backbone layer, the Backbone layer is connected to the Neck layer, the Neck layer is connected to the Head layer, and the Head layer is connected to the second output layer. The Backbone layer includes a CSPDarknet53 feature extractor, a convolutional layer, and a CSP block. The Neck layer includes a top-down FPN structure and a bottom-up feature pyramid containing two PAN structures. The Head layer includes multiple convolutional layers and three prediction scales. The second output layer outputs the target classification probability;
[0020] Obtain a second historical underwater image dataset, and divide the second historical underwater image dataset into a second training set and a second validation set;
[0021] Train and validate the initial underwater target recognition model according to the second training set and the second validation set, optimize the hyperparameters of the initial underwater target recognition model according to the results of the training and validation, and obtain an optimized underwater target recognition model after the training and validation are completed;
[0022] The output of the underwater target recognition model is:
[0023] Y2(X2) = H(N(B(X2))) + L2(Y2(X2), Y2 true (X2));
[0024] where X2 is the input of the underwater target recognition model, Y2(X2) is the classification prediction result of the underwater target recognition model, B(·) is the Backbone function, N(·) is the Neck function, H(·) is the Head function, and L2(·) is the loss function of the underwater target recognition model, is the actual classification result of the underwater target recognition model;
[0025] The loss function of the underwater target recognition model is:
[0026]
[0027] where are the weight coefficients for adjusting classification, localization, and confidence of the loss function of the underwater target recognition model respectively, is the classification loss of the loss function of the underwater target recognition model, is the localization loss of the loss function of the underwater target recognition model, is the confidence loss of the loss function of the underwater target recognition model.
[0028] Based on the above technical solutions, preferably, step S3 includes:
[0029] Using a simultaneous localization and mapping model to construct a three-dimensional map of the underwater environment, and real-time updating the target three-dimensional position and the target classification information over time, and integrating the real-time updated target three-dimensional position and the target classification information onto the three-dimensional map of the underwater environment to obtain a three-dimensional dynamic map of the underwater environment and display it on the monitoring terminal screen of the submersible.
[0030] Based on the above technical solutions, preferably, step S4 includes:
[0031] S41: Mark the current position of the submersible as the starting point on the three-dimensional dynamic map of the underwater environment, set the task target point, initialize the path planning algorithm, set the cost function and the heuristic function, perform an initial search of the path planning algorithm based on the cost function and the heuristic function, generate an initial optimal path, and the submersible initially travels along the initial optimal path;
[0032] S42: Detect the environmental changes of the current node, update the cost function of the current node based on the environmental changes of the current node, re-plan the local optimal path based on the updated cost function and the heuristic function, and drive the submersible to travel along the local optimal path;
[0033] S43: Repeat step S42 until the submersible reaches the task target point, smooth the path generated each time to ensure that the motion constraints of the submersible are met, and sequentially connect the nodes where the initial optimal path and each section of the local optimal path intersect to obtain the travel path of the submersible.
[0034] Based on the above technical solutions, preferably, the calculation formula of the cost function is:
[0035]
[0036] The calculation formula of the heuristic function is:
[0037]
[0038] The calculation formula of the updated cost function is:
[0039]
[0040] The calculation formula of the path smoothing process is:
[0041]
[0042] Where g(n) is the cost function from the starting position to the current node n, n j is the node j on the submersible path, nj+1 is the node j+1 on the path of the submersible, d(n j , n j+1 ) is the cost between node j and node j+1, h(n) is the predicted cost from the current node n to the target node, (x n , y n , z n ) and (x goal , y goal , z goal ) are the three-dimensional position coordinates of the current node n and the target node respectively, g'(n) is the cost function of the current node n after the environmental change update, c(n j ) is the additional cost of environmental change for node j, P(t) is the path curve after path smoothing, t is the time parameter, P1 and P2 are the interpolation nodes on the path curve respectively, and P0 and P3 are the curve smooth and continuous points on the path curve respectively.
[0043] In a second aspect, the present invention also provides a submersible path planning system based on visual recognition, and the system includes:
[0044] An underwater image processing module, configured to collect underwater original image data of the water area where the submersible is located, preprocess the underwater original image data, and obtain underwater image data;
[0045] An object classification and positioning module, configured to identify the underwater image data based on a three-dimensional position recognition model to obtain the three-dimensional position of the target object in the water area where the submersible is located, and identify the underwater image data based on an underwater target recognition model to obtain the classification information of the target object in the water area where the submersible is located;
[0046] An underwater dynamic map module, configured to construct a three-dimensional map of the underwater environment through a simultaneous localization and mapping model, update the three-dimensional position of the target object and the classification information of the target object in real time, and integrate them into the three-dimensional map of the underwater environment to obtain a three-dimensional dynamic map of the underwater environment;
[0047] A real-time path planning module, configured to set the starting point and the target point of the submersible according to the three-dimensional dynamic map of the underwater environment, perform path planning on the submersible through a path planning algorithm, and generate a travel path of the submersible.
[0048] In a third aspect, the present invention also provides an electronic device, which is characterized by including: at least one processor, at least one memory, a communication interface, and a bus;
[0049] Among them, the processor, the memory, and the communication interface complete mutual communication through the bus. The memory stores program instructions executable by the processor, and the processor invokes the program instructions to implement the steps of a submersible path planning method based on visual recognition.
[0050] In a fourth aspect, the present invention also provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer instructions, and the computer instructions enable a computer to implement the steps of a submersible path planning method based on visual recognition.
[0051] A submersible path planning method and system based on visual recognition of the present invention have the following beneficial effects compared with the prior art:
[0052] (1) By processing and recognizing underwater image data, the precise three-dimensional position and classification information of the target object are obtained, and this information is integrated into a dynamically updated three-dimensional map, providing comprehensive and real-time environmental information for path planning. Combining visual recognition technology with three-dimensional map construction and path planning algorithms can perceive and adapt to changes in the underwater environment in real time, reducing the computational pressure of submersible path planning, improving the processing efficiency of submersible path planning, and realizing real-time path planning of the submersible;
[0053] (2) High-precision target positioning is achieved through a three-dimensional position recognition model. By using 2D feature extraction based on ResNet-50 and depth estimation of MonoDepth2, combined with feature fusion and 3D position regression, the positioning accuracy of underwater targets is improved;
[0054] (3) High-precision target object classification is achieved through an underwater target recognition model. The YOLOv4 architecture is adopted, including a CSPDarknet53 feature extractor, FPN and PAN structures, as well as multi-scale prediction, improving the accuracy of target detection and providing a reliable perception basis for the navigation and task execution of the submersible in a complex underwater environment;
[0055] (4) By setting a cost function and a heuristic function to generate an initial optimal path, continuously detecting environmental changes during the execution process, dynamically updating the cost function, and performing local path optimization, it ensures global optimality and also has the flexibility to handle local changes. Through path smoothing, it ensures that the generated path meets the motion constraints of the submersible, improving the executability of the path and the motion efficiency of the submersible. Description of the Drawings
[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 Flowchart of a submersible path planning method based on visual recognition according to the present invention;
[0058] Figure 2 Structural diagram of a three-dimensional position recognition model according to the present invention;
[0059] Figure 3 Structural diagram of an underwater target recognition model according to the present invention;
[0060] Figure 4 Structural diagram of a submersible path planning system based on visual recognition according to the present invention. Specific implementation manners
[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in combination with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0062] Please refer to Figure 1 , the present invention provides a submersible path planning method based on visual recognition, including the following steps:
[0063] S1. Collect the original underwater image data of the water area where the submersible is located, and preprocess the original underwater image data to obtain underwater image data;
[0064] S2. Based on the three-dimensional position recognition model, recognize the underwater image data to obtain the three-dimensional positions of the target objects in the water area where the submersible is located, and based on the underwater target recognition model, recognize the underwater image data to obtain the classification information of the target objects in the water area where the submersible is located;
[0065] S3. Construct a three-dimensional map of the underwater environment through a simultaneous localization and mapping model, update the three-dimensional positions of the target objects and the classification information of the target objects in real time, and integrate them into the three-dimensional map of the underwater environment to obtain a three-dimensional dynamic map of the underwater environment;
[0066] S4. Set the starting point and target point of the submersible according to the three-dimensional dynamic map of the underwater environment, and perform path planning on the submersible through a path planning algorithm to generate the travel path of the submersible.
[0067] Specifically, in this embodiment, through the processing and recognition of underwater image data, the accurate three-dimensional position and classification information of the target object are obtained, and this information is integrated into the dynamically updated three-dimensional map, providing comprehensive and real-time environmental information for path planning. Combining visual recognition technology with three-dimensional map construction and path planning algorithms can perceive and adapt to changes in the underwater environment in real time, reduce the computational pressure of submersible path planning, improve the processing efficiency of submersible path planning, and achieve real-time path planning of the submersible.
[0068] Step S1 includes:
[0069] Use a high-resolution underwater camera to collect the original underwater image data of the water area where the submersible is located, perform noise reduction processing on the original underwater image data through Gaussian filtering to obtain the first underwater image data, perform image enhancement on the first underwater image data to obtain the second underwater image data, perform color correction on the second underwater image data to obtain the third underwater image data, perform image defogging processing on the third underwater image data to obtain the fourth underwater image data, and perform geometric correction on the fourth underwater image data to obtain the underwater image data.
[0070] In a specific embodiment, use a high-resolution underwater camera (such as a professional underwater camera with 4K resolution and 60fps frame rate) to collect the original underwater image data of the water area where the submersible is located.
[0071] Perform noise reduction processing on the original underwater image data through Gaussian filtering. Use a 5x5 Gaussian kernel with a standard deviation of 1.5 to perform convolution operation on the image to effectively remove high-frequency noise and obtain the first underwater image data.
[0072] Perform image enhancement on the first underwater image data;
[0073] Apply the Contrast Limited Adaptive Histogram Equalization (CLAHE) algorithm. Divide the image into 8x8 small blocks, and limit the contrast threshold to 3.0 to enhance local contrast;
[0074] Use the unsharp masking technique. Blur with a 3x3 Gaussian kernel (standard deviation is 0.5), and then perform weighted superposition with the original image (weight factor is 0.7) to enhance image details and obtain the second underwater image data.
[0075] Perform color correction on the second underwater image data;
[0076] Adopt the white balance algorithm, use the gray world assumption method to estimate the global illumination, and adjust the RGB channel gains;
[0077] Apply the color restoration algorithm, such as the underwater image color restoration method based on the physical model, consider the absorption and scattering characteristics of light, restore the true color, and obtain the third underwater image data.
[0078] Perform image defogging processing on the third underwater image data;
[0079] Apply the improved dark channel prior algorithm, combine with the characteristics of underwater images, estimate the transmission rate map, and optimize the atmospheric light estimation through the soft matting algorithm to achieve underwater image defogging, and obtain the fourth underwater image data.
[0080] Perform geometric correction on the fourth underwater image data;
[0081] Use the perspective transformation method based on feature point matching, combine with the IMU (Inertial Measurement Unit) data, and correct the image distortion caused by the movement of the submersible to obtain the final underwater image data.
[0082] Specifically, in this embodiment, through the comprehensive processing of multiple steps and multiple algorithms, the clarity, contrast, and color restoration degree of underwater images are significantly improved. In particular, the image enhancement using adaptive histogram equalization and unsharp masking techniques, as well as the color restoration based on the physical model, greatly improve the visual quality of underwater images and provide a high-quality data basis for subsequent target recognition and position determination.
[0083] This embodiment combines Gaussian filtering for noise reduction and the dark channel prior defogging algorithm, enabling the submersible to effectively cope with various complex underwater environments, such as turbid water quality, weak light conditions, etc., and greatly expanding the operation range and environmental adaptability of the submersible.
[0084] Through the geometric correction step that combines IMU data, the influence of the movement of the submersible on image acquisition is effectively compensated, ensuring the geometric accuracy of the image data and improving the spatial positioning accuracy of objects in the underwater environment.
[0085] Although a complex image processing process is adopted, by selecting algorithms with higher computational efficiency (such as CLAHE, improved dark channel prior algorithm, etc.) and combining with hardware acceleration technologies (such as GPU parallel computing), while ensuring the processing quality, the real-time requirements can also be met, providing support for the real-time navigation and decision-making of the submersible.
[0086] Step S2 includes:
[0087] The construction process of the three-dimensional position recognition model includes:
[0088] Construct an initial 3D position recognition model based on a convolutional neural network and a deep learning network. The initial 3D position recognition model includes a first input layer, a 2D feature extraction layer, a depth estimation layer, a feature fusion layer, a 3D position regression layer, and a first output layer. The 2D feature extraction layer is constructed based on ResNet-50, and the depth estimation layer is constructed based on MonoDepth2. The first input layer is connected to both the 2D feature extraction layer and the depth estimation layer. Both the 2D feature extraction layer and the depth estimation layer are connected to the feature fusion layer. The feature fusion layer is connected to the 3D position regression layer, and the 3D position regression layer is connected to the first output layer. The first output layer outputs 3D position coordinates;
[0089] Obtain a first historical underwater image dataset and divide the first historical underwater image dataset into a first training set and a first validation set;
[0090] Train and validate the initial 3D position recognition model according to the first training set and the first validation set, and optimize the hyperparameters of the initial 3D position recognition model according to the results of the training and validation. After the training and validation are completed, an optimized 3D position recognition model is obtained;
[0091] The output of the 3D position recognition model is:
[0092] P 3D (X1) = Regression(Fusion(ResNet 50(X1), MonoDepth2(X1)));
[0093] Where X1 is the input of the 3D position recognition model, and P 3D (X1) is the predicted 3D position coordinates of the 3D position recognition model, Regression(·) is the regression layer function, Fusion(·) is the feature fusion function, ResNet50(·) is the two-dimensional feature extracted by the 2D feature extraction layer, and MonoDepth2(·) is the depth map output by the depth estimation layer;
[0094] The loss function of the 3D position recognition model is:
[0095]
[0096] Where L1(P 3D (X1), P true ) is the loss function of the 3D position recognition model, P true is the actual 3D position coordinates of the 3D position recognition model, P 3D,i is the predicted 3D position coordinates of the i-th sample, P true,i is the actual 3D position coordinates of the i-th sample, and N is the total number of samples.
[0097] In a specific embodiment, the detailed construction steps of the initial 3D position recognition model include:
[0098] The first input layer: accepts RGB image inputs of 224x224x3.
[0099] The 2D feature extraction layer: uses the pre-trained ResNet-50, removes the last fully connected layer, retains the convolutional layers, and outputs a 2048-dimensional feature map.
[0100] The depth estimation layer: adopts an improved MonoDepth2 model, uses an encoder-decoder structure, the encoder uses ResNet-18, the decoder uses 4 upsampling blocks, and outputs a depth map of the same size as the input image.
[0101] The feature fusion layer: uses an attention mechanism to perform weighted fusion of 2D features and depth features. The specific implementation is a multi-head self-attention mechanism, with the number of heads set to 8 and the hidden layer dimension to 256.
[0102] The 3D position regression layer: uses 3 fully connected layers, each followed by a ReLU activation function and a Dropout rate of 0.5. The last layer outputs 3 values representing the horizontal, vertical, and longitudinal coordinates of the target object.
[0103] The first output layer: outputs normalized 3D coordinate values within the range of [-1, 1].
[0104] The processing steps of the first historical underwater image dataset include:
[0105] Obtain a dataset containing 100,000 underwater images with accurate 3D position annotations.
[0106] Data augmentation: random rotation (±15°), random scaling (0.8 - 1.2 times), random adjustment of brightness and contrast (±20%), random horizontal flipping.
[0107] Dataset division: 80% for the training set and 20% for the validation set.
[0108] Training and validation of the initial 3D position recognition model:
[0109] Initialization: The 2D feature extraction layer uses the pre-trained weights on ImageNet, and the other layers use He initialization.
[0110] Optimizer: Adam optimizer, with an initial learning rate of 0.0001, using a cosine annealing learning rate scheduler.
[0111] The batch size is set to 32, the number of training epochs is set to 100, and it is evaluated on the validation set in each epoch, and the best model is saved.
[0112] Hyperparameter Optimization: Use the Bayesian optimization method to optimize hyperparameters such as the learning rate, Dropout rate, and the number of attention mechanism heads. The optimization objective is set as the average Euclidean distance error on the validation set.
[0113] Specifically, in this embodiment, by combining 2D image features and depth information and using the attention mechanism for feature fusion, the model can achieve high-precision three-dimensional position recognition in complex underwater environments. Experiments show that the average positioning error on the test set is less than 5 cm, greatly improving the spatial perception ability of the submersible.
[0114] Adopt a variety of data augmentation techniques and an improved loss function to make the model more adaptable to various underwater environmental changes (such as light changes and perspective changes). Introduce the smooth L1 loss to improve the model's ability to handle outliers.
[0115] By using the pre-trained ResNet-50 as a feature extractor and adopting the improved MonoDepth2 for depth estimation, while ensuring the accuracy, the computational amount is significantly reduced. Combined with GPU acceleration, the model can achieve a near-real-time processing speed (>20 fps), meeting the requirements of real-time navigation of the submersible.
[0116] Step S2 also includes:
[0117] The construction process of the underwater target recognition model includes:
[0118] Construct an initial underwater target recognition model, which includes a second input layer, a Backbone layer, a Neck layer, a Head layer, and a second output layer. The second input layer is connected to the Backbone layer, the Backbone layer is connected to the Neck layer, the Neck layer is connected to the Head layer, the Head layer is connected to the second output layer. The Backbone layer includes a CSPDarknet53 feature extractor, a convolutional layer, and a CSP block. The Neck layer includes a top-down FPN structure and a bottom-up feature pyramid containing two PAN structures. The Head layer includes multiple convolutional layers and three prediction scales. The second output layer outputs the target classification probability.
[0119] Obtain a second historical underwater image dataset and divide the second historical underwater image dataset into a second training set and a second validation set.
[0120] Train and validate the initial underwater target recognition model according to the second training set and the second validation set, and optimize the hyperparameters of the initial underwater target recognition model according to the results of training and validation. After training and validation are completed, an optimized underwater target recognition model is obtained.
[0121] Specifically,
[0122] The output of the underwater target recognition model is:
[0123] Y2(X2) = H(N(B(X2))) + L2(Y2(X2), Y2 true (X2));
[0124] where X2 is the input of the underwater target recognition model, Y2(X2) is the classification prediction result of the underwater target recognition model, B(·) is the Backbone function, N(·) is the Neck function, H(·) is the Head function, L2(·) is the loss function of the underwater target recognition model, is the actual classification result of the underwater target recognition model;
[0125] The loss function of the underwater target recognition model is:
[0126]
[0127] where are the weight coefficients for adjusting classification, localization, and confidence of the loss function of the underwater target recognition model respectively, is the classification loss of the loss function of the underwater target recognition model, is the localization loss of the loss function of the underwater target recognition model, is the confidence loss of the loss function of the underwater target recognition model.
[0128] In a specific embodiment, the detailed construction steps of the initial underwater target recognition model include:
[0129] Second input layer: Accepts RGB image input of 416x416x3.
[0130] The Backbone layer includes:
[0131] CSPDarknet53 feature extractor, using 53 convolutional layers, including 5 CSP blocks.
[0132] Convolutional layers, using 3x3 and 1x1 convolutional kernels alternately, with strides of 1 and 2.
[0133] CSP blocks, each block contains 1 1x1 convolutional layer and multiple 3x3 convolutional layers, using residual connections.
[0134] The Neck layer includes:
[0135] FPN structure, 3 layers from top to bottom, with feature map sizes of 13x13, 26x26, and 52x52 respectively.
[0136] PAN structure, two bottom-up paths, each path containing 3 convolutional layers and 1 upsampling layer.
[0137] The Head layer includes:
[0138] 3 prediction scales, namely 13x13, 26x26, and 52x52 respectively. Each prediction scale contains 3 convolutional layers, and the number of output channels of the last layer is 3*(5 + E), where E is the number of data categories.
[0139] The second output layer: outputs the position, confidence, and class probability of each prediction box.
[0140] The processing steps of the second historical underwater image dataset include:
[0141] Obtain an underwater image dataset containing 200,000 images with target annotations, covering 20 common underwater targets.
[0142] Perform data augmentation: Mosaic augmentation, splicing 4 images into one; MixUp augmentation, mixing two images in a certain proportion, randomly rotating (±15°), randomly scaling (0.8 - 1.2 times), randomly flipping horizontally, randomly flipping vertically, randomly adjusting brightness (±30%), contrast (±20%), saturation (±20%), and randomly adding Gaussian noise and salt-and-pepper noise.
[0143] Dataset division: 80% for the training set and 20% for the validation set.
[0144] Training and validation of the initial underwater target recognition model:
[0145] Initialization: The Backbone layer uses pre-trained weights, and other layers use He initialization.
[0146] Optimizer: SGD optimizer, with an initial learning rate of 0.01, momentum of 0.937, and weight decay of 0.0005.
[0147] Learning rate scheduling: Cosine annealing strategy, with the minimum learning rate being 1 / 100 of the initial learning rate.
[0148] The sample batch size is set to 64, the number of training epochs is set to 300 epochs, and it is evaluated on the validation set every 10 epochs, and the best model is saved.
[0149] Hyperparameter optimization: Use the genetic algorithm to optimize hyperparameters such as anchor box sizes and loss function weight coefficients.
[0150] Specifically, in this embodiment, by adopting an improved YOLOv4 architecture and various optimization strategies, high-precision target recognition is achieved in a complex underwater environment. On the test set, the average precision reaches 92%, greatly improving the underwater target recognition ability of the submersible.
[0151] By using CSPDarknet53 as the backbone network and adopting a multi-scale prediction strategy, while ensuring high precision, the processing speed is significantly improved. The model can achieve a real-time processing speed of over 30 FPS, meeting the requirements of real-time navigation and decision-making of the submersible.
[0152] Step S3 includes:
[0153] Utilize the Simultaneous Localization and Mapping (SLAM) model to construct a three-dimensional map of the underwater environment. As time changes, the three-dimensional position of the target and the target classification information are updated in real time, and the real-time updated three-dimensional position of the target and the target classification information are integrated into the three-dimensional map of the underwater environment to obtain a three-dimensional dynamic map of the underwater environment and display it on the monitoring terminal screen of the submersible.
[0154] In a specific embodiment, the selection and implementation steps of the Simultaneous Localization and Mapping (SLAM) model include:
[0155] First, adopt ORB-SLAM2 as the basic framework, which is suitable for sparse feature point extraction and tracking in the underwater environment. Use the ORB feature extraction algorithm, combined with the characteristics of underwater images, to enhance the stability and matching accuracy of feature points. Combine IMU (Inertial Measurement Unit) data for Visual-Inertial Odometry (VIO) fusion to improve the positioning accuracy and robustness.
[0156] Secondly, conduct the construction and update of the three-dimensional map;
[0157] Construct an initial sparse three-dimensional map through the key frames and feature points obtained by the SLAM system. For each newly acquired frame, extract feature points in real time and match them with the features in the map to update the three-dimensional position and classification information of the target object. Use local map optimization to optimize the accuracy of the local map, and conduct global map optimization regularly to ensure the accuracy and consistency of the entire map.
[0158] Thirdly, perform the integration and display of target information;
[0159] Combine the target classification information output by the target recognition model with the three-dimensional position data and update it into the three-dimensional map. Use color coding and labels to identify different types of target objects to enhance the readability of the map. On the monitoring terminal of the submersible, use OpenGL or WebGL technology to render the three-dimensional dynamic map in real time to provide intuitive visual feedback and obtain a three-dimensional dynamic map of the underwater environment.
[0160] Finally, conduct data management and storage;
[0161] Use a database (such as SQLite) to store historical map data and target information, support backtracking and analysis, implement incremental data storage and compression, and optimize storage space and access speed.
[0162] Specifically, in this embodiment, by combining SLAM and target recognition technologies, a high-precision three-dimensional dynamic map can be constructed in a complex underwater environment. The real-time updated map provides accurate environmental perception capabilities, supports the autonomous navigation and task execution of the submersible. Using visual-inertial fusion and local / global optimization technologies, the system realizes real-time map updating and display while ensuring high precision. Even in a dynamically changing underwater environment, the system can maintain stable performance.
[0163] Through the intuitive three-dimensional map display, the submersible operator can obtain real-time environmental information and the status of target objects. This visualization ability not only improves the safety of operation but also provides support for the decision-making of the submersible.
[0164] Step S4 includes:
[0165] S41, Mark the current position of the submersible on the three-dimensional dynamic map of the underwater environment as the starting point, set the task target point, initialize the path planning algorithm, set the cost function and the heuristic function, perform the initial search of the path planning algorithm based on the cost function and the heuristic function, generate the initial optimal path, and the submersible initially travels along the initial optimal path;
[0166] S42, Detect the environmental changes of the current node, update the cost function of the current node based on the environmental changes of the current node, re-plan the local optimal path based on the updated cost function and the heuristic function, and drive the submersible to travel along the local optimal path;
[0167] S43, Repeat step S42 until the submersible reaches the task target point, smooth the path generated each time to ensure that the motion constraints of the submersible are met, and sequentially connect the nodes where the initial optimal path and each segment of the local optimal path intersect to obtain the travel path of the submersible.
[0168] Specifically, in this embodiment, by real-time detecting environmental changes and updating the cost function, the path planning strategy can be dynamically adjusted, enabling the submersible to flexibly respond to complex and changeable underwater environments, such as sudden obstacles or water flow changes, greatly improving the navigation ability and safety of the submersible in unknown or dynamic environments.
[0169] Adopting the strategy of combining initial global path planning with dynamic local path adjustment significantly reduces the computational complexity, avoids global path replanning every time the environment changes, and focuses on local optimization, thus achieving efficient real-time path planning. This method can complete local path adjustment within milliseconds, meeting the real-time requirements during the high-speed movement of the submersible.
[0170] By smoothing the generated path, it ensures that the path meets the movement constraints of the submersible, not only improving the executability of the path but also optimizing the movement trajectory of the submersible, reducing unnecessary sharp turns and acceleration changes, thereby reducing energy consumption, extending the working time of the submersible, and improving the stability of navigation at the same time.
[0171] Through the combination of initial global path planning and continuous local optimization, while ensuring global optimality, it also has the flexibility to handle local changes, enabling the submersible to follow the overall optimal strategy in a complex underwater environment and flexibly respond to local challenges, improving the efficiency and success rate of task completion.
[0172] The calculation formula of the cost function is as follows:
[0173]
[0174] The calculation formula of the heuristic function is as follows:
[0175]
[0176] The calculation formula of the updated cost function is as follows:
[0177]
[0178] The calculation formula of the path smoothing process is as follows:
[0179]
[0180] Among them, g(n) is the cost function from the starting position to the current node n, n j is the node j on the submersible path, n j+1 is the node j + 1 on the submersible path, d(n j , n j+1 ) is the cost between node j and node j + 1, h(n) is the predicted cost from the current node n to the target node, (x n , y n , z n ) and (x goal , y goal , z goal) are the three-dimensional position coordinates of the current node n and the target node respectively, g'(n) is the cost function of the current node n after the environmental change update, and c(n j ) is the additional cost of environmental change for node j, P(t) is the path curve after path smoothing, t is the time parameter, P1 and P2 are the interpolation nodes on the path curve respectively, and P0 and P3 are the curve smooth and continuous points on the path curve respectively.
[0181] Specifically, the cost function g(n) of this embodiment considers the cumulative cost from the starting position to the current node, while the heuristic function h(n) estimates the expected cost from the current node to the target. By combining the cost function and the heuristic function, it ensures that the path planning considers both known information and estimates future costs, thus achieving a better path selection globally. The heuristic function uses the Euclidean distance, providing a reasonable estimate and ensuring the optimality of the algorithm.
[0182] The updated cost function g'(n) introduces the additional cost of environmental change c(n j ), enabling it to respond quickly to environmental changes. This dynamic adjustment mechanism enables the submersible to adapt to complex and changing underwater environments in real time, such as avoiding newly emerged obstacles or taking advantage of favorable water currents, greatly improving the safety and efficiency of navigation.
[0183] The path smoothing process adopts the cubic Hermite interpolation method. By calculating P(t), it ensures the continuity and smoothness of the path, not only improving the executability of the path but also optimizing the motion characteristics of the submersible. The smooth path reduces sharp turns and sudden speed changes, thereby reducing energy consumption and extending the working time of the submersible.
[0184] Please refer to Figure 2 , the present invention also provides a submersible path planning system based on visual recognition, and the system includes:
[0185] An underwater image processing module for collecting underwater raw image data of the water area where the submersible is located, preprocessing the underwater raw image data to obtain underwater image data;
[0186] An object classification and positioning module for identifying the underwater image data based on a three-dimensional position recognition model to obtain the three-dimensional positions of target objects in the water area where the submersible is located, and identifying the underwater image data based on an underwater target recognition model to obtain the classification information of target objects in the water area where the submersible is located;
[0187] An underwater dynamic map module for constructing a three-dimensional map of the underwater environment through a simultaneous localization and mapping model, updating the three-dimensional positions of the target objects and the classification information of the target objects in real time, and integrating them into the three-dimensional map of the underwater environment to obtain a three-dimensional dynamic map of the underwater environment;
[0188] A real-time path planning module is used to set the starting point and target point of the submersible according to the three-dimensional dynamic map of the underwater environment, and perform path planning on the submersible through a path planning algorithm to generate the driving path of the submersible.
[0189] Specifically, a submersible path planning system based on visual recognition in this embodiment realizes the comprehensive processing from the original image data to the three-dimensional position and classification information of the target object through the underwater image processing module and the object classification and positioning module. Through the integrated environmental perception ability, the submersible can comprehensively and accurately understand the complex underwater environment, providing a reliable information basis for path planning. In a complex and dynamic underwater environment, this highly integrated perception ability can significantly improve the environmental adaptability and task execution efficiency of the submersible.
[0190] The underwater dynamic map module not only constructs a static three-dimensional environmental map through the simultaneous localization and mapping technology, but also can update the position and classification information of the target object in real time. This dynamic modeling ability enables the system to timely capture and respond to changes in the underwater environment, such as moving fish schools, floating objects or other dynamic obstacles. It improves the navigation safety and task adaptability of the submersible in a complex and changing underwater environment, enabling it to better cope with unknown or rapidly changing situations.
[0191] Based on the latest three-dimensional dynamic map of the underwater environment, the real-time path planning module can quickly generate and adjust the driving path of the submersible. Through this real-time planning ability based on the latest environmental information, the submersible can flexibly navigate in a complex and changeable underwater environment, effectively avoid obstacles, and optimize the path according to the task requirements. Especially in the face of emergencies or sudden changes in the environment, the system can quickly re-plan the path to ensure the safety of the submersible and the continuity of the task. This adaptive ability greatly improves the operation efficiency of the submersible in various underwater environments.
[0192] The present invention also discloses an electronic device, including: at least one processor, at least one memory, a communication interface and a bus: wherein, the processor, the memory and the communication interface complete communication with each other through the bus; the memory stores program instructions executable by the processor, and the processor calls the program instructions to implement a submersible path planning method based on visual recognition.
[0193] The present invention also discloses a computer-readable storage medium storing computer instructions, which cause the computer to implement all or part of the steps of the method for path planning of a submersible based on visual recognition according to the embodiments of the present invention. The storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory ROM, random access memory RAM, magnetic disks, or optical discs.
[0194] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A path planning method for a submersible based on visual recognition, characterized in that Including the following steps: S1. Collect the original underwater image data of the water area where the submersible is located, preprocess the original underwater image data to obtain underwater image data; S2. Identify the underwater image data based on a three-dimensional position recognition model to obtain the three-dimensional positions of target objects in the water area where the submersible is located, and identify the underwater image data based on an underwater target recognition model to obtain the classification information of target objects in the water area where the submersible is located; S3. Construct a three-dimensional map of the underwater environment through a simultaneous localization and mapping model, update the three-dimensional positions of the target objects and the classification information of the target objects in real time, and integrate them into the three-dimensional map of the underwater environment to obtain a three-dimensional dynamic map of the underwater environment; S4. Set the starting point and target point of the submersible according to the three-dimensional dynamic map of the underwater environment, perform path planning on the submersible through a path planning algorithm, and generate a driving path for the submersible.
2. The method for path planning of a submersible based on visual recognition according to claim 1, wherein Step S1 includes: Use a high-resolution underwater camera to collect the original underwater image data of the water area where the submersible is located, perform noise reduction processing on the original underwater image data through Gaussian filtering to obtain the first underwater image data, perform image enhancement on the first underwater image data to obtain the second underwater image data, perform color correction on the second underwater image data to obtain the third underwater image data, perform image dehazing processing on the third underwater image data to obtain the fourth underwater image data, and perform geometric correction on the fourth underwater image data to obtain underwater image data.
3. The method for path planning of a submersible based on visual recognition according to claim 2, wherein Step S2 includes: The construction process of the three-dimensional position recognition model includes: Construct an initial three-dimensional position recognition model based on a convolutional neural network and a deep learning network. The initial three-dimensional position recognition model includes a first input layer, a 2D feature extraction layer, a depth estimation layer, a feature fusion layer, a 3D position regression layer, and a first output layer. The 2D feature extraction layer is constructed based on ResNet-50, the depth estimation layer is constructed based on MonoDepth2, the first input layer is respectively connected to the 2D feature extraction layer and the depth estimation layer, both the 2D feature extraction layer and the depth estimation layer are connected to the feature fusion layer, the feature fusion layer is connected to the 3D position regression layer, the 3D position regression layer is connected to the first output layer, and the first output layer outputs three-dimensional position coordinates; Obtain a first historical underwater image data set, and divide the first historical underwater image data set into a first training set and a first validation set; Train and validate the initial three-dimensional position recognition model according to the first training set and the first validation set, optimize the hyperparameters of the initial three-dimensional position recognition model according to the results of training and validation, and obtain an optimized three-dimensional position recognition model after training and validation.
4. The method for path planning of a submersible based on visual recognition according to claim 3, wherein, Step S2 also includes: The construction process of the underwater target recognition model includes: Construct an initial underwater target recognition model, where the initial underwater target recognition model includes a second input layer, a Backbone layer, a Neck layer, a Head layer, and a second output layer. The second input layer is connected to the Backbone layer, the Backbone layer is connected to the Neck layer, the Neck layer is connected to the Head layer, and the Head layer is connected to the second output layer. The Backbone layer includes a CSPDarknet53 feature extractor, a convolutional layer, and a CSP block. The Neck layer includes a top-down FPN structure and a bottom-up feature pyramid containing two PAN structures. The Head layer includes multiple convolutional layers and three prediction scales. The second output layer outputs the target classification probability; Obtain a second historical underwater image dataset and divide the second historical underwater image dataset into a second training set and a second validation set; Train and validate the initial underwater target recognition model according to the second training set and the second validation set, optimize the hyperparameters of the initial underwater target recognition model according to the results of training and validation, and obtain an optimized underwater target recognition model after training and validation are completed; The output of the underwater target recognition model is: Y2(X2) = H(N(B(X2))) + L2(Y2(X2), Y2 true (X2)); Among them, X2 is the input of the underwater target recognition model, Y2(X2) is the classification prediction result of the underwater target recognition model, B(·) is the Backbone function, N(·) is the Neck function, H(·) is the Head function, and L2(·) is the loss function of the underwater target recognition model. is the actual classification result of the underwater target recognition model; The loss function of the underwater target recognition model is: Among them, are the weight coefficients for adjusting classification, localization, and confidence of the loss function of the underwater target recognition model, is the classification loss of the loss function of the underwater target recognition model, is the localization loss of the loss function of the underwater target recognition model, is the confidence loss of the loss function of the underwater target recognition model.
5. The method for path planning of a submersible based on visual recognition according to claim 4, characterized in that Step S3 includes: Use a simultaneous localization and mapping model to construct a three-dimensional map of the underwater environment, update the target three-dimensional position and the target classification information in real time as time changes, and integrate the real-time updated target three-dimensional position and the target classification information into the three-dimensional map of the underwater environment to obtain a three-dimensional dynamic map of the underwater environment and display it on the monitoring terminal screen of the submersible.
6. The method for path planning of a submersible based on visual recognition according to claim 5, wherein Step S4 includes: S41, Mark the current position of the submersible on the three-dimensional dynamic map of the underwater environment as the starting point, set the task target point, initialize the path planning algorithm, set the cost function and the heuristic function, perform the initial search of the path planning algorithm based on the cost function and the heuristic function, generate an initial optimal path, and the submersible initially travels along the initial optimal path; S42, Detect the environmental changes of the current node, update the cost function of the current node based on the environmental changes of the current node, re-plan the local optimal path based on the updated cost function and the heuristic function, and drive the submersible to travel along the local optimal path; S43, Repeat step S42 until the submersible reaches the task target point, smooth the path generated each time to ensure that the motion constraints of the submersible are satisfied, and sequentially connect the nodes where the initial optimal path and each segment of the local optimal path intersect to obtain the travel path of the submersible.
7. The method for path planning of a submersible based on visual recognition according to claim 6, wherein, The calculation formula of the cost function is: The calculation formula of the heuristic function is: The calculation formula of the updated cost function is: The calculation formula of the path smoothing process is: Among them, g(n) is the cost function from the starting position to the current node n, where n j is the node j on the submersible path, and n j+1 is the node j + 1 on the submersible path. d(n j , n j+1 ) is the cost between node j and node j + 1, h(n) is the predicted cost from the current node n to the target node, (x n , y n , z n ) and (x goal , y goal , z goal ) are the three-dimensional position coordinates of the current node n and the target node respectively. g'(n) is the cost function of the current node n after being updated due to environmental changes, c(n j ) is the additional cost of environmental changes for node j, P(t) is the path curve after path smoothing, t is the time parameter, P1 and P2 are the interpolation nodes on the path curve respectively, and P0 and P3 are the curve smooth and continuous points on the path curve respectively.
8. A submersible path planning system based on visual recognition, characterized in that, The system includes: An underwater image processing module for collecting underwater original image data in the water area where the submersible is located, preprocessing the underwater original image data, and obtaining underwater image data; An object classification and positioning module, configured to identify the underwater image data based on a three-dimensional position recognition model to obtain the three-dimensional positions of target objects in the water area where the submersible is located, and identify the underwater image data based on an underwater target recognition model to obtain the classification information of the target objects in the water area where the submersible is located; An underwater dynamic map module, configured to construct a three-dimensional map of the underwater environment through a simultaneous localization and mapping model, update the three-dimensional positions of the target objects and the classification information of the target objects in real time, and integrate them into the three-dimensional map of the underwater environment to obtain a three-dimensional dynamic map of the underwater environment; A real-time path planning module, configured to set the starting point and the target point of the submersible according to the three-dimensional dynamic map of the underwater environment, perform path planning on the submersible through a path planning algorithm, and generate a travel path of the submersible.
9. An electronic device, characterized in that, Comprising: At least one processor, at least one memory, a communication interface, and a bus; Wherein, the processor, the memory, and the communication interface complete mutual communication through the bus, the memory stores program instructions executable by the processor, and the processor invokes the program instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to implement the method according to any one of claims 1 to 7.