A side scan sonar image feature extraction method and system

By combining the LF-Net deep learning network and transfer learning with the KNN algorithm, the problems of incomplete data coverage and data dependency in side-scan sonar image matching are solved, achieving efficient and accurate underwater target matching, which is applicable to fields such as marine resource development and military exploration.

CN120411539BActive Publication Date: 2025-11-18SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510503504.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-11-18
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Existing side-scan sonar image matching methods suffer from problems such as incomplete data coverage, low efficiency and strong subjectivity of traditional methods, and excessive data dependence of deep learning methods, which make it difficult to match underwater targets.

Method used

By employing the LF-Net deep learning network combined with transfer learning and the KNN algorithm, feature point detection and description are performed through a pre-trained model, achieving end-to-end feature extraction and matching, and reducing reliance on side-scan sonar image data.

Benefits of technology

It achieves efficient and accurate side-scan sonar image matching under small sample data conditions, improves the matching quality of complex underwater scenes, reduces data acquisition and annotation costs, and is applicable to fields such as marine resource development and military exploration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411539B_ABST
    Figure CN120411539B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of image processing, and discloses a side-scan sonar image feature extraction method and system, which extracts feature points of two adjacent side-scan sonar images by using an LF-Net network, and matches the feature points extracted by the LF-Net by using a KNN algorithm. The application introduces a transfer learning mode based on a deep learning network LF-Net to solve the sonar image matching problem, detects feature points of side-scan sonar images by constructing a pre-training measured data matching model, and solves the problem of poor generalization ability caused by the small amount of side-scan sonar image data. The application realizes accurate matching of sonar image feature points by combining the KNN algorithm, reduces the influence of false matching by statistically analyzing the matching results, and improves the robustness of matching. The application can accurately match two adjacent side-scan sonar images without too much side-scan sonar image data, and provides a new solution for underwater sonar image matching in complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, and in particular relates to a method and system for extracting features from side-scan sonar images. Background Technology

[0002] With the vigorous development of various marine activities and the continuous development of coastal areas, the demand for underwater topographic maps is increasing. Side-scan sonar (SSS) detection technology, based on shipborne and submersible platforms, has become the mainstream method for detecting seabed topography and underwater targets due to its advantages such as high resolution and full coverage. However, in different real-world scenarios, due to the limitation of the scan width, a single side-scan image may not completely cover the target. This makes it difficult to fully construct the image of some larger underwater targets, hindering the reconstruction of their overall appearance and thus restricting their application in important fields such as seabed topographic mapping, seabed resource development, and AUV navigation. Therefore, accurate feature extraction and matching of side-scan sonar images is of great significance.

[0003] During operation, side-scan sonar records the intensity of backscattered echoes from the seabed and generates a grayscale image based on the intensity; this image is called a side-scan sonar image. Traditional methods of manually matching adjacent side-scan sonar images suffer from low efficiency and high subjectivity. Therefore, many scholars have conducted research on feature extraction and matching of side-scan sonar images.

[0004] The paper "Sylvie Daniel, Fabrice Le L'eannec. Side-Scan Sonar Image Matching. IEEE Journal of Oceanic Engineering, 1998, 23(3): 245-259P" proposes an algorithm for side-scan sonar image matching based on acoustic shadow features. It matches two side-scan sonar images by matching the geometric shape of shadow features in the two images and their corresponding key position information. The paper "KHATER HA, GAD AS, OMRAN EA, et al. Enhancement matching algorithms using fusion of multiple similarity metrics for sonar images[J]. World Applied Sciences Journal, 2009, 6(6): 759-763" proposes a side-scan sonar image matching strategy that integrates SUSAN and Harris corner features. Under ideal conditions where the features are stable and evenly distributed, this method shows superior matching performance. The paper "VANDRISH P, VARDYA, WALKER D, et al. Side-scan sonar image registration for AUV navigation[C] / / 2011IEEE Symposium on Underwater Technology and Workshop on Scientific Use of Submarine Cables and Related Technologies.IEEE,2011:1-7." attempts to apply the SIFT algorithm to the matching process of side-scan sonar images acquired by AUVs. The results show that the registration effect is better when the side-scan image features are obvious and the noise is small. The paper "Yang Fanlin. Multibeam and side-scan sonar data fusion and its application in seabed sediment classification[D]. Wuhan University,2003." proposes a registration method based on the same feature points of contour lines and isobaths.The paper "Zhao J, Tao W, Zhang H, et al. Study on side scan sonar image matching based on the integration of SURF and similarity calculation of typical areas[C]. Oceans 2010IEEE-Sydney.IEEE,2010:1-4P." proposes a side scan sonar image matching method based on the SURF algorithm and similarity calculation of typical areas. It first performs coarse matching of adjacent strips using the SURF algorithm, then performs geometric correction to ensure the matched image and the reference image have the same observation conditions. Finally, it completes precise matching by calculating the similarity of typical areas in the two corrected images. While these methods have achieved feature extraction and matching of side scan sonar images to varying degrees using traditional methods, they generally suffer from poor generalization ability. In cases where side scan sonar images have high noise levels, many invalid features, and low signal-to-noise ratios, they generate numerous mismatches, significantly deteriorating the matching performance.

[0005] In recent years, the rapid development of deep learning has prompted scholars to explore the potential of convolutional neural networks (CNNs) in sonar image feature extraction, aiming to obtain deep, multi-dimensional features to improve the accuracy of image matching. However, current research on the application of CNNs in sonar image matching is still scarce.

[0006] The paper "VALDENEGRO-TORO M. Improving sonar image patch matching via deep learning[C] / / 2017 European Conference on Mobile Robots(ECMR).IEEE,2017:1-6" first applied the idea of ​​deep learning to the sonar image matching problem, and proposed a method to solve the sonar image matching problem using the idea of ​​binary classification. The literature “Yang Haibo. Research on Preprocessing and Matching Methods for Side-Scan Sonar Images [D]. Harbin Engineering University, 2020” proposes a side-scan sonar image template matching method based on the U-net (Ronneberger O, Fischer P, Brox TU-net: Convolutional networks for biomedical image segmentation [C] / / Medical imagecomputing and computer-assisted intervention–MICCAI 2015:18th international conference,Munich,Germany,October 5-9,2015,proceedings,part III 18.SpringerInternational Publishing,2015:234-241.) network, which is mainly aimed at the side-scan sonar image matching and stitching scenarios on the seabed with rugged terrain.The paper “Zhou Xiaoteng. Application of Deep Learning-Based Side-Scan Sonar Image Matching and UUV-Assisted Navigation [D]. Harbin Institute of Technology, 2022. DOI:10.27061 / d.cnki.ghgdu.2022.001205” explores the simple and complex environments covered in underwater navigation tasks, and discusses the feasibility of combining deep learning matching algorithms, such as AffNet (Chen X, Fu C, Tie M, et al. AFFNet: An attention-based feature-fused network for surface defect segmentation [J]. Applied Sciences, 2023, 13(11):6428.), with HardNet (Chao P, Kao CY, Ruan YS, et al. Hardnet: A low memory traffic network [C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2019:3552-3561.), to perform feature detectors in different scenarios. The above method trains the model using a data-driven approach and achieves matching of side-scan sonar images. While this method achieves matching of side-scan sonar images, it requires a large amount of side-scan sonar image data for model training; otherwise, the matching effect is poor.

[0007] LF-Net, proposed by Yuki et al. in 2018, is an end-to-end image matching network based on unsupervised learning. Its end-to-end architecture eliminates the need for independent training of the feature point detector and descriptor networks, thus avoiding the previous limitation of separate step-by-step processing of the feature point detector and descriptor. Furthermore, it differs from supervised learning networks that rely on SIFT detectors, such as LIFT networks, which require pre-extracting initial feature points during the self-descriptor training phase and simultaneously enhancing the uniqueness and recognition ability of feature points during detector training. Similar to SIFT, LF-Net achieves matching by selecting several feature points from the two images to be matched and then describing these selected feature points. During training, it utilizes pre-given pose labels to perform recursive accuracy evaluation, thereby optimizing the learning algorithms of the feature point detector and feature point descriptor. The core of LF-Net revolves around two key components: (1) a multi-scale fully convolutional module responsible for generating the coordinates, orientation, and scale information of feature points; and (2) a feature point description module used to output the feature vectors of the feature points. These are referred to as the detection module and description module in LF-Net, respectively.

[0008] Based on the above analysis, the problems and shortcomings of the existing technology are as follows:

[0009] (1) Currently, due to the limitation of the scanning width during the side-scan sonar data acquisition process, the data has the problem that a single image cannot completely cover the target object, making it difficult to completely construct some large-scale underwater targets.

[0010] (2) The traditional method of manually matching adjacent side-scan sonar images has problems of low efficiency and strong subjectivity; at the same time, the method has poor generalization ability, generates a lot of mismatched points, and the matching effect is poor.

[0011] (3) The data set of the deep learning-based side-scan sonar image matching method is relatively small, while the model needs to be trained with a large amount of side-scan sonar image data, otherwise its matching effect is poor. Summary of the Invention

[0012] To overcome the problems existing in related technologies, the present invention discloses a method and system for extracting side-scan sonar image features, and particularly relates to a method and system for extracting side-scan sonar image features based on LF-Net. The technical solution is as follows:

[0013] This invention is implemented as follows: a side-scan sonar image feature extraction method, comprising the following steps:

[0014] S1, LF-Net Training: Construct a network architecture based on a dual-branch structure and use dual-modal loss to train the detection network and construct the loss function;

[0015] S2, Feature map generation: LF-Net uses the ResNet architecture to input the side-scan sonar image into the trained multi-scale fully convolutional network and outputs an image o with feature points;

[0016] S3, Feature point detection: Feature point detection is performed using LF-Net, and the detection network has a serial structure;

[0017] S4, Feature Description: Calculate the orientation map of each feature point using the arctan function; construct a quadruple and perform feature extraction on the image region to obtain the feature description map;

[0018] S5, KNN Corresponding Point Matching: Based on the detected feature points and the obtained feature description map, the KNN algorithm is used to match corresponding points between two adjacent sonar images.

[0019] In step S1, in constructing the dual-branch-based network architecture, the two branches share the same set of parameters, and the input contains the image I to be paired. i ,I jIn addition, depth and pose information are obtained through the structured light algorithm; branch i undertakes differentiable tasks to backpropagate gradients and drive model optimization; branch j performs non-differentiable operations, including feature point extraction and generating a supervised target for the first branch.

[0020] In step S1, the loss function of LF-Net is divided into an image-level loss function and an image patch-level loss function;

[0021] Image-level loss functions are used to optimize the feature point detection network in LF-Net by analyzing image I. j The fractional graph S calculated above j Transform to Figure I i Using the above as training samples, after a rigid transformation ω process, the fractional image S is obtained. j And extract K feature points from the fractional image, and perform a test on the fractional image S. j A Gaussian filter with a kernel of 0.5 is applied to obtain a smoothed score image. The expression for the loss function is as follows:

[0022] L im (S i ,S j )=|S i -g(ω(S j ))| 2

[0023] In the formula, S i Let S be the score map of the i-th image, representing the probability of the feature point response; j Let be the score map of the j-th image; g is a non-differentiable Gaussian filter that applies a Gaussian kernel to clean up feature points; ω is a rigid transformation.

[0024] Image block-level loss functions are used to optimize local descriptors generated by the feature description network, including matching feature point loss L. pair Geometric consistency loss L geom Triple loss L tri ;

[0025] Matching feature point loss L pair Minimize the distance between the feature vectors of the same keypoint in two images to ensure that the descriptors remain consistent after transformation; by extracting K feature points from each iteration of branch i, and using the current iteration number as the image I... i ,I j The interrelationships between them are transformed by rigid transformation of the coordinates of K feature points in image I. i Coordinate system down-transformation to image I j In the coordinate system; after completing the coordinate transformation, based on the scale and orientation information obtained from branch j, the transformation results in the new coordinate system are fused, from I jBy precisely selecting the image region, two image blocks are obtained. and Feature extraction is performed on each image patch, and the extracted features are... and k represents the number of feature point matching pairs. This is the feature descriptor for the k-th keypoint in the i-th image. Given the feature descriptor corresponding to k after geometric transformation of the j-th image, we obtain the image block-level loss function;

[0026]

[0027] LF-Net uses geometric consistency loss L geom This ensures that the scale and orientation of feature points remain consistent across different images, guaranteeing that LF-Net possesses scale invariance and rotation invariance; L geom It consists of two parts: scale loss and orientation loss. This ensures that the direction of the matching key points remains consistent. This ensures that the scale of the matching key points remains consistent. This represents the scale of the k-th keypoint in the i-th image. Indicates the direction of the k-th key point in the i-th image; This represents the scale of the corresponding keypoint in the j-th image after geometric transformation. λ represents the orientation of the corresponding keypoint in the j-th image after geometric transformation. ori and λ scale The hyperparameters Li control the effects of orientation and scale loss, respectively. geom As shown below;

[0028]

[0029] LF-Net uses a triplet loss function to optimize feature descriptors. For non-corresponding image pairs, it uses them as negative samples for progressive mining training, making the distance between matching positive samples closer and the distance between non-matching negative samples farther. The margin C hyperparameter sets a minimum interval to ensure that the distance between matching points and non-matching points is large enough. This represents the feature descriptor of the k-th keypoint in the i-th image. Let represent the positive samples of matching keypoints after geometric transformation of the j-th image. Let L represent a random negative sample in the j-th image, and C represent the Margin hyperparameter in the triplet loss, which controls the separation between positive and negative samples; L represents the triplet loss. tri The mathematical expression for is shown below;

[0030]

[0031] Where k′≠k, C=1;

[0032] After applying the three loss functions, the loss functions used for training the detection network and describing the network are obtained as follows:

[0033] L det =L im +λ pair L pair +L geom

[0034] L desc =L tri .

[0035] In step S2, the ResNet architecture comprises three modules: a convolutional module, a residual module, and a fully connected module. Each module embeds multiple sets of 5×5 convolutional kernels. Batch normalization is performed after each convolutional operation. Leaky-ReLU is used as the activation function, and the number of output channels in all convolutional layers is fixed at 16, with the output dimension consistent with the input. The convolutional module performs preliminary feature extraction, reducing the spatial size of the input data and improving computational efficiency. The residual module addresses the vanishing gradient problem in deep neural networks, enabling efficient information flow between layers and facilitating residual learning through identity mapping. The fully connected module reduces the number of parameters through global average pooling, improving generalization ability, and performs the final classification decision.

[0036] In step S3, feature point detection includes:

[0037] LF-Net performs multi-scale detection based on feature map o; the feature map o is scaled N times, with a scaling range of... N=5, Perform a convolution operation on the feature maps after N scaling operations using a 5×5 convolution kernel to obtain N fractional maps h. N (1≤n≤N), each of the N fractional images corresponds to N fractional images at different scales. Differentiable nonmaximum suppression is achieved by performing convolution operations on each of the N fractional images using a 15×15 Softmax operator, resulting in N fractional images sharpened by the Softmax operator.

[0038] LF-Net will each Resize the image to the original size; the resized fractional image is as follows. Apply an approximate softmax operator to all fractional graphs of the same scale. The data is then merged to form the final score map of the scale space;

[0039]

[0040] In the formula, ⊙ represents the Hadamard product; after scale fusion, the K highest-scoring pixels are selected from the final scale-invariant score map S as feature points and local softargmax is applied to obtain sub-pixel accuracy.

[0041] Furthermore, the Softmax operator is used to normalize the input values ​​into a probability distribution so that the sum of the output values ​​is 1;

[0042] The function definition of the softmax operator is as follows, where softmax(z) i ) represents the value of the i-th element after the Softmax transformation, and represents a probability value in the probability distribution;

[0043]

[0044] This indicates that for input z i Perform an exponential transformation to ensure the value is non-negative and amplify a larger z. i The impact of the value; This represents the sum of the exponential transformations of all input elements, used for normalization to make the sum of all outputs equal to 1, forming a probability distribution.

[0045] In step S4, the feature description includes: applying a 5×5 convolution to each pixel on the feature map o, outputting two values ​​as sine and cosine for each pixel, and calculating the orientation map θ of each feature point using the arctan function;

[0046] Descriptor extraction: By analyzing the score map S, K feature points with the highest scores are selected, and their coordinate data is obtained; the key points are combined with scale and orientation information from the scale estimation map and orientation estimation map to form a four-tuple structure p. k =(x,y,s,θ) k And perform feature extraction operations on the corresponding image regions.

[0047] Furthermore, the orientation map θ of each feature point is calculated using the arctan function. The direction calculated by arctan is a continuously changing angle value, without any abrupt changes in direction, providing a smooth direction estimate that is applicable to various scales, rotations, and perspective changes, thus improving the robustness of the feature points. Moreover, LF-Net needs to find the same keypoints in images with different viewpoints and rotational changes; the direction angle calculated by arctan ensures that the feature points still point in the same direction after rotational transformations, improving rotation invariance. Its calculation formula is as follows:

[0048]

[0049] In the formula, θ represents the direction angle of the key point, and f x f represents the gradient of a feature point in the x-direction. y This represents the gradient of the feature point in the y-direction.

[0050] In step S5, KNN corresponding point matching includes: matching corresponding points between two adjacent sonar images using the KNN algorithm based on the detected feature points and the obtained feature description map;

[0051] Before matching, the feature descriptors output by LF-Net are processed using L2 normalization.

[0052] For two side-scan sonar images to be matched, the KNN algorithm is used to match feature points. Euclidean distance is calculated for feature point matching. The KNN algorithm is used for image registration, target detection, and clustering tasks. In side-scan sonar image matching, this method accurately calculates the similarity of feature points in adjacent images, improving the accuracy of stitching and classification. K feature points and their corresponding descriptors D_A and D_B are extracted respectively. The descriptor set D_B is used as the search database, and the descriptor set D_A is used as the query set. For each descriptor obtained from any one sonar image, the K nearest neighbor descriptors are found for each descriptor obtained from the other sonar image.

[0053]

[0054] In the formula, For feature point A i Coordinates on the X-axis For feature point B j Coordinates on the X-axis; For feature point A i Coordinates on the Y-axis For feature point B j Coordinates on the Y-axis;

[0055] The nearest neighbor is determined by calculating the Euclidean distance between descriptors, and the D_A is calculated for each query descriptor. i With all descriptors D_B in the database j The similarity is used to select the two points with the highest similarity as corresponding points for matching.

[0056] Another object of the present invention is to provide a side-scan sonar image feature extraction system, which is used to control the side-scan sonar image feature extraction method, and the system includes:

[0057] The LF-Net training module is used to build a dual-branch-based network architecture and train the detection network using dual-modal loss.

[0058] The feature map generation module is used by LF-Net to input side-scan sonar images into the trained multi-scale fully convolutional network using the ResNet architecture, and output an image o with feature points;

[0059] The feature point detection module is used to detect feature points using LF-Net, and the detection network has a serial structure.

[0060] The feature description module is used to calculate the orientation map of each feature point using the arctan function; it constructs a quadruple and performs feature extraction on the image region to obtain the feature description map.

[0061] The KNN corresponding point matching module is used to match corresponding points between two adjacent sonar images using the KNN algorithm, based on the detected feature points and the obtained feature description map.

[0062] Combining all the above technical solutions, the beneficial effects of this invention are as follows:

[0063] To address the challenges of limited datasets and difficult model training in deep learning-based side-scan sonar image matching methods, this invention proposes a method that incorporates transfer learning using the LF-Net deep learning network. A pre-trained matching model based on measured data from outdoor scenes is transferred to side-scan sonar images. An automatic feature point extraction algorithm is applied to the side-scan sonar image data, and accurate feature point matching is achieved through a joint KNN matching algorithm. This method can accurately match two adjacent side-scan sonar images even with limited available data. Experiments validated the feasibility of this method and compared it with traditional methods, demonstrating that the LF-Net deep learning network can effectively perform side-scan sonar image matching, providing a novel solution for underwater sonar image matching in complex scenarios.

[0064] To address the issue of insufficient target coverage in a single image during side-scan sonar data acquisition due to the limited scan width, this invention employs transfer learning based on the LF-Net deep learning network to solve the sonar image matching problem. First, this invention constructs a pre-trained, experimentally tested data matching model for feature point detection in side-scan sonar images, overcoming the poor generalization ability caused by the limited amount of side-scan sonar image data. Second, it combines the KNN algorithm to achieve accurate matching of acoustic and image feature points, and statistical analysis of the matching results reduces the impact of mismatches, improving the robustness of the matching process.

[0065] This invention proposes a feature extraction and matching method for side-scan sonar images based on LF-Net, addressing the challenges of limited datasets and difficult model training in deep learning-based side-scan sonar image matching methods. Furthermore, this invention introduces a transfer learning approach using the LF-Net deep learning network. It employs a pre-trained matching model based on measured data from outdoor scenes, transferring this model to side-scan sonar images. An automatic feature point extraction algorithm is used on the side-scan sonar image data, and accurate feature point matching is achieved through a combined KNN matching algorithm, reducing the reliance on massive training data for deep learning methods.

[0066] This invention fully considers the current engineering scenarios that require large-scale, high-resolution measurements, where a single survey line is insufficient to cover the entire target area. It is necessary to match sonar data between two adjacent survey lines. However, existing mainstream algorithms, including the ORB algorithm and the Harris algorithm, have poor performance. This method can effectively improve the quality of sonar image matching.

[0067] This invention reduces data dependence: it eliminates the need for training models with a large number of sonar images, thus reducing data acquisition and annotation costs; it improves efficiency: it is suitable for complex underwater scenarios (such as noisy environments with low signal-to-noise ratios), improving the efficiency of applications such as seabed topography mapping, underwater target detection, and AUV navigation; it has market potential: it can be applied to fields such as marine resource development, military exploration, and shipwreck salvage, promoting the intelligent upgrading of underwater detection technology and possessing high market conversion potential.

[0068] Currently, the field of sonar image matching mainly relies on traditional algorithms (such as SIFT, ORB, and Harris) or deep learning models that require large amounts of data for training. This case introduces LF-Net combined with transfer learning into side-scan sonar image matching, addressing the following gaps: Small sample generalization ability: Traditional methods experience a sharp performance drop when the amount of data is limited, while transfer learning utilizes pre-trained models (such as those used in outdoor scenarios) to adapt to sonar images, overcoming data limitations; End-to-end optimization: LF-Net's integrated design of detection and description modules avoids the error accumulation problem of traditional step-by-step processing, filling the gap in the application of unsupervised learning in sonar image matching.

[0069] Traditional methods face two major challenges in side-scan sonar image matching: mismatches in noisy and low signal-to-noise ratio environments: traditional algorithms (such as SIFT and Harris) rely on manually generated features, resulting in high mismatch rates in complex underwater scenarios; and poor model generalization ability: deep learning requires a large amount of labeled data, while the high cost of acquiring sonar images hinders the practical application of these models. This paper addresses the model generalization problem under limited sample conditions through transfer learning, effectively suppressing mismatches by combining the KNN algorithm with statistical analysis. Experiments demonstrate that its robustness is significantly better than traditional methods, overcoming the aforementioned long-standing technical bottlenecks. Traditionally, deep learning is considered dependent on massive amounts of data: sonar image data is scarce, making direct application of deep learning difficult. This paper adapts pre-trained models to sonar images through transfer learning, verifying the feasibility of deep learning training in scenarios with scarce sonar image samples, thus overcoming technical biases. Attached Figure Description

[0070] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure;

[0071] Figure 1 This is a flowchart of the side-scan sonar image feature extraction method provided in an embodiment of the present invention;

[0072] Figure 2 This is a diagram of the LF-Net structure provided in an embodiment of the present invention;

[0073] Figure 3 This is a schematic diagram of side-scan sonar data 1 provided in an embodiment of the present invention. (a) The figure shows the sonar image of the left side of the adjacent survey line, and (b) The figure shows the sonar image of the right side of the adjacent survey line.

[0074] Figure 4 This is a schematic diagram of side-scan sonar data 2 provided in an embodiment of the present invention. (a) The figure shows the sonar image of the left side of the adjacent survey line, and (b) The figure shows the sonar image of the right side of the adjacent survey line.

[0075] Figure 5 This is a schematic diagram of the feature point detection results of the side-scan sonar data 1 provided in the embodiment of the present invention. (a) The figure shows the sonar image of the left side of the adjacent survey line, and (b) The figure shows the sonar image of the right side of the adjacent survey line.

[0076] Figure 6 This is a schematic diagram of the feature point detection results of the side-scan sonar data 2 provided in the embodiment of the present invention. (a) The figure shows the sonar image of the left side of the adjacent survey line, and (b) The figure shows the sonar image of the right side of the adjacent survey line.

[0077] Figure 7This is a directional description diagram of side-scan sonar data 1 provided in an embodiment of the present invention. (a) The diagram shows the sonar image of the left side of the adjacent survey line, and (b) The diagram shows the sonar image of the right side of the adjacent survey line.

[0078] Figure 8 The following is a directional description diagram of the side-scan sonar data provided in the embodiment of the present invention: (a) is the sonar image of the left side of the adjacent survey line, and (b) is the sonar image of the right side of the adjacent survey line.

[0079] Figure 9 The following is a 1-scale description diagram of the side-scan sonar data provided in the embodiment of the present invention: (a) is the sonar image of the left side of the adjacent survey line, and (b) is the sonar image of the right side of the adjacent survey line.

[0080] Figure 10 The following is a two-scale description diagram of the side-scan sonar data provided in the embodiment of the present invention: (a) is the sonar image of the left side of the adjacent survey line, and (b) is the sonar image of the right side of the adjacent survey line.

[0081] Figure 11 This is a matching result diagram of side-scan sonar data 1 provided in an embodiment of the present invention;

[0082] Figure 12 This is a matching result diagram of side-scan sonar data 2 provided in an embodiment of the present invention;

[0083] Figure 13 This is a diagram showing the matching results of the 1ORB algorithm for side-scan sonar data provided in this embodiment of the invention;

[0084] Figure 14 This is a diagram showing the Harris algorithm matching results of side-scan sonar data provided in this embodiment of the invention;

[0085] Figure 15 This is a diagram showing the 2ORB algorithm matching results of side-scan sonar data provided in this embodiment of the invention;

[0086] Figure 16 This is a diagram showing the matching results of the side-scan sonar data using the Harris algorithm provided in an embodiment of the present invention. Detailed Implementation

[0087] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0088] The innovation of the side-scan sonar image feature extraction method and system provided in this invention is as follows:

[0089] 1. Deep feature adaptation based on transfer learning: By combining LF-Net with transfer learning, the pre-trained natural scene model (Outdoor scene) is transferred to underwater sonar images, which solves the problems of low data volume and high annotation cost of side-scan sonar.

[0090] 2. End-to-end multi-scale feature fusion architecture: Adopting the end-to-end unsupervised architecture of LF-Net, it realizes the integrated optimization of feature detection and description, and breaks through the error accumulation problem of traditional step-by-step processing (such as SIFT staged detection and description).

[0091] To address the challenges of limited datasets and difficult model training in deep learning-based side-scan sonar image matching methods, this invention proposes a method that incorporates transfer learning using the LF-Net deep learning network. A pre-trained matching model based on measured data from outdoor scenes is transferred to side-scan sonar images. An automatic feature point extraction algorithm is applied to the side-scan sonar image data, and accurate feature point matching is achieved through a joint KNN matching algorithm. This proposed method can accurately match two adjacent side-scan sonar images even with limited available data. Experiments validated the feasibility of this method and compared it with traditional methods, demonstrating that the LF-Net deep learning network can effectively perform side-scan sonar image matching tasks, providing a novel solution for underwater sonar image matching in complex scenarios.

[0092] To address the issue of insufficient target coverage in a single image during side-scan sonar data acquisition due to the limited scan width, this invention employs transfer learning based on the LF-Net deep learning network to solve the sonar image matching problem. First, a pre-trained, experimentally-based matching model is constructed to detect feature points in side-scan sonar images, resolving the poor generalization ability caused by the limited amount of side-scan sonar image data. Second, the KNN algorithm is combined to achieve accurate matching of acoustic and image feature points. Statistical analysis of the matching results reduces the impact of mismatches and improves the robustness of the matching process.

[0093] Example 1: When side-scan sonar measures underwater targets, adjacent side-scan sonar images acquired by two adjacent survey lines will have overlapping parts. Similar parts exhibit the same image features. This invention uses an LF-Net network to extract feature points from two adjacent side-scan sonar images and employs the KNN algorithm to match the feature points extracted by the LF-Net. The specific feature matching method flow is as follows: Figure 1 As shown.

[0094] The side-scan sonar image feature extraction method provided in this embodiment of the invention includes the following steps:

[0095] S1, LF-Net Training: Construct a network architecture based on a dual-branch structure and use dual-modal loss to train the detection network and construct the loss function;

[0096] like Figure 2 As shown, LF-Net constructs a two-branch network architecture designed for efficient training. Both branches share the same set of parameters, and their input contains the image I to be paired. i ,I j The system also incorporates depth and pose information obtained through the structured light algorithm (SfM). Branch i performs differentiable tasks to backpropagate gradients and drive model optimization, while branch j performs non-differentiable operations such as feature point extraction and generating supervision targets for the first branch.

[0097] The loss function of LF-Net consists of two parts: (1) image-level loss and (2) patch-level loss. For keypoint detection networks, the image-level loss is crucial; however, it may indirectly affect the ability to extract image patches during iterative training. Therefore, LF-Net strategically employs dual-modal loss to train the detection network, while describing the network training as limited to the patch-level loss function.

[0098] Image-level loss functions are primarily used to optimize the feature point detection network in LF-Net, ensuring consistency and accuracy of feature points under different viewpoints and geometric transformations. This is achieved by optimizing the image I... j The fractional graph s calculated above j Transform to Figure I i The above is used as the training sample, and ω in the formula represents this rigid transformation. After the ω process, the transformed fractional image S can be obtained. j And extract K feature points from its fractional graph, and perform a fractional graph S j Performing a Gaussian filter with a Gaussian kernel of 0.5 yields a smooth score map. This series of operations is called g. The mathematical expression of the loss function is:

[0099]

[0100] In the formula, S i Let S be the score map of the i-th image, representing the probability of the feature point response; j Let be the score map of the j-th image; g is a non-differentiable Gaussian filter that applies a Gaussian kernel to clean up feature points; ω is a rigid transformation.

[0101] Image block-level loss functions focus on optimizing feature description networks. Their main function is to optimize the local descriptors generated by the network, improving the discriminative ability of the feature descriptors so that they can accurately match between different images. It utilizes the K feature points extracted in each iteration of branch i to optimize the image I at the current iteration number. i ,I j The interrelationships between them are transformed by rigid transformation of the coordinates of K feature points in image I. i Coordinate system down-transformation to image I j In the coordinate system, this process is similar to the transformation process in the image-level loss function, but the transformation direction is exactly opposite. After completing the coordinate transformation, based on the scale and orientation information obtained from branch j, the transformation results in the new coordinate system are fused, and the result is obtained from I... j By precisely selecting an image region, two image blocks can be obtained. and By extracting features from each of them, we can obtain the features. and Here, k represents the number of feature point matching pairs, from which we can obtain the image block-level loss function, mathematically expressed as:

[0102]

[0103] LF-Net obtains the loss function L by forcing the estimated orientation to be the same as the orientation of the transformed set of matching keypoints. geom Its mathematical expression is as follows:

[0104]

[0105] In the formula, and These represent the scale and orientation of the keypoints after orientation transformation, λ. ori and λ scale These are weight parameters.

[0106] Triple Loss Function: To better describe network learning, LF-Net considers non-corresponding image pairs as negative samples for progressive mining during training, introducing the triple loss function L. tri Its mathematical expression is as follows:

[0107]

[0108] Where k′≠k, C=1;

[0109] Based on the three loss functions described above, the loss functions used for training the detection network and describing the network can be derived as follows:

[0110] L det =L im+λ pair L pair +L geom (5)

[0111] L desc =L tri (6)

[0112] S2, Feature map generation: LF-Net uses the ResNet architecture to input the side-scan sonar image into the trained multi-scale fully convolutional network and outputs an image o with feature points;

[0113] like Figure 2 As shown, the side-scan sonar image is input into a multi-scale fully convolutional network, i.e., the detection network mentioned above, which outputs an image o with feature points, i.e., a feature map. LF-Net uses the ResNet architecture to perform this process, which contains three modules, each embedding multiple sets of 5x5 convolutional kernels. After each convolution operation, a batch normalization step is performed. Leaky-ReLU is used as the activation function, and the number of output channels of all convolutional layers is fixed at 16, and their output dimension is consistent with the input, so that the entire feature extraction part always maintains the full image resolution.

[0114] S3, Feature point detection: Feature point detection is performed using LF-Net, and the detection network has a serial structure;

[0115] Feature point detection is performed using LF-Net, which has a serial network structure. The steps are as follows: To ensure scale invariance of feature points, LF-Net performs multi-scale detection based on feature map o. They first scale feature map o N times, with a scaling range of... N=5, Then, a 5x5 convolution kernel is applied to the feature maps that have been scaled N times to obtain N score maps. N (1≤n≤N), each of these N fractional images corresponds to N fractional images at different scales. Applying a 15*15 Softmax operator to each of these N fractional images via convolution to achieve differentiable nonmaximum suppression, we can obtain N fractional images sharpened by the Softmax operator. The Softmax operator normalizes the input values ​​to a probability distribution, ensuring that the sum of the output values ​​is 1. In feature point detection, the Softmax operator enhances the saliency of keypoints, making local maxima more prominent while suppressing other response values. Through the Softmax operator, the network can more clearly distinguish keypoints from other non-keypoint regions, thereby improving the accuracy and robustness of keypoint detection. This ability to enhance saliency allows the network to better handle noise and interference in complex scenes, improving the reliability of keypoint detection. The function definition of the Softmax operator is as follows:

[0116]

[0117] This indicates that for input z i Perform an exponential transformation to ensure the value is non-negative and amplify a larger z. i The impact of the value; This represents the sum of the exponential transformations of all input elements, used for normalization to make the sum of all outputs equal to 1, forming a probability distribution.

[0118] Since the result of nonmaximum suppression is scale-dependent, LF-Net will... Resize the image to the original size; the resized fractional image is as follows. Apply an approximate softmax operator to all fractional graphs of the same scale. The process of fusing the data to form the final scale-space score map is as follows:

[0119]

[0120] In the formula, ⊙ represents the Hadamard product; after scale fusion, the K highest-scoring pixels are selected from the final scale-invariant score map S as feature points and local softargmax is applied to obtain sub-pixel accuracy.

[0121] S4, Feature Description: Calculate the orientation map of each feature point using the arctan function; construct a quadruple and perform feature extraction on the image region to obtain the feature description map;

[0122] The feature description section mainly focuses on determining the orientation. By applying a 5x5 convolution to each pixel on the feature map o, each pixel outputs two values, which can be considered as sine and cosine. The orientation map θ of each feature point is then calculated using the arctan function.

[0123] θ(x,y)=arctan(sin(x,y),cos(x,y)) (9)

[0124] The descriptor extraction method involves analyzing the score map S, selecting K feature points with the highest scores, and obtaining their coordinate data. Subsequently, these keypoints are combined with scale and orientation information derived from the scale estimation map and orientation estimation map to form a four-tuple structure p. k =(x,y,s,θ) k Based on this quadruple, detailed feature extraction operations are performed on the corresponding image regions.

[0125] S5, KNN Corresponding Point Matching: Based on the detected feature points and the obtained feature description map, the KNN algorithm is used to match corresponding points between two adjacent sonar images.

[0126] For the feature points detected by S3 and the feature descriptor map obtained by S4, the KNN algorithm is used to match the corresponding points of two adjacent sonar images. The specific process is as follows: Before matching, the feature descriptor output by LF-Net is normalized. Here we use L2 normalization to ensure that descriptors of different scales and directions are comparable and avoid matching deviations caused by differences in units.

[0127] For the two side-scan sonar images to be matched, extract K feature points and their corresponding descriptors (D). A and D B The descriptor set D in the right figure. B The descriptor set D_A in the left image serves as the query set, acting as the search database. For each descriptor obtained from any sonar image, its K nearest neighbors are found for each descriptor obtained from another sonar image.

[0128]

[0129] In the formula, For feature point A i Coordinates on the X-axis For feature point B j Coordinates on the X-axis; For feature point A i Coordinates on the Y-axis For feature point B j Coordinates on the Y-axis;

[0130] Next, the nearest neighbor is determined by calculating the Euclidean distance between the descriptors. For each query descriptor (D_A), the nearest neighbor is calculated. i ) and all descriptors in the database (D_B j The similarity between the two points is used to determine the matching criteria. The two points with the highest similarity are selected as corresponding points for matching.

[0131] Example 2: The side-scan sonar image feature extraction system provided in this embodiment of the invention includes:

[0132] The LF-Net training module is used to build a network architecture based on two branches, train the detection network using a dual-modal loss function, and describes the training of the network only as a block-level loss function.

[0133] The feature map generation module is used by LF-Net to input side-scan sonar images into the trained multi-scale fully convolutional network using the ResNet architecture, and output an image O with feature points.

[0134] The feature point detection module is used to detect feature points using LF-Net, and the detection network has a serial structure.

[0135] The feature description module is used to calculate the orientation map of each feature point using the arctan function; it constructs a quadruple and performs feature extraction operations on the corresponding image regions to obtain the feature description map.

[0136] The KNN corresponding point matching module is used to perform corresponding point matching between two adjacent sonar images using the KNN algorithm, based on the feature points obtained from the feature point detection module and the feature description map obtained from the feature description module.

[0137] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0138] To further demonstrate the positive effects of the above embodiments, the present invention conducts the following experiments based on the above technical solutions.

[0139] The paper employs two sets of experiments: Data 1 and Data 2, both targeting shipwrecks. Data 1 shows a simpler shipwreck structure, while Data 2 shows a more complex structure and higher noise levels. Experimental results demonstrate that the proposed LF-Net-based side-scan sonar image feature extraction and matching method exhibits significant advantages in high-noise, complex target scenarios (Data 2), validating its end-to-end architecture and transfer learning generalization ability. Traditional algorithms, while applicable in simple scenarios (Data 1), are ill-suited to complex environments.

[0140] 1. Introduction of Experimental Data

[0141] Experimental data such as Figure 3 and Figure 4As shown, the data comes from the SeabedObjects (Ship and Airplane) side-scan sonar dataset publicly released by Hohai University. This dataset includes data from multiple sonar equipment suppliers such as L-3Klein Associates, EdgeTech, Icocean, Hydro-techMarine, and Tritech. The dataset contains 385 ship images and 62 aircraft images. All images are taken directly from raw large-area side-scan sonar images without any processing.

[0142] Data 1 and 2 are both images of the shipwreck acquired by side-scan sonar, specifically as follows: Figure 3 and Figure 4 As shown.

[0143] This experiment used Python and the TensorFlow open-source machine learning platform on PyCharm to build the network. The computer configuration was Windows 10 operating system, 16G DDR5 memory, NVIDIA RTX 1650 graphics card, and i5-12490F CPU.

[0144] This experiment utilizes a pre-trained "Outdoor" model for transfer learning. A pre-trained model matching data from outdoor scenes is then transferred to underwater acoustic images, particularly side-scan sonar images, for feature point detection. In this test, the number of feature points selected, K, for each image is set to 500. This smaller feature point setting improves the algorithm's speed and reduces memory usage.

[0145] 2. Experimental Results and Analysis

[0146] Figure 5 and Figure 6 These are images from data 1 and data 2, respectively, after feature point detection.

[0147] As shown in Table 1, after feature point detection was performed on the two sets of data, orientation and scale descriptions were performed on the two sets of data respectively. Feature description images were obtained using a description network. Figure 7 and Figure 8 Describe the direction of the image, Figure 9 and Figure 10 Describe the image by scale.

[0148] Based on feature point detection in side-scan sonar images by constructing a pre-trained matching model using measured data, the KNN matching algorithm was used to match two sets of side-scan sonar images, such as... Figure 11 and Figure 12 As shown.

[0149] Table 1 Number of LF-Net Matching Pairs

[0150] data Number of matching points Data 1 166 Data 1 183

[0151] like Figure 5 and Figure 6 The results of sonar image feature extraction using LF-Net show that the features extracted by LF-Net are mainly concentrated in areas with significant topographic changes and the edges of ground features, with fewer features in flat areas or sonar shadow areas. In non-feature areas, due to factors such as varying reflectivity of the seabed topography, LF-Net also detected a large number of feature points in the background areas. These points are considered noise during the registration process and will be removed in subsequent matching. Figure 11 and Figure 12 As shown, the KNN algorithm is applied to the generated feature points to match the feature points of two adjacent side-scan sonar images pairwise. It can be seen that the KNN algorithm removes most of the noise, the matching point pairs are relatively dense, and the matching effect is good.

[0152] The experiments compared and analyzed the matching results of the SIFT and ORB algorithms with the proposed algorithm. Figure 13-16 As shown. Similar to the feature point detection process in LF-Net, the number of feature points selected, K, for each image is set to 500. Both algorithms perform RANSAC algorithm removal of mismatched point pairs after matching, with the RANSAC distance threshold set to 8.

[0153] Table 2 shows that using the LF-Net outdoor data pre-trained model for feature point extraction, without relying on additional side-scan sonar image data for model training, improves the matching performance by 23 and 67 pairs respectively on Data 1, and by 66 and 16 pairs respectively on Data 2, compared to the ORB and Harris algorithms. This demonstrates that using the LF-Net outdoor data pre-trained model for automatic feature point extraction from side-scan sonar image data, combined with the KNN matching algorithm, can effectively achieve side-scan sonar image matching tasks.

[0154] Table 2. Number of matching point pairs using other methods

[0155] data ORB Harris Data 1 143 99 Data 2 117 167

[0156] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention and within the spirit and principles of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for feature extraction from side-scan sonar images, characterized in that, The method includes the following steps: S1, LF-Net Training: Construct a network architecture based on a dual-branch structure and use dual-modal loss to train the detection network and construct the loss function; S2, Feature map generation: LF-Net uses the ResNet architecture to input the side-scan sonar image into the trained multi-scale fully convolutional network and outputs an image o with feature points; S3, Feature point detection: Feature point detection is performed using LF-Net, and the detection network has a serial structure; S4, Feature Description: Calculate the orientation map of each feature point using the arctan function; construct a quadruple and perform feature extraction on the image region to obtain the feature description map; S5, KNN Corresponding Point Matching: Based on the detected feature points and the obtained feature description map, the KNN algorithm is used to match corresponding points between two adjacent sonar images.

2. The side-scan sonar image feature extraction method according to claim 1, characterized in that, In step S1, in constructing the dual-branch-based network architecture, the two branches share the same set of parameters, and the input contains the image I to be paired. i ,I j In addition, depth and pose information are obtained through the structured light algorithm; branch i undertakes differentiable tasks to backpropagate gradients and drive model optimization; Branch j performs non-differentiable operations, including feature point extraction and generating a supervised target for the first branch.

3. The side-scan sonar image feature extraction method according to claim 1, characterized in that, In step S1, the loss function of LF-Net is divided into an image-level loss function and an image patch-level loss function; Image-level loss functions are used to optimize the feature point detection network in LF-Net by analyzing image I. j The fractional graph S calculated above j Transform to image I i Using the above as training samples, after a rigid transformation ω process, the fractional image S is obtained. j And extract K feature points from the fractional image, and perform a test on the fractional image S. j A Gaussian filter with a kernel of 0.5 is applied to obtain a smoothed score image. The expression for the loss function is as follows: L im (S i ,S j )=|S i -g(ω(S j ))| 2 In the formula, S i S is the score image of the i-th image. j Let g be the score map of the j-th image, g be a non-differentiable Gaussian filter, and ω be a rigid transform. Image block-level loss functions are used to optimize local descriptors generated by the feature description network, including matching feature point loss L. pair Geometric consistency loss L geom Triple loss L tri ; Matching feature point loss L pair Minimize the distance between the feature vectors of the same keypoint in two images to ensure that the descriptors remain consistent after transformation; by extracting K feature points from each iteration of branch i, and using the current iteration number as the image I... i ,I j The interrelationships between them are transformed by rigid transformation of the coordinates of K feature points in image I. i Coordinate system down-transformation to image I j In the coordinate system; after completing the coordinate transformation, based on the scale and orientation information obtained from branch j, the transformation results in the new coordinate system are fused, from I j By precisely selecting the image region, two image blocks are obtained. and Feature extraction is performed on each image patch, and the extracted features are... and k represents the number of feature point matching pairs. This is the feature descriptor for the k-th keypoint in the i-th image. Given the feature descriptor corresponding to k after geometric transformation of the j-th image, we obtain the image block-level loss function; LF-Net uses geometric consistency loss L geom This ensures that the scale and orientation of feature points remain consistent across different images, guaranteeing that LF-Net possesses scale invariance and rotation invariance; L geom It consists of two parts: scale loss and orientation loss. This ensures that the direction of the matching key points remains consistent. This ensures that the scale of the matching key points remains consistent. This represents the scale of the k-th keypoint in the i-th image. Indicates the direction of the k-th key point in the i-th image; This represents the scale of the corresponding keypoint in the j-th image after geometric transformation. λ represents the orientation of the corresponding keypoint in the j-th image after geometric transformation. ori and λ scale The hyperparameters Li control the effects of orientation and scale loss, respectively. geom As shown below; LF-Net uses a triplet loss function to optimize feature descriptors. For non-corresponding image pairs, it uses them as negative samples for progressive mining training, making the distance between matching positive samples closer and the distance between non-matching negative samples farther. The margin C hyperparameter sets a minimum interval to ensure that the distance between matching points and non-matching points is large enough. This represents the feature descriptor of the k-th keypoint in the i-th image. Let represent the positive samples of matching keypoints after geometric transformation of the j-th image. Let L represent a random negative sample in the j-th image, and C represent the Margin hyperparameter in the triplet loss, which controls the separation between positive and negative samples; L represents the triplet loss. tri The mathematical expression for is shown below; Where k'≠k, C=1; After applying the three loss functions, the loss functions used for training the detection network and describing the network are obtained as follows: L det =L im +λ pair L pair +L geom L desc L tri 。 4. The side-scan sonar image feature extraction method according to claim 1, characterized in that, In step S2, the ResNet architecture comprises three modules: a convolutional module, a residual module, and a fully connected module. Each module embeds multiple sets of 5×5 convolutional kernels. Batch normalization is performed after each convolutional operation. Leaky-ReLU is used as the activation function, and the number of output channels in all convolutional layers is fixed at 16, with the output dimension consistent with the input. The convolutional module performs preliminary feature extraction, reducing the spatial size of the input data and improving computational efficiency. The residual module addresses the vanishing gradient problem in deep neural networks, enabling efficient information flow between layers and facilitating residual learning through identity mapping. The fully connected module reduces the number of parameters through global average pooling, improving generalization ability, and performs the final classification decision.

5. The side-scan sonar image feature extraction method according to claim 1, characterized in that, In step S3, feature point detection includes: LF-Net performs multi-scale detection based on feature map o; the feature map o is scaled N times, with a scaling range of... N=5, Perform a convolution operation on the feature maps after N scaling operations using a 5×5 convolution kernel to obtain N fractional maps h. N (1≤n≤N), each of the N fractional images corresponds to N fractional images at different scales. Differentiable nonmaximum suppression is achieved by performing convolution operations on each of the N fractional images using a 15×15 Softmax operator, resulting in N fractional images sharpened by the Softmax operator. LF-Net will each Resize the image to the original size; the resized fractional image is as follows. Apply an approximate softmax operator to all fractional graphs of the same scale. The data is then merged to form the final score map of the scale space; In the formula, ⊙ represents the Hadamard product; after scale fusion, the K pixels with the highest scores are selected from the final scale-invariant score map S as feature points and local softargmax is applied to obtain sub-pixel accuracy.

6. The side-scan sonar image feature extraction method according to claim 5, characterized in that, The Softmax operator is used to normalize the input values ​​into a probability distribution such that the sum of the output values ​​is 1. The softmax operator is defined as follows, where softmax(z) i ) represents the value of the i-th element after the Softmax transformation, and represents a probability value in the probability distribution; In the formula, For input z i Perform an exponential transformation to ensure the value is non-negative and amplify a larger z. i The impact of the value; The sum of the exponential transformations of all input elements is used for normalization, so that the sum of all outputs is 1, forming a probability distribution.

7. The side-scan sonar image feature extraction method according to claim 1, characterized in that, In step S4, the feature description includes: applying a 5×5 convolution to each pixel on the feature map o, outputting two values ​​as sine and cosine for each pixel, and calculating the orientation map θ of each feature point using the arctan function; Descriptor extraction: By analyzing the score map S, K feature points with the highest scores are selected, and their coordinate data is obtained; the key points are combined with scale and orientation information from the scale estimation map and orientation estimation map to form a four-tuple structure p. k =(x,y,s,θ) k And perform feature extraction operations on the corresponding image regions.

8. The side-scan sonar image feature extraction method according to claim 7, characterized in that, The orientation map θ of each feature point is calculated using the arctan function. The direction calculated by arctan is a continuously changing angle value, without abrupt changes in direction, providing a smooth direction estimate. This is applicable to various scales, rotations, and perspective changes, improving the robustness of feature points. Furthermore, LF-Net needs to find the same keypoints in images with different viewpoints and rotational changes. The direction angle calculated by arctan ensures that feature points still point in the same direction after rotational transformations, improving rotation invariance. The calculation formula is as follows: In the formula, θ is the direction angle of the key point, and f x f is the gradient of the feature point in the x-direction. y Let be the gradient of the feature point in the y-direction.

9. The side-scan sonar image feature extraction method according to claim 1, characterized in that, In step S5, KNN corresponding point matching includes: matching corresponding points between two adjacent sonar images using the KNN algorithm based on the detected feature points and the obtained feature description map; Before matching, the feature descriptors output by LF-Net are processed using L2 normalization. For two side-scan sonar images to be matched, the KNN algorithm is used to match feature points. Euclidean distance is calculated for feature point matching. The KNN algorithm is used for image registration, target detection, and clustering tasks. In side-scan sonar image matching, this method accurately calculates the similarity of feature points in adjacent images, improving the accuracy of stitching and classification. K feature points and their corresponding descriptors D_A and D_B are extracted respectively. The descriptor set D_B is used as the search database, and the descriptor set D_A is used as the query set. For each descriptor obtained from any one sonar image, the K nearest neighbor descriptors are found for each descriptor obtained from the other sonar image. In the formula, For feature point A i Coordinates on the X-axis For feature point B j Coordinates on the X-axis; For feature point A i Coordinates on the Y-axis For feature point B j Coordinates on the Y-axis; The nearest neighbor is determined by calculating the Euclidean distance between descriptors, and the D_A is calculated for each query descriptor. i With all descriptors D_B in the database j The similarity is used to select the two points with the highest similarity as corresponding points for matching.

10. A side-scan sonar image feature extraction system, characterized in that, This system is used to control the side-scan sonar image feature extraction method according to any one of claims 1-9, and the system includes: The LF-Net training module is used to build a dual-branch-based network architecture and train the detection network using dual-modal loss. The feature map generation module is used by LF-Net to input side-scan sonar images into the trained multi-scale fully convolutional network using the ResNet architecture, and output an image o with feature points; The feature point detection module is used to detect feature points using LF-Net, and the detection network has a serial structure. The feature description module is used to calculate the orientation map of each feature point using the arctan function; it constructs a quadruple and performs feature extraction on the image region to obtain the feature description map. The KNN corresponding point matching module is used to match corresponding points between two adjacent sonar images using the KNN algorithm, based on the detected feature points and the obtained feature description map.

Citation Information

Patent Citations

  • Substation equipment defect image matching method

    CN111401384A

  • Optical character detection and recognition system based on deep learning

    CN118119983A