A sonar fish identification method based on self-supervised learning

By improving the feature decoupling and walrus optimization algorithm of the DINOv2 visual Transformer, the problems of weak echoes and feature aliasing of small targets in sonar fish recognition are solved, achieving high-precision recognition and multi-task robustness in complex waters.

CN120411756BActive Publication Date: 2025-09-12SHENYANG LIAOHAI EQUIP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510925781.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-12
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Existing sonar fish identification methods face the problems of weak echo signals of small fish targets being drowned out by noise and feature aliasing in complex waters, and lack adaptive adjustment capabilities, resulting in a decrease in accuracy when the model is applied in remote areas.

Method used

A self-supervised learning-based method is used to improve the DINOv2 visual Transformer. The sonar features are decomposed into scale sub-features, texture sub-features, and direction sub-features through the feature decoupling matrix module. Cross-scale consistency and sparsity constraints are imposed, and the walrus optimization algorithm is combined with parameter adaptive optimization to achieve multi-objective comprehensive fitness-driven feature fusion.

Benefits of technology

It significantly improved the detection rate and recognition accuracy of small fish targets, enhanced the model's generalization ability and deployment flexibility in complex waters, and improved the multi-task robustness and overall performance of fish recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411756B_ABST
    Figure CN120411756B_ABST
Patent Text Reader

Abstract

The present invention discloses a sonar fish recognition method based on self-supervised learning. The method comprises the following steps: obtaining preprocessed sonar echo image data; generating a multi-scale sonar image patch set; inputting the multi-scale sonar image patch set into an improved DINOv2 visual Transformer encoder to obtain an initial multi-scale sonar feature representation; obtaining decoupled scale sub-features, texture sub-features, and directional sub-features; outputting a multi-scale sonar decoupled feature representation that satisfies constraints; initializing a population of walrus optimization algorithm parameter triples and calculating a multi-objective comprehensive fitness value; updating the walrus optimization algorithm parameter triples and recalculating the multi-objective comprehensive fitness value until convergence conditions are met; and outputting a final fish recognition result, a fish length estimation result, and a fish school density result through a classifier when the convergence conditions are met. The improved DINOv2 visual Transformer of the present invention significantly improves the detection rate of small fish targets in complex waters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fish identification, and in particular to a sonar fish identification method based on self-supervised learning. Background Art

[0002] With the increasing demand for underwater fishery resource monitoring and intelligent environmental perception, sonar imaging systems have become an important means widely used in fish distribution surveys, ecological assessments, and aquaculture supervision. Existing technologies generally use sonar equipment to obtain underwater echo signals and use traditional image processing or convolutional neural network methods to identify fish targets. However, due to the limitations of the spatial resolution of sonar echoes, environmental background noise, and target scale differences, they face significant challenges in actual complex waters.

[0003] Existing multi-scale sonar image recognition methods generally rely on convolution-pooling structures. The shared parameters of these convolution kernels across objects of different sizes can easily cause weak echo signals from small fish targets to be smeared by noise, leading to missed and false detections. Furthermore, traditional algorithms lack the ability to adapt to changes in the distribution scale of fish schools and are unable to effectively generalize echo texture features across different devices and water environments, resulting in a sharp drop in accuracy when the models are relocated and applied remotely. Summary of the Invention

[0004] One purpose of the present invention is to propose a sonar fish recognition method based on self-supervised learning. The present invention improves the DINOv2 visual Transformer to significantly improve the detection rate of small fish targets in complex waters.

[0005] A sonar fish identification method based on self-supervised learning according to an embodiment of the present invention includes:

[0006] collecting sonar echo image data and performing preprocessing to obtain preprocessed sonar echo image data;

[0007] Perform pyramid multi-scale slicing on the pre-processed sonar echo image data to generate a multi-scale sonar image patch set;

[0008] The multi-scale sonar image patch set is fed into the improved DINOv2 visual Transformer encoder to obtain the initial multi-scale sonar feature representation.

[0009] A feature decoupling matrix module is inserted into the improved DINOv2 visual Transformer encoder to perform dimension division on the initial multi-scale sonar feature representation to obtain decoupled scale sub-features, texture sub-features, and direction sub-features.

[0010] Apply cross-scale consistency constraints and same-scale sparsity constraints to the decoupled scale sub-features, texture sub-features, and direction sub-features, and output a multi-scale sonar decoupled feature representation that meets the constraints.

[0011] Initialize the walrus optimization algorithm parameter triple population, obtain the candidate multi-scale sonar feature fusion results and input them into the classifier, output fish category, fish body length estimation and fish density, and calculate the multi-objective comprehensive fitness value;

[0012] The multi-objective comprehensive fitness value is used as the fitness function of the walrus optimization algorithm to iteratively adjust the walrus optimization algorithm parameter triple population, update the walrus optimization algorithm parameter triple and recalculate the multi-objective comprehensive fitness value until the convergence condition is met;

[0013] When the convergence conditions are met, the classifier outputs the final fish identification results, fish body length estimation results and fish school density results, and generates a confidence heat map at the same time.

[0014] Optionally, the collecting and preprocessing of sonar echo image data includes:

[0015] Collect the original sonar echo signals of the water area to be monitored through the side-scan sonar array or the forward-looking sonar array;

[0016] The original sonar echo signal sequence is band-pass filtered to eliminate non-working frequency band noise, and the reflection intensity envelope signal is extracted through envelope detection to construct the envelope intensity mapping matrix;

[0017] The envelope intensity mapping matrix is ​​beamformed and the envelope intensity information of multiple beam directions is reconstructed into a two-dimensional intensity image using a delayed weighted superposition method.

[0018] The linear normalization method is used to uniformly map the intensity value of each pixel in the two-dimensional intensity image after beamforming to a grayscale value range of 0 to 255, and the grayscale-processed two-dimensional image is output as the preprocessed sonar echo image data.

[0019] Optionally, performing pyramid multi-scale slicing includes:

[0020] An image pyramid sequence is constructed by downsampling. The pre-processed sonar echo image data is layered according to the scale scaling coefficient set to construct three scale levels: large-scale image, medium-scale image and micro-scale image.

[0021] Perform sliding window slicing operations on each scale level of the image pyramid sequence to obtain multiple sonar image patches;

[0022] Constructing a sonar image patch scale identification vector for each sonar image patch;

[0023] The sonar image patch sets at all scale levels are merged to form a multi-scale sonar image patch set, and the scale identification vectors of all sonar image patches are simultaneously merged to form a scale identification set corresponding to the multi-scale sonar image patches.

[0024] Optionally, the construction of the improved DINOv2 visual Transformer encoder includes:

[0025] The multi-scale sonar image patch set is fed into the DINOv2 visual Transformer encoder modified with multi-scale perceptual attention.

[0026] The multi-scale sonar image patch set contains multiple sonar image patches, and each sonar image patch is marked according to its index in the set;

[0027] Each sonar image patch is flattened and linearly transformed to obtain an initial feature vector of the sonar image patch. A scale embedding vector and a spatial position embedding vector are added to the initial feature vector of the sonar image patch according to the scale identification information of the sonar image patch, thereby forming a sonar image patch feature vector that contains both scale characteristics and spatial position information.

[0028] Compute the multi-scale perceptual attention matrix for any two sonar image patch feature vectors containing scale embeddings and spatial position embeddings;

[0029] Perform weighted fusion of the value vector based on the multi-scale perception attention matrix to generate a multi-scale attention feature vector after fusion of scale information;

[0030] A multi-scale residual connection structure for multi-scale sonar fish recognition tasks is introduced into the multi-scale attention feature vector, and the feature enhanced output vector is obtained through feedforward neural network mapping;

[0031] All feature-enhanced output vectors constitute the initial multi-scale sonar feature representation set.

[0032] Optionally, the insertion of the feature decoupling matrix module includes:

[0033] A feature decoupling matrix module is inserted into the improved DINOv2 visual Transformer encoder. The input of the feature decoupling matrix module is the initial multi-scale sonar feature representation set.

[0034] Through the feature decoupling matrix module, scale feature extraction, texture feature extraction and direction feature extraction operations are performed on each feature enhancement output vector to obtain scale sub-feature vector, texture sub-feature vector and direction sub-feature vector;

[0035] All scale sub-feature vectors, texture sub-feature vectors and direction sub-feature vectors are set in index order to form a scale sub-feature set, a texture sub-feature set and a direction sub-feature set respectively.

[0036] Optionally, the output of the multi-scale sonar decoupling feature representation that satisfies the constraint conditions includes:

[0037] The scale sub-feature vectors of each sonar image patch form a scale sub-feature vector sequence, the texture sub-feature vectors of each sonar image patch form a texture sub-feature vector sequence, and the direction sub-feature vectors of each sonar image patch form a direction sub-feature vector sequence;

[0038] A cross-scale consistency constraint is imposed on the scale sub-feature vector sequence. The cross-scale consistency constraint is used to measure the semantic consistency of fish target representation of sonar image patches at different scale levels, and a cross-scale loss is constructed to measure the semantic consistency of fish target representation at different scales. :

[0039] ;

[0040] in, represents the total number of sonar image patches, Indicates the The scaled sub-feature vector of a sonar image patch, Indicates the The scaled sub-feature vector of a sonar image patch, Indicates the The scale level identification of each sonar image patch, Indicates the The scale level identification of each sonar image patch, represents the cross-scale mask function, if and If the sonar image patches are from different scale levels, the value is 1, otherwise it is 0. Indicates the and The square of the Euclidean distance between the sub-feature vectors of the scale;

[0041] The texture sub-feature vector sequence is constrained to be of the same scale. The texture sparsity loss is obtained by calculating the ratio of the L1 norm to the L2 norm of the texture sub-feature vector of each sonar image patch and summing the ratios of all sonar image patches at the same scale level.

[0042] A local structure constraint is imposed on the sequence of directional sub-eigenvectors. The Euclidean distance between the directional sub-eigenvectors of each sonar image patch and all its spatially adjacent sonar image patches is calculated, and the directional consistency loss is obtained by accumulating the Euclidean distances of all adjacent patch pairs.

[0043] The scale consistency loss, texture sparsity loss and direction consistency loss are weightedly summed to obtain the overall multi-scale decoupled feature constraint loss;

[0044] Based on the overall multi-scale decoupled feature constraint loss, the scale sub-feature set, texture sub-feature set and direction sub-feature set are reversely optimized and updated to output a multi-scale sonar decoupled feature representation that meets the constraint conditions.

[0045] Optionally, the calculation of the multi-objective comprehensive fitness value includes:

[0046] Initialize a population of walrus optimization algorithm parameter triplets based on the multi-scale sonar fish recognition task characteristics. The population contains several walrus optimization algorithm parameter triplets. Each walrus optimization algorithm parameter triplet contains three parameters: scale sub-feature weight vector, classification threshold, and posterior fusion coefficient.

[0047] Based on the scale perception mechanism, the scale sub-feature weight vector of each walrus optimization algorithm parameter triplet is scale-adaptively perturbed;

[0048] The adjusted scale sub-feature weight vector is used to perform weighted fusion on the multi-scale sonar decoupling feature representation and calculate the candidate multi-scale sonar feature fusion result;

[0049] A classifier with an adaptive fusion decision-making mechanism is constructed. After the candidate multi-scale sonar feature fusion results are input into the classifier, the posterior fusion coefficient is used to weight the fish category probability distribution output by the classifier, and the fusion-enhanced fish category probability is output.

[0050] The fish category label is determined based on the fusion-enhanced fish category probability, and the classifier estimates the fish body length and fish school density, obtaining the output results of the fish category label, fish body length estimation value, and fish school density estimation value;

[0051] For each walrus optimization algorithm parameter triplet, a multi-objective comprehensive fitness function is used to evaluate it and obtain the multi-objective comprehensive fitness value.

[0052] Optionally, the updating of the multi-objective comprehensive fitness value includes:

[0053] The multi-objective comprehensive fitness value is used as the fitness function of the walrus optimization algorithm. The multi-objective comprehensive fitness values ​​corresponding to each walrus optimization algorithm parameter triple are arranged in sequence to form a fitness sequence. The walrus optimization algorithm parameter triple with the largest multi-objective comprehensive fitness value is selected as the leading walrus parameter triple.

[0054] The parameter triple population is iteratively updated based on the leader-cooperation-follower mechanism of the walrus optimization algorithm;

[0055] In the leader guidance phase, all walrus optimization algorithm parameter triplets are guided and updated according to the difference between the current parameter state and the leader walrus parameter triplet;

[0056] In the collaborative perturbation phase, each walrus optimization algorithm parameter triplet is locally and interactively updated with any other parameter triplet in the same population;

[0057] In the follow-up correction stage, the scale sub-feature weight vector is fine-grainedly corrected according to the energy change trend of the candidate multi-scale sonar feature fusion results at each scale level;

[0058] After each round of parameter update, the multi-scale sonar features are re-fused based on the updated walrus optimization algorithm parameter triples and input into the adaptive fusion classifier to recover the fish category labels, fish body length estimates, and fish school density estimates. The multi-objective comprehensive fitness value is then recalculated, completing the real-time iterative correction of the fitness values ​​of all parameter triplets.

[0059] Determine whether the current iteration meets the convergence conditions. The convergence conditions include any of the following:

[0060] The global fitness value improvement rate is lower than the preset threshold, or the population parameter change amplitude is lower than the minimum disturbance threshold, or the preset maximum number of iterations is reached;

[0061] If any convergence condition is met, the parameter update iteration ends; otherwise, the process returns to continue the optimization of population parameters and fitness calculation.

[0062] Optionally, when the convergence condition is met, the walrus optimization algorithm parameter triplet with the highest multi-objective comprehensive fitness value is selected as the optimal walrus optimization algorithm parameter triplet, and the multi-scale sonar decoupling feature representation is weightedly fused using the optimal walrus optimization algorithm parameter triplet to obtain the final multi-scale sonar feature fusion result, and the final fish recognition result, fish body length estimation result and fish school density result are output through the classifier, and a confidence heat map is generated at the same time.

[0063] The beneficial effects of the present invention are:

[0064] (1) The present invention is based on a sonar feature encoding mechanism based on feature decoupling and multi-scale consistency, which effectively alleviates the problem of weak echoes of small fish being submerged and feature aliasing, and improves the sensitivity and accuracy of fish target detection. A feature decoupling matrix module is introduced into the DINOv2 visual Transformer backbone to achieve explicit splitting of the initial multi-scale sonar features, decomposing the feature vector of each sonar image patch into scale sub-features, texture sub-features and direction sub-features. Through joint training of multi-scale consistency and sparsity constraints, the features at different scale levels are kept aligned across scales and sparse at the same scale, effectively suppressing the feature coupling problem caused by high noise and strong background interference. The improved DINOv2 visual Transformer significantly improves the detection rate of small fish targets in complex waters.

[0065] (2) The present invention combines scale-adaptive perturbation with the walrus optimization algorithm of posterior fusion decision to realize unsupervised adaptive optimization of fish recognition weight parameters and thresholds in multiple scenarios and multiple devices, thereby improving the system's generalization capability and deployment flexibility. By introducing scale-aware mechanism and energy-sensitive adaptive perturbation into the walrus optimization algorithm parameter triplet, the weights of sonar features of different scales can be dynamically adjusted to adapt to the target fish population distribution and the echo characteristics of different sonar equipment and different water environments. At the same time, the adaptive decision-making mechanism that integrates posterior probability and prior probability can dynamically adjust the classification threshold and weight parameters according to the fish distribution in the target environment.

[0066] (3) The present invention proposes a full-process evolutionary search strategy driven by multi-objective comprehensive fitness, which realizes end-to-end self-optimization of feature fusion weights, classification thresholds and fusion coefficients, and improves the overall performance of fish species discrimination, body length estimation and group density assessment. The multi-objective comprehensive fitness function is used to incorporate multiple indicators such as fish category recognition accuracy, estimation error of fish body length and fish density, and scale sensitivity into a unified evaluation system. The leader-cooperation-follower mechanism of the walrus optimization algorithm is combined to perform global and local adaptive evolution of the parameter triple population. This not only improves the multi-task robustness of fish recognition, but also can dynamically optimize the weight distribution of different tasks according to actual application requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0068] Figure 1 This is a flow chart of a sonar fish identification method based on self-supervised learning proposed by the present invention;

[0069] Figure 2This is a schematic diagram of the structure of the DINOv2 visual Transformer encoder with improved multi-scale perception attention in the sonar fish recognition method based on self-supervised learning proposed in the present invention. DETAILED DESCRIPTION

[0070] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0071] refer to Figure 1-Figure 2 , a sonar fish identification method based on self-supervised learning, comprising the following steps:

[0072] collecting sonar echo image data and performing preprocessing to obtain preprocessed sonar echo image data;

[0073] In this embodiment, sonar echo image data is collected and preprocessed, including:

[0074] Collect the original sonar echo signals of the water area to be monitored through the side-scan sonar array or the forward-looking sonar array;

[0075] The sonar echo signal includes multiple time sampling points and multiple beam directions. The sonar echo signal is sampled in the time domain using the sonar system's center frequency, beam width, and sampling period to obtain the original sonar echo signal sequence. Each signal in the original sonar echo signal sequence is characterized by an amplitude envelope function and an instantaneous phase. The amplitude envelope function is used to reflect the echo energy intensity of the beam direction at the corresponding moment, and the instantaneous phase is used to describe the instantaneous phase characteristics of the signal.

[0076] The original sonar echo signal sequence is band-pass filtered to eliminate non-working frequency band noise, and the reflection intensity envelope signal is extracted through envelope detection to construct the envelope intensity mapping matrix;

[0077] Each element of the envelope intensity mapping matrix is ​​used to reflect the acoustic reflection intensity distribution of fish or other objects at the corresponding time sampling point and beam direction. The envelope intensity mapping matrix is ​​calculated by the absolute value of the amplitude envelope function.

[0078] The envelope intensity mapping matrix is ​​beamformed and the envelope intensity information of multiple beam directions is reconstructed into a two-dimensional intensity image region using a delayed weighted superposition method.

[0079] The intensity value of each image point in the beamforming process is equal to the result of weighted superposition of the envelope intensity values ​​of multiple beam directions with corresponding weights. The envelope intensity value is used to reflect the contribution of each beam direction in the sonar reconstruction process with corresponding weights.

[0080] The intensity value of each pixel in the beamformed two-dimensional intensity image is uniformly mapped to a grayscale value range of 0 to 255 using a linear normalization method, and the grayscale-processed two-dimensional image is output as the preprocessed sonar echo image data;

[0081] The lower limit of the grayscale value is equal to the minimum intensity value in the intensity image, and the upper limit of the grayscale value is equal to the maximum intensity value in the intensity image.

[0082] Perform pyramid multi-scale slicing on the pre-processed sonar echo image data to generate a multi-scale sonar image patch set;

[0083] In this embodiment, pyramid-type multi-scale slicing is performed, including:

[0084] An image pyramid sequence is constructed by downsampling. The pre-processed sonar echo image data is layered according to the scale scaling coefficient set to construct three scale levels: large-scale image, medium-scale image and micro-scale image.

[0085] The scale scaling coefficient set includes three scale levels: original resolution, half resolution, and quarter resolution, which are used to construct three scale levels: large-scale image, medium-scale image, and micro-scale image, respectively.

[0086] Perform sliding window slicing operations on each scale level of the image pyramid sequence to obtain multiple sonar image patches;

[0087] The sliding window operation sets a fixed window size and window step size. Both the window size and window step size are set according to the scale level to which they belong. The sliding window moves stepwise across the image, cropping multiple sonar image patches. Each sonar image patch originates from an image at the corresponding scale level in the image pyramid sequence. Each sonar image patch has clear source scale information and constitutes a set of sonar image patches at the scale level.

[0088] Constructing a sonar image patch scale identification vector for each sonar image patch;

[0089] The sonar image patch scale identification vector consists of three elements: the scale scaling factor of the patch, the central horizontal coordinate and the central vertical coordinate of the patch on the image layer to which it belongs. The sonar image patch scale identification vector is used to record the scale level and image space position of each sonar image patch.

[0090] Merge the sonar image patch sets at all scale levels to form a multi-scale sonar image patch set, and simultaneously merge the scale identification vectors of all sonar image patches to form a scale identification set corresponding to the multi-scale sonar image patches;

[0091] The multi-scale sonar image patch set is used to represent local image areas at different scales, and the scale identifier set is used to indicate the scale level and spatial position of each patch.

[0092] The multi-scale sonar image patch set is fed into the improved DINOv2 visual Transformer encoder to obtain the initial multi-scale sonar feature representation.

[0093] In this implementation, the construction of the DINOv2 visual Transformer encoder is improved, including:

[0094] The multi-scale sonar image patch set is fed into the DINOv2 visual Transformer encoder modified with multi-scale perceptual attention.

[0095] The multi-scale sonar image patch set contains multiple sonar image patches, and each sonar image patch is marked according to its index in the set.

[0096] Each sonar image patch is flattened and linearly transformed to obtain an initial feature vector of the sonar image patch. A scale embedding vector and a spatial position embedding vector are added to the initial feature vector of the sonar image patch according to the scale identification information of the sonar image patch, thereby forming a sonar image patch feature vector that contains both scale characteristics and spatial position information.

[0097] The scale embedding vector is used to explicitly encode the scale-level characteristics of the sonar image patch, and the spatial position embedding vector is used to represent the position information of the sonar image patch in the original image space.

[0098] Compute the multi-scale perceptual attention matrix for any two sonar image patch feature vectors containing scale embeddings and spatial position embeddings;

[0099] ;

[0100] in, is the query vector calculated from the sonar image patch feature vector, is the key vector calculated from the sonar image patch feature vector, are the dimensions of the query vector and key vector, is the scale embedding vector, representing the The scale level to which a sonar image patch belongs is used to explicitly highlight the scale properties of the patch during the feature enhancement stage;

[0101] Multi-scale similarity weights are used to measure the correlation between sonar image patches at different scales. Adjusted by scale embedding information, these weights can distinguish and highlight cross-scale feature correlations, thereby improving the DINOv2 visual Transformer encoder's ability to perceive and identify the intrinsic relationships between sonar image patches at different scales.

[0102] Perform weighted fusion of the value vector based on the multi-scale perception attention matrix to generate a multi-scale attention feature vector after fusion of scale information;

[0103] The value vector is obtained by linear mapping the feature vector of the corresponding sonar image patch. The multi-scale attention feature vector is used to fully capture the cross-scale feature relationship between sonar image patches of different scales.

[0104] The multi-scale residual connection structure for multi-scale sonar fish recognition tasks is introduced into the multi-scale attention feature vector, and the feature enhancement output vector is defined for:

[0105] ;

[0106] in, is the multi-scale attention feature vector, is the scale feature enhancement factor, which is used to adjust the sensitivity balance between microscale, mesoscale and macroscale features. is the feedforward neural network mapping, is the sonar image patch feature vector;

[0107] All feature-enhanced output vectors Construct the initial multi-scale sonar feature representation set.

[0108] In this implementation, a multi-scale perception attention mechanism and scale-space joint embedding are introduced into the DINOv2 visual Transformer backbone, which can explicitly enhance the model's perception of feature dependencies and cross-scale information between targets of different scales in sonar images, and significantly improve the model's detection sensitivity for weak echoes of micro-scale and small-sized fish. By introducing scale embedding in the attention weight calculation, the improved DINOv2 visual Transformer encoder enhances its ability to express key target areas when processing multi-scale fish schools and complex backgrounds, while suppressing bottom reflection interference, achieving efficient cross-scale information fusion. Compared with traditional Transformers, it can effectively improve the recall rate of small target recognition and the discrimination of fish feature expression.

[0109] A feature decoupling matrix module is inserted into the improved DINOv2 visual Transformer encoder to divide the initial multi-scale sonar feature representation into dimensions according to scale sub-features, texture sub-features, and direction sub-features, thereby obtaining decoupled scale sub-features, texture sub-features, and direction sub-features.

[0110] In this embodiment, the insertion of the feature decoupling matrix module includes:

[0111] A feature decoupling matrix module is inserted into the improved DINOv2 visual Transformer encoder. The input of the feature decoupling matrix module is the initial multi-scale sonar feature representation set.

[0112] The initial multi-scale sonar feature representation set consists of multiple feature enhanced output vectors.

[0113] Through the feature decoupling matrix module, scale feature extraction, texture feature extraction and direction feature extraction operations are performed on each feature enhancement output vector to obtain scale sub-feature vector, texture sub-feature vector and direction sub-feature vector;

[0114] The scale sub-feature vector is obtained by applying the scale feature mapping operation to the feature enhancement output vector. The scale feature mapping operation is used to explicitly extract the feature components related to the scale variation of the sonar image patch from the feature enhancement output vector, thereby enhancing the model's sensitivity to micro-scale, meso-scale, and macro-scale fish targets.

[0115] The texture sub-feature vector is the result of applying a texture feature mapping operation to the feature enhancement output vector. The texture feature mapping operation is used to explicitly extract the texture feature components of the sonar echo signal on different fish targets, thereby enhancing the ability to identify the texture details of fish targets.

[0116] The directional sub-eigenvector is the result of applying the directional feature mapping operation to the feature-enhanced output vector. The directional feature mapping operation is used to explicitly extract the characteristic components related to the target orientation and spatial distribution in the sonar echo signal, thereby improving the model's ability to distinguish the directionality of fish targets.

[0117] Performing set operations on all scale sub-feature vectors, texture sub-feature vectors, and direction sub-feature vectors in index order to form a scale sub-feature set, a texture sub-feature set, and a direction sub-feature set, respectively;

[0118] The scale sub-feature set is used to specifically characterize the scale variation law of sonar image patches of different scales. The texture sub-feature set is used to specifically characterize the diversity and difference of the echo texture of fish targets. The direction sub-feature set is used to specifically characterize the spatial distribution and direction information of fish targets.

[0119] In this embodiment, the sonar feature encoding mechanism based on feature decoupling and multi-scale consistency effectively alleviates the problems of weak echoes of small fish being submerged and feature aliasing, and improves the sensitivity and accuracy of fish target detection. The feature decoupling matrix module is introduced into the DINOv2 visual Transformer backbone to realize the explicit splitting of the initial multi-scale sonar features, and decompose the feature vector of each sonar image patch into scale sub-features, texture sub-features and direction sub-features. Through joint training of multi-scale consistency and sparsity constraints, the features at different scale levels are kept aligned across scales and sparse at the same scale, effectively suppressing the feature coupling problem caused by high noise and strong background interference. The improved DINOv2 visual Transformer significantly improves the detection rate of small fish targets in complex waters.

[0120] Apply cross-scale consistency constraints and same-scale sparsity constraints to the decoupled scale sub-features, texture sub-features, and direction sub-features, and output a multi-scale sonar decoupled feature representation that meets the constraints.

[0121] In this embodiment, the output of the multi-scale sonar decoupling feature representation that satisfies the constraints includes:

[0122] The scale sub-feature vectors of each sonar image patch form a scale sub-feature vector sequence, the texture sub-feature vectors of each sonar image patch form a texture sub-feature vector sequence, and the direction sub-feature vectors of each sonar image patch form a direction sub-feature vector sequence;

[0123] The scale sub-feature vector, texture sub-feature vector and orientation sub-feature vector of each sonar image patch respectively represent the scale attribute, texture attribute and orientation attribute of the patch in the multi-scale sonar image.

[0124] A cross-scale consistency constraint is imposed on the scale sub-feature vector sequence. The cross-scale consistency constraint is used to measure the semantic consistency of fish target representation of sonar image patches at different scale levels, and a cross-scale loss function is constructed to measure the semantic consistency of fish target representation at different scales. :

[0125] ;

[0126] in, represents the total number of sonar image patches, Indicates the The scaled sub-feature vector of a sonar image patch, Indicates the The scaled sub-feature vector of a sonar image patch, Indicates the The scale level identification of each sonar image patch, Indicates the The scale level identification of each sonar image patch, represents the cross-scale mask function, if and If the sonar image patches are from different scale levels, the value is 1, otherwise it is 0. Indicates the and The square of the Euclidean distance between the sub-feature vectors of the scale;

[0127] Cross-scale consistency loss achieves cross-scale semantic alignment by comparing the Euclidean distance of scale sub-features between sonar image patches at different scale levels;

[0128] The principle is that in sonar scenarios, the same fish target may appear completely different in images of different scales. Conventional L2 regularization, cross entropy loss, and balanced weighting only focus on dense alignment of single-scale features, which can easily cause the loss of weak features of small-sized fish. This formula uses a cross-scale mask function to Selectively constraining feature alignment at different scales not only takes into account global consistency but also ensures that the representation semantics between micro-scale and large-scale features are always optimally coupled.

[0129] The cross-scale loss is used to align cross-scale features in terms of representation semantics, thereby improving the robustness of the sonar system in detecting fish targets with drastic scale changes.

[0130] The texture sub-feature vector sequence is constrained to be of the same scale. The texture sparsity loss is obtained by calculating the ratio of the L1 norm to the L2 norm of the texture sub-feature vector of each sonar image patch and summing the ratios of all sonar image patches at the same scale level.

[0131] ;

[0132] in, represents the set of all scale levels, Represents scale level The number of sonar image patches, Represents scale level Next The texture sub-feature vector of the sonar image patch, Indicates the The L1 norm of the texture sub-feature vector, Indicates the The L2 norm of the texture sub-feature vector, represents a small constant that prevents division by zero;

[0133] The same-scale sparsity constraint is used to measure the sparse expression degree of the texture sub-feature vectors of each sonar image patch at the same scale level, and the accumulated results are used to highlight the texture differences of the target echo under background interference.

[0134] The principle of texture sparsity loss is to use the ratio of the L1 norm to the L2 norm to measure the sparsity of texture sub-features at the same scale level. Unlike traditional L2 regularization and universal sparse loss, this formula limits sparsity to a set of patches at the same scale and level, encouraging the simplification of texture expression at each scale in a fine-grained manner. It can effectively highlight weak target textures, suppress clutter, and improve the system's sensitivity to target signals in complex environments.

[0135] A local structure constraint is imposed on the sequence of directional sub-eigenvectors. The Euclidean distance between the directional sub-eigenvectors of each sonar image patch and all its spatially adjacent sonar image patches is calculated, and the directional consistency loss is obtained by accumulating the Euclidean distances of all adjacent patch pairs.

[0136] ;

[0137] in, Indicates the The directional sub-feature vectors of sonar image patches, Indicates the The spatial location of the sonar image patch is adjacent to the The directional sub-feature vectors of sonar image patches, Indicates the The set of patch indices of adjacent sonar image patches in the original image space, Indicates the The first and The square of the Euclidean distance between the eigenvectors of the direction;

[0138] The local structure constraint is used to encourage the directional sub-feature vectors of spatially adjacent sonar image patches to maintain consistency, and the accumulated results are used to improve the stability of the directional features of fish targets within the spatial neighborhood.

[0139] Directional consistency loss actively strengthens the consistency of directional features of fish schools or single targets in continuous spatial areas by minimizing the local Euclidean distance of directional sub-features of adjacent patches in spatial positions. Compared with conventional global smoothing, pooling, or general graph convolution regularization terms that only consider large-scale feature consistency, directional consistency loss introduces local spatial structural information into the main feature learning process, explicitly models the spatial dependency between neighboring patches, and greatly improves the ability to maintain group direction, arrangement, and distribution information of fish in sonar images.

[0140] The scale consistency loss, texture sparsity loss and direction consistency loss are weightedly summed to obtain the overall multi-scale decoupled feature constraint loss;

[0141] The overall multi-scale decoupled feature constraint loss is used to simultaneously regulate the optimal representation states of scale sub-features, texture sub-features and orientation sub-features.

[0142] Based on the overall multi-scale decoupled feature constraint loss, the scale sub-feature set, texture sub-feature set and direction sub-feature set are reversely optimized and updated to output a multi-scale sonar decoupled feature representation that meets the constraint conditions.

[0143] Initialize the population of walrus optimization algorithm parameter triples, each of which consists of a scale sub-feature weight, a classification threshold, and a posterior fusion coefficient. For each walrus optimization algorithm parameter triple in the population of walrus optimization algorithm parameter triples, use the corresponding scale sub-feature weight to perform weighted fusion on the multi-scale sonar decoupling feature representation to obtain a candidate multi-scale sonar feature fusion result. Input the candidate multi-scale sonar feature fusion result into the classifier, output fish category, fish body length estimation, and fish school density, and calculate the multi-objective comprehensive fitness value corresponding to the walrus optimization algorithm parameter triple based on the output results.

[0144] In this embodiment, the calculation of the multi-objective comprehensive fitness value includes:

[0145] Initialize a population of walrus optimization algorithm parameter triplets based on the multi-scale sonar fish recognition task characteristics. The population contains several walrus optimization algorithm parameter triplets. Each walrus optimization algorithm parameter triplet contains three parameters: scale sub-feature weight vector, classification threshold, and posterior fusion coefficient.

[0146] The scale sub-feature weight vector is used to perform adaptive weight fusion of sonar features of different scales. The classification threshold is used to distinguish different types of fish targets. The posterior fusion coefficient is used to dynamically adjust the fusion strength of multi-scale sonar decoupling features to adapt to the differences in target scale distribution in the water area.

[0147] According to the scale perception mechanism, the scale sub-feature weight vector of each walrus optimization algorithm parameter triple is Perform scale-adaptive perturbations:

[0148] ;

[0149] in, Represents the mesoscale level of the walrus optimization algorithm parameter triple The corresponding scale sub-feature weight, is the scale perturbation factor, Scale level The sensitivity coefficient is used to reflect the importance of different scale sonar features to fish target detection. Represents scale level The average energy of all sonar image patches, Represents the scale level in the candidate multi-scale sonar feature fusion results energy distribution of the features;

[0150] The scale sub-feature weight vector is weighted and adjusted by the original weight value, scale perturbation factor, sensitivity coefficient of the scale level, average energy of all sonar image patches of the scale level, and energy distribution of scale level features in the current candidate multi-scale sonar feature fusion result, so as to realize adaptive adjustment of the weights of sonar features of different scales and improve the dynamic adaptability to the actual scale changes of fish schools.

[0151] The adjusted scale sub-feature weight vector is used to perform weighted fusion on the multi-scale sonar decoupling feature representation and calculate the candidate multi-scale sonar feature fusion result. ;

[0152] During the fusion process, the multi-scale sonar decoupling features at different scale levels are weighted according to their scale sub-feature weight coefficients. The weighted results of all scale levels are merged into the candidate multi-scale sonar feature fusion results. The candidate multi-scale sonar feature fusion results reflect the characteristic distribution of multi-scale sonar fish targets.

[0153] A classifier with an adaptive fusion decision-making mechanism is constructed. After the candidate multi-scale sonar feature fusion results are input into the classifier, the posterior fusion coefficient is used to weight the fish category probability distribution output by the classifier, and the fusion-enhanced fish category probability is output.

[0154] The fusion-enhanced fish category probability is the weighted sum of the fish category prior probability of the classifier and the posterior category probability given by the fusion result of the candidate multi-scale sonar features, which is obtained by the posterior fusion coefficient. The fusion-enhanced category probability accurately reflects the fish distribution characteristics in actual waters.

[0155] The fish category label is determined based on the fusion-enhanced fish category probability, and the classifier estimates the fish body length and fish school density, obtaining the output results of the fish category label, fish body length estimation value, and fish school density estimation value;

[0156] The classifier outputs a fish category label based on the candidate multi-scale sonar feature fusion results and the classification threshold, that is, it determines which specific fish species the current input belongs to; secondly, based on the candidate multi-scale sonar feature fusion results and the regression relationship learned by the model, it outputs a fish length estimate, that is, it makes a quantitative prediction of the length of the identified fish; at the same time, based on the candidate multi-scale sonar feature fusion results and the posterior fusion coefficient, it performs weighted aggregation on the response intensity and number of patches of the fish school, and outputs a fish density estimate, that is, a quantitative estimate of the number of fish per unit area.

[0157] For each walrus optimization algorithm parameter triplet, a multi-objective comprehensive fitness function is used to evaluate it and obtain the multi-objective comprehensive fitness value;

[0158] The multi-objective comprehensive fitness function includes the fish category recognition accuracy, fish body length estimation error, fish school density estimation error and scale sensitivity evaluation item of the scale weight vector. The multi-objective comprehensive fitness function is obtained by weighted summation of the corresponding coefficients. The fish category recognition accuracy is used to measure the consistency between the fish category label and the true label. The fish body length estimation error and the fish school density estimation error are used to measure the errors between the estimated values ​​and the true values ​​of the fish body length and fish school density, respectively. The scale sensitivity evaluation item of the scale weight vector is used to reflect the adaptive ability of the parameter triplet to detect fish targets of different scales.

[0159] In this embodiment, the walrus optimization algorithm that combines scale-adaptive perturbation and posterior fusion decision-making realizes unsupervised adaptive optimization of fish recognition weight parameters and thresholds in multiple scenarios and multiple devices, thereby improving the system's generalization capability and deployment flexibility. By introducing a scale-aware mechanism and energy-sensitive adaptive perturbation into the walrus optimization algorithm parameter triplet, the weights of sonar features at different scales can be dynamically adjusted to adapt to the size distribution of target fish schools and the echo characteristics of different sonar equipment and different water environments. At the same time, the adaptive decision-making mechanism that integrates posterior probability and prior probability can dynamically adjust the classification threshold and weight parameters according to the fish distribution in the target environment.

[0160] The multi-objective comprehensive fitness value is used as the fitness function of the walrus optimization algorithm. The leading update, cooperative update and follower update strategies are used to iteratively adjust the walrus optimization algorithm parameter triple population. In each iteration cycle, the walrus optimization algorithm parameter triple is updated according to the multi-objective comprehensive fitness value and the multi-objective comprehensive fitness value is recalculated until the convergence condition is met.

[0161] In this embodiment, the updating of the multi-objective comprehensive fitness value includes:

[0162] The multi-objective comprehensive fitness value is used as the fitness function of the walrus optimization algorithm. The multi-objective comprehensive fitness values ​​corresponding to each walrus optimization algorithm parameter triple are arranged in sequence to form a fitness sequence. The walrus optimization algorithm parameter triple with the largest multi-objective comprehensive fitness value is selected as the leading walrus parameter triple.

[0163] The leader walrus parameter triplet is used as a global optimization guide in the subsequent parameter update process of the population.

[0164] The parameter triple population is iteratively updated based on the leader-cooperation-follower mechanism of the walrus optimization algorithm;

[0165] In the leader guidance phase, all walrus optimization algorithm parameter triplets are guided and updated according to the difference between the current parameter state and the leader walrus parameter triplet;

[0166] The parameter adjustment amplitude is jointly controlled by the guidance intensity factor and the random perturbation coefficient that obeys the uniform distribution. The update direction always points to the current optimal leading walrus parameter triplet, which is used to achieve global exploration and convergence improvement.

[0167] In the collaborative perturbation phase, each walrus optimization algorithm parameter triplet is locally and interactively updated with any other parameter triplet in the same population;

[0168] The parameter adjustment range is jointly determined by the collaborative disturbance intensity factor and the disturbance control coefficient. The population diversity is introduced by randomly selecting collaborative objects to avoid falling into the local optimum.

[0169] In the follow-up correction stage, the scale sub-feature weight vector is fine-grainedly corrected according to the energy change trend of the candidate multi-scale sonar feature fusion results at each scale level;

[0170] The correction amplitude is determined by the scale direction correction coefficient. The adjustment amount is equal to the difference between the energy mean of the current scale level and the energy of the scale level feature in the candidate multi-scale sonar feature fusion result. By strengthening the low-energy scale weight, the sensitivity to weak echo targets of fish is improved.

[0171] After each round of parameter update, the multi-scale sonar features are re-fused based on the updated walrus optimization algorithm parameter triples and input into the adaptive fusion classifier to recover the fish category labels, fish body length estimates, and fish school density estimates. The multi-objective comprehensive fitness value is then recalculated, completing the real-time iterative correction of the fitness values ​​of all parameter triplets.

[0172] Determine whether the current iteration meets the convergence conditions. The convergence conditions include any of the following:

[0173] The global fitness value improvement rate is lower than the preset threshold, or the population parameter change amplitude is lower than the minimum disturbance threshold, or the preset maximum number of iterations is reached;

[0174] If any convergence condition is met, the parameter update iteration ends; otherwise, the process returns to continue the optimization of population parameters and fitness calculation.

[0175] When the convergence conditions are met, the classifier outputs the final fish identification results, fish body length estimation results and fish school density results, and generates a confidence heat map at the same time.

[0176] In this embodiment, when the convergence conditions are met, the walrus optimization algorithm parameter triplet with the highest multi-objective comprehensive fitness value is selected as the optimal walrus optimization algorithm parameter triplet, and the multi-scale sonar decoupling feature representation is weightedly fused using the optimal walrus optimization algorithm parameter triplet to obtain the final multi-scale sonar feature fusion result, and the final fish recognition result, fish body length estimation result and fish school density result are output through the classifier, and a confidence heat map is generated at the same time.

[0177] In this implementation, a full-process evolutionary search strategy driven by multi-objective comprehensive fitness is proposed to achieve end-to-end self-optimization of feature fusion weights, classification thresholds, and fusion coefficients, thereby improving the overall performance of fish species identification, body length estimation, and group density assessment. A multi-objective comprehensive fitness function is used to incorporate multiple indicators such as fish category recognition accuracy, estimation error of fish body length and fish density, and scale sensitivity into a unified evaluation system. The leader-collaborator-follower mechanism of the walrus optimization algorithm is combined to perform global and local adaptive evolution of the parameter triplet population. This not only improves the multi-task robustness of fish identification, but also dynamically optimizes the weight distribution of different tasks according to actual application requirements.

[0178] Example 1: A large-scale intelligent sonar monitoring test was carried out in the nearshore A fishing ground. The test water area was about 4 square kilometers, with a water depth of between 18 meters and 35 meters. The bottom was a mixture of sand, mud and reefs, the water was turbid, and the sonar echo signal-to-noise ratio was low.

[0179] The entire experimental system is carried on a medium-sized fishery survey vessel and is equipped with multiple types of sonar equipment, including two sets of 200kHz side-scan sonar arrays and a set of 120kHz forward-looking sonar arrays. The sonar data acquisition bandwidth is up to 10MB / s. The experimental platform performs a sonar scan every 10 seconds on each observation route, and the amount of data collected per hour is about 30GB. A total of 247 hours of sonar raw data were collected on-site by towing, and a total of 16,200 frames of standardized sonar grayscale image data were obtained. The fish species covered 12 economic species, including small mullet, yellowfin sea bream, large yellow croaker, etc. The target body length ranged from 7cm to 46cm, and the fish density ranged from 1 to 14 fish per square meter, forming a multi-scale and complex scene data set.

[0180] The staff pre-processed the original sonar echo signals and uniformly used denoising, beamforming and grayscale processes to convert the original echo signals into standardized two-dimensional grayscale images. For example, the original sonar images collected on the 28th voyage at a water depth of 28 meters and a turbidity of 15 NTU had an original envelope signal intensity range of -80dB to -43dB. After bandpass filtering and normalization, they were mapped to a grayscale value range of 0~255, retaining the details of small fish and strong background clutter.

[0181] Next, each frame of preprocessed sonar imagery is pyramid-downsampled to generate three scale levels: 256×256, 128×128, and 64×64. Sliding slices are then generated using 32×32, 24×24, and 16×16 windows to obtain sonar image patches with different scales and spatial location identifiers. For the grayscale image "F28_20250615_0932" collected on June 15th, a total of 312 patches were generated, including 94 microscale patches, 106 mesoscale patches, and 112 large-scale patches. Each patch carries the original image coordinates and scale information.

[0182] All patches are fed into the DINOv2 visual Transformer encoder, which is modified with multi-scale perceptual attention. The system automatically embeds the scale and spatial position of each patch's feature vector. This multi-scale attention mechanism models the global dependencies between features at different scales and in different spatial locations. For example, in the data set "F12_20250617_1410" collected in an area with dense fish, the model significantly increases the weight of weak echoes from small fish, thus avoiding misjudgments caused by highly reflective bottom surfaces.

[0183] After processing by the feature decoupling matrix module, the model automatically decomposes each eigenvector into scale sub-features, texture sub-features, and directional sub-features. Taking the sample "F33_20250702_0738" as an example, through matrix mapping, the scale sub-features of micro-scale fish school patches are significantly distinguished from large-scale fish schools. The texture sub-features show high sparsity in various clutter areas, and the directional sub-features show strong consistency in continuous spatial distribution areas.

[0184] Subsequently, the system imposes cross-scale consistency, same-scale sparsity and directional local structure loss on all sub-features according to the constraints. Through reverse optimization, the model ensures the semantic alignment of cross-scale features and enhances the feature expression of micro-scale fish. For example, in the scene "F41_20250620_1612" with strong clutter and low fish density, the micro-scale texture sub-features constrained by sparsity significantly stand out from the background, effectively suppressing false detection.

[0185] During the parameter optimization phase, the system initialized a population of 48 parameter triplets, each containing a scale-based sub-feature weight, a dynamic classification threshold, and a posterior fusion coefficient. The system automatically applied scale-sensitive adaptive perturbations to each weighting group and dynamically monitored energy distribution. In areas with strong bottom reflectance, the Walrus optimization algorithm increased the microscale sub-feature weights by approximately 23%, accurately separating weak fish echoes from background noise.

[0186] The classifier uses adaptive fusion decision-making to input the candidate feature fusion results and output fish category, body length estimation, and fish density. On-site calibration was performed using manually placed target fish and high-resolution underwater video. 2,122 validation samples were statistically measured, and the comparison data between the present invention and traditional methods were obtained as follows:

[0187] Table 1 Comparison data between the present invention and the traditional method

[0188]

[0189] Specific training sample examples (part):

[0190] "F12_20250617_1410": Small mullet, actual body length 13.8 cm, AI estimated 14.1 cm, school density actual 4.2 fish / m², estimated 4.5 fish / m²;

[0191] F19_20250619_1017: Large yellow croaker, actual body length 34.5 cm, AI estimated 34.2 cm, actual school density 2.1 fish / m², estimated 2.0 fish / m²;

[0192] "F28_20250615_0932": Yellowfin seabream, actual body length 23.7cm, AI estimated 23.4cm, actual population density 6.7 / m², estimated 6.9 / m²;

[0193] “F41_20250620_1612”: densely packed small fish, actual body length 7.5cm, AI estimated 7.9cm, actual school density 13.4 fish / m², estimated 12.8 fish / m².

[0194] Compared with the traditional CNN+GA method, for extremely small objects with strong clutter such as "F41_20250620_1612", the recall rate of the proposed method is 19 percentage points higher, and the errors of body length and density estimation are reduced by more than 60%.

[0195] In addition, the system was deployed off-site in nearby waters to collect and test 1,400 sets of data. The results showed that the average recognition accuracy of the present invention dropped from 92.8% to 90.6%, a decrease of only 2.4%, while the traditional method dropped to 72.3%, a decrease of 13.7%.

[0196] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A sonar fish identification method based on self-supervised learning, characterized in that: include: collecting sonar echo image data and performing preprocessing to obtain preprocessed sonar echo image data; Perform pyramid multi-scale slicing on the pre-processed sonar echo image data to generate a multi-scale sonar image patch set; The multi-scale sonar image patch set is fed into the improved DINOv2 visual Transformer encoder to obtain the initial multi-scale sonar feature representation. A feature decoupling matrix module is inserted into the improved DINOv2 visual Transformer encoder to perform dimension division on the initial multi-scale sonar feature representation to obtain decoupled scale sub-features, texture sub-features, and direction sub-features. Apply cross-scale consistency constraints and same-scale sparsity constraints to the decoupled scale sub-features, texture sub-features, and direction sub-features, and output a multi-scale sonar decoupled feature representation that meets the constraints. Initialize the walrus optimization algorithm parameter triple population, obtain the candidate multi-scale sonar feature fusion results and input them into the classifier, output fish category, fish body length estimation and fish density, and calculate the multi-objective comprehensive fitness value; The multi-objective comprehensive fitness value is used as the fitness function of the walrus optimization algorithm to iteratively adjust the walrus optimization algorithm parameter triple population, update the walrus optimization algorithm parameter triple and recalculate the multi-objective comprehensive fitness value until the convergence condition is met; When the convergence conditions are met, the classifier outputs the final fish identification results, fish body length estimation results and fish school density results, and generates a confidence heat map at the same time.

2. The sonar fish identification method based on self-supervised learning according to claim 1, characterized in that: The collecting and preprocessing of sonar echo image data includes: Collect the original sonar echo signals of the water area to be monitored through the side-scan sonar array or the forward-looking sonar array; The original sonar echo signal sequence is band-pass filtered to eliminate non-working frequency band noise, and the reflection intensity envelope signal is extracted through envelope detection to construct the envelope intensity mapping matrix; The envelope intensity mapping matrix is ​​beamformed and the envelope intensity information of multiple beam directions is reconstructed into a two-dimensional intensity image using a delayed weighted superposition method. The linear normalization method is used to uniformly map the intensity value of each pixel in the two-dimensional intensity image after beamforming to a grayscale value range of 0 to 255, and the grayscale-processed two-dimensional image is output as the preprocessed sonar echo image data.

3. The sonar fish identification method based on self-supervised learning according to claim 1 is characterized in that: The pyramid-type multi-scale slicing comprises: An image pyramid sequence is constructed by downsampling. The pre-processed sonar echo image data is layered according to the scale scaling coefficient set to construct three scale levels: large-scale image, medium-scale image and micro-scale image. Perform sliding window slicing operations on each scale level of the image pyramid sequence to obtain multiple sonar image patches; Constructing a sonar image patch scale identification vector for each sonar image patch; The sonar image patch sets at all scale levels are merged to form a multi-scale sonar image patch set, and the scale identification vectors of all sonar image patches are simultaneously merged to form a scale identification set corresponding to the multi-scale sonar image patches.

4. The sonar fish identification method based on self-supervised learning according to claim 3 is characterized in that: The construction of the improved DINOv2 visual Transformer encoder includes: The multi-scale sonar image patch set is fed into the DINOv2 visual Transformer encoder modified with multi-scale perceptual attention. The multi-scale sonar image patch set contains multiple sonar image patches, and each sonar image patch is marked according to its index in the set; Each sonar image patch is flattened and linearly transformed to obtain an initial feature vector of the sonar image patch. A scale embedding vector and a spatial position embedding vector are added to the initial feature vector of the sonar image patch according to the scale identification information of the sonar image patch, thereby forming a sonar image patch feature vector that contains both scale characteristics and spatial position information. Compute the multi-scale perceptual attention matrix for any two sonar image patch feature vectors containing scale embeddings and spatial position embeddings; Perform weighted fusion of the value vector based on the multi-scale perception attention matrix to generate a multi-scale attention feature vector after fusion of scale information; A multi-scale residual connection structure for multi-scale sonar fish recognition tasks is introduced into the multi-scale attention feature vector, and the feature enhanced output vector is obtained through feedforward neural network mapping; All feature-enhanced output vectors constitute the initial multi-scale sonar feature representation set.

5. The sonar fish identification method based on self-supervised learning according to claim 4 is characterized in that: Insertion of the feature decoupling matrix module includes: A feature decoupling matrix module is inserted into the improved DINOv2 visual Transformer encoder. The input of the feature decoupling matrix module is the initial multi-scale sonar feature representation set. Through the feature decoupling matrix module, scale feature extraction, texture feature extraction and direction feature extraction operations are performed on each feature enhancement output vector to obtain scale sub-feature vector, texture sub-feature vector and direction sub-feature vector; All scale sub-feature vectors, texture sub-feature vectors and direction sub-feature vectors are set in index order to form a scale sub-feature set, a texture sub-feature set and a direction sub-feature set respectively.

6. The sonar fish identification method based on self-supervised learning according to claim 5, characterized in that: The output of the multi-scale sonar decoupling feature representation that satisfies the constraint conditions includes: The scale sub-feature vectors of each sonar image patch form a scale sub-feature vector sequence, the texture sub-feature vectors of each sonar image patch form a texture sub-feature vector sequence, and the direction sub-feature vectors of each sonar image patch form a direction sub-feature vector sequence; A cross-scale consistency constraint is imposed on the scale sub-feature vector sequence. The cross-scale consistency constraint is used to measure the semantic consistency of fish target representation of sonar image patches at different scale levels, and a cross-scale loss is constructed to measure the semantic consistency of fish target representation at different scales. : ; in, represents the total number of sonar image patches, Indicates the The scaled sub-feature vector of a sonar image patch, Indicates the The scaled sub-feature vector of a sonar image patch, Indicates the The scale level identification of each sonar image patch, Indicates the The scale level identification of each sonar image patch, represents the cross-scale mask function, if and If the sonar image patches are from different scale levels, the value is 1, otherwise it is 0. Indicates the and The square of the Euclidean distance between the sub-feature vectors of the scale; The texture sub-feature vector sequence is constrained to be of the same scale. The texture sparsity loss is obtained by calculating the ratio of the L1 norm to the L2 norm of the texture sub-feature vector of each sonar image patch and summing the ratios of all sonar image patches at the same scale level. A local structure constraint is imposed on the sequence of directional sub-eigenvectors. The Euclidean distance between the directional sub-eigenvectors of each sonar image patch and all its spatially adjacent sonar image patches is calculated, and the directional consistency loss is obtained by accumulating the Euclidean distances of all adjacent patch pairs. The scale consistency loss, texture sparsity loss and direction consistency loss are weightedly summed to obtain the overall multi-scale decoupled feature constraint loss; Based on the overall multi-scale decoupled feature constraint loss, the scale sub-feature set, texture sub-feature set and direction sub-feature set are reversely optimized and updated to output a multi-scale sonar decoupled feature representation that meets the constraint conditions.

7. The sonar fish identification method based on self-supervised learning according to claim 6, characterized in that: The calculation of the multi-objective comprehensive fitness value includes: Initialize a population of walrus optimization algorithm parameter triplets based on the multi-scale sonar fish recognition task characteristics. The population contains several walrus optimization algorithm parameter triplets. Each walrus optimization algorithm parameter triplet contains three parameters: scale sub-feature weight vector, classification threshold, and posterior fusion coefficient. Based on the scale perception mechanism, the scale sub-feature weight vector of each walrus optimization algorithm parameter triplet is scale-adaptively perturbed; The adjusted scale sub-feature weight vector is used to perform weighted fusion on the multi-scale sonar decoupling feature representation and calculate the candidate multi-scale sonar feature fusion result; A classifier with an adaptive fusion decision-making mechanism is constructed. After the candidate multi-scale sonar feature fusion results are input into the classifier, the posterior fusion coefficient is used to weight the fish category probability distribution output by the classifier, and the fusion-enhanced fish category probability is output. The fish category label is determined based on the fusion-enhanced fish category probability, and the classifier estimates the fish body length and fish school density, obtaining the output results of the fish category label, fish body length estimation value, and fish school density estimation value; For each walrus optimization algorithm parameter triplet, a multi-objective comprehensive fitness function is used to evaluate it and obtain the multi-objective comprehensive fitness value.

8. The sonar fish identification method based on self-supervised learning according to claim 7, characterized in that: The updating of the multi-objective comprehensive fitness value includes: The multi-objective comprehensive fitness value is used as the fitness function of the walrus optimization algorithm. The multi-objective comprehensive fitness values ​​corresponding to each walrus optimization algorithm parameter triple are arranged in sequence to form a fitness sequence. The walrus optimization algorithm parameter triple with the largest multi-objective comprehensive fitness value is selected as the leading walrus parameter triple. The parameter triple population is iteratively updated based on the leader-cooperation-follower mechanism of the walrus optimization algorithm; In the leader guidance phase, all walrus optimization algorithm parameter triplets are guided and updated according to the difference between the current parameter state and the leader walrus parameter triplet; In the collaborative perturbation phase, each walrus optimization algorithm parameter triplet is locally and interactively updated with any other parameter triplet in the same population; In the follow-up correction stage, the scale sub-feature weight vector is fine-grainedly corrected according to the energy change trend of the candidate multi-scale sonar feature fusion results at each scale level; After each round of parameter update, the multi-scale sonar features are re-fused based on the updated walrus optimization algorithm parameter triples and input into the adaptive fusion classifier to recover the fish category labels, fish body length estimates, and fish school density estimates. The multi-objective comprehensive fitness value is then recalculated, completing the real-time iterative correction of the fitness values ​​of all parameter triplets. Determine whether the current iteration meets the convergence conditions. The convergence conditions include any of the following: The global fitness value improvement rate is lower than the preset threshold, or the population parameter change amplitude is lower than the minimum disturbance threshold, or the preset maximum number of iterations is reached; If any convergence condition is met, the parameter update iteration ends; otherwise, the process returns to continue the optimization of population parameters and fitness calculation.

9. The sonar fish identification method based on self-supervised learning according to claim 8, characterized in that: When the convergence condition is met, the walrus optimization algorithm parameter triplet with the highest multi-objective comprehensive fitness value is selected as the optimal walrus optimization algorithm parameter triplet, and the multi-scale sonar decoupling feature representation is weightedly fused using the optimal walrus optimization algorithm parameter triplet to obtain the final multi-scale sonar feature fusion result, and the final fish recognition result, fish body length estimation result and fish school density result are output through the classifier, and a confidence heat map is generated at the same time.

Citation Information

Patent Citations

  • Target counting method and device, terminal and computer readable storage medium

    CN116109916A

  • Intelligent fish identifying and monitoring method and system based on multi-sensor data

    CN117214904A