Underwater target classification method based on acousto-optic image fusion Transformer and related device

By adopting the acousto-optical image fusion Transformer method in the classification of underwater target image, combining the acousto-optical fusion module, channel shuffling module, self-attention mechanism and style pooling feature calibration module, the problem of insufficient capture of significance features of underwater target image is solved, and higher classification accuracy and lower misclassification rate are achieved.

CN120107655AActive Publication Date: 2025-06-06ZHEJIANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510098186.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

In the prior art, the significant feature capture of underwater target images is insufficient, resulting in the problem of high misclassification.

Method used

Using a method based on the acousto-optical image fusion Transformer, a multimodal fusion image tensor is obtained through the acousto-optical fusion module, and an acousto-optical image channel shuffling module is designed to enhance the multimodal information fusion effect. The self-attention-based Transformer architecture is used to extract the features of the fused image tensor, and a pooled calibration attention mechanism and a stylized pooled feature calibration module are introduced to improve the semantic significance of feature extraction.

Benefits of technology

The acousto-optical image fusion improves the richness of the underwater target features and improves the semantic significance of the features, thereby significantly improving the classification accuracy of the underwater target images and reducing misclassification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107655A_ABST
    Figure CN120107655A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image classification in computer vision, and discloses an underwater target classification method based on acousto-optic image fusion Transformer and a related device. The underwater target image classification method comprises the following steps: obtaining optical images to be classified and acoustic images corresponding to the optical images; and based on the acoustic and optical images to be classified, performing underwater target classification by using a pre-trained depth image classification network to obtain a classification result. According to the invention, the acousto-optic fusion module carries out acousto-optic image fusion, and the feature extraction module carries out feature extraction and feature mixing, so that the diversity and semantic richness of features are enhanced, and the underwater target image classification accuracy is improved. The technical problem of high misclassification of the underwater target image caused by insufficient single-mode image feature extraction and insufficient feature richness capture in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image classification in computer vision, and in particular relates to an underwater target classification method based on an acoustic-optical image fusion Transformer and a related device. Background Art

[0002] With the development of intelligent underwater unmanned equipment, the role of underwater target recognition technology has become increasingly prominent. Underwater target image classification refers to the technology of classifying targets in images obtained by underwater imaging equipment. It can be applied to underwater target search and rescue, underwater environment detection, ship navigation and obstacle avoidance, etc. It is a key step in the intelligentization of underwater unmanned equipment and marine environment monitoring. Explanatoryally, the marine water environment is complex and there are many environmental interferences (for example, such as pixel size, focus distance, shooting angle, environmental dirt, external light reflection factors, etc.), and the targets in the acquired images are not easy to distinguish, which will affect the accuracy of underwater target image classification tasks. Further explanatory, underwater targets have inter-class similarities and intra-class differences in morphology, which determines that the category of underwater targets is more difficult to judge. In summary, underwater target image classification is a very challenging task.

[0003] Existing deep learning-based image classification methods are mainly based on convolutional neural networks (CNNs) and vision transformers (ViTs). Convolutional neural networks use the prior knowledge of image translation invariance to aggregate pixel information layer by layer through convolution operations, thereby learning underwater target features and achieving recognition on single data source images, such as optical images collected by underwater optical sensors (such as underwater cameras) or acoustic images acquired by acoustic sensors (such as sonar). However, CNN-based image classification methods are limited by local receptive fields and cannot effectively capture long-range contextual information, resulting in low classification accuracy. In contrast, vision transformers excel in modeling long-range dependencies of features and can usually achieve higher accuracy (Top-1 Accuracy) than traditional convolutional neural networks in image classification tasks. However, the classification method of transformers with a single data type and a single feature pattern is difficult to fully describe the rich texture details and feature diversity of underwater targets due to the lack of richness of feature expression. Therefore, it is crucial to propose a method for multi-source data fusion enhanced recognition to improve the recognition accuracy of underwater targets. In addition, traditional visual Transformer image classification networks often lack feature enhancement modules, which makes it impossible to effectively step in subtle or difficult-to-distinguish features of underwater targets, thereby affecting the classification accuracy of acoustic or optical images.

[0004] Multi-source information fusion technology aims to integrate different types of data, make full use of the advantages of each data source, and extract richer and more representative information. However, existing multi-source information fusion methods usually face the problem of limited flow of multimodal information between feature channels, and traditional point-by-point convolution fusion methods are computationally expensive. In addition, existing Transformer-based image recognition methods often use ordinary self-attention mechanisms and fail to fully consider how to enhance the capture of salient information of fused feature maps. Therefore, constructing an efficient visual Transformer cognitive computing architecture based on the fusion of acoustic images and optical images can enhance the ability to extract salient features in complex underwater environments and improve the accuracy of underwater target classification, which has important theoretical and practical significance. Summary of the invention

[0005] The purpose of the present invention is to provide an underwater target classification method and related device based on acoustic-optical image fusion Transformer, aiming to solve the problem that the underwater target image in the prior art is insufficient in capturing the salient features, resulting in a high misclassification rate. In the technical solution disclosed in the present invention, an acoustic-optical fusion module is proposed to obtain a multimodal fusion image tensor; at the same time, in order to increase the information flow between multimodal data channels, the present invention designs an acoustic-optical image channel shuffling module to enhance the multimodal information fusion effect; in terms of feature extraction, the present invention uses a Transformer architecture based on self-attention to extract the features of the fused image tensor. In order to further improve the global dependency modeling capability of the self-attention mechanism, the present invention designs a pooling calibration attention mechanism and proposes a feature calibration module based on style pooling. The method of the present invention improves the richness of underwater target features through acoustic-optical image fusion, and improves the semantic significance of underwater target features through the feature enhancement self-attention module, thereby improving the classification accuracy of underwater target images.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides an underwater target classification method based on acoustic-optical image fusion Transformer, comprising the following steps:

[0008] Acquiring acoustic and optical images of underwater targets to be classified;

[0009] Based on the acoustic image and the optical image of the underwater target to be classified, a pre-trained deep image classification network is used to classify the underwater target to obtain a classification result;

[0010] in,

[0011] The deep image classification network includes an acoustic-optical fusion module, a feature extraction module and a linear classification head connected in sequence;

[0012] The acoustic-optical fusion module first performs image preprocessing on the acoustic image and the optical image, and converts them into acoustic image tensors and optical image tensors respectively; then, channel splicing and shuffling are performed on the paired acoustic image tensors and optical image tensors to obtain the acoustic-optical fusion image tensors; the feature extraction module adopts a multi-stage encoder structure, and the encoder of each stage includes a serial patch embedding and a number of encoder modules based on a pooling calibration attention mechanism;

[0013] The encoder module includes a convolution position encoding module, a first layer normalization operation, a pooling calibration attention module, a second layer normalization operation and a feedforward neural network layer connected in series, and the input of the first layer normalization operation is added to the output of the pooling calibration attention module, and the input of the second layer normalization operation is added to the output of the feedforward neural network through a skip connection layer;

[0014] The pooling calibration attention module includes a feature calibration module based on style pooling, a multi-head transposed self-attention module and a convolutional relative position encoding module; wherein the feature calibration module based on style pooling is used to input the query Q, key K, value V matrix formed by linear mapping of the output feature map of the first layer normalization operation, and output the query Q after feature calibration * , key K * , value V * Matrix; The multi-head transposed self-attention module is used to input the feature-calibrated query Q * , key K * , value V * matrix, outputs the multi-head transposed self-attention module output feature map; the convolution relative position encoding module is used to input the feature-calibrated query Q output by the style pooling-based feature calibration module * , value V * Matrix, output convolution relative position encoding module output matrix; the sum of the multi-head transposed self-attention module output feature map and the convolution relative position encoding module output matrix forms the pooled calibration attention module output feature map.

[0015] A second aspect of the present invention provides an underwater target classification system based on acoustic-optical image fusion Transformer, which is used to implement the underwater target classification method based on acoustic-optical image fusion Transformer as described in any one of the first aspects of the present invention.

[0016] In a third aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, an underwater target classification method based on an acoustic-optical image fusion Transformer as described in any one of the first aspect of the present invention is implemented.

[0017] In a fourth aspect of the present invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the underwater target classification method based on the acoustic-optical image fusion Transformer as described in any one of the first aspects of the present invention is implemented.

[0018] Compared with the prior art, the present invention has the following beneficial effects:

[0019] The present invention provides an underwater target classification method based on an acoustic-optical image fusion Transformer, which is classified through a pre-trained deep image classification network, wherein the deep image classification network comprises an acoustic-optical fusion module, a feature extraction module and a linear classification head; the acoustic-optical fusion module performs channel splicing and channel shuffling on an acoustic image tensor and an optical image tensor; the feature extraction module performs feature extraction on the acoustic-optical fusion image tensor, and inputs the outputted acoustic-optical fusion feature map into a linear classification module to obtain a classification result. In summary, in view of the problem of single feature extraction of underwater target images using a single modality in the prior art, the present invention makes full use of the multimodal features of acoustic images and optical images, enhances the diversity of features, and can effectively improve the classification accuracy of underwater targets and reduce misclassification. The present invention adopts a method of image pre-fusion and uses a single branch for feature extraction. Compared with the double-branch method of image post-fusion, the present invention has a lower computational burden and a higher reasoning speed, and introduces a channel shuffling and a feature calibration module based on style pooling to further enhance feature flow and feature extraction. The method of the present invention is based on the visual Transformer framework and introduces multimodal visual cognitive computing into the field of automatic classification, which can effectively improve the accuracy of underwater target classification. In addition, each module is simple to implement, has no excessive dependencies, and has strong applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below; obviously, the drawings described below are some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 It is a flowchart of an underwater target classification method based on acoustic-optical image fusion Transformer provided by an embodiment of the present invention.

[0022] Figure 2 It is a schematic diagram of the structure of a deep image classification network in an embodiment of the present invention.

[0023] Figure 3It is a structural diagram of an encoder module in an embodiment of the present invention.

[0024] Figure 4 It is a structural diagram of a pooling calibration attention module in an embodiment of the present invention.

[0025] Figure 5 is a schematic diagram of underwater target image classification results in a specific embodiment of the present invention; wherein, Figure 5 (a) is a schematic diagram of the classification results of reefs category image samples. Figure 5 (b) is a schematic diagram of the classification results of the shipwrecks category image sample. Figure 5 (c) is a schematic diagram of the classification results of fishes category image samples. Figure 5 (d) is a schematic diagram of the classification results of the (Autonomous Underwater Vehicles) AUVs category image samples. Figure 5 (e) is a schematic diagram of the classification results of frogmen category image samples.

[0026] Figure 6 It is a schematic diagram of an underwater target classification system based on acoustic-optical image fusion Transformer provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] See also Figure 1 and Figure 2 In an embodiment of the present invention, a method for underwater target classification based on acoustic-optical image fusion Transformer is provided, comprising the following steps:

[0029] Step 1, obtaining acoustic and optical images of underwater targets to be classified; for example, the optical image to be classified is an RGB image (i.e., a color image) taken underwater, and the image contains underwater targets to be classified; the acoustic image to be classified is generated by the optical image through a style transfer network, and this process is a conventional technology in the art and will not be described in detail in the present invention;

[0030] Step 2, based on the acoustic image and optical image of the underwater target to be classified, a pre-trained deep image classification network is used to classify the underwater target to obtain a classification result; explanatory, the classification result is output in the form of category and probability;

[0031] Among them, Figure 2 As shown, the deep image classification network includes an acoustic-optical fusion module, a feature extraction module and a linear classification head which are connected in sequence.

[0032] The acoustic-optical fusion module first performs image preprocessing on the paired input acoustic image and optical image, including two operations: first, the length and width of the acoustic image and the optical image are scaled to query 224×224 pixels, and then each pixel of the acoustic image and the optical image is normalized from 0-255 to 0-1 interval, thereby forming an acoustic image tensor and an optical image tensor. Then, the acoustic image tensor and the optical image tensor are channel-joined, and then the acoustic image tensor and the optical image tensor are orderly shuffled in the channel dimension to form alternating acoustic image tensor layers and optical image tensor layers, thereby obtaining an acoustic-optical fusion image tensor.

[0033] The feature extraction module includes stage 1, stage 2, stage 3, and stage 4 connected in sequence; wherein, stage 1 includes 4-fold patch embedding and a plurality of encoder modules connected in series in sequence; each of stage 2, stage 3, and stage 4 includes 2-fold patch embedding and a plurality of encoder modules connected in series in sequence; explanatory, the 4-fold patch embedding is used to input the acousto-optic fusion image tensor and output a 4-fold downsampled feature map; the 2-fold patch embedding is used to input the output feature map of the encoder module of the previous stage and output a 2-fold downsampled feature map; the encoder module is used to input the 4-fold patch embedding Or a 2x downsampled feature map, and an output encoder module output feature map; further specifically, by way of example, the 4x patch embedding downsamples the input acousto-optic fusion image tensor by 4x, and the downsampling operation uses 2 convolution operations with a filter size of 3×3 and a step size of 2, and performs 1 depthwise separable convolution operation with a filter size of 3×3 and a step size of 1; the 2x patch embedding downsamples the output feature map of the encoder module of the previous stage by 2x, and the downsampling operation uses 1 depthwise separable convolution operation with a filter size of 3×3 and a step size of 2.

[0034] like Figure 3As shown, the encoder module includes convolution position coding, a first layer normalization operation, a pooling calibration attention module, a second layer normalization operation and a feedforward neural network layer, and the input of the first layer normalization operation is added to the output of the pooling calibration attention module, and the input of the second layer normalization operation is added to the output of the feedforward neural network through a jump connection layer; specifically, the convolution position coding is used to input the input feature map of the encoder module, and output the convolution position coding output feature map; the first layer normalization operation is used to input the convolution position coding output feature map, and output the first layer normalization operation. The first layer of normalization operation output feature map; the pooling calibration attention module is used to input the first layer of normalization operation output feature map, and output the pooling calibration attention module output feature map; the second layer of normalization operation is used to input the sum of the pooling calibration attention module output feature map and the convolution position encoding output feature map, and output the second layer of normalization operation output feature map; the feedforward neural network layer is used to input the second layer of normalization operation output feature map, and output the feedforward neural network output feature map; the encoder module output feature map is the sum of the feedforward neural network output feature map and the second layer of normalization operation input feature map.

[0035] like Figure 4 As shown, the pooling calibration attention module includes a feature calibration module based on style pooling, a multi-head transposed self-attention module and a convolutional relative position encoding module; the present invention adopts a feature calibration module based on style pooling to enhance the multi-head transposed self-attention module; the feature calibration module based on style pooling adopts three feature calibration networks to perform feature calibration operations based on style pooling on the input query Q, key K, and value V matrices respectively, where the query Q, key K, and value V matrices are obtained by linear mapping the output feature maps of the first layer normalization operation; each feature calibration network includes a trunk network and a branch network, and the output of the feature calibration network is the output of the trunk network and the output of the branch network The broadcast-based Hadamard product is a Hadamard product; wherein the operation of the backbone network is identity mapping, and the operation of the branch network includes sequentially connected global average pooling, global maximum pooling, one-dimensional convolution, and sigmoid activation function operations; through average pooling and maximum pooling, the branch network can effectively extract the channel correlation of the query Q, key K, and value V matrix formed by the feature map of the acoustic-optical fusion image, and mix the information related to adjacent channels through one-dimensional convolution, and then obtain the weights of each acoustic image channel and optical image channel of the underwater target through the Sigmoid function and act on the query Q, key K, and value V matrix in the backbone network, thereby forming a query Q after feature calibration * , key K * , value V * matrix; feature-calibrated query Q * , key K * , value V *The matrix enhances the saliency of features by suppressing or enhancing the multi-source heterogeneous feature information channels; it forms a subsequent pooled calibration attention to capture the enhanced long-distance dependencies, thereby improving the accuracy of acoustic-optical image fusion and classification. The multi-head transposed self-attention module inputs the query Q after feature calibration * , key K * , value V * Matrix, output multi-head transposed self-attention module output feature map; convolutional relative position encoding module is used to input feature-calibrated query Q * , value V * Matrix, output convolution relative position encoding module output matrix; the output feature map of the pooling calibration attention module is the sum of the output feature map of the multi-head self-attention module and the output matrix of the convolution relative position encoding module.

[0036] In the embodiment of the present invention, the cognitive computing technology based on visual transformer is applied to the field of underwater target automatic recognition. The acoustic image tensor and the optical image tensor are channel-joined and channel-shuffled through the acoustic-optical fusion module, and an improved feature extraction module is proposed to extract acoustic-optical image features; the acoustic-optical fusion module is used to fuse the acoustic image features and the optical image features, which enhances the diversity and richness of the features, and an image classification network is constructed based on this to solve the underwater target image classification task. Further explanatory, the feature extraction module and feature fusion module disclosed in the embodiment of the present invention are based on basic convolution and matrix operations, and the input and output data forms are consistent, without too much dependence, which is convenient for application in various underwater target depth image classification models, and has a wide range of application prospects. In summary, the underwater target classification method based on acoustic-optical image fusion transformer provided in the embodiment of the present invention uses a method based on acoustic image and optical image fusion and feature extraction to improve the diversity of underwater target features, which can enhance the diversity and semantic richness of the extracted features and improve the accuracy of underwater target image classification. Therefore, the technical solution disclosed in the embodiment of the present invention can solve the technical problem of misclassification of underwater target images caused by insufficient feature extraction of single modality images and insufficient consideration of feature richness in the above-mentioned prior art.

[0037] In one embodiment of the present invention, the acoustic-optical fusion module performs channel splicing and channel shuffling on the acoustic image tensor and the optical image tensor; inputs the acoustic image and the optical image, and outputs the acoustic-optical fusion image tensor. The acoustic-optical fusion module may be:

[0038]

[0039] In the formula, X a is the acoustic image tensor, X o is the optical image tensor, X aois the concatenated tensor of the acoustic image tensor and the optical image tensor, X is the acoustic-optical fusion image tensor, Cat is channel concatenation; Shuffle is channel shuffling, X a1 ,X a2 ,X a3 Respectively represent channel 1, channel 2, and channel 3 of the acoustic image tensor; X o1 ,X o2 ,X o3 Represent channel 1, channel 2, and channel 3 of the optical image tensor respectively.

[0040] In one embodiment of the present invention, the training process of the deep image classification network includes:

[0041] 1) Data preprocessing: The acoustic images and optical images in the training set and the acoustic images and optical images in the validation set are scaled to a size of 224×224 and converted into tensors. The images in the training set and the validation set are paired acoustic images and optical images. After channel concatenation and channel shuffling, an acoustic-optical fusion image tensor is formed.

[0042] 2) Data loading and batch processing: According to the set batch size of 64, the data of the training set and the validation set are loaded in batches;

[0043] 3) Model initialization: Initialize the defined optimizer and cross entropy loss function based on the acoustic-optical image fusion Transformer;

[0044] 4) Model training: For each training cycle (epoch), the acoustic-optical fusion image tensor is input into the network through forward propagation, and the prediction result is obtained through the feature extraction module and the linear classification head; the difference between the prediction result and the actual label is calculated using the cross entropy loss function; the gradient is calculated and the network parameters are updated to minimize the loss through the error back propagation algorithm;

[0045] 5) Model validation: After each training cycle, the model performance is evaluated by calculating the cross entropy loss and accuracy on the validation set;

[0046] 6) Model saving: If the accuracy on the validation set improves, save the current optimal model.

[0047] In a specific embodiment of the present invention, the process of the underwater target classification method including the network training process specifically includes the following steps:

[0048] Step 1: Construct an optical image dataset of the underwater target to be identified; the image satisfies the three color channels of RGB (i.e., color images) and contains the corresponding manual classification results; the length and width pixel sizes of the dataset range from 480 to 1920; use the style transfer network to generate the corresponding acoustic image dataset; randomly divide the dataset into a training set and a validation set; and transmit the collected images in pairs to the computer that executes the algorithm.

[0049] Step 2: Construct an image classification network, including a sequentially connected sound-light fusion module, a feature extraction module, and a linear classification head;

[0050] In the embodiment of the present invention, the acoustic-optical fusion module performs channel splicing on the acoustic image tensor and the optical image tensor, inputs the acoustic image and the optical image, and outputs the acoustic-optical fusion image tensor. The acoustic-optical fusion module is:

[0051]

[0052] In the formula, X a is the acoustic image tensor, X o is the optical image tensor, X ao is the sound-light fusion image tensor, Cat is channel splicing; Shuffle is channel shuffling, X a1 ,X a2 ,X a3 Represents channel 1, channel 2, channel 3, X of the acoustic image tensor respectively. o1 ,X o2 ,X o3 Represent channel 1, channel 2, and channel 3 of the optical image tensor respectively.

[0053] The feature extraction module includes stage 1, stage 2, stage 3, and stage 4 connected in sequence; the stage 1 includes 4x patch embedding and 2 encoder modules connected in series in sequence; the stage 2, the stage 3, and the stage 4 all include 2x patch embedding connected in series in sequence, including 2, 6, and 2 encoder modules respectively; the 4x patch embedding module includes the following two convolution operations with a stride of 2 and a filter size of 3×3 and one depthwise separable convolution operation with a stride of 1 and a filter size of 3×3:

[0054]

[0055] In the formula, X 1 The acoustic-optical fusion image tensor input to the feature extraction module, is the feature map output by the convolution operation, Output feature map of depth-separable convolution, Conv 3×3 For a convolution operation with a filter size of 3×3, DWC 3×3It is a depth-wise separable convolution operation with a filter size of 3×3; after each convolution operation, batch normalization and Hardswish activation function processing are performed; the two convolution operations downsample the resolution of the input image by 4 times and expand the number of channels of the input image to 40; the depth-wise separable convolution operation does not change the image resolution and the number of channels; the 2x patch embedding module of stages 2, 3, and 4 contains a depth-wise separable convolution operation with a stride of 2 and a filter size of 3×3, which downsamples the input feature map resolution by 2 times and expands the number of input feature map channels to 60, 80, and 120 respectively.

[0056] In an embodiment of the present invention, the encoder module includes a convolution position encoding, a first layer normalization operation, a pooling calibration attention module, a second layer normalization operation and a feedforward neural network layer connected in series in sequence, and the input of the first layer normalization operation is added to the output of the pooling calibration attention module, and the input of the second layer normalization operation is added to the output of the feedforward neural network through a skip connection layer;

[0057] Among them, the convolution position encoding operation is:

[0058]

[0059] In the formula, X 2 is the 4x patch embedding or 2x patch embedding in the feature extraction module or the output feature map of the encoder module, Output feature map for convolution position encoding, Conv 3×3 is a convolution operation with a filter size of 3×3.

[0060] In an embodiment of the present invention, the pooling calibration attention module includes a feature calibration module based on style pooling, a multi-head transposed self-attention module and a convolutional relative position encoding module; wherein the feature calibration module based on style pooling is used to input the query Q, key K, value V matrix formed by linear mapping of the output feature map of the first layer normalization operation, and output the query Q after feature calibration * , key K * , value V * Matrix; The multi-head transposed self-attention module is used to input the feature-calibrated query Q * , key K * , value V * matrix, outputs the multi-head transposed self-attention module output feature map; the convolution relative position encoding module is used to input the feature-calibrated query Q output by the style pooling-based feature calibration module * , value V *Matrix, output convolution relative position encoding module output matrix; Pooling calibration attention The sum of the multi-head transposed self-attention module output feature map and the convolution relative position encoding module output matrix forms the pooling calibration attention module output feature map. In the pooling calibration attention module,

[0061] The feature calibration module based on style pooling is:

[0062]

[0063] In the formula, Q * , K * 、V * is the query, key, and value matrix after feature calibration; Q, K, and V are the query, key, and value matrices formed by linear mapping of the output feature map of the first layer normalization operation; σ(·) is the Sigmoid activation function; Conv(·) is a one-dimensional convolution operation with a convolution kernel size; [Ave, Max] # It is the style pooling operation in the channel dimension, including global average pooling Ave and global maximum pooling Max; # indicates different operation matrices; It is Hadamard.

[0064] The multi-head transposed self-attention module is,

[0065]

[0066] Where Att is the output feature map of the multi-head transposed self-attention module; Q * , K * 、V * is the query, key, and value matrix after feature calibration; d is the scaling factor; Softmax(·) is the exponential normalization function; (·) T Represents matrix transpose;

[0067] The convolutional relative position encoding module is,

[0068]

[0069] Where ReP is the output matrix of the convolutional relative position encoding module; DConv 3,5 Perform depthwise convolution operations with filter sizes of 3×3 and 5×5 for each channel; represents the Hadamard product.

[0070] The output feature map of the pooled calibration attention module, which is composed of the output feature map of the multi-head transposed self-attention module and the output matrix of the convolutional relative position encoding module, is:

[0071]

[0072] Where ConvAtt represents the output feature map of the pooled calibrated attention module.

[0073] In the embodiment of the present invention, the classification module uses a linear classification head to generate target categories; the linear classification head is a commonly used module in the field of image classification.

[0074] Step 3: Data preprocessing. Before the acoustic image and optical image are input into the network training in pairs, each image is first scaled to 224 pixels in width and height, so that the image size is 224×224 pixels. Finally, the image pixels are normalized from 0-255 to 0-1 to form the acoustic image tensor and optical image tensor, respectively.

[0075] Step 4: The paired acoustic image tensors and optical image tensors are formed into an acoustic-optical fusion image tensor after channel concatenation and channel shuffling, and then the acoustic-optical fusion feature map is obtained through the feature extraction module, and then the underwater target image classification result is output through the linear classification head module;

[0076] In addition, illustratively, in each step of training, back propagation is performed starting from the loss function value; the AdamW optimizer is used to optimize the network parameters according to the gradient information obtained by back propagation, thereby guiding the neural network to achieve accurate image classification results according to the input image.

[0077] In the following specific embodiments of the present invention: the optical image dataset used contains a total of 3745 optical images taken in underwater environments in 5 categories, and the acoustic image dataset used contains a total of 3745 acoustic images in 5 categories. The optical images are generated by a style transfer network, and the width and height of the images are 480-1920 pixels; the style transfer network uses a style transfer network based on a convolutional neural network; the acoustic dataset and the optical dataset are randomly divided into a training set consisting of 2452 pictures and a validation set consisting of 1293 pictures, respectively, and the images are first scaled to 224×224 pixels in the preprocessing stage; the initialization method of the image classification network in the embodiment of the present invention is random initialization; the operating environment is a computer with a framework such as PyTorch, which can read a given picture and complete the construction and training of the model of this method. The training time of the embodiment of the present invention on a Gold6626R@2.90GHz CPU, 8G memory and NVIDIA GeForce RTX3090 GPU is about 1 hour.

[0078] In an embodiment of the present invention, the specific implementation steps include: first setting relevant training parameters, setting the optimizer used for network update in the present invention to the AdamW optimizer, setting its momentum value to 0.05, setting the initial learning rate to 0.001, and setting the weight decay coefficient to 0.00001. Set the learning rate adjustment strategy to linear preheating and cosine simulated annealing, and the preheating cycle is 5. The deep image classification network of the embodiment of the present invention is composed of an acoustic-optical fusion module, a feature extraction module, and a linear classification head. The four stages of each feature extraction network include 2, 2, 6, and 2 encoder modules respectively; the input of the network is an RGB image (i.e., a color image), which is processed into an acoustic-optical fusion image tensor in the acoustic-optical fusion module, and the texture and abstract semantic information of the acoustic-optical fusion image tensor are extracted by the feature extraction module. The feature map corresponding to the output is subjected to multiple stages of feature extraction to obtain an acoustic-optical fusion feature map, which is then passed to the linear classification head module to obtain a classification result image containing the target category and probability. Next, when using the divided data set for network training, 64 pictures are randomly selected from the training set and input into the network each time, and the selected AdamW optimizer is used to update the parameters. The training is completed after 30 rounds of iteration on the data set. Finally, the pictures in the validation set are input into the trained network for classification, and the results of the embodiment of the invention are obtained, as shown in FIG. Figure 5 As shown, Figure 5 (a) is a schematic diagram of the classification results of reefs category image samples. Figure 5 (b) is a schematic diagram of the classification results of the shipwrecks category image sample. Figure 5 (c) is a schematic diagram of the classification results of fishes category image samples. Figure 5 (d) is a schematic diagram of the classification results of the (Autonomous Underwater Vehicles) AUVs category image samples. Figure 5 (e) is a schematic diagram of the classification result of frogmen category image samples. From the categories and probabilities of underwater targets given in the embodiment, it can be seen that the method proposed in the present invention achieves a good underwater target image classification result. Among the five underwater target image samples given, the underwater category is correctly given with a high probability, which reflects the effectiveness of the present invention in the task of accurately classifying underwater images.

[0079] In summary, an embodiment of the present invention discloses an underwater target classification method based on an acoustic-optical image fusion Transformer, in which the acoustic-optical fusion module is used for multimodal image fusion, and the feature extraction module is used for multimodal image tensor feature extraction and feature mixing; by introducing the fusion of acoustic features and optical features, the diversity of features is enhanced. The method of the present invention introduces visual cognitive computing into the field of automatic classification, which can effectively improve the accuracy of underwater target image classification, and each module is simple to implement, without excessive dependence, and has strong applicability. The underwater target classification method based on the acoustic-optical image fusion Transformer provided by the present invention has an implementation result of the calibration of the target type in the image, which can solve the technical problem of underwater target image misclassification caused by insufficient feature extraction of a single modality image and insufficient consideration of feature richness in the above-mentioned prior art.

[0080] The following are device embodiments of the present invention, which can be used to implement the method embodiments of the present invention. For details not disclosed in the device embodiments, please refer to the method embodiments of the present invention.

[0081] See also Figure 6 In an embodiment of the present invention, a system for underwater target classification based on acoustic-optical image fusion Transformer is provided, comprising:

[0082] A data acquisition module, which is used to obtain acoustic images and optical images of underwater targets to be classified;

[0083] A deep image classification network underwater target classification module is used to classify underwater targets based on the acoustic image and optical image of the underwater target to be classified using a pre-trained deep image classification network to obtain a classification result;

[0084] in,

[0085] The deep image classification network includes an acoustic-optical fusion module, a feature extraction module and a linear classification head connected in sequence;

[0086] The acoustic-optical fusion module first performs image preprocessing on the acoustic image and the optical image, and converts them into acoustic image tensors and optical image tensors respectively; then, channel splicing and shuffling are performed on the paired acoustic image tensors and optical image tensors to obtain the acoustic-optical fusion image tensors; the feature extraction module adopts a multi-stage encoder structure, and the encoder of each stage includes a serial patch embedding and a number of encoder modules based on a pooling calibration attention mechanism;

[0087] The encoder module includes a convolution position encoding module, a first layer normalization operation, a pooling calibration attention module, a second layer normalization operation and a feedforward neural network layer connected in series, and the input of the first layer normalization operation is added to the output of the pooling calibration attention module, and the input of the second layer normalization operation is added to the output of the feedforward neural network through a skip connection layer;

[0088] The pooling calibration attention module includes a feature calibration module based on style pooling, a multi-head transposed self-attention module and a convolutional relative position encoding module; wherein the feature calibration module based on style pooling is used to input the query Q, key K, value V matrix formed by linear mapping of the output feature map of the first layer normalization operation, and output the query Q after feature calibration * , key K * , value V * Matrix; The multi-head transposed self-attention module is used to input the feature-calibrated query Q * , key K * , value V * matrix, outputs the multi-head transposed self-attention module output feature map; the convolution relative position encoding module is used to input the feature-calibrated query Q output by the style pooling-based feature calibration module * , value V * Matrix, output convolution relative position encoding module output matrix; the sum of the multi-head transposed self-attention module output feature map and the convolution relative position encoding module output matrix forms the pooled calibration attention module output feature map.

[0089] In one embodiment of the present invention, a computer device is provided, the computer device comprising a processor and a memory, the memory is used to store a computer program, the computer program comprises program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, which are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used to perform the operation of the underwater target classification method based on the acoustic-optical image fusion Transformer.

[0090] In one embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the underwater target classification method based on the acoustic-optical image fusion Transformer in the above embodiment.

[0091] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0092] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0093] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. An underwater target classification method based on acoustic-optical image fusion Transformer, characterized in that: The following steps are involved: Acquiring acoustic and optical images of underwater targets to be classified; Based on the acoustic image and the optical image of the underwater target to be classified, a pre-trained deep image classification network is used to classify the underwater target to obtain a classification result; in, The deep image classification network includes an acoustic-optical fusion module, a feature extraction module and a linear classification head connected in sequence; The acoustic-optical fusion module first performs image preprocessing on the acoustic image and the optical image, and converts them into acoustic image tensors and optical image tensors respectively; then, channel splicing and shuffling are performed on the paired acoustic image tensors and optical image tensors to obtain the acoustic-optical fusion image tensors; the feature extraction module adopts a multi-stage encoder structure, and the encoder of each stage includes a serial patch embedding and a number of encoder modules based on a pooling calibration attention mechanism; The encoder module includes a convolution position encoding module, a first layer normalization operation, a pooling calibration attention module, a second layer normalization operation and a feedforward neural network layer connected in series, and the input of the first layer normalization operation is added to the output of the pooling calibration attention module, and the input of the second layer normalization operation is added to the output of the feedforward neural network through a skip connection layer; The pooling calibration attention module includes a feature calibration module based on style pooling, a multi-head transposed self-attention module and a convolutional relative position encoding module; wherein the feature calibration module based on style pooling is used to input the query Q, key K, value V matrix formed by linear mapping of the output feature map of the first layer normalization operation, and output the query Q after feature calibration * , key K * , value V * Matrix; The multi-head transposed self-attention module is used to input the feature-calibrated query Q * , key K * , value V * matrix, outputs the multi-head transposed self-attention module output feature map; the convolution relative position encoding module is used to input the feature-calibrated query Q output by the style pooling-based feature calibration module * , value V * Matrix, output convolution relative position encoding module output matrix; the sum of the multi-head transposed self-attention module output feature map and the convolution relative position encoding module output matrix forms the pooled calibration attention module output feature map.

2. According to claim 1, the underwater target classification method based on acoustic-optical image fusion Transformer is characterized in that: The acoustic-optical fusion module performs channel splicing and channel shuffling on the acoustic image tensor and the optical image tensor of the same size and the same number of channels to obtain an acoustic-optical fusion image tensor; wherein the channel shuffling is to orderly shuffle the acoustic image tensor and the optical image tensor according to the channels, so that the channels of the acoustic-optical fusion image tensor are alternating acoustic image tensor layers and optical image tensor layers.

3. The underwater target classification method based on acoustic-optical image fusion Transformer according to claim 2 is characterized in that: In the deep image classification network, the calculation of the sound and light fusion module is expressed as: Where, X a is the acoustic image tensor, X o is the optical image tensor, X ao is the concatenated tensor of the acoustic image tensor and the optical image tensor, Cat is channel concatenation; Shuffle is channel shuffle; X is the acoustic-optical fusion image tensor; X a1 ,X a2 ,X a3 Represents channel 1, channel 2, channel 3, X of the acoustic image tensor respectively. o1 ,X o2 ,X o3 Represent channel 1, channel 2, and channel 3 of the optical image tensor respectively.

4. The underwater target classification method based on acoustic-optical image fusion Transformer according to claim 1 is characterized in that: The style pooling-based feature calibration module includes three feature calibration networks for query Q, key K, and value V matrices based on style pooling, respectively; each feature calibration network includes a trunk network and a branch network, the operation of the trunk network is identity mapping, and the operations of the branch network are global average pooling, global maximum pooling, one-dimensional convolution, and sigmoid activation function operations in sequence; the broadcast-based Hadamard product of the trunk network output and the branch network output is used as the output of the feature calibration network; The outputs of the three feature calibration networks are the feature-calibrated query Q * , key K * , value V * matrix.

5. The underwater target classification method based on acoustic-optical image fusion Transformer according to claim 4 is characterized in that: The calculation of the pooling calibration attention module is expressed as: in, In the formula, Q * , K * 、V * is the query, key, and value matrix after feature calibration; Q, K, and V are the query, key, and value matrices formed by linear mapping of the output feature map of the first layer normalization operation before feature calibration; σ(·) is the Sigmoid activation function; Conv(·) is a one-dimensional convolution operation with a convolution kernel size; [Ave, Max] Q 、[Ave,Max] K 、[Ave,Max] V They are style pooling operations for the channel dimensions of the Q, K, and V matrices, including global average pooling Ave and global maximum pooling Max; is the Hadamard product; DConv 3,5 The depth convolution operation with filter size of 3×3 and 5×5 is performed for each channel; d is the scaling factor; Softmax(·) is the exponential normalization function; (·) T is the matrix transpose; ConvAtt is the output of the pooled calibration attention module.

6. The underwater target classification method based on acoustic-optical image fusion Transformer according to claim 1 is characterized in that: The feature extraction module includes stage 1, stage 2, stage 3, and stage 4 connected in sequence; the stage 1 includes 4-fold patch embedding and a plurality of encoder modules connected in series in sequence; the stage 2, the stage 3, and the stage 4 each include 2-fold patch embedding and a plurality of encoder modules connected in series in sequence; The 4x patch embedding is used to downsample the input acousto-optic fusion image tensor by 4x and output a 4x downsampled feature map; The 2x patch embedding is used to downsample the output feature map of the encoder module of the previous stage by 2x, and output a 2x downsampled feature map; The encoder module is used to encode the 4x or 2x downsampled feature map, and the output encoder module outputs the feature map.

7. The underwater target classification method based on acoustic-optical image fusion Transformer according to claim 1 or 6, characterized in that: In the encoder module: The convolution position encoding module is used to input the input feature map of the encoder module and output the convolution position encoding output feature map; The first layer normalization operation is used to input the convolution position encoding output feature map, and output the first layer normalization operation output feature map; The pooling calibration attention module is used to input the first layer normalization operation output feature map, and output the pooling calibration attention module output feature map; The second layer normalization operation is used to input the sum of the output feature map of the pooling calibration attention module and the output feature map of the convolution position encoding, and output the output feature map of the second layer normalization operation; The feedforward neural network layer is used to input the second layer normalization operation output feature map, and output the feedforward neural network output feature map; The encoder module output feature map is the sum of the feedforward neural network output feature map and the second layer normalization operation input feature map.

8. An underwater target classification system based on acoustic-optical image fusion Transformer, characterized in that: include: A data acquisition module, which is used to obtain acoustic images and optical images of underwater targets to be classified; A deep image classification network underwater target classification module is used to classify underwater targets based on the acoustic image and optical image of the underwater target to be classified using a pre-trained deep image classification network to obtain a classification result; in, The deep image classification network includes an acoustic-optical fusion module, a feature extraction module and a linear classification head connected in sequence; The acoustic-optical fusion module first performs image preprocessing on the acoustic image and the optical image, and converts them into acoustic image tensors and optical image tensors respectively; then, channel splicing and shuffling are performed on the paired acoustic image tensors and optical image tensors to obtain the acoustic-optical fusion image tensors; the feature extraction module adopts a multi-stage encoder structure, and the encoder of each stage includes a serial patch embedding and a number of encoder modules based on a pooling calibration attention mechanism; The encoder module includes a convolution position encoding module, a first layer normalization operation, a pooling calibration attention module, a second layer normalization operation and a feedforward neural network layer connected in series, and the input of the first layer normalization operation is added to the output of the pooling calibration attention module, and the input of the second layer normalization operation is added to the output of the feedforward neural network through a skip connection layer; The pooling calibration attention module includes a feature calibration module based on style pooling, a multi-head transposed self-attention module and a convolutional relative position encoding module; wherein the feature calibration module based on style pooling is used to input the query Q, key K, value V matrix formed by linear mapping of the output feature map of the first layer normalization operation, and output the query Q after feature calibration * , key K * , value V * Matrix; The multi-head transposed self-attention module is used to input the feature-calibrated query Q * , key K * , value V * matrix, outputs the multi-head transposed self-attention module output feature map; the convolution relative position encoding module is used to input the feature-calibrated query Q output by the style pooling-based feature calibration module * , value V * Matrix, output convolution relative position encoding module output matrix; the sum of the multi-head transposed self-attention module output feature map and the convolution relative position encoding module output matrix forms the pooled calibration attention module output feature map.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the underwater target classification method based on the acoustic-optical image fusion Transformer according to any one of claims 1 to 8 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the underwater target classification method based on acoustic-optical image fusion Transformer according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Emotion recognition method and device, equipment and medium

    CN114168823A

  • Microalgae image classification method based on feature calibration Transform and related device

    CN118657999A

  • Road drivable area detection method and system capable of learning depth position code guidance

    CN118675128A

  • Sensor fusion

    US20230237783A1

  • Coarse-to-fine heterologous image matching method based on edge guidance

    WO2024148969A1