An underwater image super-resolution and target recognition method

Through super-resolved target recognition model and innovative training strategies, the problems of low underwater image quality and difficult target recognition are solved, and image clarity and recognition performance are improved, which are suitable for underwater unmanned detection technology.

CN119067848BActive Publication Date: 2025-07-29UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411154805.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2025-07-29
Estimated Expiration
2044-08-22

AI Technical Summary

Technical Problem

The quality of underwater image has been severely reduced, the processing efficiency of traditional methods is low, the identification of underwater biological targets is difficult, the data set is scarce, the algorithm accuracy and speed are insufficient, the robustness is poor, and the generalization ability is limited.

Method used

The super-resolution target recognition model, including encoder, decoder and feature pyramid network, is adopted, combined with orthogonal bidirectional attention module and upper scaler, through convolutional layer and depth feature extraction, using training strategies of similar graph consistency loss and cross-label supervision, to improve image clarity and recognition performance.

Benefits of technology

It significantly improves the signal-to-noise ratio and clarity of underwater images, enhances the target recognition performance, alleviates the scarcity of data sets, and improves the robustness and computing efficiency of the algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119067848B_ABST
    Figure CN119067848B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of underwater imaging detection, and discloses an underwater image super-resolution and target recognition method, which specifically includes: The encoder includes a convolutional layer and a deep feature extraction part; The convolutional layer is used to downsample the input underwater image to extract initial features, and then the features are sequentially input into the deep feature extraction part composed of multiple orthogonal bidirectional attention modules; The features are sequentially input into an upscaler and a convolutional layer to obtain a reconstructed image of the underwater image; The features output by the encoder are input into a feature pyramid network; After connecting multiple aligned regions of interest, they are input into a convolutional layer to obtain the classification result of the target and the bounding box of the target. It can significantly improve the signal-to-noise ratio and clarity of the image, laying a foundation for the improvement of target recognition performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater imaging detection, and in particular to an underwater image super-resolution and target recognition method. Background Art

[0002] Underwater imaging detection technology has become a research hotspot amidst the rapid growth of underwater resource development, environmental monitoring, and maritime military applications. Breakthroughs in deep learning have injected new vitality into this field. As a key technology for intelligent underwater detection, research on underwater image target detection is increasingly active and widely applied in a variety of tasks, including aquatic biological detection, water environment exploration, seabed modeling, salvage and rescue, submarine pipeline detection, mine countermeasures, and anti-submarine warfare. This research not only broadens the boundaries of underwater detection but also provides strong technical support for further development in related fields. However, due to significant differences between underwater and terrestrial environments, particularly the absorption and scattering of light by water, the quality of underwater images captured by electronic devices is severely degraded, manifested by defects such as severe color distortion, loss of detail, reduced contrast, and blurring, significantly hindering accurate information acquisition. Traditional underwater image processing methods, such as enhancement, restoration, and super-resolution reconstruction algorithms, are generally limited by their reliance on degradation models and low processing efficiency, limiting their application scope and real-time performance.

[0003] Currently, a major constraint on the large-scale application of unmanned underwater detection technology is the performance shortcomings of detection algorithms and the lack of underwater image datasets. In most cases, human intervention is still required to assist in processing. Furthermore, underwater biological targets are small and densely distributed, and the presence of overlapping occlusions further complicates target identification. Therefore, improving algorithm accuracy and speed, expanding underwater image datasets, enhancing robustness in complex environments, broadening the algorithm's generalization capabilities, and optimizing the model's computational efficiency have become key challenges that need to be overcome in this field. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides an underwater image super-resolution and target recognition method.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0006] A method for underwater image super-resolution and target recognition, wherein the super-resolution target recognition model adopted includes an encoder, a decoder, and a feature pyramid network; specifically, the following steps are included:

[0007] Step 1: The encoder includes a convolutional layer and a deep feature extraction part. The convolutional layer is used to downsample the input underwater image I to extract the initial feature F0, and then F0 is successively input into the deep feature extraction part composed of A orthogonal bidirectional attention modules. The A orthogonal bidirectional attention modules respectively output features F1, F2, …, F A , and F0, F1, F2, …, F A are concatenated to obtain the feature C1 output by the encoder;

[0008] Step 2: The decoder includes an upscaler and a convolutional layer. The feature C1 is successively input into the upscaler and the convolutional layer to obtain the reconstructed image of the underwater image

[0009] Step 3: The feature C1 output by the encoder is input into the feature pyramid network. The feature pyramid network includes a bottom-up path, a top-down path, and a lateral connection path. The bottom-up path uses a deep residual network. The region proposal network is used to extract regions of interest from each feature map in the top-down path and align them. After connecting multiple aligned regions of interest, they are input into a convolutional layer to obtain the classification result of the target and the bounding box of the target.

[0010] Furthermore, the orthogonal bidirectional attention module includes multiple alternating convolutional layers and attention modules, which are used to alternately perform convolutional operations and attention operations on the input feature F;

[0011] The attention module includes two bidirectional long short-term memory networks. The first bidirectional long short-term memory network scans each pixel row by row along the feature F from left to right and from right to left to obtain two feature maps. The first bidirectional long short-term memory network concatenates the two feature maps and outputs them to combine the left and right context information of each pixel. The second bidirectional long short-term memory network scans each column of the feature map output by the first bidirectional long short-term memory network in a bottom-up and top-down order and obtains the overall feature map through concatenation. A convolutional layer is used to convert the overall feature map into D channels to form the feature weight α, where D = W × H, and W and H are the width and height of the input feature F respectively. For the feature vector p w,h ∈R D at the position (w, h) in the feature weight α, a convolutional layer is used to change the dimension to C × W × H and then normalize it using the softmax function to obtain a new feature weight α w,h ∈R D . Then, by weighting the new feature weight α w,h generated for each pixel, the context feature F att is obtained:

[0012] The new feature weight of the i-th channel is obtained by the following formula:

[0013]

[0014] is the feature vector of the pixel at position (w, h) in the feature weight α of the i-th channel;

[0015] For the context feature F of the same size as F participated by the pixel at position (w, h) att is obtained by the following formula:

[0016]

[0017] where f w,h ∈R D is the convolutional feature at position (w, h) in F, is the feature at position (w, h) in F att .

[0018] Furthermore, the upscaler includes a plurality of convolutional groups, and each convolutional group contains a normal convolution and a sub-pixel convolution; Let Conv(n c , k, n f ) be a convolutional layer, where n c , k, n f are the number of filter channels, kernel size, and number of filters respectively; For the input feature containing n c feature maps, the normal convolution Conv(n c , 3, r 2 ·n f ) increases the number of feature maps of the input feature to r 2 ·n f , obtaining a feature tensor of size H × W × r 2 ·n f , where r represents the upscaling factor (the magnification can be customized); H and W respectively represent the height and width of the feature maps of the input feature; The sub-pixel convolution PixelShuffle(r) rearranges the elements in the feature tensor of size H × W × r 2 ·n f into a feature tensor of size (r·H) × (r·W) × n f .

[0019] Furthermore, the training strategy of the super-resolution target recognition model includes:

[0020] Pairwise similarity Figure 1 consistency loss:

[0021] Use two super-resolution object recognition models with the same structure but different initialization methods. The encoder and decoder constitute an autoencoder network. Define the autoencoder networks of the two super-resolution object recognition models as the first autoencoder network A(θ1) and the second autoencoder network A(θ2) respectively. Denote the feature pyramid networks of the two super-resolution object recognition models as the first feature pyramid network F(θ1) and the second feature pyramid network F(θ2) respectively; A(θ1) and F(θ1) form the first super-resolution object recognition model, and A(θ2) and F(θ2) form the second super-resolution object recognition model;

[0022] In the autoencoder network, for the feature map P from the encoder e , that is, the feature C1, and the feature map P output by the decoder d , that is, the reconstructed image calculate the similarity Figure 1 consistency loss;

[0023] For a given batch size B, the shapes of P e and P d are B×C×W×H and B×C'×W'×H' respectively, where H and W are the width and height of the feature map P e respectively, and C is the corresponding number of feature channels; W' and H' are the width and height of the feature map P d respectively, and C' is the corresponding number of feature channels; To calculate the similarity Figure 1 consistency loss, first obtain the similarity map matrix, that is, the Gram matrix, and return the similarity values between two pairs of inputs according to the outputs of the autoencoder networks A(θ1), A(θ2); To reduce the influence of the variation between the similarity values, perform L2 normalization on the similarity map matrix to obtain the pairwise similarity map matrix ψ with the shape of B×B, that is, for the i-th super-resolution object recognition model:

[0024]

[0025] ‖·‖2 represents the L2 norm, i∈{1,2}, and ψ1 and ψ2 are the similarity map matrices generated by the autoencoder networks A(θ1), A(θ2) respectively;

[0026] The proposed similarity Figure 1 consistency loss l sim is:

[0027]

[0028] L D (·,·) is the root mean square error, respectively represent the similarity map matrices of the feature map P e in A(θ1), A(θ2), Feature maps P output by A(θ1) and A(θ2) respectively d Similarity graph matrix of

[0029] Cross-label supervision:

[0030] Take the classification results output by the first super-resolution target recognition model and the second super-resolution target recognition model as pseudo-labels Y1 and Y2 respectively, and use the pseudo-labels Y1 and Y2 as supervision signals: use the pseudo-label Y2 as the supervision of the first super-resolution target recognition model, use the pseudo-label Y1 as the supervision of the second super-resolution target recognition model, and jointly constrain with the cross-entropy loss function and the Dice loss function to obtain the loss of the first super-resolution target recognition model And the loss of the second super-resolution target recognition model

[0031]

[0032] where α is a hyperparameter that balances the contributions of the two loss terms; l ce (·,·) is the cross-entropy loss function; l dice (·,·) is the Dice loss function; i ∈ {1, 2}; then the total cross-label supervision loss l Y is:

[0033]

[0034] The total loss function of the training strategy described above is:

[0035]

[0036] Compared with the prior art, the beneficial technical effects of the present invention are:

[0037] The present invention proposes a novel and efficient underwater image super-resolution and target recognition method, which can significantly improve the signal-to-noise ratio and clarity of images, laying a foundation for the improvement of target recognition performance. At the same time, the present invention also designs an innovative model training strategy to effectively alleviate the problem of scarcity of underwater image datasets, thus meeting the urgent needs of practical applications. These series of innovative measures can lay a solid foundation for the wide application of underwater unmanned detection technology. Brief Description of the Drawings

[0038] Figure 1 Schematic diagram of the architecture of the super-resolution target recognition model proposed by the present invention;

[0039] Figure 2 Schematic diagrams of the orthogonal bidirectional attention module and the attention module in the embodiments of the present invention;

[0040] Figure 3 Schematic diagram of the upscaler in the embodiment of the present invention;

[0041] Figure 4 Principle diagram of the upscaler in the embodiment of the present invention;

[0042] Figure 5 Schematic diagram of the training strategy in the embodiment of the present invention. Detailed implementation manners

[0043] A preferred implementation manner of the present invention will be described in detail below with reference to the accompanying drawings.

[0044] 1. Overall architecture:

[0045] The architecture of the super-resolution target recognition model proposed by the present invention is as Figure 1 shown. The super-resolution target recognition model of the present invention is built on the classical autoencoder structure, adopts a dense upsampling framework, performs initial feature extraction and deep feature extraction in the encoder part, and performs image reconstruction in the decoder.

[0046] Taking a fish school as an underwater target as an example, for the input underwater image I containing a fish school, first use a convolutional layer (Conv) to downsample the image to extract the initial feature F0, and then send F0 to the deep feature extraction part composed of A orthogonal bidirectional attention modules (Orthogonal Bidirectional Attention, OBA). The features extracted at different levels include F0, F1, F2, …, F A , are all transferred to the decoder for image super-resolution reconstruction. The decoder consists of an upscaler Upscaler and a convolutional layer Conv, which is the reconstructed image output by the decoder.

[0047] The feature C1 output by the encoder is sent into the bottom-up path in the feature pyramid network (FPN). The bottom-up path uses a deep residual network (ResNet), and a region proposal network (RPN) is used in the top-down path to form region of interest extraction and then connection. Then, the connected features are fused through a convolutional layer to obtain the classification of the target and the bounding box regression of the target, realizing target recognition.

[0048] 2. Orthogonal bidirectional attention module (OBA):

[0049] As Figure 2 shown, in the orthogonal bidirectional attention module, convolutional operations and attention operations alternate to further extract features from the input feature F.

[0050] In the attention module, the bidirectional long short-term memory network (LSTM) scans each pixel row by row from left to right and from right to left along each row of F respectively, and then combines the two obtained feature maps to combine the left context information and right context information of each pixel. Similarly, another bidirectional LSTM is used to scan each column of the obtained feature map in a bottom-up and top-down order. After obtaining the overall feature map through scanning in four directions, a convolutional layer is used to convert the feature map into D channels, where D = w × h, to form the feature weight α, and then the final attention weight F is obtained by weighting the weights generated for each pixel. att For the feature vector p of the pixel at position (w, h) w,h ∈R D , after normalization using the softmax function, an attention weight α w,h ∈R D is obtained. Therefore, for the context feature F of the same size as F participated by the pixel at position (w, h) att , it can be obtained through the following formula:

[0051]

[0052] where f w,h ∈R D is the convolutional feature at position (w, h) in F, is the feature at position (w, h) in f att .

[0053] 3. Upscaler:

[0054] As shown in Figure 3 and Figure 4 , the upscaler consists of multiple convolutional groups, and each convolutional group contains a normal convolution and a sub-pixel convolution. For the input feature containing N c feature maps, the normal convolution Conv(n c , 3, r 2 ·n f ) increases the number of feature maps to R 2 ·N f . The sub-pixel convolution PixelShuffle(r) rearranges the elements in the H×W×r 2 ·n f feature tensor into a feature tensor of size (r·H)×(r·W)×n f .

[0055] 4. Training strategy:

[0056] 4.1. Pairwise similarity Figure 1 consistency loss:

[0057] like Figure 5 As shown in the figure, the present invention uses two super-resolution target recognition models with the same structure but different random initialization methods. For the sake of simplicity, the lower autoencoder network is defined as A(θ1), A(θ2), and the upper feature pyramid network is defined as F(θ1), F(θ2). In the autoencoder network, based on the feature map P from the encoder, e , i.e., feature C1, and the decoder output P d , that is, reconstructing the image Calculating Similarity Figure 1 Consistency loss. For a given batch size B, these outputs have shapes B×C×W×H and B×C'×W'×H', where H and W are the spatial dimensions of the feature map and C is the number of feature channels. The similarity map matrix is a Gram matrix that returns the similarity value between two pairs of inputs based on the network output. To reduce the impact of changes between values, the similarity map is L2 normalized. This calculation results in a pairwise similarity map matrix (ψ) of shape B×B. That is, for each network:

[0058]

[0059] Its training loss is

[0060]

[0061] L D is the root mean square error, ψ1 and ψ2 are the similarity graph matrices generated by networks A(θ1) and A(θ2), respectively.

[0062] resemblance Figure 1 The role of consistency training is equivalent to adding multiple perturbations to the features learned by the encoder and forcing the output of the decoder to be consistent with the output before the perturbation. It is used to cope with complex underwater environments, which in disguise increases the sample size and improves the robustness of the algorithm.

[0063] 4.2, Cross-label Supervision:

[0064] like Figure 5 As shown in the figure, when the two networks output the classification results, they obtain corresponding classification labels Y1 and Y2, which are used as supervision signals. Classification label Y2 is used as supervision for network 1, and Y1 is used as supervision for network 2, and the cross entropy loss function is used for constraint.

[0065] It is apparent to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention, and any reference signs in the claims should not be construed as limiting the claims involved.

[0066] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment contains only one independent technical solution. This narrative manner of the specification is merely for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in the various embodiments can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. An underwater image super-resolution and target recognition method, characterized in that The super-resolution object recognition model adopted includes an encoder, a decoder, and a feature pyramid network; specifically, it includes the following steps: Step 1, the encoder includes a convolutional layer and a deep feature extraction part; the convolutional layer is used to downsample the input underwater image I to extract the initial feature F0, and then F0 is sequentially input into the deep feature extraction part composed of A orthogonal bidirectional attention modules; the A orthogonal bidirectional attention modules respectively output the features F1, F2, …, F A , for F0, F1, F2, …, F A are concatenated to obtain the feature C1 output by the encoder; Step 2: The decoder includes an upscaler and a convolutional layer; the feature C1 is sequentially input into the upscaler and the convolutional layer to obtain the reconstructed image of the underwater image Step 3: Input the feature C1 output by the encoder into the feature pyramid network; the feature pyramid network includes a bottom-up path, a top-down path, and a lateral connection path; the bottom-up path uses a deep residual network; use a region selection network to extract regions of interest from each feature map in the top-down path and perform alignment; after connecting multiple aligned regions of interest, input them into a convolutional layer to obtain the classification result of the object and the bounding box of the object; The orthogonal bidirectional attention module includes multiple alternating convolutional layers and attention modules, and is used to alternately perform convolutional operations and attention operations on the input feature F; The attention module includes two bidirectional long short-term memory networks. The first bidirectional long short-term memory network scans each pixel row by row along the feature F, from left to right and from right to left, to obtain two feature maps. The first bidirectional long short-term memory network splices the two feature maps and outputs them to combine the left context information and the right context information of each pixel. The second bidirectional long short-term memory network scans each column of the feature map output by the first bidirectional long short-term memory network in a bottom-up and top-down order, and obtains the overall feature map through splicing; the convolutional layer is used to convert the overall feature map into D channels to form the feature weight α, where D = W × H, and W and H are the width and height of the input feature F respectively; for the feature vector p of the pixel at the position (w, h) in the feature weight α w,h ∈R D , the convolutional layer is used to change the dimension to C × W × H and then the softmax function is used for normalization to obtain a new feature weight α w,h ∈R D , and then by weighting the new feature weight α w,h generated for each pixel, the context feature F att is obtained: The new feature weight of the i-th channel Obtained by the following formula: is the feature vector of the pixel at position (w, h) in the feature weight α of the i-th channel; The context feature F of the same size as F participated by the pixel at the position (w, h) att is obtained by the following formula: where, f w,h ∈R D is the convolutional feature at position (w, h) in F, is F att the feature at position (w, h) therein; The training strategy of the super-resolution object recognition model includes: (1) Calculate the pairwise similarity map consistency loss: Define the autoencoder networks of two super-resolution object recognition models as the first autoencoder network A(θ1) and the second autoencoder network A(θ2) respectively. The structures of the two super-resolution object recognition models are the same but the initialization methods are different. The encoder and decoder constitute the autoencoder network. The feature pyramid networks of the two super-resolution object recognition models are respectively denoted as the first feature pyramid network F(θ1) and the second feature pyramid network F(θ2); A(θ1) and F(θ1) form the first super-resolution object recognition model, and A(θ2) and F(θ2) form the second super-resolution object recognition model; In the autoencoder network, for the feature map P from the encoder e , that is, the feature C1, and the feature map P output by the decoder d , that is, the reconstructed image calculate the similarity map consistency loss; For a given batch size B, P e and P d are shaped as B×C×W×H and B×C'×W'×H' respectively, where H and W are the width and height of the feature map P e respectively, C is the corresponding number of feature channels; W' and H' are the width and height of the feature map P d respectively, C' is the corresponding number of feature channels; to calculate the similarity map consistency loss, first, a similarity map matrix, i.e., the Gram matrix, needs to be obtained, and the similarity values between two pairs of inputs are returned according to the outputs of the autoencoder networks A(θ1) and A(θ2); the similarity map matrix is L2-normalized to obtain a pairwise similarity map matrix ψ of shape B×B, i.e., for the i-th super-resolution target recognition model: ‖·‖2 represents the second norm, i ∈ {1, 2}, and ψ1 and ψ2 are the similarity map matrices generated by the autoencoder networks A(θ1) and A(θ2) respectively; The proposed similarity graph consistency loss l sim is as follows: L D (·,·) is the root mean square error, respectively representing the similarity graph matrices of the feature maps P in A(θ1) and A(θ2), e respectively representing the similarity graph matrices of the feature maps P output by A(θ1) and A(θ2); d ​​ (2) Perform cross-label supervision: Take the classification results output by the first super-resolution target recognition model and the second super-resolution target recognition model as the pseudo-labels Y1 and Y2 respectively, and use the pseudo-labels Y1 and Y2 as supervision signals: use the pseudo-label Y2 as the supervision of the first super-resolution target recognition model, use the pseudo-label Y1 as the supervision of the second super-resolution target recognition model, and jointly constrain them with the cross-entropy loss function and the Dice loss function to obtain the loss of the first super-resolution target recognition model and the loss of the second super-resolution target recognition model where α is a hyperparameter that balances the contributions of the two loss terms; l ce (·,·) is the cross-entropy loss function; l dice (·,·) is the Dice loss function; i ∈ {1, 2}; then the total cross-label supervision loss e Y is: The total loss function of the training strategy is as follows:

2. The underwater image super-resolution and target recognition method according to claim 1, characterized in that The upscaler includes a plurality of convolutional groups, and each convolutional group contains a normal convolution and a sub-pixel convolution; let Conv(n c , k, n f ) be a convolutional layer, where n c , k, n f are the number of filter channels, kernel size, and number of filters respectively; for an input feature containing n c feature maps, the normal convolution Conv(n c , 3, r 2 ·n f ) increases the number of feature maps of the input feature to r 2 ·n f , obtaining a feature tensor of size H×W×r 2 ·n f , where r represents the upscaling factor; H and W represent the height and width of the feature maps of the input feature respectively; the sub-pixel convolution PixelShuffle(r) rearranges the elements in the feature tensor of size H×W×r 2 ·n f into a feature tensor of size (r·H)×(r·W)×n f .

Citation Information

Patent Citations

  • Remote sensing image super-resolution reconstruction method and system based on multi-scale network, and medium

    CN118071602A

  • Apparatus for processing image super resolution using graph information and its method

    KR102682543B1