A Complex Object Recognition Method Based on Deep Twin Self-Attention Network

Through the deep twin self-attention network method, combined with sliding window self-attention transformation and residual connection, the feature balance problem of convolutional neural network in complex image scenarios is solved, and more efficient semantic segmentation and recognition are achieved.

CN119399545BActive Publication Date: 2025-07-18BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411572324.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-07-18
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

In complex image scenarios, convolutional neural networks are difficult to balance features of different scales, resulting in low recognition accuracy, inaccurate local positioning, insufficient exchange of global context information, and unbalanced classification recognition when processing large sample data.

Method used

A method based on a deep twin self-attention network is adopted, and feature extraction and fusion are performed through the twin encoder and decoder structure, combined with the sliding window self-attention transformation mechanism and residual connection, and the adaptive image scale processing and fusion loss function are used to achieve long-distance semantic information interaction and feature balance.

Benefits of technology

It improves the generalization and recognition accuracy of the network, balances local and global features, enhances semantic spatial correlation, solves the recognition challenges of convolutional neural networks in complex image scenarios, and achieves more efficient semantic segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399545B_ABST
    Figure CN119399545B_ABST
Patent Text Reader

Abstract

The present invention discloses a complex object recognition method based on a deep twin self-attention network, which relates to the technical field of object recognition methods. The present invention applies a twin network based on a sliding window multi-head self-attention transformation mechanism, which enhances the global modeling ability of the feature extraction network and the acquisition efficiency of long-distance semantic information, improves the underlying feature extraction ability and the abstraction effect of high-level semantic information; the twin network performs personalized modeling, improving the feature segmentation effect; the overall network forms a U-shaped structure, with a locally symmetric encoder-decoder structure, which supplements the global context information and strengthens the spatial association of semantic information; residual links are added to efficiently train deep networks and enhance the recognition accuracy; and by adding adaptive image feature scale input processing, the input image is subjected to adaptive scale transformation, achieving an accurate end-to-end semantic segmentation result and realizing multi-scale input to enhance the generalization ability and accuracy of the network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition methods, and in particular, to a complex target recognition method based on a deep twin self-attention network. Background Art

[0002] In complex image scenes, due to the phenomenon of same-spectrum different objects or same-object different spectra, the spectral characteristics of the same type of ground objects vary greatly, and the spectral characteristics of different types of targets overlap with each other. This makes the within-class variance increase and the between-class variance decrease, which will cause the confusion between image feature details and high-level semantic information and cannot be solved only by relying on the recognition of expert human eyes. Traditional algorithms such as color clustering methods cannot mine deeper high-level semantic information behind images, resulting in a shallow understanding of the local and global aspects of images and low task completion efficiency. Therefore, it is necessary to draw on artificial intelligence deep learning algorithms to complete tasks and identify the high-level information contained in images.

[0003] At the same time, with the development of digital vision technology, images in various fields are moving towards large data volumes and diverse types of information. The tasks of image recognition and semantic segmentation have great application value in many fields such as land resource surveys, geographic information mapping, autonomous driving navigation, and medical image detection. There is a practical need to accurately and effectively identify precise detail features from images and at the same time extract accurate semantic information and balance the types of information.

[0004] Artificial intelligence algorithms can extract and analyze specific explicit features, abstract and generalize them into high-level semantic information, and systematize the process of specific extraction and abstraction, enhancing efficiency and accuracy. At the same time, feature details will be lost during the process. The main task of the visual space is to segment and locate feature details, and the main task of the semantic space is to summarize and learn high-level concept information, which is where the contradiction lies.

[0005] In a large number of complex image scenes, problems such as the imbalance between the underlying features and high-level semantic information of images, confusion in the recognition of approximate features, inaccurate local positioning and segmentation positions, neglect of small-scale features, insufficient context information exchange due to multi-scale changes in the global context, and processing of large sample data volumes will be faced. These problems make the semantic segmentation task more challenging in grasping detail features and the balance of semantic information at multiple scales.

[0006] Common complex object recognition methods based on convolutional neural networks on the market currently use cascaded convolutional and pooling operations for feature extraction. Since the size of the convolutional kernel is usually small, such as 3×3 or 5×5, the receptive field of the convolutional operation is limited, which is equivalent to performing operations on a very small area of the image or feature map each time. In the shallower convolutional layers, only small-scale features can be extracted, and effective long-distance dependencies cannot be established. In the deeper convolutional layers, global semantic information can be obtained, but some details are lost. To solve the above problems, the present invention proposes a complex object recognition method based on a deep twin self-attention network. Summary of the Invention

[0007] The purpose of the present invention is to propose a complex object recognition method based on a deep twin self-attention network to solve the following problems existing in the prior art:

[0008] (1) There are different-scale features in complex image scenes, and convolutional neural networks have difficulty in balancing large-scale and small-scale features, ultimately affecting the accuracy of complex object recognition tasks;

[0009] (2) The imbalance between the underlying features and high-level semantic information of the image leads to inaccurate local positioning and segmentation positions, and detailed features are ignored;

[0010] (3) The classification and recognition imbalance in the image global due to insufficient exchange of context information and processing of large sample data volumes.

[0011] To achieve the above purpose, the present invention adopts the following technical solutions:

[0012] A complex object recognition method based on a deep twin self-attention network specifically includes the following steps:

[0013] S1. First, separate the high-dimensional image information into two-channel dimensions, and each channel is respectively input into the sub-encoding network for feature extraction. First, the features of each channel are input into the convolutional module of the twin encoding network to initially extract the image features. Then, input into the encoder (RSwinT-Encoder) multiple residual module structures Down modules containing residual connections, and the input picture first performs adaptive scale adaptive;

[0014] S2. Input the image processed in S1 into a series of sliding window self-attention transformation module (Swin Transformer) blocks to further extract and process the features, and then restore the size of the output feature map after processing by the image size adaptive module (Resize);

[0015] S3. At each stage, the hierarchical encoder of the Siamese network simultaneously generates feature maps of different depths through downsampling, and for the sub-encoding network with input channels containing richer information, records the features in each encoding stage and the features of the original image resolution.

[0016] S4. At the end of the encoding stage, fuse the channels of the two sub-network branches, and fuse and splice the feature maps extracted from each channel.

[0017] S5. Input the fused image obtained in S4 into multiple Up modules of the residual module structure containing residual connections in the decoder (RSwinT-Decoder). First, after the image is upsampled, it is gradually restored to the feature map of the input size.

[0018] S6. Use the feature fusion technology at each decoding stage to combine the high-resolution features of the hierarchical encoder in S3 and the upsampled features of the decoder (RSwinT-Decoder) in S5.

[0019] S7. At the decoding stage, after the feature fusion in S6, repeat the operations described in S1-S2.

[0020] S8. Input the feature map of the last stage obtained in S7 into the segmentation head, classify each pixel, divide the pixels into pre-determined categories, and obtain the semantic segmentation recognition result.

[0021] S9. Calculate the semantic segmentation recognition result obtained in S8 using the fusion loss function, and backpropagate to improve the network parameters.

[0022] Preferably, the specific content of S1 includes the following:

[0023] Separate the input high-dimensional image into two channels, input the same sub-encoding network to form a Siamese network. In the sub-encoding network, first input the image into 2 layers of CBL modules for preliminary feature extraction. The encoder (RSwinT-Encoder) includes multiple Down modules of the residual module structure containing residual connections. First, perform an adaptive scaleadaptive operation on the input image, complete the adaptive interpolation calculation, and transform the image size into the least common multiple of 2 and 7 to be suitable for the subsequent sliding window self-attention mechanism calculation.

[0024] Preferably, the specific content of S2 includes the following:

[0025] S2.1. Divide the input image into multiple small patches, then flatten the patches and convert them into vectors of a fixed dimension that can be processed through a linear layer, and add positional information to the feature vectors of each patch to ensure that the model can understand the spatial relationships between different patches; the feature vectors after linear embedding are further divided into multiple non-overlapping windows, the size of the windows is fixed, and self-attention calculations are performed independently within each window;

[0026] S2.2. Keep the sequence with positional encoding in S2.1 as the residual quantity and input it into the normalization layer LN , for each element of the input sequence, perform W - MSA multi-head self-attention calculation and add it to the residual quantity. The specific formula is expressed as:

[0027] (1)

[0028] In formula (1), the output sequence, x l-1 represents the input sequence;

[0029] S2.3. Keep the output quantity obtained in S2.2 as the residual and input it into the normalization layer LN , the multi-layer perceptron network MLP , and then add it to the residual quantity. The specific formula is expressed as:

[0030] (2)

[0031] In formula (2), x l represents the output sequence;

[0032] S2.4. Keep the output sequence in S2.3 as the residual and input it into the normalization layer LN , for each element of the input sequence, perform SW - SWA shifted window self-attention calculation and add it to the residual quantity. The specific formula is expressed as:

[0033] (3)

[0034] In formula (3), represents the output sequence;

[0035] S2.5. Use the output quantity obtained in S2.4 as the residual and the input quantity, and the process is the same as S2.3. The specific formula is expressed as:

[0036] (4)

[0037] In formula (4), x l+1 represents the output sequence;

[0038] The calculation method of self-attention inside the window is shown as follows:

[0039] (5)

[0040] In formula (5), for each element of the input sequence, a query vector matrix, a key vector matrix, and a value vector matrix are obtained by multiplying with different matrices, Q , K , V ∈ , B The value of is taken from the bias matrix ∈ (2M+1)(2M-1) , M 2 represents the number of patches of the window, d represents Q or K the dimension of;

[0041] S2.6. After a series of sliding window self-attention transformation modules (Swin Transformer), interpolation calculation is performed using the image size adaptive module (Resize) that restores the size, and the image size is adaptively restored to the input size to facilitate the design of the hierarchical feature extraction structure.

[0042] Preferably, the S3 specifically includes the following:

[0043] Feature downsampling is implemented through a convolutional layer with a stride of 2 and a convolution kernel of 3×3, including a regularization layer and an activation function LeakRelu , and the hierarchical encoder generates four different-depth feature maps of 1 / 2, 1 / 4, 1 / 8, and 1 / 16 through downsampling at each stage. And for the channels containing richer information separated in S1, the 1 / 2, 1 / 4, 1 / 8 features and the original image resolution features in each encoding stage of its encoding network are recorded as skip connections to supplement the high-resolution features of the decoder.

[0044] Preferably, the S4 specifically includes the following:

[0045] On the feature level, the feature maps of the two channel branches from the siamese network are merged by splicing, and the similarity characteristics learned by the siamese network are used to improve the segmentation performance, and it has high interpretability when processing highly similar images;

[0046] Preferably, the S5 specifically includes the following:

[0047] Feature upsampling is implemented through a convolutional layer with a step size of 2 and a convolutional kernel of 3×3, including a regularization layer and an activation function LeakRelu , and every time it goes to a shallower layer, its width and height are both twice that of the previous layer, and its number of channels is 1 / 2 of the previous layer, restoring the feature maps of 1 / 8, 1 / 4, and 1 / 2 times. Among them, the feature maps with restored resolution contain both rich detailed feature information and high-level semantic information extracted from the deep feature maps.

[0048] Preferably, S6 specifically includes the following content:

[0049] Deep features Figure 1 / 16 is restored to the size of the previous layer 1 / 8 through upsampling and fused with the high-resolution features at the corresponding encoding stage Figure 1 / 8, deep features Figure 1 / 8 is restored to the size of the previous layer 1 / 4 through upsampling and fused with the high-resolution features at the corresponding encoding stage Figure 1 / 4, deep features Figure 1 / 4 is restored to the size of the previous layer 1 / 2 through upsampling and fused with the high-resolution features at the corresponding encoding stage Figure 1 / 2, deep features Figure 1 / 2 is restored to the original resolution size 1 of the image through upsampling and fused with the selected and recorded high-resolution feature maps at the encoding stage. This skip connection uses the splicing method rather than simple addition to ensure the integrity and richness of information during feature transmission, improving the accuracy of segmentation and edge details.

[0050] Preferably, the output resolution obtained after the segmentation head processing in S8 is consistent with the original Figure 1 , and the value of each pixel is its class probability.

[0051] Preferably, S9 specifically includes the following content:

[0052] Fusing the calculations of two loss functions can improve the robustness and generalization ability of the model, reduce overfitting, provide more accurate gradient signals, and facilitate the model to better learn and adjust parameters. At the same time, because Lovasz Loss has differentiable properties in mathematics, it is more convenient to combine with other loss functions, thereby further improving the model performance; at the same time, when fusing the calculations of the two losses, the weights between them can be adjusted to balance the weights between different objectives; in this way, the optimization degree of the model on different task metrics can be flexibly controlled to meet the actual needs;

[0053] The loss function uses Soft Cross Entropy Loss and Lovasz Loss, which are fused with a weight ratio of 1:1. Among them, the soft cross entropy, that is, the cross entropy with label smoothing, can improve generalization. The specific formula for the soft cross entropy is as follows:

[0054] (6)

[0055] In formula (6), L SCE represents the value of the Soft Cross Entropy Loss function; y i represents the label of the i-th category in the true label, which is set as a soft label, y i ∈(0,1); P ( x i ) represents the probability of the i-th category predicted by the model;

[0056] The meaning of Lovasz Loss is to gradually optimize the prediction results to make them closer to the true labels. It not only focuses on correctly classified samples but also considers boundary samples and misclassified samples. Therefore, it can handle the problems of class imbalance and uneven difficulty of samples. The specific formula is as follows:

[0057] (7)

[0058] (8)

[0059] (9)

[0060] (10)

[0061] In formula (7), Δ Jc represents the loss function to be optimized; y* represents the ground truth; c represents the set of mispredicted pixels; M c represents the set where the network segmentation result does not match the label, M c The domain is {0,1} p , p represents the number of pixels;

[0062] In formula (8), the lovasz extension is used for smooth extension and is specifically implemented in multi-class segmentation; f i ( c ) represents thec Probability values after softmax-like; using a scoring function f i ( c ) to construct a pixel errors m i ( c ) vector;

[0063] In Equation (9), use the errors m ( c ) vector to construct an alternative Δ Jc of the loss function;

[0064] In Equation (10), in order to optimize the mIoU metric for evaluating all classes, average the above loss ( f ( c ))

[0065] Compared with the prior art, the present invention provides a complex target recognition method based on a deep twin self-attention network, having the following beneficial effects:

[0066] Compared with the common image semantic segmentation methods on the market, the present invention uses a new U-shaped symmetric encoder-decoder structure constructed based on the residual connection deep sliding window self-attention mechanism transformation, with scale adaptive embedded, and combines the fusion loss calculation, realizing long-term long-distance semantic information interaction, supplementing context information, effectively balancing the low-level features and high-level features, greatly improving the prediction efficiency of the network, not only solving the problem of the input size limitation of the Swin Transformer sliding window self-attention mechanism, but also improving the generalization, accuracy and class balance metrics of the network, enhancing the semantic space correlation while not losing the details of the feature space, and balancing the global and local recognition accuracies BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 is the grid structure diagram mentioned in Embodiment 1 of the present invention;

[0068] Figure 2 is the schematic diagram of the important module of the network structure mentioned in Embodiment 1 of the present invention;

[0069] Figure 3 is the schematic diagram of the self-attention mechanism calculation mentioned in Embodiment 1 of the present invention;

[0070] Figure 4 is the diagram of the comparative qualitative experimental results mentioned in Embodiment 1 of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0071] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0072] Embodiment 1:

[0073] Based on the summary of the existing complex target recognition methods based on convolutional neural networks and symmetric encoding-decoding structure networks, the present invention proposes an image semantic segmentation method based on a deep self-attention transformation twin network. More specifically, a complex target recognition method based on a deep twin self-attention network is proposed. The overall framework adopts a U-shaped structure, and the twin encoder network, decoder box and improved Swin Transformer are combined for feature extraction. The network comprehensively considers the global information and local feature relationships of the input data, uses the characteristics of the twin network to strengthen personalized modeling, further supplements the target location details, establishes long-distance connections and dependencies between features, so that a larger range of context semantic information can be explored and has stronger segmentation ability. The specific contents are as follows:

[0074] As Figure 1 、 Figure 2 shown, the twin deep self-attention residual U-shaped network is mainly divided into the Down module in the encoder (RSwinT-Encoder) and the Up module in the decoder (RSwinT-Decoder). And in the encoder (RSwinT-Encoder), a twin network is innovatively used to extract high-dimensional image features. Each stage is based on the improved Swin Transformer for feature extraction, and feature fusion is performed at the end of the encoding stage and at the skip connection part. Finally, it is input into the segmentation head, and the result is to label the semantic information for each pixel. Starting from the input picture, the specific steps are as follows:

[0075] Step 1: First, separate the high-dimensional image information into two-channel dimensions. Each channel image is first input into a 2-layer CBL module, that is, preliminary feature extraction is performed through a convolutional layer with a convolution kernel of 3×3, a regularization layer, and an activation layer. The input picture in the sub-encoding network then enters the encoder Down module. There are multiple residual module structures Down modules with residual connections in the encoder (RSwinT-Encoder). First, perform an adaptive scale adaptive operation on the picture. The adaptive interpolation calculation must transform the image size into a common multiple of 2 and 7, which meets the fixed size requirements of the Swin Transformer and is applicable to the subsequent sliding window self-attention mechanism calculation.

[0076] Step 2: Since the input of Swin Transformer is a fixed one-dimensional sequence rather than a two-dimensional image, the image needs to be converted into an input sequence. Meanwhile, to reduce the computational complexity and avoid too long input image sequences, the input image is divided into several small patches, which are flattened and then converted into vectors with a fixed dimension that can be processed through a linear layer. And position information is added to the feature vectors of each patch to ensure that the model can understand the spatial relationships between different patches. The feature vectors after linear embedding and position embedding are further divided into multiple non-overlapping windows, the sizes of which are fixed, and self-attention calculations are independently performed inside each window. As Figure 1 shown, meanwhile, to design the structural residual, the sequence with positional encoding is reserved as the residual quantity and input into the normalization layer , for each element of the input sequence, multi-head self-attention calculation is performed and added to the residual quantity, which can be expressed by the formula:

[0077] (1)

[0078] In formula (1), the output sequence, x l-1 represents the input sequence;

[0079] Furthermore, the output quantity is reserved as the residual and successively input into the normalization layer LN , the multi-layer perceptron network MLP , and then added to the residual quantity. The specific formula is expressed as:

[0080] (2)

[0081] In formula (2) represents the output sequence.

[0082] Furthermore, x l the output sequence is reserved as the residual and input into the normalization layer LN , for each element of the input sequence, SW - MSA shifted window self-attention calculation is performed and added to the residual quantity. The specific formula is expressed as:

[0083] (3)

[0084] In formula (3), represents the output sequence.

[0085] Furthermore, the output quantity is reserved as the residual, and the input process is also to successively input into the normalization layerLN , the multi - layer perceptron network MLP , and then added to the residual amount, and the specific formula is expressed as:

[0086] (4)

[0087] In formula (4), x l+1 represents the output sequence.

[0088] The calculation method of the above self - attention within the window is shown as follows:

[0089] (5)

[0090] In formula (5), for each element of the input sequence, the query vector matrix, key vector matrix, and value vector matrix are obtained by multiplying with different matrices, Q , K , V ∈ B The value of is taken from the bias matrix ∈ (2M+1)(2M-1) , M 2 represents the number of patches of the window, d represents Q or K the dimension of, and the self - attention calculation mechanism is as Figure 3 shown. After a series of improved Swin Transformers, through the Resize module for restoring the size, interpolation calculation is performed to adaptively restore the image size to the input size, which is beneficial to the design of the hierarchical feature extraction structure.

[0091] Step 3: Use the feature extraction network to extract feature maps with four different depths of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 for feature map fusion. Feature downsampling is implemented through a convolutional layer with a stride of 2 and a convolution kernel of 3×3, which includes a regularization layer and an activation function LeakRelu , and the hierarchical encoder generates feature maps with four different depths of 1 / 2, 1 / 4, 1 / 8, and 1 / 16 through downsampling at each stage. In the encoding stage, features of 1 / 2, 1 / 4, and 1 / 8 are cascaded. For the sub - encoding network with more information - rich channels, the feature maps of each depth will be retained to serve as skip connections to supplement the high - resolution features of the decoder.

[0092] Step 4: At the feature level, the feature maps of the two channel branches from the siamese encoding network are merged by concatenation. The similarity characteristics learned by the siamese network are used to improve the segmentation performance, and it has high interpretability when dealing with highly similar images;

[0093] Step 5: The deepest 1 / 16 feature map after fusion in Step 4 is subjected to feature upsampling through a convolutional layer with a stride of 2 and a convolution kernel of 3×3, including a regularization layer and an activation function LeakRelu , and then upsampling is performed layer by layer. Every time it reaches a shallower layer, its width and height are 2 times that of the previous layer, and its number of channels is 1 / 2 times that of the previous layer, restoring the feature maps of 1 / 8, 1 / 4, and 1 / 2 times, and also retaining the fused feature maps at each scale. The restored-resolution feature maps contain both rich detailed feature information and high-level semantic information extracted from the deep feature maps.

[0094] Step 6: The deep feature Figure 1 / 16 is upsampled to the size of the previous layer 1 / 8 and fused with the selected high-resolution feature Figure 1 / 8 at the corresponding encoding stage. The deep feature Figure 1 / 8 is upsampled to the size of the previous layer 1 / 4 and fused with the selected high-resolution feature Figure 1 / 4 at the corresponding encoding stage. The deep feature Figure 1 / 4 is upsampled to the size of the previous layer 1 / 2 and fused with the selected high-resolution feature Figure 1 / 2 at the corresponding encoding stage. The deep feature Figure 1 / 2 is upsampled to the original resolution size 1 of the image and fused with the high-resolution feature map at the corresponding encoding stage. This skip connection uses the concatenation method instead of simple addition to ensure the integrity and richness of information during the feature transfer process, so as to improve the accuracy and edge details of segmentation;

[0095] Step 7: After layer-by-layer upsampling, enter the same improved Swin Transformer layer as in Step 2, which includes adaptive size adjustment, a series of sliding window self-attention calculations, a series of residual connections, and image size restoration, for feature and semantic information extraction;

[0096] Step 8: Input the feature map of the last stage in Step 6 into the convolutional layer segmentation head, and the output resolution of the network is Figure 1 consistent with the original, and the value of each pixel is its class probability.

[0097] Step 9: Calculate the loss function for the network output for backpropagation to update the network parameters, including MLP parameters, self-attention matrix weight parameters, etc. Fusing the calculations of the two loss functions can improve the robustness and generalization ability of the model, reduce overfitting, provide more accurate gradient signals, and facilitate better learning and parameter adjustment of the model. At the same time, since LovaszLoss has differentiable properties in mathematics, it is more convenient to be combined with other loss functions to further improve the model performance. Meanwhile, when fusing the calculations of the two losses, the weights between them can be adjusted to balance the weights of different objectives. This can flexibly control the optimization degree of the model on different task metrics to meet the actual needs.

[0098] The loss function uses Soft Cross Entropy Loss and Lovasz Loss fused with a weight ratio of 1:1. Among them, the soft cross entropy, that is, the cross entropy using label smoothing, will increase the generalization ability. The specific formula for the soft cross entropy is as follows:

[0099] (6)

[0100] In Equation (6), L SCE The value of the Soft Cross Entropy Loss function; y i The label of the i-th category in the true label, set as a soft label, y i ∈(0,1); P ( x i ) represents the probability of the i-th category predicted by the model.

[0101] The meaning of Lovasz Loss is to gradually optimize the prediction results to make them closer to the true labels. It not only focuses on correctly classified samples but also considers boundary samples and misclassified samples. Therefore, it can handle the problems of class imbalance and uneven difficulty of samples. The specific formula is as follows:

[0102] (7)

[0103] (8)

[0104] (9)

[0105] (10)

[0106] In Equation (7), Δ Jc oss function; y* round truth; c represents the set of predicted error pixels; M c the set where the network segmentation result does not match the label, M c {0,1} p , p represents the number of pixels;

[0107] In Equation (8), the lovasz extension is used for smooth extension and is specifically implemented in multi-class segmentation; f i ( c ) represents the probability value after the softmax of the c th class; The scoring function f i ( c ) is used to construct a pixel errors m i ( c ) vector;

[0108] In Equation (9), the errors m ( c ) Δ Jc loss function;

[0109] In Equation (10), in order to optimize the mIoU metric for evaluating all classes, average the above loss ( f ( c ))

[0110] In summary, compared with the previous convolutional feature extraction network and U-shaped network, the semantic segmentation efficiency of the present invention has been improved, and the accuracy of segmentation and recognition results has been enhanced. It can be widely applied to fields such as medical treatment and remote sensing that require semantic segmentation. The siamese network has model interpretability. By explaining the correlation between the input and output, the essence of the segmented image can be understood. Moreover, the siamese network can be personalized modeled, the feature differences can be understood, and personalized predictions can be made based on these differences. By applying the sliding window multi-head self-attention transformation mechanism, the problem that the shallow texture information and the deep semantic information cannot be taken into account at the same time is improved, the global modeling ability of the feature extraction network and the acquisition efficiency of long-distance semantic information are enhanced, and the underlying feature extraction ability and the abstraction effect of high-level semantic information are improved. The overall network forms a U-shaped structure, with a locally symmetric encoder-decoder structure to supplement the global context information and strengthen the spatial correlation of semantic information. In addition, residual links are added to train the deep network with high efficiency and enhance the recognition accuracy. By adding adaptive image feature scale input processing, the input image is adaptively scaled to solve the redundant image preprocessing process caused by the network input limitation problem, achieve accurate end-to-end semantic segmentation results, and implement multi-scale input to enhance the generalization ability and accuracy of the network model. The network uses a fusion loss function to improve both the segmentation accuracy and the recognition accuracy. In application, for the semantic segmentation task of complex scenes, this network solves the problems of inaccurate and unclear local positioning segmentation, neglect of small-scale feature recognition, and low global category information recognition efficiency caused by multi-scale features in the image and the imbalance between underlying features and high-level semantic information therein.

[0111] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.

Claims

1. A complex target recognition method based on a deep twin self-attention network, characterized in that, The deep twin self-attention network includes: The twin encoding network can interpret high-dimensional image information to enhance individual feature modeling, improve segmentation performance, and enhance model generalization; The sliding window self-attention transformation module based on residual connection is used to enhance the network model's recognition ability for objects of different sizes and enhance the relevance of the global long-distance semantic information space; The image size adaptive module is used to solve the fixed limitation based on the sliding window self-attention mechanism; The locally symmetric hierarchical deep sliding window self-attention mechanism transformation encoder-decoder structure network, which includes a twin encoding network and a decoder. The twin encoding network includes Down modules at different levels, and the decoder includes Up modules at different levels; Multiple skip connections are used to directly connect the corresponding layers of the encoder and decoder to retain and transfer high-resolution feature information; The fusion loss function is used to improve both recognition accuracy and classification balance; The method specifically includes the following steps: S1. First, separate the high-dimensional image information into two channel dimensions, and each channel is input into two sub-encoding networks of the twin encoding network for feature extraction; first, the features of each channel are input into the convolutional module in the twin encoding network to initially extract image features; then input into multiple residual module structures Down modules with residual connections in the twin encoding network, and the input image first undergoes adaptive scale adaptive; S2. Input the image processed in S1 into a series of sliding window self-attention transformation modules based on residual connection to further extract and process features, and then restore the size of the output feature map after being processed by the image size adaptive module; S3. At each stage of the twin encoding network, feature maps of different depths are generated through downsampling, and for the sub-encoding network with a more information-rich input channel, record the features and the original image resolution features in each encoding stage; S4. At the end of the encoding stage, fuse the branch channels of the two sub-encoding networks, that is, fuse and splice the feature maps extracted from each channel in the twin channels of the twin encoding network; S5. Input the fused image obtained in S4 into multiple residual module structures Up modules with residual connections in the decoder. After the image is upsampled, it is gradually restored to the feature map of the input size; S6. In each decoding stage, use the feature fusion technology to combine the high-resolution features of the hierarchical encoder in S3 and the upsampled features of the decoder in S5; S7. In the decoding stage, after the feature fusion in S6, repeat S1-S2; S8. Input the feature map of the last stage obtained in S7 into the segmentation head, classify each pixel, and divide the pixels into pre-determined categories to obtain the semantic segmentation recognition result; S9. Use the fusion loss function to calculate the semantic segmentation recognition result obtained in S8, and backpropagate to improve the network parameters.

2. The complex target recognition method based on a deep twin self-attention network according to claim 1, characterized in that The function of the residual module structure Down module in S1 is image size adjustment, using the sliding window self-attention transformation module for feature extraction and downsampling; S1 specifically includes the following content: S1.

1. Separate the input high-dimensional image into two channels and input the same sub-encoding network to form a twin network; S1.

2. Scale adapt the input image to meet the fixed size requirements of the sliding window self-attention transformation module; S1.

3. Input the image processed in S1.2 into the subsequent sliding window self-attention transformation module for calculation.

3. A complex target recognition method based on a deep twin self-attention network according to claim 1, characterized in that, The specific content of S2 is as follows: S2.

1. Divide the image into multiple small blocks, flatten the small blocks and convert them into vectors of a fixed dimension through a linear layer, and add position information to the feature vectors of each small block to ensure that the model understands the spatial relationship between different blocks; S2.

2. After being processed by S2.1, the sequence with positional encoding information is retained as the residual quantity and input into the normalization layer LN , for each element of the input sequence, perform multi-head self-attention calculation and add it to the residual quantity. The specific formula is expressed as: (1) In formula (1), represents the output sequence, x l-1 represents the input sequence; S2.

3. Retain the output obtained in S2.2 as the residual and input it into the normalization layer successively LN , the multi-layer perceptron network MLP , and then add it to the residual. The specific formula is expressed as: (2) In formula (2), x l represents the output sequence; S2.

4. Keep the output sequence in S2.3 as the residual and input it into the normalization layer LN , for each element of the input sequence, perform moving window self-attention calculation and add it to the residual quantity. The specific formula is expressed as: (3) In formula (3), represents the output sequence; S2.

5. Use the output obtained in S2.4 as the residual and the input, and repeat S2.

3. The specific formula is expressed as: (4) In formula (4), x l+1 represents the output sequence; S2.

6. After passing through a series of sliding window self-attention transformation modules, use the image size adaptive module with the restored size to restore the image size to the input size to facilitate the design of the hierarchical feature extraction structure.

4. The complex target recognition method based on the deep twin self-attention network according to claim 1, characterized in that The specific content of S3 is as follows: The sub-encoding network simultaneously performs feature downsampling through a convolutional layer with a stride of 2 and a convolutional kernel of 3×3. Every time it goes to a deeper layer, its width and height are both 1 / 2 of the previous layer, and the number of channels is 2 times that of the previous layer, forming four different-depth feature maps of 1 / 4, 1 / 8, 1 / 16, and 1 / 32; for the image with more channels separated from the high-dimensional image in S1, record its input sub-network layer-by-layer sampling features, where the shallow feature maps contain detailed feature information, and the deep feature maps contain high-level semantic information.

5. A complex target recognition method based on a deep twin self-attention network according to claim 1, characterized in that, The specific content of S4 is as follows: At the feature level, merge the feature maps of the two branches from the siamese network by splicing, and use the similar feature characteristics learned by the siamese network to improve the segmentation performance and ensure its high interpretability when processing highly similar images.

6. A complex target recognition method based on a deep twin self-attention network according to claim 1, characterized in that, The specific content of S5 is as follows: After fusing the two-channel image in S4, perform feature upsampling through a convolutional layer with a stride of 2 and a convolutional kernel of 3×3. Every time it goes to a shallower layer, its width and height are both 2 times that of the previous layer, and its number of channels is 1 / 2 of the previous layer, restoring three-depth feature maps of 1 / 8, 1 / 4, and 1 / 2; Among them, the feature maps with the restored resolution contain both detailed feature information and high-level semantic information extracted from the deep feature maps.

7. A complex target recognition method based on a deep twin self-attention network according to claim 1, characterized in that, The specific content of S6 is as follows: The deep feature map 1 / 16 after channel fusion is upsampled to the size 1 / 8 of the previous layer and fused with the corresponding high-resolution feature map 1 / 8 recorded in the selected channel of S3 to ensure the integrity and richness of information during the feature transfer process and improve the accuracy and edge details of the segmentation.

8. A complex target recognition method based on a deep twin self-attention network according to claim 1, characterized in that, The specific content of S9 is as follows: Use Soft Cross Entrophy Loss and Lovza Loss to fuse the semantic segmentation recognition results with a weight of 1:1, perform backpropagation, and update the network parameters to improve the image segmentation accuracy and classification accuracy.

Citation Information

Patent Citations

  • Transform-based twin network image denoising method and system, medium and equipment

    CN114359109A

  • Remote sensing image change detection method based on improved Transform twin network

    CN115984700A