Radar image segmentation method and device based on THAM-ResUNet
Through the THAM-ResUNet network, the THAM hybrid attention mechanism and residual structure are used to solve the problem of weak edge features in radar image segmentation, achieving higher accuracy image segmentation and faster training efficiency.
Patent Information
- Application Number
- CN202510626097.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Radar images are difficult to accurately segment due to significant spot noise, low signal-to-noise ratio and weak edge characteristics.
Using the THAM-ResUNet network, the target defocus and texture characteristics are captured by introducing the THAM hybrid attention mechanism, combining phase attention, channel attention and spatial attention, and combining residual structure to achieve accurate segmentation of radar images.
It significantly improves the segmentation accuracy of weakly scattered areas in radar images, suppresses background noise interference, and improves the generalization ability and training efficiency of the model.
Smart Images

Figure CN120495314A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of ISAR image processing and relates to a radar image segmentation method and equipment based on a neural network model. Background Art
[0002] Inverse Synthetic Aperture Radar (ISAR) is a key branch of synthetic aperture radar (SAR). It can acquire detailed images of non-cooperative moving targets (such as aircraft, ships, and missiles) at all times of day and in all weather conditions, and at long distances, thus possessing significant application value. In recent years, segmentation techniques for ISAR images have played a crucial role in target feature extraction and recognition. However, compared to optical images, radar images suffer from significant speckle noise, low signal-to-noise ratio, and weak edge features, making accurate segmentation difficult. Summary of the Invention
[0003] The present invention aims to solve the problem that radar images have weak edge features compared to optical images, which makes it difficult to accurately segment them.
[0004] A radar image segmentation method based on THAM-ResUNet, which first obtains an ISAR space target image and then feeds it into a THAM-ResUNet network for radar image segmentation for segmentation; the THAM-ResUNet network is built based on a UNet network model, in which a THAM hybrid attention mechanism is set after each convolution module of the UNet network model. The output of the last THAM attention mechanism is passed through a fully connected layer to obtain an output vector of the THAM-ResUNet, thereby achieving radar image segmentation;
[0005] The processing of the THAM attention mechanism includes:
[0006] Split the input feature map G(a,x,y)=A(a,x,y)·exp(jφ(a,x,y)) into the amplitude spectrum A(a,x,y) and the phase spectrum φ(a,x,y), where a is the channel dimension, x is the distance dimension, and y is the azimuth dimension; calculate the phase change of adjacent frequency components:
[0007] Δφ(a,x,y)=φ(a,x,y-1)-φ(a,x,y)+φ(a,x,y+1)-φ(a,x,y)
[0008] Then generate the phase attention features:
[0009] F φ =γ·Conv3D(Stack[φ(a,x,y),Δφ(a,x,y)])+b
[0010] Among them, Stack[·] is splicing along the channel dimension, Conv3D represents three-dimensional convolution, b is the bias term; γ is the weight parameter;
[0011] Then F φ The feature F is obtained by concatenating it with the input feature map G(a,x,y), and the HAM attention processing is performed on the feature F to obtain the THAM output feature.
[0012] Furthermore, the process of performing HAM attention processing on feature F includes:
[0013] For feature F, average pooling and maximum pooling are performed to obtain and Based on the parameters α and β, we get
[0014]
[0015] Then Perform one-dimensional fast convolution to obtain a channel refined feature map; then introduce a channel separation ratio parameter λ to divide the channel refined feature map into important channel group F1′ and less important feature F2′;
[0016] Perform average pooling and maximum pooling on F1′, and splice the two pooling results in the channel dimension to obtain a set of output features; perform the same processing on F2′ to obtain a set of output features; for the two sets of output features, convolution is performed through a shared convolution layer with a convolution kernel size of 7×7 to generate two sets of feature maps, which are normalized and activated to obtain tensors. and The two tensors are multiplied with F1′ and F2′ respectively to obtain the spatial refined features F1″ and F2″, and the two are added element by element to obtain the THAM output features.
[0017] Further, During the one-dimensional fast convolution, the convolution kernel size is in Indicates the closest odd number.
[0018] Furthermore, the THAM-ResUNet network processing process based on the UNet network model is as follows:
[0019] The input feature map first passes through a 7×7 convolution module, and the output of the 7×7 convolution module is processed by the THAM attention mechanism;
[0020] The output of the 7×7 convolutional module is fed into the bottleneck residual module. There are at least four bottleneck residual modules, and the THAM attention mechanism is introduced after each bottleneck residual module. Different bottleneck residual convolution blocks include multiple bottleneck convolution residual units. Each bottleneck residual convolution unit has three convolution layers, and the input and the output of the last convolution layer are connected with residuals. The residual connection in each bottleneck residual convolution unit is then processed by the ReLU activation layer.
[0021] The output of the last bottleneck residual module is fed into the basic residual convolution module. There are at least three basic residual convolution modules, each of which has two convolution layers. The residual connection in each basic residual convolution module is processed by the ReLU activation layer. The THAM attention mechanism is introduced after each basic residual convolution block. The jump connection dimension is adjusted by upsample after the THAM attention mechanism of the penultimate basic residual convolution module and the penultimate third basic residual convolution module.
[0022] The THAM attention mechanism set after the 7×7 convolution module and the first two bottleneck residual modules is skip-connected with the last three basic residual convolution blocks in the decoder.
[0023] Furthermore, the 7×7 convolution module is provided with a convolution layer with a convolution kernel size of 7×7, followed by a batch normalization layer, a maximum pooling layer and a ReLU activation layer.
[0024] Furthermore, in the 7×7 convolution module, the convolution step size of the 7×7 convolution layer is 2 pixels, and the step size of the maximum pooling layer is 2 pixels.
[0025] Furthermore, the number of the bottleneck residual modules is set to 4, and the number of the basic residual convolution modules is set to 3;
[0026] The four bottleneck residual modules are denoted as bottleneck residual module 1-neck residual module 4; bottleneck residual module 1 has 3 bottleneck convolution residual units, and only the first convolution unit is down-sampled; bottleneck residual module 2 has 4 bottleneck convolution residual units, and only the first convolution unit is down-sampled; bottleneck residual module 3 has 6 bottleneck convolution residual units, and only the first convolution unit is down-sampled; bottleneck residual module 4 has 3 bottleneck convolution residual units.
[0027] Furthermore, the structure of the bottleneck residual convolution unit is as follows:
[0028] When the bottleneck residual convolution unit is set to downsample, the convolution kernel size in the first convolution layer is 1×1 pixel and the stride is 1 pixel; the convolution kernel size in the second convolution layer is 3×3 pixels and the stride is 2 pixels when downsampling. The residual connection dimension is adjusted by the downsample branch; the convolution kernel size in the third convolution layer is 1×1 pixel and the stride is 1 pixel. Each convolution layer is followed by a batch normalization layer.
[0029] When the bottleneck residual convolution unit does not set downsampling processing, the convolution kernel size in the first convolution layer is 1×1 pixel and the stride is 1 pixel; the convolution kernel size in the second convolution layer is 3×3 pixels and the stride is 1 pixel; the convolution kernel size in the third convolution layer is 1×1 pixel and the stride is 1 pixel; each convolution layer is followed by a batch normalization layer.
[0030] Furthermore, the two convolutional layers of the basic residual convolution module are as follows:
[0031] The convolution kernel size in the first convolution layer is 3×3 pixels and the stride is 1 pixel;
[0032] The convolution kernel size in the second convolution layer is 3×3 pixels and the stride is 1 pixel;
[0033] Each convolutional layer is followed by a batch normalization layer.
[0034] A radar image segmentation device based on THAM-ResUNet, the device comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the radar image segmentation method based on THAM-ResUNet.
[0035] Beneficial effects of the present invention: The image segmentation framework of THAM-ResUNet proposed in the present invention can realize the segmentation of ISAR images of space targets. First, the method designs THAM (triple hybrid attention mechanism) according to the characteristics of radar images, innovatively designs phase attention to capture the target defocus features, and then uses channel attention to screen key features, spatial attention to strengthen the target edge response, and combines the residual structure to effectively capture the detailed distribution characteristics of the scattering points, thereby achieving better segmentation effects. In addition, the present invention also proposes a single-cycle strategy training optimization method, which can speed up the training efficiency of the network, reduce the possibility of model overfitting, and improve the generalization ability of the model.
[0036] The present invention utilizes deep learning technology to better meet the requirements of radar image segmentation under conditions such as significant speckle noise, low signal-to-noise ratio and weak edge features. The network maintains a certain stability to the perspective changes and noise of the input image. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram of generating ISAR images of space targets and corresponding annotated images;
[0038] Figure 2 This is the overall structure diagram of the THAM-ResUNet framework proposed in this invention;
[0039] Figure 3 It is the optimization effect of single-cycle strategy training. DETAILED DESCRIPTION
[0040] In response to the problems in the background technology, this patent proposes a radar image segmentation method based on THAM-ResUNet. This method introduces THAM to optimize feature selection, captures target defocus features and target texture features through phase attention, filters key features through channel attention, strengthens target edge response through spatial attention, and combines the residual structure to suppress gradient vanishing. It uses a deep network to effectively capture the detailed distribution characteristics of scattering points, significantly improving the segmentation accuracy of weak scattering areas in radar images while suppressing the interference of background noise. In addition, the jump connection structure of UNet further enhances the fusion of multi-scale features and improves segmentation accuracy. Finally, by using an innovative single-cycle strategy training optimization method, the model generalization ability is effectively improved and the network training efficiency is improved.
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work are not included.
[0042] Specific implementation method 1: Combination Figures 1 to 2 To explain this embodiment,
[0043] The radar image segmentation method based on THAM-ResUNet described in this embodiment includes the following steps:
[0044] S1. Based on the three-dimensional model of the space target, an ISAR space target image dataset is obtained. Then, the corresponding label information is obtained through labelme. The sample dataset consisting of the ISAR image and the labeled image is divided into a training sample set and a test sample set.
[0045] In the process of obtaining the ISAR space target image dataset, the ISAR imaging part uses the scattering point information of the space target three-dimensional model and performs ISAR simulation imaging according to the set radar parameters to obtain the space target ISAR image dataset.
[0046] Annotate an ISAR image dataset using LabelMe and generate a corresponding label information map. To do this, open LabelMe via the command line or an application, import the ISAR image dataset, draw the annotation regions on the images, and enter the label names. This generates a JSON-formatted annotation file containing the label information. Convert the JSON file to a PNG-formatted label information map.
[0047] The ISAR image sequence dataset is then divided into a training set and a test set. The training set contains a label information map, in which each pixel has a label, indicating whether the pixel belongs to the target. The test set does not have a label information map.
[0048] S2. Construct a radar image segmentation network based on the THAM-ResUNet framework, use ISAR images and corresponding label information graphs as training samples to train THAM-ResUNet to obtain a training model; then use the trained segmentation model to perform segmentation testing on the test ISAR image data to obtain the segmentation results of the test ISAR image data.
[0049] The ISAR image data was fed into the THAM-ResUNet network for processing, and the label information graph was used as the label for training the THAM-ResUNet network. The ISAR image data size was 3×512×512, and the label information graph data size was 1×512×512. Target image segmentation training and testing were performed based on the THAM-ResUNet network architecture.
[0050] Construct a THAM-ResUNet network based on THAM triple hybrid attention mechanism for radar image segmentation.
[0051] S201, THAM hybrid attention mechanism:
[0052] First, phase attention processing is performed. The input feature map G(a,x,y) = A(a,x,y) · exp(jφ(a,x,y)) can be split into an amplitude spectrum A(a,x,y) and a phase spectrum φ(a,x,y), where a is the channel dimension, x is the range dimension, and y is the azimuth dimension. The phase spectrum reflects the phase shift caused by the target's Doppler effect. This shift manifests as azimuth defocus in the image, making the ISAR image edges more irregular. Severe defocusing can also cause target geometric distortion. In addition, high-frequency phase changes correspond to the microstructure of the scattering points, which is conducive to capturing the target's texture features. Defocus features and target texture features can be captured by calculating the phase changes of adjacent frequency components.
[0053] Δφ(a,x,y)=φ(a,x,y-1)-φ(a,x,y)+φ(a,x,y+1)-φ(a,x,y) (1)
[0054] As shown in the above formula, only the azimuth dimension difference is used to extract the phase change while avoiding the introduction of irrelevant noise, and the boundaries are symmetrically filled. In order to solve the influence of different defocus phenomena on the image segmentation effect, the trainable parameter γ is used to generate the phase attention feature:
[0055] F φ =γ·Conv3D(Stack[φ(a,x,y),Δφ(a,x,y)])+b (2)
[0056] Among them, Stack[·] is splicing along the channel dimension, Conv3D uses convolution with a 1×3×3 three-dimensional convolution kernel, and b is the bias term.
[0057] For the weight parameter γ, when γ→0, phase difference is disabled and degenerates to the baseline model; 0<γ<1, the phase attention weight is weakened; γ>1, the phase attention weight is enhanced.
[0058] Then perform channel attention processing: F φ The feature F is obtained by concatenating it with the input feature map G(a,x,y). The feature F is obtained by average pooling and maximum pooling respectively. and Average pooling and maximum pooling also play different roles in different stages of image feature extraction, and an adaptive mechanism is designed accordingly. and Multiply the trainable parameters α and β by 1 / 2 and add them together, and finally use element-by-element summation to get
[0059]
[0060] Then Perform one-dimensional fast convolution to capture the relationship between channels, and the convolution kernel size is in Indicates the closest An odd number of convolution is performed to obtain a channel refined feature map. Then a channel separation ratio parameter λ is introduced to divide the channel refined feature map into important channel groups F1′ and less important features F2′.
[0061] Then, spatial attention is calculated based on F1′ and F2′, and average pooling and maximum pooling are performed on F1′ and F2′ respectively. The two pooling results are concatenated in the channel dimension to obtain two sets of output features. These two sets of features are then convolved by a shared convolution layer with a convolution kernel size of 7×7, generating two sets of feature maps of size H×W×1. After normalization and activation function, the tensors are obtained. and The two tensors are multiplied with F1′ and F2′ respectively to obtain the spatial refined features F1″ and F2″, and the two are added element by element to obtain the final refined features, namely the THAM output features.
[0062] S202. Build a network model framework:
[0063] like Figure 2 As shown in the figure, the THAM-ResUNet network contains three convolution modules with different structures, namely 7×7 convolution module, bottleneck residual module and basic residual convolution module.
[0064] The input image first passes through a 7×7 convolutional module. As part of the encoder, the 7×7 convolutional module has one convolutional layer to expand the receptive field and capture coarse-grained features. The convolution kernel size is 7×7 pixels with a stride of 2 pixels. This is followed by a batch normalization layer, a max pooling layer, and a ReLU activation layer. The max pooling layer has a stride of 2 pixels. The THAM attention mechanism is set after the 7×7 convolutional module.
[0065] The output of the 7×7 convolution module is fed into the bottleneck residual module. The bottleneck residual module is set to multiple, and in this embodiment, it is set to 4, as part of the encoder. The THAM attention mechanism is introduced after each bottleneck residual module to dynamically calibrate the features. Figure 2As shown in the figure, different bottleneck residual convolution blocks include multiple bottleneck convolution residual units. Each bottleneck residual convolution unit is equipped with 3 convolution layers for downsampling the encoding part. The input and the output of the last convolution layer (including the batch normalization layer) are residually connected. The convolution kernel size in the first convolution layer is 1×1 pixel and the step size is 1 pixel. The convolution kernel size in the second convolution layer is 3×3 pixels and the step size is 2 pixels during downsampling. The residual connection dimension is adjusted by the downsample branch. The convolution kernel size in the third convolution layer is 1×1 pixel and the step size is 1 pixel. Each convolution layer is followed by a batch normalization layer. The residual connection in each bottleneck residual convolution unit is processed by the ReLU activation layer.
[0066] In this embodiment, different numbers of bottleneck convolution residual units are set in different bottleneck residual modules. In this embodiment, bottleneck residual module 1 has 3 bottleneck convolution residual units, and only the first convolution unit is used for downsampling; bottleneck residual module 2 has 4 bottleneck convolution residual units, and only the first convolution unit is used for downsampling; bottleneck residual module 3 has 6 bottleneck convolution residual units, and only the first convolution unit is used for downsampling; bottleneck residual module 4 has 3 bottleneck convolution residual units.
[0067] The output of the last bottleneck residual module is sent to the basic residual convolution module. The basic residual convolution modules are set to multiple, and in this embodiment, they are set to 3, which serve as decoders; each basic residual convolution module has 2 convolution layers for upsampling the decoding part, and the input and the last layer of convolution output are residually connected. The convolution kernel size in the first convolution layer is 3×3 pixels, and the step size is 1 pixel. The convolution kernel size in the second convolution layer is 3×3 pixels, and the step size is 1 pixel. Each convolution layer is followed by a batch normalization layer; the residual connection in each basic residual convolution module is processed by a ReLU activation layer; the THAM attention mechanism is introduced after each basic residual convolution block to adaptively fuse multi-scale features; the THAM attention mechanism of the basic residual convolution module 1 and the basic residual convolution module 2 are both adjusted by upsample to adjust the jump connection dimension.
[0068] A skip connection is designed between the downsampling encoding part and the upsampling decoding part, effectively fusing the multi-scale features of the encoding and decoding parts and mitigating the information loss caused by downsampling. In this implementation, the THAM attention mechanism set after the 7×7 convolutional module and the first two bottleneck residual modules of the encoder is skip-connected to the last three basic residual convolution blocks in the decoder.
[0069] The output of the last THAM attention mechanism passes through a fully connected layer with an output dimension of 2 to obtain the output vector of THAM-ResUNet, thereby realizing the radar image segmentation method.
[0070] In this embodiment, a residual network is used to construct a model. In a conventional neural network, each layer maps the input to the output by learning, while in residual learning, the goal of each layer is to learn the residual between the input mapping and the expected output, that is, the network tries to fit the difference between the input and the expected output. Compared with ordinary networks, ResNet adds a short-circuit mechanism between every two or three convolutional layers to form a residual learning structure. When there is a residual, it means that new features can continue to be learned, thereby enriching the feature information. Conversely, the network can be kept in a good state and the degradation problem can be alleviated. The basic residual block structure consists of two 3×3 convolutional layers, and residual learning is achieved through jump connections; deep networks use a "bottleneck" structure residual block, and each residual block uses 1×1, 3×3, and 1×1 convolutional layers to reduce the number of parameters and computational complexity.
[0071] The overall network architecture is similar to that of the U. The encoder's main structure consists of four downsampling modules with matching Relu activation functions, batch normalization layers, and max pooling layers, enabling high-level image feature extraction and data dimensionality reduction. The decoder's upsampling module consists of an upsampling convolutional layer, a skip structure, and a Relu activation function. The upsampling convolutional layer gradually restores the image resolution. The resulting feature map is concatenated with the feature information extracted by the corresponding layer during the downsampling process via skip connections, outputting a feature map with the same size as the original image.
[0072] In some embodiments, for the THAM-ResUNet network, its specific structural parameters are as follows:
[0073] In the downsampling encoding part, the output size of the 7×7 convolution module is 64×256×256 pixels;
[0074] After being input into the bottleneck residual module 1, it is passed to the first convolution layer of the first bottleneck residual convolution unit;
[0075] The first convolutional layer processes the input data and outputs a feature map of 64×256×256 pixels, which is then processed by the batch normalization layer and fed into the second convolutional layer.
[0076] The second convolutional layer performs downsampling with a step size of 2 pixels and outputs a feature map of 64×128×128 pixels. After being processed by the batch normalization layer, it is sent to the third convolutional layer.
[0077] The third convolutional layer performs dimensionality reorganization with a stride of 1 pixel and outputs a feature map of 256×128×128 pixels. This feature map is fused with the residual connection through the downsample branch, processed by the batch normalization layer and the ReLu layer, and then passed to the remaining two bottleneck residual convolution units of the bottleneck residual module 1.
[0078] The remaining two bottleneck residual convolution units of the bottleneck residual module 1 differ from the first bottleneck residual convolution unit only in the second convolution layer, which does not require downsampling and has a stride of 1 pixel. Finally, the THAM attention mechanism outputs a feature map of 256×128×128 pixels and passes it to the bottleneck residual module 2.
[0079] Compared with bottleneck residual module 1, bottleneck residual module 2 has only one more bottleneck residual convolution unit, and finally outputs a feature map of 512×64×64 pixels, which is passed to bottleneck residual module 3;
[0080] Bottleneck residual module 3 has only three more bottleneck residual convolution units than bottleneck residual module 1, and finally outputs a feature map of 1024×32×32 pixels, which is passed to bottleneck residual module 4;
[0081] Compared with bottleneck residual module 1, bottleneck residual module 4 does not perform downsampling in the second convolution layer of the first bottleneck residual convolution unit, has a step size of 1 pixel, and finally outputs a feature map of 1024×64×64 pixels through the THAM attention mechanism and passes it to the upsampling decoding part.
[0082] In the upsampling decoding part, the input of each basic residual convolution module is the jump connection fusion result of the output of the previous layer and the corresponding downsampling encoding part;
[0083] The input of the first basic residual convolution module is the 1024×64×64 pixel feature map of the previous layer plus the 512×64×64 pixel feature map before downsampling by the bottleneck residual module 3, which is passed to the first convolution layer of the first basic residual convolution module;
[0084] The first convolutional layer performs upsampling with a step size of 1 pixel and outputs a feature map of 256×64×64 pixels, which is then processed by the batch normalization layer and fed into the second convolutional layer.
[0085] The second convolutional layer processes the input data and outputs a feature map of 1024×64×64 pixels. After processing by the batch normalization layer and the ReLu layer, it outputs a feature map of 512×128×128 pixels through upsample to the second basic residual convolution module;
[0086] The input of the second basic residual convolution module is the 512×128×128 pixel feature map of the first basic residual convolution module plus the 256×128×128 pixel feature map before downsampling by the bottleneck residual module 2. The other steps are the same as the first basic residual convolution module, and the output feature map of 256×256×256 pixels is sent to the third basic residual convolution module.
[0087] The input of the third basic residual convolution module is the 256×256×256 pixel feature map of the second basic residual convolution module plus the 64×256×256 pixel feature map before downsampling by the bottleneck residual module 1. The other steps are the same as the first basic residual convolution module, and the output feature map of 64×512×512 pixels is sent to the fully connected layer.
[0088] The output size of the fully connected layer is 2×512×512.
[0089] Combine Figure 3 It shows that the single-cycle strategy training optimization method of the present invention can effectively accelerate the training efficiency of the network. The training process is divided into four parts, namely two 30-epoch freeze trainings and two 30-epoch thaw trainings. Since the network structure of the downsampling part is the most complex, the downsampling part is frozen in the freezing stage, and the downsampling parameters of the freezing training process are kept unchanged to reduce the complexity of the model; after multiple learning rate test experiments, the learning rate is selected as 0.001 in the freezing stage, and the learning rate is selected as 0.0001 in the thaw stage; since the network is relatively complex and the network architecture is deep, a weight decay value of 0.000001 is selected in both the freezing and thaw stages; in addition, in order to verify the convergence of the model after four trainings, an additional 10-epoch convergence test stage is set. This stage is after the second thaw training, so the learning rate and weight decay are consistent with the thaw training to test whether the model is overfitting. Although the THAM-ResUNet model based on the single-cycle strategy optimization training method did not rise as fast as the THAM-ResUNet model in the early stage, the THAM-ResUNet model based on the single-cycle strategy optimization training method has stabilized after 120 epochs. It can be seen that the single-cycle training optimization method is beneficial to accelerating the model training efficiency.
[0090] S3. After the training is completed, a trained THAM-ResUNet model is obtained and used for actual image segmentation. In this embodiment, the ISAR image sequence in the test data set is input into the model to obtain the image segmentation result. Specific implementation method two:
[0092] This embodiment is a radar image segmentation device based on THAM-ResUNet, which includes a processor and a memory. It should be understood that the device includes any device including a processor and a memory described in the present invention, and may also include other units and modules that perform display, interaction, processing, control, and other functions through signals or instructions;
[0093] At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the radar image segmentation method based on THAM-ResUNet.
[0094] It should be understood that instructions include computer program products, software, or computerized methods corresponding to any method described herein; the instructions can be used to program a computer system or other electronic device. Computer storage media can include readable media on which instructions are stored, and can include but are not limited to magnetic storage media and optical storage media; magneto-optical storage media include read-only memory (ROM), random access memory (RAM), erasable programmable memory (e.g., EPROM and EEPROM), and flash memory layers, or other types of media suitable for storing electronic instructions. Those skilled in the art will also understand that the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0095] The present application is described with reference to the flowcharts and / or block diagrams of the methods, systems, and computer program products according to the embodiments of the present application, and can also be used for corresponding devices. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0096] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0098] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0099] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
[0100] The above examples are merely illustrative of the calculation model and process of the present invention and are not intended to limit the embodiments of the present invention. Persons skilled in the art will readily appreciate that other variations or modifications based on the above description are possible. This list of embodiments is not exhaustive; however, any obvious variations or modifications derived from the technical solution of the present invention remain within the scope of protection of the present invention.
Claims
1. A radar image segmentation method based on THAM-ResUNet, characterized in that: First, an ISAR space target image is obtained and then fed into a THAM-ResUNet network for radar image segmentation. The THAM-ResUNet network is built based on a UNet network model. Each convolutional module of the UNet network model is followed by a THAM hybrid attention mechanism. The output of the last THAM attention mechanism is passed through a fully connected layer to obtain the output vector of the THAM-ResUNet, thereby achieving radar image segmentation. The processing of the THAM attention mechanism includes: Split the input feature map G(a,x,y)=A(a,x,y)·exp(jφ(a,x,y)) into the amplitude spectrum A(a,x,y) and the phase spectrum φ(a,x,y), where a is the channel dimension, x is the distance dimension, and y is the azimuth dimension; calculate the phase change of adjacent frequency components: Δφ(a,x,y)=φ(a,x,y-1)-φ(a,x,y)+φ(a,x,y+1)-φ(a,x,y) Then generate the phase attention features: F φ =γ·Conv3D(Stack[φ(a,x,y),Δφ(a,x,y)])+b Among them, Stack[·] is splicing along the channel dimension, Conv3D represents three-dimensional convolution, b is the bias term; γ is the weight parameter; Then F φ The feature F is obtained by concatenating it with the input feature map G(a,x,y), and the HAM attention processing is performed on the feature F to obtain the THAM output feature.
2. The radar image segmentation method based on THAM-ResUNet according to claim 1, characterized in that: The process of HAM attention processing for feature F includes: For feature F, average pooling and maximum pooling are performed to obtain and Based on the parameters α and β, we get Then Perform one-dimensional fast convolution to obtain a channel refined feature map; then introduce a channel separation ratio parameter λ to divide the channel refined feature map into important channel group F1′ and less important feature F2′; Perform average pooling and maximum pooling on F1′, and splice the two pooling results in the channel dimension to obtain a set of output features; perform the same processing on F2′ to obtain a set of output features; for the two sets of output features, convolution is performed through a shared convolution layer with a convolution kernel size of 7×7 to generate two sets of feature maps, which are normalized and activated to obtain tensors. and The two tensors are multiplied with F1′ and F2′ respectively to obtain the spatial refined features F1″ and F2″, and the two are added element by element to obtain the THAM output features.
3. The radar image segmentation method based on THAM-ResUNet according to claim 2, characterized in that: right During the one-dimensional fast convolution, the convolution kernel size is in Indicates the closest odd number.
4. The radar image segmentation method based on THAM-ResUNet according to any one of claims 1 to 3, characterized in that: The THAM-ResUNet network processing process based on the UNet network model is as follows: The input feature map first passes through a 7×7 convolution module, and the output of the 7×7 convolution module is processed by the THAM attention mechanism; The output of the 7×7 convolutional module is fed into the bottleneck residual module. There are at least four bottleneck residual modules, and the THAM attention mechanism is introduced after each bottleneck residual module. Different bottleneck residual convolution blocks include multiple bottleneck convolution residual units. Each bottleneck residual convolution unit has three convolution layers, and the input and the output of the last convolution layer are connected with residuals. The residual connection in each bottleneck residual convolution unit is then processed by the ReLU activation layer. The output of the last bottleneck residual module is fed into the basic residual convolution module. There are at least three basic residual convolution modules, each of which has two convolution layers. The residual connection in each basic residual convolution module is processed by the ReLU activation layer. The THAM attention mechanism is introduced after each basic residual convolution block. The jump connection dimension is adjusted by upsample after the THAM attention mechanism of the penultimate basic residual convolution module and the penultimate third basic residual convolution module. The THAM attention mechanism set after the 7×7 convolution module and the first two bottleneck residual modules is skip-connected with the last three basic residual convolution blocks in the decoder.
5. The radar image segmentation method based on THAM-ResUNet according to claim 4, characterized in that: The 7×7 convolution module is provided with a convolution layer with a convolution kernel size of 7×7, followed by a batch normalization layer, a maximum pooling layer, and a ReLU activation layer.
6. The radar image segmentation method based on THAM-ResUNet according to claim 5, characterized in that: In the 7×7 convolution module, the convolution step size of the 7×7 convolution layer is 2 pixels, and the step size of the maximum pooling layer is 2 pixels.
7. The radar image segmentation method based on THAM-ResUNet according to claim 4, characterized in that: The number of the bottleneck residual modules is set to 4, and the number of the basic residual convolution modules is set to 3; The four bottleneck residual modules are denoted as bottleneck residual module 1-neck residual module 4; bottleneck residual module 1 has 3 bottleneck convolution residual units, and only the first convolution unit is down-sampled; bottleneck residual module 2 has 4 bottleneck convolution residual units, and only the first convolution unit is down-sampled; bottleneck residual module 3 has 6 bottleneck convolution residual units, and only the first convolution unit is down-sampled; bottleneck residual module 4 has 3 bottleneck convolution residual units.
8. The radar image segmentation method based on THAM-ResUNet according to claim 7, characterized in that: The structure of the bottleneck residual convolution unit is as follows: When the bottleneck residual convolution unit is set to downsample, the convolution kernel size in the first convolution layer is 1×1 pixel and the stride is 1 pixel; the convolution kernel size in the second convolution layer is 3×3 pixels and the stride is 2 pixels when downsampling. The residual connection dimension is adjusted by the downsample branch; the convolution kernel size in the third convolution layer is 1×1 pixel and the stride is 1 pixel. Each convolution layer is followed by a batch normalization layer. When the bottleneck residual convolution unit does not set downsampling processing, the convolution kernel size in the first convolution layer is 1×1 pixel and the stride is 1 pixel; the convolution kernel size in the second convolution layer is 3×3 pixels and the stride is 1 pixel; the convolution kernel size in the third convolution layer is 1×1 pixel and the stride is 1 pixel; each convolution layer is followed by a batch normalization layer.
9. The radar image segmentation method based on THAM-ResUNet according to claim 7, characterized in that: The two convolutional layers of the basic residual convolution module are as follows: The convolution kernel size in the first convolution layer is 3×3 pixels and the stride is 1 pixel; The convolution kernel size in the second convolution layer is 3×3 pixels and the stride is 1 pixel; Each convolutional layer is followed by a batch normalization layer.
10. A radar image segmentation device based on THAM-ResUNet, characterized in that: The device includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the radar image segmentation method based on THAM-ResUNet according to any one of claims 1 to 9.
Citation Information
Patent Citations
Medical image segmentation method based on multiple scales and attention
CN114359292A
Structured light illumination microscope reconstruction method of amplitude phase channel attention
CN115689889A
Solder paste stirring intelligent temperature control system and method based on temperature sensor
CN119200711A
Object detection model and method for detecting object occupying fire escape route, and use
WO2023207163A1
Three-dimensional point-cloud semantic segmentation method based on multi-level boundary enhancement for unstructured environment
WO2024230038A1