High-frequency sensing semi-supervised electromagnetic shielding light window image segmentation method based on multitask and transformation consistency learning

By combining the HAMTC-Net network with fully enhanced transformation consistency, equivariant loss and high-frequency perception modules, the problems of transformation consistency and insufficient learning ability of multi-task encoders in existing methods are solved, and the accuracy and generalization performance of electromagnetic shielding window image segmentation are improved.

CN120707848APending Publication Date: 2025-09-26SHAANXI UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510725833.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing semi-supervised image segmentation methods fail to fully utilize the global learning capabilities of transformation consistency and multi-task reinforcement encoders, resulting in limited understanding of complex scenes. In particular, in electromagnetic shielding window images, high-frequency features are not fully utilized, affecting the generalization performance of the model.

Method used

A high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning is adopted. By constructing the HAMTC-Net network, combining fully enhanced transformation consistency, equivariant loss and high-frequency perception modules, and using wavelet transform to extract high-frequency features, the segmentation accuracy and generalization ability of the model are enhanced.

Benefits of technology

The segmentation accuracy and generalization ability of electromagnetic shielding window images have been significantly improved, the model's efficiency in utilizing unlabeled data has been improved, the dependence on large amounts of labeled data has been reduced, the ability to understand complex scenes has been enhanced, and uncertainty has been reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707848A_ABST
    Figure CN120707848A_ABST
Patent Text Reader

Abstract

The invention discloses a high-frequency perception semi-supervised electromagnetic shielding light window image segmentation method based on multitask and transformation consistency learning, and the method comprises the steps: 1, constructing an electromagnetic shielding light window OM image data set, carrying out the shooting of an electromagnetic shielding light window, selecting an original image, generating a corresponding same-name label image, and dividing the same-name label image into a training set, a test set, and a verification set; 2, constructing a semi-supervised electromagnetic shielding light window image segmentation network model HAMTC-Net; 3, a network model HAMTC-Net is trained; 4, evaluating the performance of the HAMTC-Net network model by using the verification set and optimizing parameters; and 5, inputting an electromagnetic shielding light window image to be segmented into the trained HAMTC-Net network model, and outputting a segmentation result. According to the method, the global learning ability of a network encoder is enhanced through a multi-task method while the transformation consistency is fully utilized, and the high-frequency characteristics of data are fully utilized to guide training, so that the segmentation precision and generalization ability of the model in a complex electromagnetic shielding light window image are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of image processing and computer vision, and specifically relates to a high-frequency perception semi-supervised electromagnetic shielding light window image segmentation method based on multi-task and transformation consistency learning. Background Art

[0002] Electromagnetic shielding optical windows effectively protect against electromagnetic radiation, ensuring the proper functioning of optoelectronic instruments and equipment. They are currently widely used in precision electronic and electrical equipment, computers, aerospace, and military electronic equipment and equipment. Electromagnetic shielding windows primarily consist of mesh regions and cracks (the gaps between meshes). Their structural parameters (including the number, size, and distribution of meshes, crack width, and crack duty cycle) directly determine the structural characteristics of the resulting metal grid after metallization, thus playing a crucial role in the electromagnetic shielding performance of the optical window. Therefore, it is necessary to statistically analyze the structural parameters of electromagnetic shielding windows to indirectly establish a correlation between experimental conditions and window performance.

[0003] To improve the quality of optical windows, researchers typically use optical microscopes (OMs) to image electromagnetic shielding windows and estimate their performance by analyzing the structural parameters of sample images. Optical microscopes are instruments that use optical principles to observe tiny objects. By leveraging the refraction, scattering, and magnification of light, OMs enable researchers to observe and study the minute structures of electromagnetic shielding windows.

[0004] Currently, the statistical analysis of the structural parameters of electromagnetic shielding windows relies primarily on manual measurement, which has significant limitations. First, an OM image of an electromagnetic shielding window typically contains hundreds of irregular grids. Manually measuring the shape parameters of each grid is not only time-consuming and laborious, but also results in inaccurate results due to irregular grid sizes and shapes, and uneven crack widths. Furthermore, manual measurement can only provide representative sampling of specific parameters and lacks statistical analysis, often leading to large errors in the analysis results. Therefore, a feasible approach is to first segment the OM image of the electromagnetic shielding window into a binary image and then use digital image processing methods to quickly obtain the structural parameters.

[0005] Traditional image segmentation methods are time-consuming and labor-intensive, with low detection accuracy and poor results. In recent years, deep learning has rapidly developed, with its accuracy and automation far exceeding traditional methods, and has now become the mainstream method in the field of image segmentation.

[0006] With the rapid development of deep learning, convolutional neural networks (CNNs) have made significant contributions to image segmentation. Ronneberger et al. proposed the symmetrical encoder-decoder network U-Net, which significantly improved image segmentation accuracy.

[0007] CNN-based methods, such as CE-Net, SA-Net, SGU-Net, and PHNet, utilize hierarchical representations to capture local features of images. These methods typically incorporate functional modules, such as pyramid feature fusion, attention mechanisms, and depthwise separable convolutions, to enhance the network's feature representation capabilities. As the Transformer model demonstrates excellent performance in natural language processing tasks, researchers have applied it to image segmentation tasks and achieved outstanding performance. Transformer-based methods, such as TransUNet, SwinUnet, ConvFormer, and FCT, utilize self-attention mechanisms to process global image information, enabling the network to better capture long-range dependencies between pixels, thereby improving the accuracy and robustness of medical image segmentation.

[0008] The aforementioned segmentation networks typically require large amounts of labeled data to achieve better generalization performance. However, electromagnetic shielding windows are difficult to label, typically requiring pixel-by-pixel comparison by experts, a time-consuming and expensive process. To address this issue, semi-supervised methods require only a small amount of labeled data and a large amount of unlabeled data to achieve superior detection performance. To fully utilize existing unlabeled data to solve image segmentation problems and reduce reliance on large amounts of labeled training data, a growing number of researchers are exploring the application of semi-supervised learning methods to image segmentation tasks.

[0009] Consistency learning is widely used in the field of semi-supervised learning. The core idea of ​​these methods is to implement pixel-level consistency in the output of the network to cope with different perturbations. For example, Yu et al. proposed an uncertainty-aware self-ensemble model (UA-MT) based on the mean teacher model (MT), which uses Monte-Carlo Dropout to exclude unreliable predictions. In addition, Li et al. proposed a transformation consistency self-ensemble model (TCSM_v2), which uses transformations such as rotation, flipping and scaling for consistency learning. Luo et al. achieved multi-task consistency constraint (DTC) by using an additional regression head to perform task conversion with the original segmentation head. Zhou et al. extracted high-frequency and low-frequency features of the image through wavelet transform to perform feature fusion, and imposed consistency constraints on their different fusion results.

[0010] While existing semi-supervised image segmentation methods have achieved promising results, they currently fail to fully exploit transformation consistency, limiting the model's ability to learn data invariance. Furthermore, existing methods fail to leverage multi-task learning to enhance the encoder's global learning capabilities, limiting their ability to understand complex scenes. Furthermore, for datasets with prominent high-frequency features, existing methods lack consideration of how to fully exploit these high-frequency features in sample images to improve the model's generalization performance on large amounts of unlabeled data, limiting performance gains in unsupervised or semi-supervised learning scenarios. Summary of the Invention

[0011] The purpose of the present invention is to provide a high-frequency perception semi-supervised electromagnetic shielding light window image segmentation method based on multi-task and transformation consistency learning. While making full use of transformation consistency, the global learning ability of the network encoder is enhanced through a multi-task method, and the high-frequency features of the data are fully utilized to guide training, thereby effectively improving the segmentation accuracy and generalization ability of the model in complex electromagnetic shielding light window images.

[0012] In order to achieve the above object, the present invention provides the following technical solutions:

[0013] A high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning includes the following steps:

[0014] Step 1: Construct an electromagnetic shielding window OM image dataset. Take photos of the electromagnetic shielding window, select the original image and generate the corresponding labeled images with the same name. Divide the original image and the corresponding labeled images into a training set, a test set, and a validation set. The validation set and the test set use the same set of images.

[0015] Step 2: Construct a semi-supervised electromagnetic shielding window image segmentation network model HAMTC-Net;

[0016] The network model HAMTC-Net includes network A (student model), network B (teacher model), fully enhanced transformation consistency module, equivariant loss module and high-frequency perception module. Both network A and network B are Unet networks, and network A and network B use the same encoder f θ and decoder f γ , set a classification decoder p after the encoder of network A φ ;

[0017] Step 3: Train the network model HAMTC-Net;

[0018] Step 3.1, initialize the parameters of the network model HAMTC-Net;

[0019] Step 3.2: Input the training set into network A (student model) and network B (teacher model) respectively. The training set images input into network A (student model) are first transformed and then enter the codec to output the corresponding feature map. The training set images input into network B (teacher model) are first transformed into the feature map output by the codec, and then the same transformation as that of network A (student model) is applied to the feature map. Finally, the consistency of the inference results of network A (student model) and network B (teacher model) is maintained.

[0020] Step 3.3, use the equivariant loss module to calculate the loss of the feature map of the input network A (student model), and input the feature map result of the input network A (student model) into the classification decoder p φ , get the predicted transformation method p φ (f θ (π i (x))) and with the actual transformation method π i Perform loss calculation to improve the utilization efficiency of unlabeled data and the generalization ability of the model;

[0021] Step 3.4: For the image of input network A (student model), after transformation π i The result is decomposed by wavelet to obtain the high-frequency components H of the image. f and low-frequency component L f , and get the normalized result H norm , using H norm As a mask, train the model; calculate the consistency loss between the prediction results of network A (student model) and the output results of network B (teacher model) after the transformation in step 3.2 under high-frequency perception constraints, and complete the model training;

[0022] Step 3.5: Calculate the total loss

[0023] Step 3.6: Update the parameters of the network model;

[0024] Step 4: Use the validation set to evaluate the performance of the HAMTC-Net network model and optimize the parameters;

[0025] Step 5: Input the electromagnetic shielding window image to be segmented into the trained HAMTC-Net network model and output the segmentation result.

[0026] Furthermore, the generation process of the label image with the same name in step 1 is: manually annotating the selected original image using labelme software, generating a json file containing the coordinates of the target area corresponding to the original image, processing the original image and the json file, and generating a label image with the same name corresponding to the original image.

[0027] Furthermore, in step 3.2, the difference between the prediction results of the teacher model and the student model is minimized by the mean square error loss (MSE). The specific process is expressed as follows:

[0028]

[0029] in, represents the difference between the prediction results of the student model and the teacher model, z i and Represent the prediction results of the student model and the teacher model, z i =π i (θ s (x i )), where π i To randomly perform Cutmix, Mixup, rotation and flip operations, θ s and θ t Represent the student model and teacher model respectively.

[0030] Furthermore, the specific operation of step 3.3 includes: defining the segmentation model network A in encoder-decoder form: f(x i )=f γ (f θ (x i ), where f represents the segmentation model network A, f θ is the encoder, f γ For the decoder;

[0031] For the input image x i , when it undergoes the transformation π i After that, the corresponding segmentation results will also change, namely:

[0032] f(π i (x i ))=π i (f(x i ))

[0033] And infer that:

[0034] f θ (π(x i ))≠f θ (x i )

[0035] Add a classification head p after the encoder φ (·) to predict the type of data transformation π applied to the image i ;

[0036] Equivariant loss function It is expressed by the following formula:

[0037]

[0038] Among them, π i represents one of Cutmix, Mixup, Rotation and Flip, so C = 4; CE represents the cross entropy loss.

[0039] Furthermore, the normalized result H in step 3.4 is norm The following steps are included: using wavelet transform to decompose the original image into low-frequency component LL, horizontal high-frequency component HL, vertical high-frequency component LH and diagonal high-frequency component HH, and the image low-frequency component L f and high frequency component H f They are defined as:

[0040] L f =LL

[0041] H f =HL+LH+HH

[0042] The high-frequency feature map H obtained by wavelet decomposition of each sample f Pixel-by-pixel normalization, the normalization calculation method can be expressed by the following formula:

[0043]

[0044] Among them, x, y represent the pixel coordinates in the image, H fmin H f The minimum value of the pixels in H fmax H f The maximum value of the pixels in H norm is the normalized result.

[0045] Furthermore, the consistency loss between the prediction result of network A (student model) in step 3.4 and the output result of network B (teacher model) after transformation in step 3.2 under high-frequency perception constraints is defined as high-frequency perception consistency loss It is expressed by the following formula:

[0046]

[0047] in, is the indicator function, is the result of the full enhancement transformation consistency between the predictions of the teacher model and the student model at the p-th pixel, u p is the value of the normalized high-frequency feature map at the p-th pixel, HT is the upper threshold, and LT is the lower threshold;

[0048] The Gaussian warm-up function is used to adjust the values ​​of LT and HT. The process is expressed as follows:

[0049] LT=LT max -LT max *λ(t)

[0050] HT=HT min +(1-HT min )*λ(t)

[0051] Among them, λ(t) represents the Gaussian warm-up function, LT max =0.15 and HT min =0.65 are the maximum value of the lower threshold and the minimum value of the upper threshold respectively. The Gaussian warm-up function is used to gradually increase the upper threshold to 1 and gradually reduce the lower threshold to 0.

[0052] Furthermore, the total loss of step 3.5 is Defined as:

[0053]

[0054] in represents the supervision loss, Unsupervised loss, λ(t) represents the Gaussian warm-up function, where t is the current training step number, t max is the maximum number of training steps;

[0055] Monitoring losses It can be expressed as:

[0056]

[0057] For the labeled training branch, the labeled data is input into the student model, and the cross entropy (CE) loss and Dice loss are used as the supervision loss. To measure the label y l and prediction results the differences between;

[0058] Unsupervised loss function It can be expressed as:

[0059]

[0060] Where γ = 0.1, They are high-frequency perceptual consistency loss and equivariant loss;

[0061] For the unlabeled training branch, first, the unlabeled data is passed through π iAfter transformation, it is input into the student model, and the prediction results of the teacher model are transformed in the same way. The consistency regularization is fully utilized to encourage the model to produce consistent predictions when the results of the transformation and then prediction of the same sample are consistent with the results of the prediction and then transformation. Then, high-frequency pixels are filtered out at the beginning of training through high-frequency perception to obtain Finally, the classification head after the student model encoder improves the model's perception of image transformation and obtains the equivariant loss

[0062] Furthermore, the total loss of step 3.5 is Defined as:

[0063]

[0064] in represents the supervision loss, Unsupervised loss, λ(t) represents the Gaussian warm-up function, where t is the current training step number, t max is the maximum number of training steps;

[0065] Monitoring losses It can be expressed as:

[0066]

[0067] For the labeled training branch, the labeled data is input into the student model, and the cross entropy (CE) loss and Dice loss are used as the supervision loss. To measure the label y l and prediction results the differences between;

[0068] Unsupervised loss function It can be expressed as:

[0069]

[0070] Where γ = 0.1, They are high-frequency perceptual consistency loss and equivariant loss;

[0071] For the unlabeled training branch, first, the unlabeled data is passed through π i After transformation, it is input into the student model, and the prediction results of the teacher model are transformed in the same way. The consistency regularization is fully utilized to encourage the model to produce consistent predictions when the results of the transformation and then prediction of the same sample are consistent with the results of the prediction and then transformation. Then, high-frequency pixels are filtered out at the beginning of training through high-frequency perception to obtain Finally, the classification head after the student model encoder improves the model's perception of image transformation and obtains the equivariant loss

[0072] Compared with the prior art, the present invention has the following beneficial effects:

[0073] (1) A fully enhanced transformation consistency training strategy is proposed. In traditional consistency constraints, due to the computational characteristics of convolutional networks on two-dimensional images, the unsupervised regularization effect of segmentation networks on randomly transformed data is limited. This limitation becomes more obvious when processing light window images with irregular shape features, as traditional transformation consistency constraints usually use simple image enhancement methods. Fully enhanced transformation consistency significantly enhances the model's learning ability for randomly transformed data by simultaneously introducing multiple enhancement methods (including CutMix, MixUp, rotation, and flipping) into the transformation consistency constraint strategy during the training process. This strategy not only enriches the transformation patterns of unlabeled data, but also enhances the constraint effect of regularization on the model, enabling the segmentation network to more fully explore the potential information in unlabeled data.

[0074] (2) A multi-task based equivariant loss is proposed to enhance the equivariance of feature representation in segmentation tasks. In segmentation tasks, the effective feature representation required by the network should be equivariant to different transformation methods, but some enhancement transformations may not conform to the prior knowledge of the segmentation task, thereby affecting the model performance. To address the above problems, the present invention introduces a classifier at the end of the encoder to complete the classification task and predict which enhancement method is applied to the image, thereby enhancing the encoder's ability to learn global features. In addition, the equivariant loss and the full enhancement transformation promote each other. This is because during the training process, when complex enhancement methods are applied to training samples, the classification task gives the model the ability to recognize the full enhancement transformation, further improving the regularization effect of transformation consistency.

[0075] (3) A new high-frequency perception module is proposed. In the existing confidence-based pixel reliability screening method, although high confidence is usually regarded as a reliable prediction result, there are still cases where pixels are evaluated as high confidence but predicted incorrectly. Such errors often have a serious negative impact on the learning of the model. To address this problem, the present invention introduces wavelet transform to extract the high-frequency feature map of the image, and uses the high-frequency feature map to screen the reliability of pixels. Specifically, relatively reliable low-frequency areas are learned at the beginning of training. As the number of training rounds increases, the segmentation ability of the model is enhanced, and the learning target gradually shifts to complex high-frequency details. This learning strategy enables the model to gradually transition from areas with high certainty to areas with high uncertainty, thereby effectively reducing the uncertainty of the model and improving the segmentation effect of the model. In addition, the high-frequency perception module also works in conjunction with the full enhancement transform consistency. When the transformation of the image data is too strong, the high-frequency feature map can effectively screen the enhanced complex pixels, thereby avoiding the overfitting problem.

[0076] (4) The present invention is comprehensively verified experimentally on an electromagnetic shielding window OM image dataset. The results show that HAMTC-Net has better segmentation performance than the current mainstream semi-supervised learning method. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is the overall flow chart of the present invention;

[0078] Figure 2 This is a diagram showing part of the electromagnetic shielding window OM data set;

[0079] Figure 3 It is the result of wavelet decomposition of the electromagnetic shielding window OM image;

[0080] Figure 4 These are the segmentation results of different methods on the electromagnetic shielding window dataset using 10% labeled data. DETAILED DESCRIPTION

[0081] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0082] like Figure 1 As shown, this paper proposes a new semi-supervised training strategy. Its main contribution lies in introducing more complex CutMix and Mixup enhancement methods to the transformation consistency method. On this basis, it proposes an equivariant loss, which introduces the classification task into the encoder of the segmentation network. Through a multi-task training strategy, the network's feature learning ability is improved, complementing the transformation consistency regularization. Finally, a high-frequency perception method is proposed. It extracts high-frequency features of the image through wavelet transform. The high-frequency feature map is used to guide the model training at the pixel level to improve segmentation quality.

[0083] The high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning described in this embodiment has the following specific steps:

[0084] Step 1: Construct an electromagnetic shielding window OM image dataset. Take a picture of the electromagnetic shielding window, select the original image and generate the corresponding labeled image with the same name. Divide the original image and the corresponding labeled image into a training set, a test set, and a validation set. The validation set and the test set use the same set of images.

[0085] This paper constructs an electromagnetic shielding window OM image dataset. The electromagnetic shielding windows were photographed in an electromagnetic shielding window preparation laboratory. The resolution of the images taken by an optical microscope (Olympus BX53M) is 2592×1944. The electromagnetic shielding window OM images have a large number of grids, uneven sizes and irregular shapes, fine cracks and irregular curvature, and low contrast. For example, some samples have Figure 2 shown

[0086] Representative images were selected from the collected data and manually annotated using labelme software. A JSON file containing the coordinates of the target area corresponding to the original image was generated. The original image and JSON file were processed to generate a labeled image with the same name as the original image. After the above processing, a total of 88 original images of electromagnetic shielding light window OM images were obtained. In the experiment of the present invention, the original images and the corresponding labeled images were divided into training and test sets, and they were cropped to a uniform size of 256×256. The final result was a total of 1890 training set images and 4270 test set images. The same set of images was used for the validation and test sets. The training set was divided into [5%, 95%] and [10%, 90%] according to labeled and unlabeled.

[0087] Step 2: Construct a semi-supervised electromagnetic shielding window image segmentation network model HAMTC-Net

[0088] The network model HAMTC-Net includes network A (student model), network B (teacher model), fully enhanced transformation consistency module, equivariant loss module and high-frequency perception module. Both network A and network B are Unet networks, and network A and network B use the same encoder f θ and decoder f γ , there is a classification decoder p after the encoder of network A φ .

[0089] Step 3: Train the network model HAMTC-Net

[0090] Step 3.1: Initialize the parameters of the network model HAMTC-Net

[0091] Step 3.2: Input the training set into network A (student model) and network B (teacher model) respectively. The training set image input into network A (student model) is first transformed by the transformation method π i Then enter the feature map corresponding to the codec output, and the training set image of the input network B (teacher model) first passes through the codec output to obtain the feature map, and then the same π as the network A (student model) is applied to the feature map. i Transformation, and finally maintain the consistency of the inference results of network A (student model) and network B (teacher model)

[0092] In the classic MT model, a consistency loss is constructed by perturbing the difference between the teacher and student models on unlabeled data, and a supervision loss is constructed on labeled data for joint learning. MT first performs supervised learning on labeled data while using the teacher model's predictions as pseudo-labels for the unlabeled data. Using different regularization methods, the student and teacher models are constrained to maintain consistency in their predictions under different perturbations. Finally, feedback from the supervision loss and consistency loss is used to update the student model.

[0093] The student model updates its parameters through backpropagation, and the teacher model's parameters are the exponential moving average (EMA) of the student model's weights. This operation enables the teacher model to continuously accumulate the network's historical prediction information for unlabeled data. The specific process of EMA can be expressed as follows:

[0094]

[0095] Among them, θ t represents the teacher model parameters, θ s represents the student model parameters, i represents the number of iterations, and α = 0.99.

[0096] In image segmentation tasks, if the input sample image undergoes transformations (such as cutmix, mixup, rotation, and flipping), the segmentation results of the corresponding sample image obtained by the segmentation network will also be transformed in the same way. However, the feature maps obtained by convolution calculation during the segmentation process do not necessarily have the same transformation effect. This phenomenon limits the unsupervised regularization effect of the segmentation network on randomly transformed data.

[0097] This paper proposes a fully enhanced variation consistency method, which simultaneously introduces Cutmix, Mixup, rotation and flipping into the training strategy of transformation consistency. Transformation consistency can enhance the regularization constraint effect and more effectively utilize unlabeled data in segmentation tasks.

[0098] For an unlabeled input image x i , in the student model, π i Applied to x i , while in the teacher model, the transformation π i After applying to the output. i During the application, random perturbation Gaussian noise is applied to network A and network B. The difference between the prediction results of the teacher model and the student model is minimized by the mean square error loss (MSE). The network is thus regularized and has transformation consistency, and its generalization ability and robustness are also improved. The process can be expressed as:

[0099]

[0100] in, represents the difference between the prediction results of the student model and the teacher model, z i and Represent the prediction results of the student model and the teacher model, z i =π i (θ s (x i )), where π i To randomly perform Cutmix, Mixup, rotation and flip operations, θ s and θ t Represent the student model and teacher model respectively.

[0101] Through the above steps, the transformation consistency is extended to more complex enhancement types. i After processing, the consistency of the inference results of the teacher model and the student model is maintained, which enhances the consistency regularization effect of semi-supervised learning.

[0102] Step 3.3, use the equivariant loss module to calculate the loss of the feature map of the input network A (student model), and input the feature map result of the input network A (student model) into the classification decoder p φ , get the predicted transformation method p φ (f θ (π i (x))) and with the actual transformation method π i Perform loss calculation to improve the utilization efficiency of unlabeled data and the generalization ability of the model

[0103] In the method of constructing consistency through transformation, the effective feature representation required for the segmentation task is the same for different transformations. However, some transformations do not conform to the prior knowledge of the segmentation task, such as geometric transformations and more complex strong enhancement methods. Therefore, the effective feature representation required for the segmentation task should be equivariant (or discriminative) for different geometric transformations and strong enhancement methods.

[0104] To address the above problems, the present invention proposes to add equivariant loss to the segmentation model to enable the encoder to learn global information. Specifically, the segmentation model network A is defined as an encoder-decoder form: f(x i )=f γ (f θ (x i ), where f represents the segmentation model network A, f θ is the encoder, f γ For the decoder. For the input image x i , when it undergoes the transformation π i After that, the corresponding segmentation results will also change, namely:

[0105] f)π i (x i ))=π i (f(x i )) (3)

[0106] Furthermore, it can be inferred that:

[0107] f θ (π(x i ))≠f θ (x i ) (4)

[0108] Therefore, we can strengthen the θ For the transformation information π i By adding a classification head p φ (·) to predict the type of random transformation, equivariant loss function It can be expressed by the following formula:

[0109]

[0110] Among them, π i represents one of Cutmix, Mixup, Rotation and Flip, so C = 4; CE represents the cross entropy loss.

[0111] Step 3.4: For the image of input network A (student model), after transformation π i The result is decomposed by wavelet to obtain the high-frequency components H of the image. f and low-frequency component L f , and get the normalized result H norm , using H norm As a mask, train the model; calculate the consistency loss between the prediction result of network A (student model) and the output result of network B (teacher model) after the transformation in step 3.2 under the high-frequency perception constraint, and complete the model training

[0112] In the frequency domain, the high-frequency components of the image correspond to information such as the edges, textures, and details of the image in the spatial domain, which is also a challenging area in the image segmentation task. Curriculum learning is a training strategy that aims to imitate the human learning process. It advocates that the model start learning from easy data and gradually advance to complex data. Therefore, the present invention proposes a high-frequency perception method. For high-frequency areas of unlabeled data, the prediction results of the student model are more likely to be unreliable. The present invention pays more attention to the low-frequency components of the image in the early stages of training. This is to allow the model to learn relatively reliable data in the early stages of training. As the number of training rounds increases, the model's segmentation ability gradually increases, so we let the model pay more attention to more complex high-frequency details.

[0113] Specifically, given a training image, the student model branch first transforms the sample image. It then uses a wavelet transform to obtain a high-frequency feature map of the transformed sample, and applies this to the model to generate a prediction map. Finally, the model is optimized using a consistency loss guided by the high-frequency feature map. Initially, the student model focuses on relatively reliable low-frequency features, guided by the high-frequency feature map. The detailed process is described below.

[0114] 2D images are essentially 2D discrete non-stationary signals containing different frequency ranges and spatial position information. Wavelet transform is used to decompose the original image into low-frequency components, horizontal high-frequency components, vertical high-frequency components and diagonal high-frequency components (LL, HL, LH and HH). f and high frequency component H f They are defined as:

[0115] L f =LL (6)

[0116] H f =HL+LH+HH (7)

[0117] The electromagnetic shielding window OM image is decomposed by wavelet as follows Figure 3 As shown, Figure 3 (a) is the original image, (b) is the result of wavelet decomposition, and (c) is the low-frequency component L of the image. f , (d) is the high frequency component H of the image f We can see that H f Focuses mainly on image details, while L f Focus on the overall semantics of the image.

[0118] Therefore, the high-frequency feature map H obtained by wavelet decomposition of each sample f Pixel-by-pixel normalization, the normalization calculation method can be expressed by the following formula:

[0119]

[0120] Among them, x, y represent the pixel coordinates in the image, H fmin H f The minimum value of the pixels in H fmax H f The maximum value of the pixels in H norm is the normalized result.

[0121] After obtaining the normalized high-frequency feature map, relatively unreliable predictions are filtered out, and only relatively reliable predictions are selected as the learning targets of the student model, and the high-frequency perceptual consistency loss is The pixel-level mean square error (MSE) loss designed for the teacher and student models can be expressed as follows:

[0122]

[0123] in, is the indicator function, is the result of the full enhancement transformation consistency between the predictions of the teacher model and the student model at the p-th pixel, u p is the value of the normalized high-frequency feature map at the p-th pixel, HT is the upper threshold, and LT is the lower threshold. By perceiving the high frequency of the image during training, both the student and teacher models can learn more reliable knowledge, thereby reducing the overall uncertainty of the model. The values ​​of LT and HT are adjusted using a Gaussian warmup function. This process can be expressed as follows:

[0124] LT=LT max -LT max *λ(t) (10)

[0125] HT=HT min +(1-HT min )*λ(t) (11)

[0126] Among them, λ(t) represents the Gaussian warm-up function, LT max =0.15 and HT min = 0.65 are the maximum and minimum values ​​of the lower and upper thresholds, respectively. A Gaussian warmup function is used to gradually increase the upper threshold to 1 and decrease the lower threshold to 0. As training progresses, our method filters out less and less data, allowing the student model to gradually learn from relatively certain situations to uncertain ones.

[0127] For the teacher model, the same transformation is applied to the network output to obtain pseudo-labels. By calculating the consistency loss between the pseudo-labels and the student model predictions under high-frequency perceptual constraints, the model parameters are optimized to complete model training.

[0128] Step 3.5: Calculate the total loss

[0129] Total loss Defined as:

[0130]

[0131] in represents the supervision loss, Unsupervised loss, λ(t) represents the Gaussian warm-up function, where t is the current training step number, t maxis the maximum number of training steps. Therefore, as the number of training steps increases, the unsupervised loss weight gradually increases and reaches a maximum of 1.

[0132] For the labeled training branch, the labeled data is input into the student model, and the cross entropy (CE) loss and Dice loss are used as the supervision loss. To measure the label y l and prediction results The difference between the supervised loss It can be expressed as:

[0133]

[0134] For the unlabeled training branch, first, the unlabeled data is passed through π i After transformation, it is input into the student model, and the prediction results of the teacher model are transformed in the same way. The consistency regularization is fully utilized to encourage the model to produce consistent predictions when the results of the transformation and then prediction of the same sample are consistent with the results of the prediction and then transformation. Then, high-frequency pixels are filtered out at the beginning of training through high-frequency perception to obtain Finally, the classification head after the student model encoder improves the model's perception of image transformation and obtains the equivariant loss Unsupervised loss function It can be expressed as:

[0135]

[0136] Where γ = 0.1, They are high-frequency perceptual consistency loss and equivariant loss respectively.

[0137] Step 3.6: Update the parameters of the network model

[0138] Step 4: Use the validation set to evaluate the performance of the network model HAMTC-Net and optimize the parameters

[0139] Step 5: Input the electromagnetic shielding window image to be segmented into the trained network model HAMTC-Net and output the segmentation result

[0140] The effects of the present invention can be further illustrated by the following experiments.

[0141] In order to verify the effect of the present invention on electromagnetic shielding window OM image segmentation, the development environment used in the experiment was PyCharm 2023.1, the deep learning framework was PyTorch 1.11.0, and training was performed on a GeForce RTX3090Ti GPU with 24GB VRAM. During the training phase, the present invention uses Daubechies 2 (Db2) wavelet as the basis wavelet for image wavelet decomposition. The Db2 wavelet is a compactly supported orthogonal wavelet with good smoothness and time-frequency localization characteristics. It is commonly used in signal processing and image analysis, and helps to enhance image feature extraction capabilities. The student model uses the Adam optimizer, the learning rate is set to 0.001, the batch size is 8, and the total number of training rounds is 150. The present invention evaluates the model performance by calculating the following indicator parameters, which are:

[0142] (1) The Dice coefficient (DI) mainly measures the similarity between two segmentation images and is the most important measurement indicator in the field of medical image segmentation. The specific calculation method is as follows:

[0143]

[0144] (2) The Jaccard similarity coefficient (JA) indicates the degree of overlap between samples, also known as the intersection-to-union ratio, which represents the ratio of the intersection to the union between the true result and the predicted result. The specific calculation method is as follows:

[0145]

[0146] (3) Sensitivity (SE) indicates the ability of the algorithm to correctly predict positive examples. The specific calculation method is as follows:

[0147]

[0148] The ranges of DI, JA and SE are all between [0,1]. The closer the DI value is to 1, the better the segmentation result is. The closer the JA value is to 1, the higher the overlap between the segmentation result and the true label is. The closer the SE value is to 1, the higher the proportion of correct predictions is.

[0149] Among them, TP represents the number of positive samples correctly predicted by the model as positive, FP represents the number of negative samples incorrectly predicted by the model as positive, and FN represents the number of positive samples incorrectly predicted by the model as negative.

[0150] To demonstrate the superiority of our invention, we compared it with mainstream methods, including MT, UA-MT, TCSMv2, CPS, DTC, and MC-Net. The experimental results are shown in Table 1.

[0151] Table 1 Quantitative comparison with different methods on the electromagnetic shielding window OM image dataset

[0152]

[0153] As can be seen, our method consistently outperforms other methods with varying amounts of labeled data. In particular, with 5% real labeled data, our method surpasses the TCSM_V2 method by 0.88% in the Dice coefficient. With 10% real labeled data, our method surpasses the TCSM_V2 method by 0.7% in the Dice coefficient.

[0154] To further verify the effectiveness of each module in the present invention, an ablation experiment was conducted on a dataset of electromagnetic shielding window OM images with only 5% annotated images. The experimental results are shown in Table 2.

[0155] Table 2 Quantitative analysis of ablation experiments on electromagnetic shielding window OM image dataset

[0156]

[0157] Without any semi-supervised techniques, the basic segmentation network Unet can obtain a Dice indicator result of 81.20% when trained with only 5% labeled data (SupOnly).

[0158] Ablation experiments on transformation consistency: Comparing the experimental results of Schemes 1, 2, and 3, we can see that compared to the other two modules, transformation consistency alone has the greatest improvement in semi-supervised learning, with a 2.23% increase in the Dice coefficient compared to the MT model. This shows that after introducing more complex transformations such as CutMix and Mixup, transformation consistency in regularization can enable the network to learn more robust features of samples, thereby improving segmentation performance.

[0159] Ablation experiments on equivariant loss and high-frequency perception: Comparing the experimental data in Table 2 shows that in Scheme 5, which simultaneously optimizes the equivariant loss and high-frequency perception, the Dice coefficient increases from 86.23% to 86.92% compared to Scheme 2. This demonstrates that high-frequency perception complements the equivariant loss, addressing the potential for overly strong random transformations to inhibit the learning performance of the student model. Comparing Schemes 1 and 6, we see that by adding a classification task to construct the equivariant loss, the Dice coefficient increases from 86.45% to 87.16%. This demonstrates that the equivariant loss enables the segmentation model encoder to identify the type of image transformation, enhancing the model's global perception. While the model regularizes transformation consistency, the equivariant loss optimizes the segmentation model encoder parameters, complementing transformation consistency and improving the model's segmentation performance.

[0160] In order to more intuitively reflect the superiority of the present invention, Figure 4 The visualization segmentation results of the test set of the above models are shown on the electromagnetic shielding window OM image dataset using only 10% of the labeled data. Figure 4 As can be seen in the figure, the detection results of the present invention are closer to the labeled image, with clearer boundaries. This is due to the fact that the present invention utilizes more complex transformation consistency while using equivariant loss, thereby enhancing the network's feature extraction capabilities. The present invention exhibits particularly improved segmentation performance for small light spots on electromagnetic shielding window OM images. This is due to the high-frequency perception module's ability to extract high-frequency information from the sample image and guide the model training from highly reliable pixels using the sample's high-frequency feature map, thereby improving the model's segmentation results.

Claims

1. A high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning, characterized by: The following steps are involved: Step 1: Construct an electromagnetic shielding window OM image dataset. Take photos of the electromagnetic shielding window, select the original image and generate the corresponding labeled images with the same name. Divide the original image and the corresponding labeled images into a training set, a test set, and a validation set. The validation set and the test set use the same set of images. Step 2: Construct a semi-supervised electromagnetic shielding window image segmentation network model HAMTC-Net; The network model HAMTC-Net includes network A (student model), network B (teacher model), fully enhanced transformation consistency module, equivariant loss module and high-frequency perception module. Both network A and network B are Unet networks, and network A and network B use the same encoder f θ and decoder f γ , set a classification decoder p after the encoder of network A φ ; Step 3: Train the network model HAMTC-Net; Step 3.1, initialize the parameters of the network model HAMTC-Net; Step 3.2: Input the training set into network A (student model) and network B (teacher model) respectively. The training set images input into network A (student model) are first transformed and then enter the codec to output the corresponding feature map. The training set images input into network B (teacher model) are first transformed into the feature map output by the codec, and then the same transformation as that of network A (student model) is applied to the feature map. Finally, the consistency of the inference results of network A (student model) and network B (teacher model) is maintained. Step 3.3, use the equivariant loss module to calculate the loss of the feature map of the input network A (student model), and input the feature map result of the input network A (student model) into the classification decoder p φ , get the predicted transformation method p φ (f θ (π i (x))) and with the actual transformation method π i Perform loss calculation to improve the utilization efficiency of unlabeled data and the generalization ability of the model; Step 3.4: For the image of input network A (student model), after transformation π i The result is decomposed by wavelet to obtain the high-frequency components H of the image. f and low-frequency component L f , and get the normalized result H norm , using H norm As a mask, train the model; calculate the consistency loss between the prediction results of network A (student model) and the output results of network B (teacher model) after the transformation in step 3.2 under high-frequency perception constraints, and complete the model training; Step 3.5: Calculate the total loss Step 3.6: Update the parameters of the network model; Step 4: Use the validation set to evaluate the performance of the HAMTC-Net network model and optimize the parameters; Step 5: Input the electromagnetic shielding window image to be segmented into the trained HAMTC-Net network model and output the segmentation result.

2. The high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning according to claim 1 is characterized in that: The process of generating the labeled image with the same name in step 1 is as follows: manually annotate the selected original image using labelme software, generate a json file containing the coordinates of the target area corresponding to the original image, process the original image and the json file, and generate a labeled image with the same name corresponding to the original image.

3. The high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning according to claim 2 is characterized in that: In step 3.2, the difference between the prediction results of the teacher model and the student model is minimized by the mean square error loss (MSE). The specific process is expressed as follows: in, represents the difference between the prediction results of the student model and the teacher model, z i and Represent the prediction results of the student model and the teacher model, z i =π i (θ S (x i )), where π i To randomly perform Cutmix, Mixup, rotation and flip operations, θ S and θ t Represent the student model and teacher model respectively.

4. The high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning according to claim 3 is characterized in that: The specific operations of step 3.3 include: defining the segmentation model network A in encoder-decoder form: f(x i )=f γ (f θ (x i ), where f represents the segmentation model network A, f θ is the encoder, f γ For the decoder; For the input image x i , when it undergoes the transformation π i After that, the corresponding segmentation results will also change, namely: f)p i (x i ))=π i (f(x i )) And infer that: f θ (π(x i ))≠f θ (x i ) Add a classification head p after the encoder φ (·) to predict the type of data transformation π applied to the image i ; Equivariant loss function l e It is expressed by the following formula: Among them, π i represents one of Cutmix, Mixup, Rotation and Flip, so C = 4; CE represents the cross entropy loss.

5. The high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning according to claim 4 is characterized in that: The normalized result H in step 3.4 norm The following steps are included: using wavelet transform to decompose the original image into low-frequency component LL, horizontal high-frequency component HL, vertical high-frequency component LH and diagonal high-frequency component HH, and the image low-frequency component L f and high frequency component H f They are defined as: L f =LL <h2 style=";text-align:left;direction:ltr">H<h2 style=";text-align:left;direction:ltr"> f <h2 style=";text-align:left;direction:ltr"> =HL+LH+HH The high-frequency feature map H obtained by wavelet decomposition of each sample f Pixel-by-pixel normalization, the normalization calculation method can be expressed by the following formula: Among them, x, y represent the pixel coordinates in the image, H fmin H f The minimum value of the pixels in H fmax H f The maximum value of the pixels in H norm is the normalized result.

6. The high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning according to claim 5 is characterized in that: The consistency loss between the prediction result of network A (student model) in step 3.4 and the output result of network B (teacher model) after transformation in step 3.2 under high-frequency perception constraints is defined as high-frequency perception consistency loss It is expressed by the following formula: in, is the indicator function, is the result of the full enhancement transformation consistency between the predictions of the teacher model and the student model at the p-th pixel, u p is the value of the normalized high-frequency feature map at the p-th pixel, HT is the upper threshold, and LT is the lower threshold; The Gaussian warm-up function is used to adjust the values ​​of LT and HT. The process is expressed as follows: LT=LT max -LT max *λ(t) HT=HT min +(1-HT min )*λ(t) Among them, λ(t) represents the Gaussian warm-up function, LT max =0.15 and HT min =0.65 are the maximum value of the lower threshold and the minimum value of the upper threshold respectively. The Gaussian warm-up function is used to gradually increase the upper threshold to 1 and gradually reduce the lower threshold to 0.

7. The high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning according to claim 6 is characterized in that: The total loss of step 3.5 Defined as: in represents the supervision loss, Unsupervised loss, λ(t) represents the Gaussian warm-up function, where t is the current training step number, t max is the maximum number of training steps; Monitoring losses It can be expressed as: For the labeled training branch, the labeled data is input into the student model, and the cross entropy (CE) loss and Dice loss are used as the supervision loss. To measure the label y l and prediction results the differences between; Unsupervised loss function It can be expressed as: Where γ = 0.1, l e They are high-frequency perceptual consistency loss and equivariant loss; For the unlabeled training branch, first, the unlabeled data is passed through π i After transformation, it is input into the student model, and the prediction results of the teacher model are transformed in the same way. The consistency regularization is fully utilized to encourage the model to produce consistent predictions when the results of the transformation and then prediction of the same sample are consistent with the results of the prediction and then transformation. Then, high-frequency pixels are filtered out at the beginning of training through high-frequency perception to obtain Finally, the classification head after the student model encoder improves the model's perception of image transformation and obtains the equivariant loss l e .

8. The high-frequency perception semi-supervised electromagnetic shielding window image segmentation method based on multi-task and transformation consistency learning according to claim 7 is characterized in that: In step 4, the Dice coefficient (DI), Jaccard similarity coefficient (JA) and sensitivity (SE) are used to evaluate the performance of the network model HAMTC-Net, specifically: Among them, TP represents the number of positive samples correctly predicted by the model as positive, FP represents the number of negative samples incorrectly predicted by the model as positive, and FN represents the number of positive samples incorrectly predicted by the model as negative. The ranges of DI, JA and SE are all between [0,1]. The closer the DI value is to 1, the better the segmentation result is. The closer the JA value is to 1, the higher the overlap between the segmentation result and the true label is. The closer the SE value is to 1, the higher the proportion of correct predictions is.

Citation Information

Cited By

  • Semi-supervision-based astronomical time-frequency image target detection method and system

    CN120764617A

  • A semi-supervised based astronomical time-frequency image target detection method and system

    CN120764617B