Double-student teacher network semi-supervised crack segmentation method based on uncertain sensing, medium and equipment
By adopting uncertain perception of dual-student teacher network and dynamic uncertainty patch adjustment technology in crack detection, the existing method relies on large amounts of labeled data and is susceptible to noise, achieving efficient crack segmentation, especially excellent performance in complex contexts.
Patent Information
- Application Number
- CN202510326789.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-27
AI Technical Summary
The existing crack detection methods rely on a large amount of labeled data, and the model is susceptible to noise, resulting in a decrease in detection accuracy, especially in complex backgrounds, which is difficult to maintain excellent segmentation performance.
A semi-supervised crack segmentation method based on uncertain perception is proposed. By constructing a dual-student teacher network (DSMT), a high-quality new training sample is generated using dynamic uncertainty patch adjustment (DUPA) technology, and crack details are retained through the hierarchical feature fusion module (HFF).
It effectively reduces the model's dependence on annotated data, and can maintain excellent segmentation performance even with a small amount of annotated data, improving the segmentation accuracy of the model in complex contexts.
Smart Images

Figure CN120219744A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and deep learning, and particularly relates to a crack image segmentation method and system based on semi-supervised learning, which is especially applicable to the scenario of civil engineering structure health monitoring. Background Art
[0002] Cracks are fractures or gaps on the surfaces of building structures such as concrete and masonry caused by natural weathering, material deformation, or external forces. The development of cracks may lead to structural instability, resulting in major safety accidents such as bridge collapses and dam damages. Traditional crack detection relies on manual visual inspection, and measures the crack size through professional instruments (such as rulers and magnifying glasses), which has problems such as low efficiency, strong subjectivity, and poor consistency. In special environments (such as underwater facilities), manual detection also needs to face challenges such as safety risks and high operation difficulties.
[0003] Early automated detection technologies (such as digital image processing and traditional machine learning) have partially achieved semi-automatic detection, but require complex parameter adjustment and preprocessing, and are easily affected by the environment, making it difficult to guarantee the detection accuracy. Methods based on deep learning have improved the efficiency through automatic feature learning, but still require a large amount of labeled data to train the model. Especially for crack detection methods based on semantic segmentation, although pixel-level positioning information can be obtained, the labeling cost is high, and labeling errors are likely to occur in regions with blurred boundaries.
[0004] Semi-supervised learning (SSL) alleviates the labeling dependence by combining a small amount of labeled data with a large amount of unlabeled data. Existing methods mainly include:
[0005] Adversarial training: Optimizes the segmentation result through the confrontation between the generator and the discriminator, but the training is unstable and the computational cost is high;
[0006] Self-training: Generates pseudo-labels using the model prediction, but it is difficult to screen out noise and errors are likely to accumulate;
[0007] Consistency learning: Adds perturbations to the model input or features and constrains the output consistency, but the overly coupled teacher-student model (such as Mean Teacher) will cause error propagation, and the knowledge singularity of a single student model limits the performance improvement.
[0008] In addition, existing segmentation networks (such as UNet) are prone to losing detailed information in the encoder-decoder interaction, and redundant features further reduce the robustness of the model, exacerbating the problem of pseudo-label noise. Therefore, developing a method that can effectively achieve pixel-level segmentation of cracks to avoid the decrease in crack detection accuracy caused by the high similarity between the teacher model and a single student model, and the sensitivity of the network to noise and the resulting incorrect pseudo-labels, is the problem to be solved by the present invention. Summary of the Invention
[0009] Aiming at the problems existing in the prior art, the purpose of the present invention is to propose a semi-supervised crack segmentation method based on an uncertain perception dual student-teacher network.
[0010] The technical solution for achieving the above object of the present invention is: a semi-supervised crack segmentation method based on uncertainty guidance, characterized by including the following steps:
[0011] S1: Construct a dual student-teacher network (DSMT), including two independently initialized student models and one teacher model, and both the student model and the teacher model are CrackUnet segmentation models based on an encoder-decoder structure;
[0012] S2: Calculate the supervised loss using the labeled data The is a linear combination of the weighted Dice loss and the cross-entropy loss; perform cross-pseudo-supervision training and consistency training on the unlabeled data, and calculate the unsupervised loss and
[0013] S3: Update the parameters of the student model according to the total loss where λ(t) is a dynamic weight coefficient that changes with the training iteration number t;
[0014] S4: Perform dynamic uncertainty patch adjustment (DUPA) to generate new training samples, including:
[0015] S401: Based on the prediction results of the student model and the teacher model, calculate two information entropy maps as uncertainty maps;
[0016] S402: Divide each uncertainty map into r×r pixel blocks (patches), and select the top K blocks with the highest entropy values;
[0017] S403: Mark the regions corresponding to the K blocks in the original sample as the background to generate a denoised new sample;
[0018] S5: Only use the new samples to retrain the network, update the parameters of the teacher model to the exponential moving average (EMA) of the parameters of the two student models, output the segmentation result of the teacher model for the input crack image, and generate a binary segmentation mask.
[0019] Further, the decoder of the CrackUnet model in step S1 includes a skip-fusion module, and the skip-fusion module performs the following operations: use a dynamic upsampler (DySample) to increase the resolution of the decoder features; input the upsampled features and the encoder features into a hierarchical feature fusion module (HFF), and generate channel weights through channel concatenation, global pooling, convolution, and Sigmoid activation; calibrate the encoder features based on the channel weights and perform spatial dimension fusion with the decoder features.
[0020] Further, the formula for generating the channel weights of the HFF module is:
[0021] w c = Sigmoid(Conv1(MaxPool([f1; f2]) + AvgPool([f1; f2])))
[0022] where f1 is the encoder feature and f2 is the decoder feature.
[0023] Further, the calculation formula for the dynamic weight coefficient λ(t) in step S3 is:
[0024]
[0025] where λ m represents the maximum value of the weighting function, t represents the current training iteration, and t m represents the maximum iteration during the training process.
[0026] Further, the formula for updating the teacher model parameters in step S5 is:
[0027] where α is used to balance the weight ratios of the two student models, β is used to control the speed of parameter update, and t represents the current training iteration.
[0028] Further, the supervision loss in step 2 is defined as:
[0029]
[0030] L seg = βL dice + (1 - β)L ce
[0031] where L dice is the dice loss function, L ce is the cross-entropy loss function, β is the weight coefficient of the dice loss, p i is the probability output result of the model, and y iis the corresponding true label, |D L | represents the number of labeled data, and H and W represent the height and width of the image respectively.
[0032] Furthermore, the unsupervised loss in step S2 and are defined as:
[0033]
[0034] where and are the probability outputs generated by the student f(θ1) and f(θ2) and the teacher model respectively, and represent the segmentation masks obtained by performing one-hot encoding on the probability outputs generated by the student models f(θ1) and f(θ2) respectively, |D U | represents the number of unlabeled data.
[0035] According to another aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the steps described in any one of the above methods are implemented.
[0036] According to another aspect of the present invention, there is provided an electronic device, the electronic device includes a processor, a memory, and a computer program stored on the memory, and the processor is used to implement the steps described in any one of the above methods when executing the computer program.
[0037] The present invention provides a semi-supervised crack segmentation method based on uncertainty-aware dual student-teacher network. Compared with traditional methods, this method can effectively reduce the dependence of the model on labeled data, and can maintain excellent segmentation performance even when there is only a small amount of labeled data. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings incorporated into the specification and constituting a part of the specification illustrate embodiments of the present invention and, together with the related written description, are used to explain the principles of the present invention. In these drawings, like reference numerals are used to represent like elements. The drawings in the following description are some embodiments of the present invention, not all embodiments. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 Shows a flowchart of a semi-supervised crack segmentation method based on uncertainty-guided dual student-teacher network provided by an embodiment of the present invention.
[0040] Figure 2The schematic diagram of the hierarchical feature fusion module (HFF) in an embodiment method of the present invention is shown.
[0041] Figure 3 The block diagram of a computer device according to an exemplary embodiment is shown. Detailed implementation manners
[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. It should be noted that, without conflict, the embodiments and features in the embodiments of the present application can be combined with each other arbitrarily.
[0043] Fully supervised crack segmentation technology can improve the accuracy of civil structure health monitoring, but highly depends on a large amount of labeled data. To reduce the model's dependence on labeled data, the present invention proposes a semi-supervised crack segmentation based on uncertainty-guided dual student-teacher network. By adopting the method of collaborative training based on the dual student-teacher (DSMT) network, multi-view learning is realized by using the divergence between sub-networks to obtain a richer and more comprehensive feature representation, thereby alleviating the coupling problem between the teacher model and the student model. At the same time, to balance the contradiction between the sub-network divergence and the quality of pseudo-labels, the uncertainty patch adjustment (DUPA) technology based on sub-network prediction is proposed. In addition, the CrackUnet model is proposed to retain crack detail information through hierarchical feature fusion, thereby further improving the accuracy of semi-supervised crack segmentation.
[0044] An embodiment of the present invention provides a semi-supervised crack segmentation method based on uncertainty perception for a dual student-teacher network, and its implementation steps can be more intuitively understood through the Figure 1 flowchart of the semi-supervised crack segmentation method based on uncertainty perception for a dual student-teacher network provided in the appendix, including:
[0045] Step S1: Construct a dual student-teacher network (DSMT), including two independently initialized student models and one teacher model, and both the student model and the teacher model are CrackUnet segmentation models based on the encoder-decoder structure.
[0046] Since the parameters of the teacher model and the student model will gradually tend to be the same as the training progresses, this embodiment solves the coupling problem of the traditional Mean Teacher (MT) framework through the following design: First, two independent student models are used for collaborative training to learn crack features from different perspectives, avoiding the teacher model being overly similar to a single student model due to weight coupling, thus solving the problem that it is difficult to improve the final segmentation accuracy of the network. Different from the conventional MT framework, the core segmentation structure in DSMT consists of three CrackUnet models. In order to retain more crack detail information, the embodiment of the present invention uses CrackUnet as the basic segmentation model of the network. CrackUnet is improved on the basis of the traditional UNet. By introducing residual connections and attention mechanisms, the model's ability to capture crack features in complex backgrounds is enhanced. Commonly used segmentation networks in the prior art, such as DeepLab or PSPNet, often have difficulty effectively distinguishing cracks from similar textures in the background when processing crack images, resulting in insufficient segmentation accuracy. However, CrackUnet, through its unique encoder-decoder structure and multi-scale feature fusion mechanism, can more accurately locate crack regions, especially performing well in the detection of crack edges and thin cracks. In addition, the lightweight design of CrackUnet enables it to operate efficiently in application scenarios with limited computing resources, meeting the real-time requirements of civil engineering structure health monitoring. These three models have the same architecture, including an encoder (pre-trained weights of ResNet34) and a decoder (integrated skip-fusion module), but are initialized with different parameters to ensure the diversity of the initial weights. One of them is used as the teacher model, and the other two are used as student models. The two student models are independent of each other and are trained using their respective paths, so they can learn information from different perspectives. Since the student models are independent of each other and the parameters of the teacher model are updated using the weighted values of the two student parameters, it can effectively prevent the teacher model from overly relying on a single student model and thus accumulating error information that affects the segmentation accuracy.
[0047] In one implementation, optionally, the decoder of the CrackUnet model in this step includes a skip-fusion module, and the skip-fusion module performs the following operations: Use a dynamic upsampler (DySample) to increase the resolution of the decoder features; Input the upsampled features and the encoder features into a hierarchical feature fusion module (HFF), and generate channel weights through channel concatenation, global pooling, convolution, and Sigmoid activation; Calibrate the encoder features based on the channel weights and perform spatial dimension fusion with the decoder features.
[0048] Different from the traditional UNet, in the embodiment of the present invention, a skip-fusion module is used for upsampling to promote the interaction between the encoder and the decoder, screen the features of the skip connection part, and at the same time use a more efficient upsampling method for upsampling instead of using traditional bilinear interpolation.
[0049] The skip fusion module consists of two parts: a hierarchical feature fusion module (HFF) and a dynamic upsampler DySample. Through dynamic upsampling and hierarchical feature fusion, the problem of information loss caused by multiple downsamplings in the traditional UNet is effectively solved. The traditional UNet uses a simple skip connection between the encoder and the decoder. Although it can transfer some detailed information, it cannot effectively fuse the features of different levels, resulting in the easy loss of detailed information. The skip-fusion module solves this problem in the following way:
[0050] Specifically, in the skip-fusion module, DySample is used to upsample the features. Then, in order to promote the interaction between the encoder-decoder and screen out important features, the upsampled features are cross-level fused with the features of the encoding part through the hierarchical feature fusion module (HFF) to retain the crack details and avoid information loss caused by multiple downsamplings in the conventional UNet. The fused features are then concatenated with the upsampled features to obtain the final feature map for subsequent decoder decoding prediction. In this way, the skip-fusion module can more effectively fuse the features of different levels, retain more crack detail information, and thus improve the segmentation accuracy of the model.
[0051] The core idea of the hierarchical feature fusion module (HFF) is to utilize the global information of two different-level features in the encoding-decoding process, screen out important features, and reduce redundant crack information. Existing methods such as simple feature concatenation or element-wise addition often have difficulty effectively utilizing the information of different levels when fusing features, resulting in insufficient feature representation ability. The HFF module solves this problem in the following way:
[0052] Such as Figure 2 , HFF has two branches, which are respectively used to obtain the global spatial information and global channel information of the two features. Specifically, the two features are concatenated along the channels, such as the Concat operation in Figure 2 , which is used for the concatenation step before feature fusion to obtain the channel information of the two features.
[0053] The Skip-Fusion module achieves a balance in reducing information loss, suppressing background noise, and enhancing detail reconstruction through dynamic upsampling and hierarchical feature fusion, enabling CrackUnet to achieve pixel-level precise segmentation in complex engineering scenarios and providing high-quality initial feature representations for semi-supervised training.
[0054] In one implementation, optionally, the channel weight generation formula of the HFF module is:
[0055] w c = Sigmoid(Conv1(MaxPool([f1; f2]) + AvgPool([f1; f2])))
[0056] where f1 is the encoder feature and f2 is the decoder feature.
[0057] To enhance the description of important features while reducing channels, as Figure 2 shown, global pooling operations (including global average pooling GlobalAvgPool and global max pooling GlobalMaxPool, corresponding to the AvgPool operation and MaxPool operation in the formula) are performed using the core components of pooling branches such as GAP and GMP to extract the global information of the feature map, and the weight representation w of the global channel information is obtained through convolution and the Sigmoid activation function c . And the feature is calibrated using w c to increase the weight of important features and suppress irrelevant features, thereby obtaining the fused feature map F t , and this process is expressed by the formula:
[0058] w c = Sigmoid(Conv1(MaxPool([f1; f2]) + AvgPool([f1; f2])))
[0059]
[0060] In addition, to enhance the spatial information correlation of the two features to be fused, in order to further correct the spatial information of the feature F t , the global spatial information w of the feature is obtained by element addition s , and the formula is:
[0061]
[0062] And the final fused feature is calculated:
[0063]
[0064] Therefore, compared with other fusion methods, the HFF module can more effectively fuse features at different levels, retain more detailed information, and suppress background noise, thereby improving the segmentation performance of the model.
[0065] Step S2: Calculate the supervised loss using the labeled data The is a linear combination of the weighted Dice loss and the cross-entropy loss; perform cross-pseudo-supervised training and consistency training on the unlabeled data, and calculate the unsupervised loss and
[0066] The training process of DSMT mainly includes the following two aspects: (1) fully supervised training on the labeled data; (2) consistency training and cross-pseudo-supervised training on the unlabeled data. Specifically, under the constraint of the supervised loss , DSMT effectively utilizes the labeled data and thus learns in the correct direction. For the unlabeled data, constraints are imposed on the two student models through the Lcps loss, which not only increases the utilization of complementary information but also prevents the negative impact on the final segmentation accuracy due to excessive differences between the models. In addition, through the L ts loss, consistency constraints of the teacher model on the two student models are achieved, thereby guiding the student models to learn more abundant knowledge.
[0067] In one implementation, optionally, the supervised loss in this step is defined as:
[0068]
[0069] L seg = βL dice + (1 - β)L ce
[0070] where L dice is the dice loss function, L ce is the cross-entropy loss function, β is the weight coefficient of the dice loss, p i is the probability output result of the model, y i is the corresponding ground truth label, |D L | represents the number of labeled data, and H and W respectively represent the height and width of the image.
[0071] To enable the model to pay more attention to the crack area and improve the accuracy of the shape and boundary of the segmented area, the embodiment of the present invention introduces a weighted dice loss to balance the importance of different classes. At the same time, in view of the fact that the cross-entropy loss can accurately capture pixel-level predictions and improve the fine-grained segmentation ability of the model. Therefore, a supervised data segmentation loss that combines the cross-entropy loss and the weighted dice loss is designed.
[0072] In one implementation, to promote the mutual learning of student models and constrain the consistency between models, the cross-supervision loss between student models and the consistency loss of the teacher guiding the students are used to represent the unsupervised loss. Optionally, the unsupervised loss in this step and are defined as:
[0073]
[0074] where, and are the probability outputs generated by the student models f(θ1) and f(θ2) and the teacher model respectively, and represent the segmentation masks obtained by one-hot encoding the probability outputs generated by the student models f(θ1) and f(θ2) respectively, and |D U | represents the number of unlabeled data.
[0075] Step S3: Update the parameters of the student model according to the total loss where λ(t) is a dynamic weight coefficient that changes with the training iteration number t;
[0076] To comprehensively consider the influence of unlabeled information in different training stages, as well as the evaluation of class balance and fine-grained segmentation ability, a training total loss composed of the supervised loss and the unsupervised loss and is designed.
[0077] In one implementation, optionally, the calculation formula of the dynamic weight coefficient λ(t) in this step is:
[0078]
[0079] where, λ m represents the maximum value of the weighting function, t represents the current training iteration, and t m represents the maximum iteration in the training process, that is, the total number of iterations.
[0080] In this embodiment, a method based on the change of the Gaussian ramp function is designed to adjust the proportion of the unsupervised loss in the total loss. At the initial stage of training, λ(t) has a small value. As the number of training epochs increases, the value of λ gradually increases until it stabilizes. This ensures that the proportion of the unlabeled loss is small at the initial stage of training, avoiding misleading the model to generate too many incorrect pseudo-labels. At the same time, as the training progresses, the proportion of the unsupervised loss is continuously increased to learn more comprehensive unlabeled information and improve the model segmentation performance.
[0081] It should be noted that the setting of the total number of iterations t m is crucial for the training of the model. It directly affects the change trend of the dynamic weight coefficient λ(t), thereby affecting the proportion of the unsupervised loss in the total loss. For datasets with a large scale and high complexity, such as civil engineering structure images containing a large number of complex backgrounds and fine cracks, the total number of iterations needs to be increased accordingly to ensure that the model has enough time to learn the internal characteristics and laws of the data. If a higher segmentation accuracy of the model is required and more fine-tuning of the model parameters is needed, a larger total number of iterations can be set to allow the model to be optimized for a longer time. Therefore, by reasonably setting the total number of iterations t m and combining the change of the dynamic weight coefficient λ(t), the roles of labeled data and unlabeled data in the training process can be effectively balanced, thereby improving the segmentation performance of the model in the semi-supervised learning scenario.
[0082] Step S4: Perform Dynamic Uncertainty Patch Adjustment (DUPA) to generate new training samples, including:
[0083] S401: Calculate two information entropy maps as uncertainty maps based on the prediction results of the student model and the teacher model;
[0084] S402: Divide each uncertainty map into r×r pixel blocks (patches), and select the top K blocks with the highest entropy values;
[0085] S403: Mark the regions corresponding to the K blocks in the original sample as the background to generate a denoised new sample;
[0086] DSMT learns complementary information from different perspectives through the interaction of two student models. However, this interaction may cause the incorrect information generated by one student model to mislead the other student model, leading it to learn in the wrong direction. Especially in some regions with complex backgrounds, significant disagreements are likely to occur between the two models, thus misleading each other and resulting in poor-quality pseudo-labels generated finally, which in turn reduces the segmentation accuracy of the model in these regions. To solve this problem, the embodiment of the present invention creatively proposes the Dynamic Uncertainty Patch Adjustment (DUPA) method. By integrating the prediction results of the student model and the teacher model in the Dual Student Teacher Network (DSMT), an information entropy map (i.e., uncertainty map) reflecting the prediction confidence of the model is generated, and low-confidence regions are screened based on this map for sample reconstruction. This method improves the overall quality of samples by suppressing noise (i.e., regions with higher uncertainty), thereby reducing the generation of incorrect pseudo-labels.
[0087] The core idea of Dynamic Uncertainty Patch Adjustment (DUPA) is to mark the k regions with the lowest confidence in the original sample as the image background, thereby constructing a new sample for retraining the network. To obtain a new sample with less noise information for retraining the model to obtain higher-quality pseudo-labels, the prediction results of the student and the teacher are integrated through steps 401-403 to obtain the uncertainty map, patches with lower confidence are obtained based on the uncertainty map, and patch reset is performed based on the original samples with low confidence. Existing methods such as self-training and adversarial training often have difficulty effectively distinguishing noise and real features when generating pseudo-labels, resulting in poor-quality pseudo-labels and thus affecting the model performance. Through the embodiment of the present invention, the DUPA method can effectively suppress noise, improve the sample quality, thereby generating higher-quality pseudo-labels and finally improving the segmentation accuracy of the model in complex backgrounds.
[0088] Specifically, the prediction results of the student model and the teacher model in DSMT are averaged, and the information entropy is calculated to obtain two uncertainty maps (i.e., UncertaintyMap1 and UncertaintyMap2) for evaluating the reliability of the model's prediction results. The smaller the information entropy value, the smaller the uncertainty of the prediction result, that is, the higher the confidence. Calculate the uncertainty map M through the prediction results of the student model and the teacher model t :
[0089]
[0090] where P i is the probability that the sample is predicted as class i, and is the number of classes (such as two classes of cracks and background).
[0091] Further, in order to obtain high-uncertainty noise regions (such as complex background textures or blurred edges), in the embodiments of the present invention, these two uncertainty maps are divided into multiple patches of size r×r pixels. By calculating the average value of all elements in each patch as the uncertainty index of the patch, and based on this index, the k patches with the lowest confidence in the uncertainty map are obtained. Divide the uncertainty map into patches and calculate the uncertainty index U of each patch ij :
[0092]
[0093] wherein, represents the patch at the i-th row and j-th column in the uncertainty map M t in, and A ij [x,y] represents the element value at the x-th row and y-th column in this patch.
[0094] In order to reduce the regions with high uncertainty in the original sample X, thereby generating high-quality pseudo-labels and improving the accuracy of model prediction, the regions with low confidence are marked as the background. First, based on Uncertain map1, the k regions with the lowest confidence in the original sample X (i.e., the regions corresponding to the blue patches) are marked as the background, and at the same time, the same processing is performed on the original sample X based on Uncertain map2. Then, the network is retrained using the new sample after patch adjustment. This strategy aims to reduce the interference of noise on model training by removing high-uncertainty regions. Traditional methods such as mean filtering or median filtering often blur image details and cause loss of useful information when denoising. However, in the embodiments of the present invention, through the DUPA method, only the regions with the highest model prediction uncertainty are marked as the background, which not only removes the noise but also retains the detailed information of the image. Compared with other denoising methods, this strategy can more effectively improve the sample quality, thereby improving the segmentation performance of the model in complex backgrounds. In specific implementation, the patch division parameter r and the screening quantity K can be dynamically adjusted according to the characteristics of the data set (for example, r = 8 for high-resolution images, and k is adaptively selected according to 10%-20% of the total number of patches) to ensure the robustness of the method in different engineering scenarios. Experimental results show that after adopting this denoising strategy, both the accuracy and recall rate of the model in the crack detection task have been significantly improved, especially in engineering scenarios with complex backgrounds
[0095] Step S5: Retrain the network using the new sample, update the teacher model parameters as the exponential moving average (EMA) of the two student model parameters, and output the segmentation result of the teacher model for the input crack image to generate a binary segmentation mask.
[0096] In this step, only the new training samples generated by S4 are used to replace the original Initial sample Iteratively train the dual student-teacher network, including: inputting the new samples into the student model, updating the parameters of the student model according to the total loss formula, then updating the parameters of the teacher model according to the current parameters of the student model according to the exponential moving average (EMA) rule, and finally outputting the segmentation result of the teacher model for the input crack image to generate a binary segmentation mask. Step S5, through EMA parameter fusion and iterative training of high-quality new samples, while ensuring the inference efficiency, significantly improves the stability, noise resistance, and segmentation accuracy of the model. Its core effect is reflected in the parameter smoothing of the teacher model, enhanced noise suppression ability, and comprehensive optimization of detailed features, providing a reliable final output guarantee for semi-supervised crack segmentation. In traditional methods, the teacher model usually directly copies the parameters of the student model, resulting in the teacher model being easily affected by the noise and errors in the student model. The EMA method gradually integrates the parameters of the student model into the teacher model through weighted averaging, enabling the teacher model to more stably learn the useful information in the student model while suppressing the influence of noise and errors.
[0097] In one implementation, optionally, the formula for updating the parameters of the teacher model in this step is: Where α is used to balance the weight ratios of the two student models, β is used to control the speed of parameter update, and t represents the current training iteration.
[0098] Finally, calculate the weighted sum of all the above loss values as the final total loss value, and use the AdamW optimizer to update the parameters of the two student models. The weights of the teacher model are updated by the weighted values of the parameters of the two student models, and the teacher parameter update process can be formulated as:
[0099]
[0100] In the formula, α is used to balance the weight ratios of the two student models, β is used to control the speed of parameter update, and t represents the current training iteration. In this way, the teacher model can more effectively guide the learning of the student model, thereby improving the overall performance of the model. It should be noted that although two student models are used for training in the training stage, only the teacher model is used for inference testing in the testing stage.
[0101] Specifically, in the inference stage, only the teacher model is used to perform forward calculation on the input crack image, output a pixel-level probability map, and generate a binary segmentation mask through thresholding (e.g., threshold = 0.5). The specific process includes: the input image extracts multi-scale features through the encoder, and fuses the encoder-decoder features through the Skip-Fusion module of the decoder; the final output layer generates a probability map through the Sigmoid activation function; and binaryzation is performed based on the probability map to obtain the mask result of the crack area.
[0102] In summary, the embodiment of the present invention proposes an Uncertainty-aware Dual Student Mean Teacher (U-DSMT) network that combines consistency regularization and pseudo-labels. By independently learning information from different perspectives through sub-networks, the problem that the teacher and students are too similar due to coupling is effectively alleviated. On the other hand, while maintaining the inconsistency of the sub-networks as much as possible, the confidence of the pseudo-labels is guaranteed by dynamically adjusting the uncertain patches. In addition, by improving the basic segmentation model, information at different levels of cracks is captured and screened, and information loss caused by ordinary upsampling is reduced. The framework proposed in the embodiment of the present invention is more suitable for crack segmentation than existing semi-supervised algorithms. With only a small number of labeled samples, it can achieve quite high accuracy. Experimental results show that this method effectively reduces the dependence of the model on labeled data, and can maintain excellent segmentation performance even when there is only a small amount of labeled data.
[0103] Generally speaking, the beneficial effects of the embodiment of the present invention mainly include:
[0104] (1) A dual student teacher network (DSMT) is proposed. This network jointly updates the weight parameters of the teacher network through two independent student models, effectively alleviating the problem that it is difficult to improve the segmentation performance caused by the teacher model being highly similar to a single student model, thereby enhancing the performance of semi-supervised crack segmentation.
[0105] (2) A dynamic uncertainty-guided patch adjustment strategy (DUPA) is proposed. This strategy evaluates the uncertainty of sample patches, regards regions with higher uncertainty (i.e., noise regions) as the background, and thus generates new samples with less interference information for retraining the network to generate more confident pseudo-labels, thereby improving the segmentation ability of the model in complex background regions.
[0106] (3) A crack segmentation model (CrackUnet) with hierarchical feature fusion is proposed. By integrating features at different levels through the skip fusion module, the problem of loss of crack detail information caused by multiple downsamplings is alleviated, and the features of the skip connections are screened to reduce the impact of redundant information on the interaction between the encoder and the decoder.
[0107] An embodiment of the present invention further provides a readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps described in any one of the above detection methods are implemented.
[0108] An embodiment of the present invention further provides an electronic device, as Figure 3 shown, the electronic device 300 includes a processor 301, a memory 302, and a computer program stored on the memory. The processor 301 is configured to implement the steps described in any one of the above detection methods when executing the computer program.
[0109] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A semi-supervised crack segmentation method based on uncertainty-guided dual-student teacher network, characterized in that: The following steps are involved: S1: Construct a dual student teacher network (DSMT), including two independently initialized student models and one teacher model, both of which are CrackUnet segmentation models based on an encoder-decoder structure; S2: Calculate supervision loss using labeled data Said It is a linear combination of weighted Dice loss and cross entropy loss; cross pseudo-supervised training and consistency training are performed on unlabeled data to calculate the unsupervised loss and S3: Based on total loss Update the student model parameters, where λ(t) is a dynamic weight coefficient that changes with the number of training iterations t; S4: Perform dynamic uncertainty patch adjustment (DUPA) to generate new training samples, including: S401: Based on the prediction results of the student model and the teacher model, two information entropy maps are calculated as uncertainty maps; S402: Divide each uncertainty map into r×r pixel patches, and select the first K patches with the highest entropy values; S403: Mark the area corresponding to K blocks in the original sample as background to generate a new sample after denoising; S5: Retrain the network using the new sample, update the teacher model parameters to the exponential moving average (EMA) of the two student model parameters, output the segmentation result of the teacher model for the input crack image, and generate a binary segmentation mask.
2. The method according to claim 1, characterized in that The decoder of the CrackUnet model in step S1 includes a skip-fusion module, which performs the following operations: A dynamic upsampler (DySample) is used to improve the resolution of decoder features. The upsampled features and encoder features are input into a hierarchical feature fusion module (HFF), and channel weights are generated through channel splicing, global pooling, convolution, and sigmoid activation. The encoder features are calibrated based on the channel weights and fused with the decoder features in the spatial dimension.
3. The method according to claim 2, characterized in that The channel weight generation formula of the HFF module is: <h2 style=";text-align:left;direction:ltr">w<h2 style=";text-align:left;direction:ltr"> c <h2 style=";text-align:left;direction:ltr"> =Sigmoid(Conv1(MaxPool([f1;f2])+AvgPool([f1;f2]))) Among them, f1 is the encoder feature and f2 is the decoder feature.
4. The method according to claim 1, characterized in that: The calculation formula of the dynamic weight coefficient λ(t) in step S3 is: where λ m represents the maximum value of the weighted function, t represents the current training iteration, and t m Indicates the maximum number of iterations during training.
5. The method according to claim 1, characterized in that The teacher model parameter updating formula in step S5 is: Among them, α is used to balance the weight ratio of the two student models, β is used to control the speed of parameter update, and t represents the current training iteration.
6. The method according to claim 1, characterized in that The supervised loss in step S2 Defined as: L seg =βL dice +(1-β)L ce Among them, L dice is the dice loss function, L ce is the cross entropy loss function, β is the weight coefficient of dice loss, p i is the probability output of the model, y i is the corresponding true label, |D L | represents the number of labeled data, H and W represent the height and width of the image respectively.
7. The method according to claim 1, characterized in that The unsupervised loss in step S2 and Defined as: in, and are the probability outputs generated by the student f(θ1) and f(θ2) and the teacher model, and They represent the segmentation masks obtained by one-hot encoding the probability outputs generated by the student models f(θ1) and f(θ2), respectively, U | represents the number of unlabeled data.
8. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer device, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory, wherein when the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Semi-supervised medical image decoupling comparison segmentation method based on uncertainty guidance
CN121353864A
Semi-supervised medical image decoupled contrast segmentation method based on uncertainty guidance
CN121353864B