A method for generating a training dataset without a segmenter for abnormal image segmentation
By introducing attention difference maximizing energy function in the diffusion model and optimizing the generation process, the label drift problem in the abnormal image segmentation training dataset is solved, and high-quality and diverse training data are generated, which improves the performance of downstream tasks.
Patent Information
- Application Number
- CN202510407252.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The prior art has mask label drift problems when generating abnormal images to segment training data sets, and relying on external cutters leads to high computational overhead, making it difficult to generate high-quality training data in data scarce scenarios.
By introducing attention difference maximizing energy functions, the attention map in the diffusion model is aligned with the reference truth mask, and the generation process is optimized to generate anomaly images matching the reference truth mask and their mask label pairs to avoid relying on external cutters.
It effectively alleviates the problem of label drift, improves the consistency between the generated image and the reference truth mask, reduces calculation costs, significantly improves the quality and diversity of generated images, and supports improvement in downstream tasks.
Smart Images

Figure CN119919758B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and particularly to a method for generating a training dataset without a segmenter for abnormal image segmentation. Background Art
[0002] In the field of computer vision and image processing, the combination of image generation and segmentation tasks has become a key research direction, especially in scenarios where samples are scarce or annotation is difficult. With the rapid development of deep learning technology, generative adversarial networks (GANs) and diffusion models have been increasingly widely used in image generation, enhancement, and annotation, especially in fields such as medicine and industry. Generating high-quality images and their corresponding label data has become the core technology to solve the problem of data scarcity.
[0003] Generative adversarial networks (GANs), as one of the early image generation methods, generate realistic images through adversarial training between a generator and a discriminator. However, GANs are prone to the problem of mode collapse, resulting in a lack of diversity in the generated samples and difficulties in image quality control. Nevertheless, GANs are still widely used in fields such as medical image generation and image inpainting. In recent years, diffusion models (such as DDPM) have gradually become the mainstream method in the field of image generation. Diffusion models generate high-quality images by introducing noise during the generation process and gradually denoising. They have strong controllability and stability and can effectively avoid the common mode collapse problem of GANs, so they are widely used in high-quality image generation and reconstruction tasks.
[0004] Abnormal images cover various types, such as medical images, ground dirt images, etc. The goal of the segmentation training data synthesis task is to generate a large number of images and their corresponding mask labels simultaneously. These image-mask pairs serve as additional training datasets for downstream segmentation tasks, especially suitable for tasks with sparse samples or difficult annotation. The prior art usually adopts the following process: input a ground truth (GT) mask (representing the abnormal area) into a pre-trained diffusion model to generate an image containing abnormal content. Among them, the GT mask serves as both the guidance for the image generation process and the segmentation label of the generated image. However, there is a problem of mask label drift in this process, that is, the GT mask does not match the actual abnormal area in the generated image. Due to the lack of clear error feedback, this problem is inevitable in the traditional process.
[0005] To solve the label drift problem, a direct strategy is to input the generated image with the drift problem into an additional frozen pre-trained segmentation model (segmenter), such as Figure 1As shown. This segmentation model can ensure that its mask prediction exactly matches the abnormal regions in the generated image, but this approach may cause the mask prediction itself to be affected by the drift problem. Finally, the diffusion model is optimized by calculating the binary cross-entropy (BCE) loss between the GT mask and the drifted predicted mask, forcing the generated abnormal regions to be closer to the GT mask, thereby alleviating the mask label drift problem. However, this strategy highly depends on the performance of the pre-trained segmentation model, which is difficult to achieve in the real world, especially in data-sparse fields. It faces huge challenges to ensure that the segmentation model performs well on limited training samples. In addition, if the segmentation model performs perfectly, there is no need to generate more training samples, which goes against the original intention of generating the dataset. At the same time, introducing an additional segmentation model will significantly increase the computational overhead and have a negative impact on the efficiency of the entire framework. Summary of the Invention
[0006] To this end, an embodiment of the present invention provides a method for generating a training dataset without a segmenter for abnormal image segmentation, aiming to solve the mask label drift problem that occurs when generating abnormal images, and without relying on an additional segmenter, improving the accuracy of the generated images by optimizing the generation process, making the abnormal regions of the generated images better aligned with the ground truth (GT) mask, thereby generating higher-quality training data for downstream abnormal image segmentation tasks.
[0007] To solve the above problems, an embodiment of the present invention provides a method for generating a training dataset without a segmenter for abnormal image segmentation, the method comprising:
[0008] Taking the ground truth mask as a conditional input into the diffusion model, and generating an image containing abnormal regions through a random noise and a step-by-step denoising process;
[0009] During the sampling process of the diffusion model, introducing an attention difference maximization energy function, which dynamically adjusts the alignment degree between the abnormal regions of the generated image and the ground truth mask by comparing the attention maps generated by the diffusion model with the ground truth mask;
[0010] By iteratively optimizing the attention difference maximization energy function, generating abnormal images and their corresponding mask label pairs that match the ground truth mask as the training dataset.
[0011] Preferably, the attention difference maximization energy function includes a first energy function and a second energy function , where:
[0012] The first energy function By maximizing the difference between the average attention of the target region and the average attention of the non-target region, the localization accuracy of the abnormal region in the generated image is enhanced;
[0013] The second energy function By penalizing the lowest K attention values in the target region and the highest K attention values in the non-target region, the regional imbalance problem between the target region and the non-target region is solved.
[0014] Preferably, the first energy function The calculation formula of is:
[0015] ;
[0016] In the formula, Represents the ground truth mask, Represents the attention map calculated when the diffusion model predicts noise at time step , Is the number of pixels in the target region, Is the number of pixels in the non-target region, Represents element-wise multiplication, Represents the sum of all elements in the matrix, Represents the absolute value function.
[0017] Preferably, the second energy function The calculation formula of is:
[0018] ;
[0019] In the formula, Represents the ground truth mask, Represents the attention map calculated when the diffusion model predicts noise at time step , And Respectively represent selecting the K highest values and the K lowest values in the non-target region and the target region, Represents element-wise multiplication, Represents the sum of all elements in the matrix, Represents the absolute value function.
[0020] Preferably, the expression of the attention difference maximization energy function is:
[0021] ;
[0022] In the formula, Represents the attention difference maximization energy function, Represents the ground truth mask, Represents the diffusion model at time step The attention map calculated when predicting noise is a weight factor used to balance the first energy function and the second energy function in terms of their influence.
[0023] Preferably, the sampling process of the diffusion model includes two stages: a contour generation stage and a detail refinement stage. The attention difference maximization energy function is only applied in the contour generation stage of the diffusion model, and the contour generation stage corresponds to the interval of 80% to 100% of the sampling time steps of the diffusion model.
[0024] Preferably, a repetition strategy is adopted in the sampling process of the diffusion model, that is, at each time step, the currently generated intermediate image is resampled as the output of the previous time step, and the control intensity of the attention difference maximization energy function is enhanced through multiple repeated samplings.
[0025] Preferably, the attention map is generated by the cross-attention layer of the U-Net in the diffusion model, and the specific calculation method is as follows:
[0026] ;
[0027] In the formula, represents the attention mechanism function, represents the activation function, represents the attention map calculated when the diffusion model predicts noise at time step The attention map calculated when predicting noise represents the text information in the latent space, represents the intermediate space features, is the dimension of the projection key and query, represents the matrix transpose operation.
[0028] An embodiment of the present invention further provides a training dataset generation system without a segmenter for abnormal image segmentation. This system is used to implement the above-mentioned method for generating a training dataset without a segmenter for abnormal image segmentation, and specifically includes:
[0029] An image generation module, which is used to input the ground truth mask as a condition into the diffusion model, and generate an image containing abnormal regions through random noise and a step-by-step denoising process;
[0030] An alignment adjustment module, which is used to introduce the attention difference maximization energy function during the sampling process of the diffusion model. The attention difference maximization energy function dynamically adjusts the alignment between the abnormal regions of the generated image and the ground truth mask by comparing the attention map generated by the diffusion model with the ground truth mask;
[0031] A dataset generation module, which is used to generate a pair of abnormal images and their corresponding mask labels that match the ground truth mask as a training dataset by iteratively optimizing the attention difference maximization energy function.
[0032] An embodiment of the present invention also provides a computer storage medium, which stores a computer software product. The computer software product includes several instructions for causing a computer device to execute the above-mentioned method for generating a training dataset without a segmenter for abnormal image segmentation.
[0033] It can be seen from the above technical solutions that the present invention application has the following beneficial effects:
[0034] (1) Effectively alleviate the label drift problem: By using the attention difference maximization energy function based on the attention map and the ground truth mask, the target region of the generated image is adjusted in each sampling step. Effectively avoid the problem of mismatched target regions caused by label drift.
[0035] (2) Do not rely on an external segmenter: Traditional diffusion models rely on an external segmenter to predict the mask of the generated image and then compare it with the ground truth mask for optimization. This method is highly dependent on the segmenter, and when the data is scarce or the segmenter is imperfect, it may affect the quality of the generated image. Compared with the existing method that relies on an external model, the present invention directly optimizes the generation process through the embedded attention difference maximization energy function, making the generation of the target region more accurate, significantly improving the consistency between the generated image and the ground truth mask, and greatly reducing the computational cost and improving the stability of the training process.
[0036] (3) No additional training data is required during the training process: The present invention adopts a "training-free guidance" method, uses an existing diffusion model as a baseline model, and carefully designs an attention-based energy function for guidance without an additional training process. This method is very beneficial in scenarios where data is scarce. Without a large amount of labeled data, it can optimize the quality of the image, thus better supporting downstream tasks.
[0037] (4) Solve the imbalance problem between the target region and the non-target region: The present invention significantly solves the problem of abnormal attention values in the generated image in the case where the proportion of the target region is small and the proportion of the non-target region is large by designing an additional energy function (the second energy function ). By adjusting the top K attention values, the generation quality of the target region is effectively improved, and label drift is further reduced.
[0038] (5) Improved the quality of generated data: The present invention significantly improves the quality of generated data. In the quantitative comparison, the images generated by the method of the present invention are significantly better than other SOTA models in terms of FID and KID metrics, indicating that the generated images are closer to real data in terms of visual quality and data distribution. In the LPIPS metric, the method of the present invention also shows high visual consistency. Qualitative comparison shows that the method of the present invention can generate clearer and more detailed images, avoiding the problems of blurring and artifacts in traditional methods. In addition, the t-SNE visualization results show that the images generated by the method of the present invention are more consistent with real images in terms of data distribution, successfully overcoming the data bias caused by the additional loss function and the auxiliary segmenter. Generally speaking, the present invention can generate images with higher quality, more details and more accurate data distribution, improving the quality of generated data.
[0039] (6) Wide application potential: The training data generated by using the method of the present invention can significantly improve the performance of downstream task models. For example, on the ETIS dataset, the mDice of the SAnet model increased by 4.4% and the mIoU increased by 5.3%. The technical solution of the present invention can be applied in multiple fields such as medical imaging and environmental detection by generating high-quality and diverse synthetic data, demonstrating wide application potential. Brief Description of the Drawings
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly describe the drawings required to be used in the embodiments. By referring to the drawings, the features and advantages of the present invention will be more clearly understood. The drawings are schematic and should not be construed as limiting the present invention in any way. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:
[0041] Figure 1 It is a schematic diagram of a strategy for solving the label drift problem provided in the background art;
[0042] Figure 2 It is a flowchart of a method for generating a training dataset without a segmenter for abnormal image segmentation provided in the embodiment;
[0043] Figure 3 It is a flowchart of the method of the present invention in the embodiment;
[0044] Figure 4 It is a schematic diagram of the action mechanism of the ADM energy function in the method of the present invention in the embodiment;
[0045] Figure 5 It is a schematic diagram of the structure of the ADM energy function in the embodiment;
[0046] Figure 6Schematic diagram of the repetition strategy algorithm in the embodiment;
[0047] Figure 7 Comparison result graph of the method of the present invention and different generation models in the embodiment;
[0048] Figure 8 Comparison result graph for the polyp dataset in the embodiment;
[0049] Figure 9 Comparison result graph of the segmentation performance on the floor dirt dataset in the embodiment;
[0050] Figure 10 Block diagram of a training dataset generation system for anomaly image segmentation without a segmenter provided in the embodiment. Detailed implementation manners
[0051] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Embodiment 1
[0052] To solve the problem of mask label drift that occurs when generating anomaly images without relying on an additional segmenter. As Figure 2 shown, the embodiment of the present invention proposes a method for generating a training dataset for anomaly image segmentation without a segmenter, and the method includes:
[0053] Taking the ground truth mask as a condition and inputting it into the diffusion model, and generating an image containing an anomaly region through a random noise and a step-by-step denoising process;
[0054] During the sampling process of the diffusion model, introducing an attention difference maximization energy function, which dynamically adjusts the alignment degree between the anomaly region of the generated image and the ground truth mask by comparing the attention map generated by the diffusion model with the ground truth mask;
[0055] By iteratively optimizing the attention difference maximization energy function, generating an anomaly image and its corresponding mask label pair that match the ground truth mask as a training dataset.
[0056] As can be seen from the above technical solution, the present invention proposes a method for generating a training dataset without a segmenter for abnormal image segmentation. This method optimizes the sampling process by guiding the generation process and avoids the problems brought by introducing an additional segmenter in traditional methods. In the method of the present invention, the attention map in the diffusion model is used to replace the mask prediction of the segmenter. Then, the proposed attention difference maximization loss (function) is calculated using the attention map and the GT mask to optimize the image. This operation enables the diffusion model to focus on the regions in the ground truth mask when generating abnormal content, thus alleviating the label drift problem.
[0057] In this embodiment, first, the ground truth (GT) mask is input into the diffusion model as a condition, and an image containing abnormal regions is generated through a random noise and a step-by-step denoising process, as Figure 3 shown.
[0058] Specifically, the GT mask M is input as a condition into the diffusion model, and the diffusion model generates an image containing abnormal regions through a random noise and a step-by-step denoising process . The GT mask M is used as an annotation for the generated image and guides the diffusion model to generate the regions where the abnormal content is located.
[0059] The above diffusion model includes two processes: a forward process and a reverse process. The forward process converts a clear image into pure Gaussian noise by continuously adding noise, with the time ranging from 0 to T. On the contrary, the reverse process reconstructs a clear image from pure Gaussian noise, with the time steps ranging from T to 0. denotes the intermediate state of the data at time step , and denotes the clear image.
[0060] The forward process adds noise to according to a certain schedule, expressed as:
[0061] ;
[0062] In the formula, is a predefined parameter, is random noise.
[0063] The reverse process can be expressed as:
[0064] ;
[0065] In the formula, is the distribution of , , is obtained through the training of the diffusion model. In practice, the model usually only learns one of them, and the other can be predefined.
[0066] A neural network is usually used to fit . Since can be obtained from the noise, we only need the network to learn the noise . Among them, the training loss function is defined as:
[0067] ;
[0068] In the formula, represents the mathematical expectation, is the square of the norm.
[0069] The basic form of the energy-guided score-based diffusion model is:
[0070] ;
[0071] In the formula, is a predefined parameter, is the score function of the gradient in the diffusion model based on the Tweedie formula, is the introduced energy function, which is used to measure the difference between the target and the generated image.
[0072] Furthermore, in the sampling process of the diffusion model, the present invention introduces the Attention Difference Maximization (ADM) energy function. The ADM energy function dynamically adjusts the alignment degree of the abnormal region of the generated image with the GT mask by comparing the attention map generated by the diffusion model with the GT mask, as Figure 4 shown.
[0073] Since the method of the present invention aims to enhance the alignment degree between the activation region of the attention map and the basic GT mask, so is replaced with , where is the ADM energy function proposed by the present invention. By introducing the ADM energy function of the present invention, the generated image is updated along a combined direction at each time step, thereby alleviating label drift, and the new formula becomes:
[0074] ;
[0075] Among them, represents the attention map calculated when the diffusion model predicts the noise at time step , which is generated by the U-Net in the diffusion model. The calculation method is as follows:
[0076] ;
[0077] In the formula, Represents the attention mechanism function, Represents the activation function, Represents the attention map calculated when the diffusion model predicts noise at time step Represents the text information in the latent space, Represents the intermediate spatial features, Is the dimension of the projected key and query, Represents the matrix transpose operation.
[0078] Furthermore, the detailed design of the first energy function is as follows:
[0079] ;
[0080] In the formula, is the number of pixels in the target region, is the number of pixels in the non-target region, Represents element-wise multiplication, Represents the sum of all elements in the matrix, Represents the absolute value function. At the same time, is normalized by maximum and minimum normalization so that the range of each activation value is from 0 to 1. Therefore, as Figure 5 shown, selects the activation values of the target region and encourages the mean of these values to be close to 1, represents the mean of in the non-target region, and promotes its value to be close to 0.
[0081] As Figure 5 shown, in the process of using the GT mask, when sampling, the first energy function updates the image by encouraging large values in the target region of the attention map and small values in the non-target region. However, after statistical analysis of the dataset used, it is found that the regional imbalance ratio is relatively high. This situation makes it difficult for the first energy function to effectively adjust the abnormal attention values, that is, the minimum value in the target region and the maximum value in the non-target region. The attention map of the intermediate result is a good example.
[0082] Abnormal images often exhibit the characteristics of small target objects and occupy a small area in the image. Medical images are typical examples. This characteristic makes the regional imbalance problem more prominent. In this case, abnormal attention values such as large attention values in the non-target region and low attention values in the target region cannot be accurately punished by the first energy function , which further leads to the difficulty of achieving the optimization goal. Even though the above first energy function There is a certain control ability over the position and shape of the target object, but the problem of regional imbalance still limits its optimization effect.
[0083] To solve this problem, the present invention proposes a second energy function , which is used to penalize the top K abnormal attention values. With the help of the second energy function , the abnormal attention values are further adjusted, and the formula is as follows:
[0084] ;
[0085] In the formula, and respectively represent the top K highest values and the top K lowest values selected from the non-target region and the target region.
[0086] All in all, the first energy function enhances the localization accuracy of the abnormal region in the generated image by maximizing the difference between the attention mean of the target region and the attention mean of the non-target region. The second energy function solves the problem of regional imbalance between the target region and the non-target region by penalizing the lowest K attention values in the target region and the highest K attention values in the non-target region.
[0087] Finally, the expression of the ADM energy function is as follows:
[0088] ;
[0089] In the formula, is a weight factor used to balance the influence of the first energy function and the second energy function .
[0090] The intermediate image updated by the ADM energy function of the present invention can be optimized in the direction of alleviating label drift.
[0091] At the same time, we notice that it is not reasonable to use the ADM energy function to update during the entire sampling process. The diffusion sampling process includes two stages: the contour generation stage and the detail refinement stage. Since in the refinement stage, the difference between the predicted intermediate image and the final clear image is not significant, using the method of the present invention may not effectively solve the label drift problem, but still affects the reconstruction of the clear image. Therefore, the method of the present invention is only introduced in the contour generation stage, where the contour generation stage corresponds to the interval of 80% to 100% of the diffusion model sampling time steps.
[0092] In addition, to enhance the control strength of the energy function, the present invention adopts a repetition strategy, that is, at each time step the current model output Resampled to the output of the previous time step and used as the model input. By repeating the sampling process multiple times, the control strength of the ADM energy function is further enhanced, thus more effectively guiding the precision of the generated image. The specific approach is as Figure 6 shown.
[0093] Finally, in terms of training the diffusion model, the pre-trained Stable Diffusion (SD) V2.1 and ControlNet are adopted as the baseline models, where the real ground truth mask M is introduced as the condition for the diffusion model. In addition, by fine-tuning ControlNet, SD can generate abnormal images. For the inference process of the diffusion model, following the denoising diffusion implicit model, 100 sampling time steps are used. The guidance scale λ is set to 9, exactly the same as ControlNet. The energy function is introduced at time steps between 80 and 100; the repetition strategy is executed three times; the weight factor is set to 0.5. For the attention map, in this embodiment, the output of the final cross-attention operation in the U-Net of SD is selected, with a size of 64×64. The reason is that the final attention map captures all the effects of conditional control on features, and the features used in this layer are closer to the final model output. The number of abnormal pixel points, denoted as K, is set to 20% of the total pixels in the target area.
[0094] The advantages of the method of the present invention are further illustrated below with specific experiments.
[0095] 1. Image quality evaluation
[0096] The quality of the generated images is evaluated in this experiment by the following metrics:
[0097] FID: Used to measure the difference between the generated image and the real image. The lower the FID value, the more similar the generated image is to the real image.
[0098] KID: Similar to FID, but more stable when dealing with fewer samples and does not assume that the data follows a Gaussian distribution.
[0099] LPIPS: Evaluates the diversity of the generated images. The higher the LPIPS value, the richer the diversity of the generated images.
[0100] As Figure 7 shown, compared with different generation models, the pictures generated by the method of the present invention perform best in most evaluation metrics. The optimal scores are in bold, and the sub-optimal ones are underlined.
[0101] 2. Evaluation of the effectiveness of downstream model data
[0102] In this experiment, experiments were conducted on five publicly available polyp segmentation datasets and one floor dirt dataset.
[0103] 2.1 Dataset Preparation
[0104] a) Polyp segmentation datasets: ETIS, CVC-ClinicDB / CVC612, CVC-ColonDB, EndoScene, Kvasir.
[0105] The image-mask pairs for training were from Kvasir-SEG (900 training samples) and CVC-ClinicDB (550 training samples), a total of 1450 image-mask pairs as the "train-real" dataset.
[0106] The test sets included: EndoScene (60 test samples), CVC-ClinicDB (62 test samples), CVC-ColonDB (380 test samples), ETIS-LaribPolypDB (196 test samples), and Kvasir (100 test samples).
[0107] b) Floor dirt dataset: It contains two types of anomalies: floor stains (500 images) and pet feces (458 images). 3 / 5 of the anomaly images were used for training, and 2 / 5 were used for validation and testing.
[0108] By duplicating the 1450 image and mask pairs in "train-real", 2900 pairs of data were created, resulting in the "train-real + train-real" dataset.
[0109] Combined the real images in "train-real" and 1450 image-mask pairs generated by the ArSDM method to obtain the "train-real + train-ArSDM" dataset.
[0110] We combined the real images (1450 pairs) with 1450 image and mask pairs generated by the method of the present invention. The "train-ours" dataset was obtained.
[0111] 2.2 Performance Evaluation and Analysis
[0112] Two different segmentation models, SANet and Polyp-PVT, were used to train the above four datasets, and the results are as Figure 8As shown. Two metrics, mIoU (%) and mDice (%), are used, where higher values indicate more accurate segmentation by the model. The results with the highest scores are highlighted in bold, and the second-highest results are underlined. "+" indicates the combination of the "train-real" dataset and other datasets. "train-ArSDM" refers to the dataset generated by the ArSDM generation model. "AVE" represents the average results of five test sets.
[0113] For the floor dirt dataset, in this experiment, 1000 pure synthetic data of each class were used to train the SegFormer model, and the SegFormer model was trained with the data generated by the method of the present invention and the data generated by previous methods (DFMGAN, Anomaly Diffusion). Two metrics, IoU (%) and F-measure (%), were used. The results are as Figure 9 shown.
[0114] Obviously, by using the training data generated by the method of the present invention, the performance of the downstream segmentation task is significantly improved. Example Two
[0115] As Figure 10 shown, the present invention provides a segmentation-free training dataset generation system for anomaly image segmentation, which is used to implement the segmentation-free training dataset generation method for anomaly image segmentation in the above-mentioned Example One, and specifically includes:
[0116] An image generation module, which is used to input the ground truth mask as a condition into the diffusion model, and generate an image containing an abnormal area through a random noise and a step-by-step denoising process;
[0117] An alignment adjustment module, which is used to introduce an attention difference maximization energy function during the sampling process of the diffusion model. The attention difference maximization energy function dynamically adjusts the alignment between the abnormal area of the generated image and the ground truth mask by comparing the attention map generated by the diffusion model with the ground truth mask;
[0118] A dataset generation module, which is used to generate a pair of abnormal images and their corresponding mask labels that match the ground truth mask through iterative optimization of the attention difference maximization energy function, as the training dataset.
[0119] A segmentation-free training dataset generation system for anomaly image segmentation in this embodiment is used to implement the foregoing segmentation-free training dataset generation method for anomaly image segmentation. Therefore, the specific implementation manners in the segmentation-free training dataset generation system for anomaly image segmentation can be seen in the embodiment part of the foregoing segmentation-free training dataset generation method for anomaly image segmentation. To avoid redundancy, they will not be elaborated here. Embodiment 3
[0120] An embodiment of the present invention provides a computer storage medium. The computer storage medium stores a computer software product. The computer software product includes a number of instructions for causing a computer device to execute the above-mentioned method for generating a training data set without a segmenter for abnormal image segmentation.
[0121] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0122] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0123] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0124] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.
Claims
1. A method for generating a training dataset without a segmenter for abnormal image segmentation, characterized in that, Including: Taking the ground truth mask as a condition and inputting it into the diffusion model, generating an image containing abnormal regions through random noise and a step-by-step denoising process; During the sampling process of the diffusion model, introducing an attention difference maximization energy function, which dynamically adjusts the alignment degree between the abnormal region of the generated image and the ground truth mask by comparing the attention map generated by the diffusion model with the ground truth mask; The attention difference maximization energy function includes a first energy function F l and a second energy function F e , where the first energy function F l enhances the localization accuracy of abnormal regions in the generated image by maximizing the difference between the attention mean of the target region and the attention mean of the non-target region; the second energy function F e solves the problem of regional imbalance between the target region and the non-target region by penalizing the lowest K attention values in the target region and the highest K attention values in the non-target region; The expression of the attention difference maximization energy function is: F(M,A t ) = αF l (M,A t ) + (1 - α)F e (M,A t ); In the formula, F represents the attention difference maximization energy function, M represents the ground truth mask, and A t represents the attention map calculated by the diffusion model when predicting noise at time step t, and α is a weight factor used to balance the first energy function F l and the second energy function F e 's influence; By iteratively optimizing the attention difference maximization energy function, generating an abnormal image and its corresponding mask label pair that match the ground truth mask as the training dataset.
2. The method for generating a training dataset without a segmenter for abnormal image segmentation according to claim 1, wherein, The first energy function F l is calculated by the following formula: where M represents the ground truth mask, A t represents the attention map calculated by the diffusion model when predicting noise at time step t, N is the number of pixels in the target region, is the number of pixels in the non-target region, · represents element-wise multiplication, ∑() represents the sum of all elements in the matrix, and abs() represents the absolute value function.
3. The method for generating a training dataset without a segmenter for abnormal image segmentation according to claim 1, wherein, The second energy function F e is calculated as follows: where M represents the reference true value mask, and A t represents the attention map calculated by the diffusion model when predicting noise at time step t. topK(·, K) and lowK(·, K) respectively represent selecting the K highest values and the K lowest values in the non-target region and the target region. · represents element-wise multiplication, ∑() represents the summation of all elements in the matrix, and abs() represents the absolute value function.
4. The method for generating a training dataset without a segmenter for abnormal image segmentation according to claim 1, wherein The sampling process of the diffusion model includes two stages: a contour generation stage and a detail refinement stage. The attention difference maximization energy function is only applied in the contour generation stage of the diffusion model, and the contour generation stage corresponds to the interval of 80% to 100% of the sampling time steps of the diffusion model.
5. The method for generating a training dataset without a segmenter for abnormal image segmentation according to claim 1, wherein A repetition strategy is adopted in the sampling process of the diffusion model, that is, at each time step, the currently generated intermediate image is resampled as the output of the previous time step, and the control intensity of the attention difference maximization energy function is enhanced through multiple repeated samplings.
6. The method for generating a training dataset without a segmenter for abnormal image segmentation according to claim 1, characterized in that, The attention map is generated by the cross-attention layer of the U-Net in the diffusion model, and the specific calculation method is: Wherein, Attention() represents the attention mechanism function, Softmax() represents the activation function, and A t represents the attention map calculated by the diffusion model when predicting noise at time step t, and Q t represents the text information in the latent space, K t represents the intermediate space feature, d is the dimension of the projected key and query, and T represents the matrix transpose operation.
7. A training dataset generation system for abnormal image segmentation without a segmenter, characterized in that, The system is used to implement the method for generating a training dataset without a segmenter for abnormal image segmentation according to any one of claims 1 to 6, specifically including: An image generation module, configured to take the ground truth mask as a condition and input it into the diffusion model, and generate an image containing abnormal regions through random noise and a step-by-step denoising process; An alignment degree adjustment module, configured to introduce an attention difference maximization energy function during the sampling process of the diffusion model, and the attention difference maximization energy function dynamically adjusts the alignment degree between the abnormal region of the generated image and the ground truth mask by comparing the attention map generated by the diffusion model with the ground truth mask; A dataset generation module, configured to generate an abnormal image and its corresponding mask label pair that match the ground truth mask as the training dataset by iteratively optimizing the attention difference maximization energy function.
8. A computer storage medium, characterized in that, The computer storage medium stores a computer software product, and the computer software product includes several instructions for causing a computer device to execute the method for generating a training dataset without a segmenter for abnormal image segmentation according to any one of claims 1 to 6.
Citation Information
Patent Citations
Document image transmission removal method and device based on fuzzy diffusion model
CN118229569A
Model training method and device, bearing detection method and device and electronic equipment
CN119048862A