A few-shot X-ray defect intelligent detection method based on diffusion generative model

CN122156210BActive Publication Date: 2026-09-11STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610633907.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-09
Publication Date
2026-09-11
Estimated Expiration
2046-05-09

AI Technical Summary

Technical Problem

[0005]本发明的目的在于克服现有技术的不足,提供一种基于扩散生成模型的少样本X射线缺陷智能检测方法,其核心流程包括数据预处理、缺陷条件建模、扩散生成模型训练以及生成-检测协同优化等步骤,用于解决工业场景中缺陷正样本稀缺、样本分布失衡、传统增强效果差、生成与检测无协同优化导致的缺陷检测精度低、泛化能力弱的问题

Benefits of technology

[0032] 1. This invention constructs a conditional diffusion generation model with physical constraints, which can generate multi-morphological, high-fidelity X-ray defect samples, covering complex types such as microcracks and low-contrast defects, and fundamentally solves the industry pain points of scarce positive defect samples and sample imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122156210B_ABST
    Figure CN122156210B_ABST
Patent Text Reader

Abstract

This invention discloses a few-sample intelligent X-ray defect detection method based on a diffusion-generative model, belonging to the fields of industrial non-destructive testing and artificial intelligence. The method first acquires and preprocesses an X-ray image dataset, constructing a defect condition embedding that preserves physical properties. It then builds a conditional diffusion-generative model optimized by incorporating rectified flow reparameterization and a consistency model distillation mechanism, training it in stages to generate high-fidelity defect samples. A multi-scale Transformer detection network is constructed, and through mixed-sample training, dynamic loss optimization, curriculum learning, and feature consistency constraints, intermediate features from the detection network are fed back to the generative model to adjust the defect condition weights, achieving joint training of the generative model and the detection network. Finally, the input image to be detected completes defect localization and classification. This invention can generate multiple types of realistic defect samples, eliminating sample imbalance and domain bias, and achieving both high detection accuracy and real-time performance under few-sample conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of industrial nondestructive testing and artificial intelligence, specifically to a method for intelligent detection of X-ray defects with few samples based on a diffusion generation model. Background Technology

[0002] X-ray nondestructive testing (NDT) technology is a core technology for identifying internal defects in high-end manufacturing fields such as power equipment, rail transportation, and aerospace. It can effectively detect hidden dangers such as welding defects, structural cracks, and foreign object inclusions. Existing testing technologies include three categories: radiographic testing, digital radiographic testing, and industrial computed tomography (CT). Among them, digital radiographic testing, with its advantages of real-time imaging and digital management, is widely used in automated industrial inspection scenarios.

[0003] In recent years, deep learning-based target detection algorithms have become the mainstream solution for intelligent X-ray defect identification, automatically extracting defect features and completing localization and classification. However, these algorithms heavily rely on large-scale, balanced labeled datasets, while defects are low-frequency events in real industrial scenarios, resulting in extremely scarce positive defect samples and a severely imbalanced sample distribution. This leads to class bias during model training, causing missed detections of critical defects.

[0004] Existing technologies mostly use traditional data augmentation methods such as rotation and flipping to expand samples, which can only achieve pixel-level transformations and cannot generate defect samples with real physical imaging characteristics. At the same time, data generation and defect detection are independent modules, and there is no collaborative optimization between the generated samples and the detection task, resulting in significant distribution bias. This cannot fundamentally solve the problem of performance degradation of detection models under conditions of few samples. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for intelligent detection of X-ray defects with few samples based on a diffusion generation model. Its core process includes data preprocessing, defect condition modeling, diffusion generation model training, and generation-detection co-optimization. This method is used to solve the problems of low defect detection accuracy and weak generalization ability caused by the scarcity of positive defect samples, unbalanced sample distribution, poor traditional enhancement effect, and lack of co-optimization between generation and detection in industrial scenarios.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A method for intelligent detection of X-ray defects using a few samples based on a diffusion generation model includes the following steps:

[0008] S1: Data Acquisition and Preprocessing: Acquire an X-ray image dataset of industrial equipment, which includes defect-free negative samples, a small number of labeled defect positive samples, and projection data of physical simulation structures; normalize and denoise all images to construct a training base dataset;

[0009] S2: Defect Category Embedding Construction: Geometric edges, gray-level gradients, and local texture multi-scale features of positive defect samples are extracted through a hierarchical feature encoding network and mapped to a unified conditional embedding space; contrast constraints are introduced to optimize the embedding space while preserving the physical characteristics of defect X-ray imaging.

[0010] S3: Construction and Training of Conditional Diffusion Generation Model: Based on the conditional diffusion probability model, a U-Net multi-scale denoising network with cross-scale attention fusion is constructed to integrate defect category embedding and location encoding; rectified flow reparameterization and consistency model distillation are introduced to optimize the sampling path and reduce the number of inference steps; a phased training mechanism is used to train the generation model and output highly realistic defect synthetic samples generated based on the conditional diffusion model.

[0011] S4: Generation-Detection Joint Collaborative Training: Construct a multi-scale Transformer object detection network, and mix real samples and synthetic samples in proportion as the training set; introduce dynamic label allocation, focus loss and generalized intersection-union loss to optimize the detection network, and adopt curriculum learning fine-tuning and feature distribution consistency constraints to achieve closed-loop joint optimization of the generation model and the detection network;

[0012] S5: Intelligent Defect Detection Inference: Input the X-ray image to be detected into the trained detection network, output the location coordinates and classification results of the defect, and complete the intelligent detection of defects in industrial X-ray images with few samples.

[0013] As a further aspect of the present invention, in step S1, the industrial equipment is a GIS combined electrical appliance; the defect types include four types of defects: loose bolts, misaligned contacts, cracked basin insulators, and foreign metal objects, and the annotation form is pixel-level segmentation annotation or bounding box annotation.

[0014] As a further aspect of the present invention, in step S3, the training objective function of the conditional diffusion probability model is:

[0015]

[0016] in, This is the intermediate state after adding noise. Embed for defect categories, For noise prediction networks, This is real noise;

[0017] The objective function for stream matching training of the rectified stream is:

[0018]

[0019] in To predict the flow field function, This represents the temporal gradient of the image state.

[0020] As a further aspect of the present invention, in step S3, the phased training mechanism specifically comprises:

[0021] In the first stage, unconditional pre-training was performed using defect-free negative samples and defective positive samples to learn the global structural distribution of X-ray images.

[0022] In the second stage, defect category conditions are introduced for fine-tuning, and mask loss constraints are applied only to defect areas.

[0023] In the third stage, a physical consistency regularization term is added to constrain the grayscale distribution of the generated image to match the statistical characteristics of the simulated projection data.

[0024] As a further aspect of the present invention, in step S3, the high-resolution branch of the multi-scale denoising network integrates dense residual blocks to enhance the ability to generate details of fine-grained defects such as cracks; the consistency model achieves rapid sampling and generation of defect samples in a single step or with fewer steps through cross-step consistency constraints.

[0025] As a further aspect of the present invention, in step S4, the target detection network is an improved DeformableDETR or RT-DETR network, which introduces a multi-scale feature fusion structure into the backbone network and combines it with a cross-attention module to perform feature interaction, fuse multi-resolution features, and enhance the detection sensitivity of small-scale, low-contrast defects.

[0026] As a further aspect of the present invention, in step S4, the overall loss function of the detection network is:

[0027]

[0028] in, This is a weighted focus loss used to mitigate sample class imbalance. The generalized intersection-union loss is used to optimize the regression accuracy of small target defect boxes; These are the balancing weighting coefficients.

[0029] As a further aspect of the present invention, in step S4, the course learning fine-tuning specifically involves: prioritizing the training of the detection network with real samples and gradually increasing the proportion of synthetic samples; combining adversarial domain alignment to reduce the difference in feature distribution between real samples and synthetic samples; and simultaneously constructing dedicated sub-model branches for different defect categories to achieve multi-task joint optimization.

[0030] As a further aspect of the present invention, in step S4, the closed-loop joint optimization specifically involves: extracting semantic features from the intermediate layer of the detection network, constraining the consistency of feature distribution between the synthesized samples and the real samples; and based on the low-confidence region of the detection network, strengthening the weights of the corresponding defect conditions in the generation model, and generating difficult samples to expand the training set.

[0031] This invention proposes a few-sample intelligent X-ray defect detection method based on a diffusion generation model, which has the following advantages and beneficial effects:

[0032] 1. This invention constructs a conditional diffusion generation model with physical constraints, which can generate multi-morphological, high-fidelity X-ray defect samples, covering complex types such as microcracks and low-contrast defects, and fundamentally solves the industry pain points of scarce positive defect samples and sample imbalance.

[0033] 2. By introducing rectified flow reparameterization and consistency model distillation, the number of sampling steps is significantly reduced while ensuring the quality of the generated product, thus meeting the real-time deployment requirements of industrial online detection.

[0034] 3. A generation-detection closed-loop collaborative optimization mechanism is adopted. Through feature distribution constraints and adaptive generation of difficult samples, the synthesized samples accurately serve the detection task and eliminate domain distribution bias.

[0035] 4. By adopting a multi-scale Transformer detection network combined with a dynamic loss optimization strategy, the detection recall rate of small-scale, low-contrast defects is significantly improved. It maintains high detection accuracy and robustness even under extremely low sample conditions, making it suitable for non-destructive testing scenarios of high-end industrial equipment such as GIS combined electrical appliances. Attached Figure Description

[0036] Figure 1 This is a schematic diagram of the X-ray image defect detection prediction results of the method of the present invention;

[0037] Figure 2 This is a block diagram of the overall structure of the small-sample X-ray defect detection system of the present invention. Detailed Implementation

[0038] The present invention will be further described below with reference to the embodiments. It should be noted that the following content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the claims, they should all be considered to fall within the protection scope of the present invention.

[0039] Please see Figures 1 to 2As shown, this invention is an intelligent detection method for X-ray defects with few samples based on a diffusion generation model. The overall technical approach includes: constructing a diffusion generation model with controllable defect conditions to achieve high-fidelity defect map synthesis; accelerating sampling and distilling the consistency model using rectified flow to achieve efficient inference; constructing a defect detection network based on multi-scale feature reconstruction; and improving the model's targeted specialization and generalization capabilities through collaborative training with generated data and real data, thereby achieving stable training and high-quality inference even when defect samples are extremely scarce.

[0040] At the data level, the data used in this invention includes three categories: the first category is real-collected X-ray images of GIS equipment, most of which are defect-free negative samples in high-resolution grayscale; the second category is a small number of defect samples annotated by experts, including categories such as loose bolts, misaligned contacts, cracks in basin insulators, and metallic foreign objects, which are segmented at the pixel level or annotated with bounding boxes; the third category is structural projection data constructed based on physical simulation software, which is used to assist the generation model in learning the geometric consistency relationship between defects and structures.

[0041] In the generative modeling stage, this invention first constructs a conditional representation of defects. For a small number of real defect samples, a hierarchical feature encoding network is used to extract their geometric edge structure, gray-level gradient distribution, and local texture response, and this multi-scale information is mapped to a unified conditional embedding space. This embedding not only describes the appearance features of the defect but also preserves its attenuation pattern and structural continuity information in X-ray imaging. To avoid embedding space degradation, this invention introduces contrast constraints in the conditional encoding stage, ensuring that the defect features maintain a clustered structure in the latent space, thereby providing stable guidance for subsequent generation.

[0042] In terms of generating the backbone framework, this invention employs a Conditional Diffusion Probabilistic Model (CDPP) as the basic framework for high-dimensional image distribution modeling. The CDPP constructs a forward path from the real image to a Gaussian distribution by progressively adding noise, and learns its reverse recovery process. The core training objective in the reverse process is to minimize the noise prediction error, expressed as:

[0043]

[0044] in This is the intermediate state after adding noise. Embed for defect categories, This is a noise prediction network. The network employs an encoder-decoder architecture with a cross-scale attention fusion structure, enabling the model to maintain overall image consistency while recovering local defect structures. To improve generation quality and structural consistency, this invention uses a multi-scale denoising network based on the U-Net backbone and introduces a cross-attention mechanism to fuse the defect category embedding vector with the structural location encoding, ensuring that the generated result matches the physical structure semantically. Simultaneously, dense residual blocks are introduced at high resolution to enhance detail representation and prevent fine-grained defects such as cracks from being over-smoothed.

[0045] Meanwhile, standard diffusion models typically require numerous iterations during the sampling phase, which is unfavorable for industrial online deployment. Therefore, this invention introduces the concept of rectified flow onto the diffusion framework, performing deterministic reparameterization of the original random diffusion path. Specifically, the original discrete noise evolution process is transformed into continuous-time flow field modeling. By learning the optimal transmission direction from the noise distribution to the true distribution, the generated trajectory becomes smoother and approximately linear. This continuous flow field training objective can be expressed as:

[0046]

[0047] in To predict the flow field function, this flow matching mechanism eliminates the need for numerous random sampling steps in path generation. Instead, it gradually approximates the target distribution along the learned optimal direction, significantly reducing inference complexity while maintaining generation quality. To further improve sampling efficiency, this invention introduces a Consistency Model mechanism, imposing cross-step consistency constraints on the generation results at different time steps. The core idea is to construct a shared mapping function that allows inputs at any time step to directly predict the final clear image, thereby reducing error accumulation. This consistency constraint enables the generation model to maintain structural stability even with a small number of sampling steps, making it particularly suitable for real-time defect data augmentation scenarios.

[0048] Regarding the training strategy, this invention employs a phased training mechanism. The first phase utilizes all negative samples and a small number of positive samples for unconditional pre-training, enabling the model to learn the overall structural distribution of X-ray spectra. The second phase introduces category conditions for conditional fine-tuning, applying mask loss constraints only to defective sample regions. The third phase introduces a physical consistency regularization term, constraining the grayscale distribution of the generated image to be consistent with the statistical characteristics of the structural projection simulation results, thus avoiding the generation of artifacts that do not conform to the X-ray attenuation law.

[0049] The generated samples are selected through an automatic quality assessment mechanism and mixed proportionally with real samples for training the detection model. For the detection model, this invention employs a multi-scale target detection network based on the Transformer architecture with hierarchical feature representation capabilities. Preferably, improved Deformable DETR and RT-DETR structures are used as the basic framework, fusing high-resolution features through a multi-scale feature pyramid. For slender targets such as cracks, this invention introduces a cross-attention module into the backbone network to enhance long-range dependency modeling capabilities.

[0050] Feature maps of different resolutions are output through the backbone network and fused using a feature pyramid network to enhance sensitivity to small-scale cracks and low-contrast defects. During the training phase, to alleviate the severe imbalance between positive and negative samples, a dynamic label allocation strategy is introduced, allowing the model to adaptively adjust the weights of positive and negative samples based on prediction confidence and spatial overlap. Simultaneously, a focus loss mechanism is used to suppress the influence of easily classified samples. A joint optimization of classification and regression losses is employed, and the overall loss function can be expressed as:

[0051]

[0052] Classification loss We employ weighted focus loss to mitigate class imbalance and locate the loss. Generalized Intersection over Union (IoU) is used to enhance the regression accuracy of small target boxes, making the model pay more attention to hard-to-get sample regions.

[0053] To improve the model's adaptability to generated data, this invention employs a curriculum fine-tuning strategy. Initial training primarily uses real samples, gradually increasing the proportion of generated samples. Simultaneously, a domain discriminant network is introduced to perform binary classification training between generated and real samples, and adversarial loss is used to align feature distributions. This makes it difficult for the model to distinguish between real and generated images in the feature space, thereby improving generalization ability. Furthermore, this invention constructs sub-model branches for different defect categories during the specialized training phase. For example, for metal foreign object detection, high-frequency texture response is enhanced; for insulation crack detection, edge-sensitive convolutional kernels are added; and for contact offset detection, a structural keypoint regression auxiliary task is introduced to achieve multi-task joint optimization.

[0054] During training, generated samples and real samples are input into the detection network to extract semantic features from the intermediate layers. Distribution consistency constraints are applied to both to ensure that generated samples closely approximate real defects in the discrimination space. Furthermore, hard sample reinforcement is performed on regions with low prediction confidence from the detection model. Corresponding conditional weights are added during the generation stage, prompting the model to generate more challenging defect samples more frequently and improving its adaptability to complex structures.

[0055] The optimization objective of the entire system is a weighted combination of diffusion loss, flow field matching loss, detection loss, and consistency constraint loss. Through joint training, the generative model gradually approximates the true defect probability distribution, while the detection model learns a more stable decision boundary with the support of expanded data. After training, the generative model can be used to continuously supplement defect sample data, while the detection model is used to complete defect localization and classification in the actual inference stage.

[0056] By introducing diffusion path modeling, rectified flow reparameterization, and a consistent mapping mechanism step by step, this invention significantly improves inference efficiency while maintaining generation stability. Furthermore, through generation-detection closed-loop optimization, it ensures that generated samples truly contribute to improved detection performance. In industrial scenarios where defective samples are extremely scarce, this method maintains a high recall rate and a low false positive rate, demonstrating good deployment feasibility and scalability.

[0057] The above is an exemplary description of the invention. Obviously, the specific implementation of the invention is not limited to the above-described manner. Any non-substantial improvement made using the inventive concept and technical solution of the invention, or the direct application of the inventive concept and technical solution to other situations without modification, is within the protection scope of the invention.

Claims

1. A method for intelligent detection of X-ray defects with few samples based on a diffusion generation model, characterized in that, Includes the following steps: S1: Data Acquisition and Preprocessing: Acquire an X-ray image dataset of industrial equipment, which includes defect-free negative samples, a small number of labeled defect positive samples, and projection data of physical simulation structures; normalize and denoise all images to construct a training base dataset; S2: Defect Category Embedding Construction: Geometric edges, gray-level gradients, and local texture multi-scale features of positive defect samples are extracted through a hierarchical feature encoding network and mapped to a unified conditional embedding space; contrast constraints are introduced to optimize the embedding space while preserving the physical characteristics of defect X-ray imaging. S3: Construction and Training of Conditional Diffusion Generation Model: Based on the conditional diffusion probability model, a U-Net multi-scale denoising network with cross-scale attention fusion is constructed to integrate defect category embedding and location encoding; rectified flow reparameterization and consistency model distillation are introduced to optimize the sampling path and reduce the number of inference steps; a phased training mechanism is used to train the generation model and output highly realistic defect synthetic samples based on the conditional diffusion generation model. S4: Generation-Detection Joint Collaborative Training: Construct a multi-scale Transformer object detection network, and mix real samples and synthetic samples in proportion as the training set; introduce dynamic label allocation, focus loss and generalized intersection-union loss to optimize the detection network, and adopt curriculum learning fine-tuning and feature distribution consistency constraints to achieve closed-loop joint optimization of the generation model and the detection network; S5: Intelligent Defect Detection Inference: Input the X-ray image to be detected into the trained detection network, output the location coordinates and classification results of the defect, and complete the intelligent detection of defects in industrial X-ray images with few samples.

2. The intelligent detection method for few-sample X-ray defects based on a diffusion generation model according to claim 1, characterized in that, In step S1, the industrial equipment is a GIS combined electrical appliance; the defect positive sample includes four types of defects: loose bolts, misaligned contacts, cracked basin insulators, and foreign metal objects, and the annotation form is pixel-level segmentation annotation or bounding box annotation.

3. The intelligent detection method for few-sample X-ray defects based on a diffusion generation model according to claim 1, characterized in that, In step S3, the training objective function of the conditional diffusion probability model is: in, It is about random variables The mathematical expectation operator is used to calculate the mean loss under multivariate conditions. This is the intermediate state after adding noise. For diffusion time step, For defect condition embedding, For noise prediction networks, This is real noise; The objective function for stream matching training of the rectified stream is: in To predict the flow field function, For the temporal gradient of the image state, It is about random variables The mathematical expectation operator.

4. The intelligent detection method for few-sample X-ray defects based on a diffusion generation model according to claim 1, characterized in that, In step S3, the phased training mechanism specifically includes: In the first stage, unconditional pre-training was performed using defect-free negative samples and defective positive samples to learn the global structural distribution of X-ray images. In the second stage, defect category conditions are introduced for fine-tuning, and mask loss constraints are applied only to defect areas. In the third stage, a physical consistency regularization term is added to constrain the grayscale distribution of the generated image to match the statistical characteristics of the simulated projection data.

5. The intelligent detection method for few-sample X-ray defects based on a diffusion generation model according to claim 1, characterized in that, In step S3, the high-resolution branch of the multi-scale denoising network integrates dense residual blocks to enhance the ability to generate details of fine-grained defects such as cracks; the consistency model achieves rapid sampling and generation of defect samples in a single step or with fewer steps through cross-step consistency constraints.

6. The intelligent detection method for few-sample X-ray defects based on a diffusion generation model according to claim 1, characterized in that, In step S4, the target detection network is an improved Deformable DETR or RT-DETR network. A multi-scale feature fusion structure is introduced into the backbone network, and a cross-attention module is used for feature interaction to fuse multi-resolution features and enhance the detection sensitivity of small-scale, low-contrast defects.

7. The intelligent detection method for few-sample X-ray defects based on a diffusion generation model according to claim 1, characterized in that, In step S4, the overall loss function of the detection network is: in, This is a weighted focus loss used to mitigate sample class imbalance. The generalized intersection-union loss is used to optimize the regression accuracy of small target defect boxes; These are the balancing weighting coefficients.

8. The intelligent detection method for few-sample X-ray defects based on a diffusion generation model according to claim 1, characterized in that, In step S4, the course learning fine-tuning specifically involves: prioritizing training the detection network with real samples and gradually increasing the proportion of synthetic samples; combining adversarial domain alignment to reduce the difference in feature distribution between real and synthetic samples; and simultaneously constructing dedicated sub-model branches for different defect categories to achieve multi-task joint optimization.

9. The intelligent detection method for few-sample X-ray defects based on a diffusion generation model according to claim 1, characterized in that, In step S4, the closed-loop joint optimization specifically involves: extracting semantic features from the intermediate layer of the detection network and constraining the consistency of feature distribution between the synthesized samples and the real samples; based on the low-confidence region of the detection network, strengthening the weights of the corresponding defect conditions in the generation model, and generating difficult samples to expand the training set.

Citation Information

Patent Citations

  • Chip defect weak supervision semantic segmentation method based on YOLO and diffusion model

    CN120125824A

  • Track defect detection method and device based on multi-modal large model

    CN120997487A