An image data enhancement method for target detection tasks
Generate high-quality X-ray contraband images through diffusion model network and image conversion technology, solving the problem of insufficient training data, achieving efficient data enhancement, and improving the detection performance and public safety guarantee capabilities of the X-ray security system.
Patent Information
- Application Number
- CN202510779647.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing X-ray security inspection system faces the problems of limited training data scale and insufficient sample diversity, which leads to insufficient generalization capabilities of detection models and attenuation of detection accuracy. Especially in the field of public safety, it is difficult to obtain images of real contraband, which limits the widespread application of data enhancement technology.
The diffusion model network is used to randomly generate various types of X-ray contraband images, convert the pseudo-color images into dual-energy X-ray RAW images through the image conversion network, and synthesize the X-ray contraband image using the dual-energy X-ray imaging principle, and combine the pseudo-color imaging technology to generate high-quality contraband image.
It enriches the diversity and complexity of the training data, improves the generalization ability and accuracy of the detection model, reduces data acquisition costs, shortens project cycles, and improves the efficiency and reliability of the security inspection system.
Smart Images

Figure CN120298279B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an image data enhancement method for target detection tasks. Background Art
[0002] In the field of computer vision, breakthroughs in object detection technology have driven rapid development in its applications across multiple scenarios. However, existing algorithms generally face the core constraints of limited training data and insufficient sample diversity, which directly lead to key issues such as insufficient model generalization and reduced detection accuracy. In the public safety sector, the development of automatic X-ray contraband detection systems has become a focus of both academia and industry. This technology, which uses X-ray transmission imaging to non-invasively inspect the internal structure of packages, holds irreplaceable strategic value in preventing public safety risks.
[0003] Current X-ray security inspection systems face two challenges in practical deployment: First, due to the lack of a unified standard training system, detection performance is often limited by the quality of security inspection images and the subjective experience of operators, resulting in significant individual differences in the reliability of detection results. Second, in peak traffic scenarios such as transportation hubs, the demand for real-time inspection of massive packages places enormous pressure on traditional manual screening models, which is prone to the risk of missed detections and misjudgments. To address these pain points, building an automated inspection system based on deep learning has become an inevitable choice, and overcoming the bottleneck of data augmentation technology is the core path to achieving system optimization.
[0004] To overcome these challenges, numerous researchers have explored in-depth technical approaches to synthesizing high-quality X-ray security inspection images. [Thomas WR, et al. Threat Image Projection (TIP) into X-ray Images of CargoContainers for Training Humans and Machines. ICCST 2016:1-7] pioneered a new framework for integrating Threat Image Projection (TIP) technology into cargo container X-ray images. This approach aims to train human observers and machine learning models to identify threat objects and enhance the robustness of machine learning algorithms. This technique cleverly exploits the approximate multiplication property of X-ray images to extract images from a threat object library and project them onto X-ray images of actual cargo. By incorporating multiple real-world scene variations, including translation, scaling, rotation, noise addition, illumination changes, volume and density adjustments, and occlusion, this approach significantly increases the diversity and challenge of the training data, thereby enhancing the adaptability of machine learning algorithms to complex environments. Experimental data strongly demonstrates that there is no significant difference in qualitative and quantitative evaluation between real threat images and TIP-synthesized images. Further application practice shows that the framework can effectively improve the performance of machine learning algorithms under controlled conditions in tasks such as cargo detection, concealed vehicle detection, and small metal threat detection. However, its practical application faces a significant challenge: the extremely difficult acquisition of real contraband images. In real-world scenarios, the probability of contraband appearing is extremely low, which means that collecting a sufficient number and variety of real contraband images for TIP technology is not only costly but also inefficient, thus limiting the widespread adoption and in-depth application of the technology.
[0005] In addition, [Dongming L, Jianchang L, Peixin Y, Feng Y, et al. A Data Augmentation Method for Prohibited Item X-Ray Pseudocolor Images in X-Ray Security Inspection Based on Wasserstein Generative Adversarial Network and Spatial-and-Channel Attention Block[J], Computational Intelligence and Neuroscience, 2022, 2022: 1-14.] proposed an innovative data augmentation method based on a Wasserstein generative adversarial network (WGAN) and a spatial-and-channel attention block, specifically designed for pseudo-color images of prohibited X-ray items. Through a carefully constructed framework, a unique spatial-and-channel attention block, a novel basic block, and a clever synthesis strategy, this method not only effectively achieves data augmentation for pseudo-color images of prohibited X-ray items, but also greatly enriches the diversity of the dataset and closely simulates the complex overlapping relationships between items in real scenes, thereby significantly enhancing the model's generalization ability and practical application results. However, training generative adversarial networks (GANs) is not easy, often facing challenges such as mode collapse, convergence difficulties, and unstable image quality. Furthermore, while this method has achieved significant progress in image generation, the quality of the generated X-ray security inspection images still needs to be further improved, particularly in terms of detail representation, texture realism, and the simulation of overlapping relationships between objects. Summary of the Invention
[0006] This paper proposes an image data enhancement method for target detection tasks, specifically for the field of X-ray contraband image detection. The core of this method lies in the use of an advanced diffusion model network, which, through a gradual denoising process, can generate a variety of X-ray contraband images of unprecedented quality. These images are not only diverse in variety but also have distinct features, significantly enriching the image library used for training and providing a solid foundation for synthesizing high-fidelity and highly practical X-ray images of contraband packages. Compared with the widely used generative adversarial networks (GANs), the diffusion model network exhibits greater stability during training, effectively avoiding problems such as mode collapse and convergence difficulties. Furthermore, the images generated by the diffusion model network are of higher quality, with finer details and more realistic textures, providing more reliable data support for subsequent target detection tasks. This innovative solution directly addresses key issues existing in the current technological landscape, such as the lack of image data and the insufficient generalization capability of detection models. It not only improves the accuracy and reliability of X-ray security inspection systems, but also provides strong support for ensuring public safety and improving security inspection efficiency.
[0007] The technical solution of the present invention is as follows: an image data enhancement method for target detection tasks, which randomly generates X-ray contraband images of different geometric shapes through a random generation module; collects X-ray pseudo-color images and dual-energy X-ray RAW images, uses the X-ray pseudo-color images as input values, and the dual-energy X-ray RAW images as true values, and jointly trains an image conversion network; generates a dual-energy X-ray contraband RAW image from the trained image conversion network for the X-ray contraband images, and generates a dual-energy X-ray wrapped RAW image from the trained image conversion network for the X-ray wrapped images; designs a mathematical model based on the principle of dual-energy X-ray imaging, and the mathematical model establishes an imaging relationship between the dual-energy X-ray contraband RAW image and the dual-energy X-ray wrapped RAW image, thereby synthesizing the X-ray contraband wrapped RAW image, and performs pseudo-color imaging to obtain the X-ray contraband wrapped image. The random generation module is a diffusion model network, including forward diffusion and reverse denoising based on the self-attention mechanism; the forward diffusion process adds noise to the original X-ray contraband image, and the reverse denoising process based on the self-attention mechanism predicts the noise at each time step 𝑘, removes the predicted noise from the original X-ray contraband image with the added noise, and generates an X-ray contraband image;
[0008] Specifically, the mathematical model is as follows: based on the linear relationship between image grayscale value and X-ray energy, a relationship formula between the grayscale value of the dual-energy X-ray contraband RAW image and the initial X-ray intensity is obtained, a relationship formula between the grayscale value of the dual-energy X-ray package RAW image and the initial X-ray intensity, and a relationship formula between the grayscale value of the X-ray contraband package RAW image and the X-ray energy formed by the X-ray transmission through the package containing the contraband is obtained;
[0009] Based on the relationship between the X-ray energy generated by X-ray transmission through a package containing contraband and the initial intensity of the X-ray, as well as the above three equations, a mathematical model for synthesizing a dual-energy X-ray RAW image of a contraband package is obtained.
[0010] In the forward diffusion process, according to the time step k Step by step processing of the original X-ray contraband image C 0 Add noise that satisfies the standard Gaussian distribution, and the time step range is: 1≤ k ≤1000; when k =1000, C k is a pure noise image; the forward diffusion process is shown in formula (1):
[0011] (1)
[0012] Where, Represents the time step k The weight of the added noise.
[0013] The reverse denoising based on the self-attention mechanism is implemented based on the denoising network; the denoising network is a ResNet encoding and decoding denoising network based on the self-attention mechanism, and the self-attention mechanism realizes the re-weighted combination of the noisy image features at each time step through the query Q, key K and value V; the input of the denoising network is the first k Step-added noise image C k , the output is k -1 step noise image C k-1 .
[0014] The denoising network uses ResNet as the network backbone, including a downsampling module, an intermediate feature processing module and an upsampling module;
[0015] The downsampling module includes a 2D convolution module and a self-attention mechanism module; the 2D convolution module performs preliminary feature extraction on the noisy image to obtain local texture and structural information of the noisy image; the self-attention mechanism module further explores the long-distance dependency between the local texture and structural information of the noisy image; the intermediate feature processing module includes a 2D convolution module and a self-attention mechanism module; the intermediate feature processing module is responsible for further refining and enhancing the output of the downsampling module; the intermediate feature processing module further extracts local features from the output of the downsampling module through a 3×3 two-dimensional convolution layer; the self-attention mechanism module re-weights and fuses local features; the upsampling module includes a nearest neighbor upsampling and a self-attention mechanism module; upsampling is performed based on the fused features extracted by the intermediate feature processing module to gradually predict noise;
[0016] At time step k The noise introduced in the forward diffusion process of is optimized by designing a loss function so that the predicted noise gradually approaches the introduced noise during the training process. The reverse denoising loss function of each time step 𝑘 is defined as:
[0017] Among them, represents the real noise, represents the predicted noise;
[0018] The variance and mean of the random generation module are calculated based on the predicted noise at each time step, as shown in formula (3):
[0019]
[0020] represents the weight of adding noise at time step k-1;
[0021] According to formula (4), the time step is obtained by the mean and variance k To time step k -1 image distribution, when k = 1, a randomly generated X-ray contraband image is obtained:
[0022]
[0023] The X-ray pseudo-color image is obtained by security inspection equipment; the dual-energy X-ray RAW image acquisition process is as follows: when X-rays pass through a package containing contraband, the energy of the X-rays attenuates according to the Bell-Lambert law, and the process is expressed by formula (5):
[0024]
[0025] in, I 0 represents the initial intensity of X-rays, I is the X-ray intensity received by the detector after attenuation, σ is the attenuation coefficient, T is the thickness of the object inside the package that the X-ray passes through;
[0026] When the detector receives the transmitted X-ray information, a dual-energy X-ray RAW image is formed. The dual-energy X-ray RAW image is a 16-bit high- and low-energy dual-energy X-ray image. The high-energy X-ray intensity I high and low-energy X-ray intensity I low It is expressed as follows:
[0027]
[0028] In formula (6), is the attenuation coefficient of high-energy X-rays, is the attenuation coefficient of low-energy X-rays.
[0029] An image conversion network is designed to convert an X-ray pseudo-color image into a dual-energy X-ray RAW image; an X-ray contraband image is converted into a dual-energy X-ray contraband RAW image by the image conversion network, and an X-ray wrapping image is converted into a dual-energy X-ray wrapping RAW image by the image conversion network;
[0030] The image conversion network mainly consists of a downsampling feature extractor and an upsampling generator;
[0031] The downsampling feature extractor includes a convolution layer and a maximum pooling layer; the convolution layer consists of a two-dimensional convolution module with a size of 3×3 and a stride of 1, a batch normalization module and a ReLU activation function; the upsampling generator includes a deconvolution layer and a convolution layer, and uses jump connections to channel-concatenate the features of each downsampling feature extractor with the deconvolution layer of the corresponding level upsampling generator, ultimately obtaining a dual-energy X-ray RAW image.
[0032] The synthesis process of the X-ray contraband package RAW image is as follows:
[0033] The generated dual-energy X-ray contraband RAW image and dual-energy X-ray package RAW image are used as data basis to perform image synthesis based on the dual-energy X-ray imaging principle;
[0034] The grayscale value of the dual-energy X-ray contraband RAW image is G con and the grayscale value of the dual-energy X-ray wrapped RAW image G back , expressed as:
[0035]
[0036] Where, α and β is the weight; is the attenuation coefficient of X-rays when passing through contraband; is the attenuation coefficient of X-rays when passing through a package; T con The thickness of the contraband; T back is the thickness of the package;
[0037] When X-rays penetrate a package containing contraband, the X-ray energy received by the detector is:
[0038]
[0039] The grayscale representation of the X-ray contraband package RAW image is:
[0040]
[0041] According to formulas (7), (8) and (9), the grayscale value of the synthesized X-ray contraband package RAW image is obtained: G syn Dual-energy X-ray contraband RAW image grayscale value G con and dual-energy X-ray wrapping RAW image grayscale value G back The mathematical relationship is:
[0042]
[0043] Formula (10) is a mathematical model constructed based on the principle of X-ray imaging, where G0 is the energy intensity received by the detector when the X-ray is unloaded. The RAW image of the X-ray contraband package is synthesized according to Formula (10), and then the final X-ray contraband package image is obtained through pseudo-color imaging technology.
[0044] The X-ray contraband package image is generated by performing pseudo-color imaging on the X-ray contraband package RAW image;
[0045] The dual energy value R reflects the recognition ability at low energy and high atomic number, and its calculation formula is shown in formula (11):
[0046]
[0047] When X-rays penetrate an object, assuming that the material of the object is evenly distributed, the material of the object is expressed in terms of effective relative atomic number. Z eff Expressed as follows, by formula (7) and the known RZ eff The curve is fitted, Z eff The calculation is shown in formula (12):
[0048]
[0049] in, and is a constant, i =1,2; X-ray contraband package RAW image is based on the effective relative atomic number Z eff The distribution of the artificial lookup table is used for coloring, and finally the X-ray contraband package image is obtained.
[0050] Beneficial effects of the present invention:
[0051] This invention uses a random generation module to randomly generate X-ray contraband images of varying geometric shapes, significantly enriching the morphological features of contraband in the dataset. In actual X-ray contraband detection scenarios, the shapes of contraband vary greatly. The method of this invention can simulate a wide range of possible shapes, making the training data more realistic. This diverse data helps the target detection model learn more comprehensive features, avoiding overfitting to contraband of specific shapes due to a single dataset, thereby significantly improving the model's generalization ability across different scenarios.
[0052] The present invention trains an image conversion network to convert randomly generated X-ray contraband images and X-ray package images into dual-energy X-ray contraband RAW images and dual-energy X-ray package RAW images, respectively. This process not only increases the data dimension but also integrates information from different image types. Different types of X-ray images contain features of contraband and packages under different imaging modes. The image conversion network can generate richer and more diverse data, providing a data foundation for the subsequent synthesis of X-ray contraband package images.
[0053] This invention utilizes the principles of dual-energy X-ray imaging to design a mathematical model, establishing an imaging relationship between dual-energy X-ray contraband RAW images and dual-energy X-ray package RAW images, thereby synthesizing an X-ray contraband package RAW image. This physics-based imaging simulation method accurately simulates the interaction between contraband and package during actual X-ray imaging, as well as the resulting imaging effect. Compared to traditional simple data augmentation methods, the images generated by this invention more closely resemble real-world X-ray images, providing more realistic and reliable training data for target detection models.
[0054] This method synthesizes large amounts of high-quality X-ray contraband package image data, significantly reducing reliance on actual data collection. This not only reduces data acquisition costs but also shortens project cycles, enabling object detection models to be put into practical application more quickly. It can effectively improve the performance of object detection models in X-ray contraband detection tasks, thus possessing broad application prospects and significant practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is the overall technical roadmap of the present invention;
[0056] Figure 2 Schematic diagram of the denoising network in the random generation module;
[0057] Figure 3 Schematic diagram of image conversion network;
[0058] Figure 4Schematic diagram of the dual-energy contraband package RAW image synthesis and X-ray contraband package image pseudo-color imaging;
[0059] Figure 5 Comparison of the X-ray contraband image generation results of the present invention with those of other excellent methods; (a) is method one; (b) is method two; (c) is method three; (d) is method four; and (e) is the method of the present invention.
[0060] Figure 6 Comparison of the X-ray contraband package image synthesis results of the present invention with those of other excellent methods; (a) is the X-ray package image; (b) is method one; (c) is method two; (d) is method three; (e) is method four; and (f) is the method of the present invention.
[0061] Figure 7 This is a subjective comparison diagram. DETAILED DESCRIPTION
[0062] The present invention relates to an image data enhancement method for target detection tasks, the method flow is as follows Figure 1 As shown in the figure, an X-ray contraband image is randomly generated through a diffusion model. The image conversion network trained on the X-ray contraband image generates a dual-energy X-ray contraband RAW image. The image conversion network trained on the X-ray wrapping image generates a dual-energy X-ray wrapping RAW image. A mathematical model is designed based on the principle of dual-energy X-ray imaging. The mathematical model establishes the imaging relationship between the dual-energy X-ray contraband RAW image and the dual-energy X-ray wrapping RAW image, thereby synthesizing the X-ray contraband wrapping RAW image. Pseudo-color imaging is then performed to obtain the X-ray contraband wrapping image.
[0063] (1) Random generation of X-ray contraband images
[0064] In reality, the probability of contraband appearing is low, resulting in data loss. To solve this problem, a random generation module is proposed. The essence of the random generation module is an unconditionally constrained diffusion model network, which can randomly generate X-ray contraband images of different postures. The random generation module consists of forward diffusion and reverse denoising based on the self-attention mechanism. In the forward diffusion process, according to the time k Step by step to give the original X-ray contraband image C 0 adds noise that satisfies the standard Gaussian distribution, k Satisfy 1≤ k ≤ K ,when k = K hour, C K Close to pure noise image, this process is shown in formula (1):
[0065]
[0066] Where, Represents the time step k The weight of the added noise.
[0067] The reverse denoising based on the self-attention mechanism is to predict the noise of each time step 𝑘 through the denoising network. The present invention proposes a ResNet encoding and decoding denoising network based on the self-attention mechanism, whose structure is as follows Figure 2 As shown, the input of the denoising network is k Step-added noise image C k , the output is k -1 step noise image C k-1 .
[0068] The denoising network uses ResNet as the network backbone, including a downsampling module, an intermediate feature processing module and an upsampling module;
[0069] The downsampling module includes a 2D convolution module and a self-attention mechanism module; the 2D convolution module performs preliminary feature extraction on the noisy image to obtain local texture and structural information of the noisy image; the self-attention mechanism module further explores the long-distance dependency between the local texture and structural information of the noisy image; the intermediate feature processing module includes a 2D convolution module and a self-attention mechanism module; the intermediate feature processing module is responsible for further refining and enhancing the output of the downsampling module; the intermediate feature processing module further extracts local features from the output of the downsampling module through a 3×3 two-dimensional convolution layer; the self-attention mechanism module re-weights and fuses local features; the upsampling module includes a nearest neighbor upsampling and a self-attention mechanism module; upsampling is performed based on the fused features extracted by the intermediate feature processing module to gradually predict noise;
[0070] At time step k The noise introduced in the forward diffusion process of is optimized by designing a loss function so that the predicted noise gradually approaches the introduced noise during the training process. The reverse denoising loss function of each time step 𝑘 is defined as:
[0071]
[0072] in, represents the real noise, represents the prediction noise;
[0073] Predict the noise at each time step Calculate the variance of the random generation module and mean , as shown in formula (3):
[0074]
[0075] represents the weight of adding noise at time step k-1;
[0076] According to formula (4), the time step is obtained by the mean and variance k To time step k -1 image distribution, when k = 1, a randomly generated X-ray contraband image is obtained:
[0077]
[0078] (2) Generation of dual-energy X-ray contraband RAW images and dual-energy X-ray package RAW images
[0079] This paper proposes an image synthesis strategy that aims to inject an X-ray image of contraband into an image of a package, thereby synthesizing the X-ray image of the contraband package required in the dataset. This process requires a RAW image of the X-ray image. When X-rays pass through a package containing contraband, the energy of the X-rays is attenuated according to the Beer-Lambert law. The process can be expressed as follows:
[0080]
[0081] in, I 0 represents the initial intensity of X-rays, I is the X-ray intensity received by the detector after attenuation, σ is the attenuation coefficient, and T is the thickness of the object in the package that the X-ray passes through. When the detector receives the transmitted X-ray information, a RAW image is formed. The RAW image is a 16-bit high- and low-energy dual-energy X-ray image. The intensities of high-energy and low-energy X-rays are expressed as I high and I low express:
[0082]
[0083] RAW images are important data for synthesizing X-ray images of contraband packages, but their disadvantage is that they are difficult to preserve and easy to lose. When RAW images are missing, the present invention designs an image conversion network to realize the conversion of X-ray pseudo-color images into dual-energy X-ray RAW images. Figure 3As shown in Figure 1, it consists of a downsampling feature extractor and an upsampling generator. First, the X-ray pseudo-color image is input to the downsampling feature extractor for feature extraction. The extracted feature module then generates high-energy RAW images and low-energy RAW images through the upsampling generator. The downsampling feature extractor includes a convolutional layer and a maximum pooling layer. The convolutional layer consists of a two-dimensional convolution module of size 3×3 and stride 1, a BatchNorm2d module, and a ReLU activation function. The upsampling generator includes a deconvolutional layer and a convolutional layer. The convolutional layers of the two feature extractors are connected to the deconvolutional layers through skip connections.
[0084] (3) X-ray contraband package RAW image synthesis
[0085] After the dual-energy contraband RAW image and package image are generated, they are synthesized using the dual-energy X-ray imaging principle. The grayscale value G of the RAW image can be linearly expressed as:
[0086]
[0087] Since the imaging strategies for high energy and low energy are the same, taking single energy as an example, the grayscale values of the X-ray foreground image, i.e., the dual-energy X-ray contraband RAW image, and the background image, i.e., the dual-energy X-ray package RAW image, are respectively G con and G back , expressed as:
[0088]
[0089] Where, α and β is the weight; is the attenuation coefficient of X-rays when passing through contraband; is the attenuation coefficient of X-rays when passing through a package; T con The thickness of the contraband; T back is the thickness of the package;
[0090] When X-rays penetrate a package containing contraband, the X-ray energy received by the detector is:
[0091]
[0092] The grayscale representation of the X-ray contraband package RAW image is:
[0093]
[0094] According to formulas (8), (9) and (10), the synthesized X-ray contraband package image grayscale value G is obtained syn The gray value G of the dual-energy X-ray contraband RAW imagecon and dual-energy X-ray wrapped RAW image grayscale value G back The mathematical relationship is:
[0095]
[0096] Formula (11) is a mathematical model constructed based on the principle of X-ray imaging, where G0 is the energy intensity received by the detector when the X-ray is unloaded. Based on Formula (11), a RAW image of the X-ray contraband package can be synthesized.
[0097] Figure 4 The synthesis process of X-ray contraband package RAW images is demonstrated. This process combines the dual-energy X-ray contraband RAW image with the dual-energy X-ray package RAW image to generate the X-ray contraband package RAW image, and then uses pseudo-color imaging technology to obtain the final X-ray contraband package image.
[0098] Therefore, it is necessary to utilize the high- and low-energy RAW image characteristics of the X-ray contraband package RAW image to generate a pseudo-color image. In dual-energy X-ray transmission technology, the dual energy value R is an important parameter that reflects the recognition ability at low energy and high relative atomic number. Its calculation formula is shown in formula (12):
[0099]
[0100] When X-rays penetrate an object, we assume that the material is evenly distributed, and the effective relative atomic number can be used. Z eff To express it, through formula (8) and the traditional RZ eff The curve is fitted, Z eff The calculation is shown in formula (13):
[0101]
[0102] in, ( i =1,2) and ( i =1,2) is a constant, depending on the high and low energy values. The RAW image of the X-ray contraband package is based on the effective relative atomic number Z eff The distribution of the artificial lookup table is used for coloring, and finally the X-ray contraband package image is obtained.
[0103] The dataset for generating X-ray contraband images and dual-energy X-ray RAW images consists of three parts: X-ray contraband images, X-ray package images, and their corresponding dual-energy X-ray RAW images. All images are sized uniformly at 512×512 to ensure data consistency and comparability. The X-ray contraband images cover eight common categories of contraband, including pistols, knives, javelins, liquid contraband, fireworks, scissors, pliers, and lighters. Each category of contraband includes 20,000 pseudo-color and RAW images. This data size not only meets training requirements but also provides strong support for the algorithm's generalization capabilities.
[0104] This invention cleverly utilizes advanced diffusion modeling technology to generate X-ray contraband images in a randomized, high-quality manner. This innovative approach significantly enriches the diversity and complexity of the data. In practical applications, the relative scarcity of contraband samples often results in incomplete training datasets, which in turn limits the performance of detection models. However, the diffusion model of this invention effectively overcomes this limitation, generating contraband images with diverse morphologies and distinct features. This significantly expands the training sample library and lays a solid foundation for improving the generalization capabilities of detection models.
[0105] Furthermore, this invention proposes a novel solution to the difficult storage and transmission of RAW image data. Utilizing multidimensional feature extraction technology and an upsampling network, we successfully convert X-ray pseudo-color images into dual-energy X-ray RAW image data. This method not only addresses the difficulty of preserving RAW image data but also provides valuable data resources for subsequent X-ray image synthesis of contraband packages. Through precise feature extraction and upsampling, we can generate data that is highly similar to the actual RAW image, ensuring the accuracy and reliability of the subsequent synthesized image.
[0106] In terms of image synthesis, this invention fully leverages the principles of dual-energy X-ray imaging to establish a precise mathematical model. This model effectively combines dual-energy X-ray contraband RAW images with dual-energy X-ray package RAW images to generate an X-ray contraband package RAW image containing the contraband. This technology not only more accurately reflects the physical characteristics of the contraband within the package, such as density and thickness, but also provides high-quality input data for subsequent target detection. By applying the principles of dual-energy imaging, we further enhance the realism and credibility of the synthesized image.
[0107] Finally, to enhance the visibility and identification of contraband in the image, this paper applies pseudo-color imaging to the synthesized X-ray contraband package RAW image. This process not only improves the visual quality of the image but also makes the contraband more prominent and easier to identify. By applying pseudo-color imaging technology, we can provide more intuitive and richer visual information for subsequent target detection, further improving detection accuracy and efficiency.
[0108] In summary, the entire algorithmic framework of this invention, from generation and combination to synthesis and imaging, forms a complete and efficient data augmentation process. This process not only addresses the shortage of contraband samples in real-world data but also overcomes the difficulty in storing RAW image data, providing strong support for training and improving the performance of X-ray contraband detection models. Ultimately, we successfully constructed an efficient and realistic training dataset, injecting new vitality into the development of intelligent and precise X-ray security inspection systems.
[0109] We qualitatively compare the algorithm proposed in this paper with several popular algorithms, and select Method 1 "Yue Z, Yutao Z, Haigang Z, Jinfeng Y, Zihao Z, et al. Data Augmentation ofX-Ray Images in Baggage Inspection Based on Generative Adversarial Networks[J], IEEE Access, 2020, 8: 86536-86544", Method 2 "Li D, Hu X, Zhang H, Yang J,et al. A GAN Based Method for Multiple Prohibited Items Synthesis of X-raySecurity Image[C], Chinese Conference on Pattern Recognition, 2021, 17(2):112-117.", Method 3 "Dongming L, Jianchang L, Peixin Y, Feng Y, et al. A DataAugmentation Method for Prohibited Item X-Ray Pseudocolor Images in X-RaySecurity Inspection Based on Wasserstein Generative Adversarial Network andSpatial-and-Channel Attention Block[J], Computational Intelligence andNeuroscience, 2022, 2022: 1-14." and method four "Jian Liu, Tim H. Lin. A Frameworkfor the Synthesis of X-Ray Security Inspection Images Based on GenerativeAdversarial Networks.[J], ISSS journal of micro and smart systems, 2023, 11:63751-63760.". These four algorithms all involve the generation of X-ray contraband images and the synthesis of X-ray contraband package images. In the X-ray contraband image generation experiment, the comparison results are as follows: Figure 5As shown. All five algorithms are able to generate X-ray contraband images with relatively complete outlines, but after detailed comparison, we found that the quality of the generation results of method one is not stable enough. For example, the generated images of scissors, knives and pistols are accompanied by a certain degree of distortion and low imaging quality. Relatively speaking, the X-ray contraband images generated by methods two, three and four are of higher and more stable quality, but there are still some defects, such as the noise of pistols and knives, and the distortion of the shapes of punches and scissors. After comparison, the images generated by the random generation network of X-ray contraband of the present invention have the smallest defects and the highest imaging quality, which fully demonstrates the significant superiority of the method of the present invention.
[0110] We inject randomly generated X-ray contraband images into package images that do not contain contraband to synthesize X-ray contraband security inspection images. In the experiment, we selected three types of randomly generated contraband: knives, lighters, and pistols for synthesis. The synthesis results are shown in Figure 2. Figure 6 As shown, the figure includes an overall rendering, a zoomed-in view of a section containing contraband, and an image of the injected contraband. Method 1 is able to fully synthesize an X-ray contraband security inspection image and demonstrate the hierarchical relationship between the contraband and the background package. However, the image's color contrast is low, and the overall visual quality needs improvement. In contrast, methods 2 and 3 produce higher-quality images, with the zoomed-in view showing a good overlap between the contraband and the package. However, a detailed comparison reveals that the overlapping area of the contraband produced by method 2 is relatively smooth, with some loss of texture information. The contraband image produced by method 3 contains noise, resulting in a decrease in the quality of the synthesized image. The image synthesized by method 4 also exhibits distortion, severely impacting image quality and subjective visual quality. In contrast, this method effectively synthesizes the image while generating high-quality X-ray contraband images. Comparison shows that the image synthesized by this method has high contrast, clearly depicts the overlapping area of the contraband and the package, and demonstrates the object's position and texture details, resulting in excellent visual quality. Therefore, the X-ray security inspection image synthesis algorithm of the present invention demonstrates significant superiority.
[0111] The subjective effect of human vision is an important basis for measuring image quality. We invited 20 experienced subway security personnel to conduct a subjective evaluation of the X-ray contraband security inspection images synthesized by each algorithm and the algorithm in this paper. The evaluation uses a scoring system, with 1 point representing "very poor" and 5 points representing "excellent". The evaluators need to score all five algorithms, and the scores of each algorithm cannot be repeated. The scoring results are as follows: Figure 7As shown in the figure, the algorithm evaluation scores for Methods 1 and 2 were primarily concentrated between 1 and 2, primarily due to the low contrast of the images synthesized by these two algorithms and color distortion in overlapping areas. In contrast, the scores for Methods 4 and 3 were relatively balanced, primarily concentrated between 3 and 4, indicating that the evaluators were relatively satisfied with the images synthesized by these two algorithms. It is worth noting that the images synthesized by the method of the present invention received the most highest scores and no lowest scores, demonstrating that the method of the present invention performed the best in the subjective evaluation.
[0112] Synthesized X-ray contraband inspection images are not only useful for routine training of security inspectors, but also effectively address the challenges of insufficient data or limited data types in contraband object detection. Data augmentation using the synthesized X-ray contraband inspection images generated by our proposed method demonstrates the importance of this method in the contraband object detection task. In experiments, we used real-world and synthetic X-ray contraband inspection images, preparing 800 images for each type of contraband. The synthesized image dataset was combined with the real-world image dataset to form an augmented dataset. A YOLOv11 object detection model was trained on the real-world, synthetic, and augmented datasets, and then detected real-world X-ray contraband images. The detection results for each dataset are shown in Table 1. Although the synthetic images still differ slightly from the real-world X-ray contraband inspection images, resulting in slightly lower detection performance than the real data, the augmented dataset shows improvements in all detection metrics. This demonstrates that data augmentation using the synthesized X-ray contraband inspection images generated by our proposed algorithm effectively improves the model's contraband detection performance.
[0113] Table 1 Detection results of different datasets
[0114] .
Claims
1. An image data enhancement method for target detection tasks, characterized in that: Randomly generate X-ray contraband images of different geometric shapes using a random generation module; collect X-ray pseudo-color images and dual-energy X-ray RAW images, using the X-ray pseudo-color images as input values and the dual-energy X-ray RAW images as true values, and jointly train an image conversion network; generate dual-energy X-ray contraband RAW images using the trained image conversion network for the X-ray contraband images, and generate dual-energy X-ray wrapped RAW images using the trained image conversion network for the X-ray wrapped images; design a mathematical model based on the principle of dual-energy X-ray imaging, the mathematical model establishing an imaging relationship between the dual-energy X-ray contraband RAW images and the dual-energy X-ray wrapped RAW images, thereby synthesizing the X-ray contraband wrapped RAW images, and performing pseudo-color imaging to obtain the X-ray contraband wrapped images; Specifically, the mathematical model is as follows: based on the linear relationship between image grayscale value and X-ray energy, a relationship formula between the grayscale value of the dual-energy X-ray contraband RAW image and the initial X-ray intensity is obtained, a relationship formula between the grayscale value of the dual-energy X-ray package RAW image and the initial X-ray intensity, and a relationship formula between the grayscale value of the X-ray contraband package RAW image and the X-ray energy formed by the X-ray transmission through the package containing the contraband is obtained; Based on the relationship between the X-ray energy generated by X-ray transmission through a package containing contraband and the initial intensity of the X-ray, as well as the three equations mentioned above, a mathematical model for synthesizing a dual-energy X-ray RAW image of a contraband package is obtained; The synthesis process of the X-ray contraband package RAW image is as follows: The generated dual-energy X-ray contraband RAW image and dual-energy X-ray package RAW image are used as data basis to perform image synthesis based on the dual-energy X-ray imaging principle; The gray value of the dual-energy X-ray contraband RAW image is G con and the grayscale value G of the dual-energy X-ray wrapped RAW image back , expressed as: Where α and β are weights; σ con is the attenuation coefficient of X-rays when passing through contraband; σ back is the attenuation coefficient of X-rays when passing through the package; T con is the thickness of the contraband; T back is the thickness of the package; When X-rays penetrate a package containing contraband, the X-ray energy received by the detector is: I syn =I0exp(-σ con T con -s back T back (8) The grayscale representation of the X-ray contraband package RAW image is: G syn =αI syn +β(9) According to formulas (7), (8) and (9), the synthetic X-ray contraband package RAW image gray value G is obtained syn The gray value G of the dual-energy X-ray contraband RAW image con and dual-energy X-ray wrapped RAW image grayscale value G back The mathematical model is: G syn =(G con -b)·(G back -β) / G0+β(10) Formula (10) is a mathematical model constructed based on the principle of X-ray imaging, where G0 is the energy intensity received by the detector when the X-ray is unloaded. The RAW image of the X-ray contraband package is synthesized according to Formula (10), and then the final X-ray contraband package image is obtained through pseudo-color imaging technology.
2. The image data enhancement method for target detection tasks according to claim 1, characterized in that The random generation module is a diffusion model network, including forward diffusion and reverse denoising based on the self-attention mechanism; During the forward diffusion process, noise is added to the original X-ray contraband image. The backward denoising process based on the self-attention mechanism predicts the noise at each time step k, removes the predicted noise from the original X-ray contraband image with the added noise, and generates the X-ray contraband image.
3. The image data enhancement method for target detection tasks according to claim 2, characterized in that: In the forward diffusion process, noise ξ that satisfies the standard Gaussian distribution is gradually added to the original X-ray contraband image C0 according to the time step k, and the range of the time step is: 1≤k≤1000; when k=1000, C k is a pure noise image; the forward diffusion process is shown in formula (1): Where, β k represents the weight of adding noise at time step k.
4. The image data enhancement method for target detection tasks according to claim 2, characterized in that: The reverse denoising based on the self-attention mechanism is implemented based on a denoising network; the denoising network is a ResNet encoding and decoding denoising network based on the self-attention mechanism, and the self-attention mechanism realizes the re-weighted combination of the noisy image features at each time step through the query Q, key K and value V; the input of the denoising network is the k-th step noisy image C k , the output is the noise image C of the k-1th step k-1 .
5. The image data enhancement method for target detection tasks according to claim 4, characterized in that: The denoising network uses ResNet as the network backbone, including a downsampling module, an intermediate feature processing module and an upsampling module; The downsampling module includes a 2D convolution module and a self-attention mechanism module; the 2D convolution module performs preliminary feature extraction on the noisy image to obtain local texture and structural information of the noisy image; the self-attention mechanism module further explores the long-distance dependency between the local texture and structural information of the noisy image; the intermediate feature processing module includes a 2D convolution module and a self-attention mechanism module; the intermediate feature processing module is responsible for further refining and enhancing the output of the downsampling module; the intermediate feature processing module further extracts local features from the output of the downsampling module through a 3×3 two-dimensional convolution layer; the self-attention mechanism module re-weights and fuses local features; the upsampling module includes a nearest neighbor upsampling and a self-attention mechanism module; upsampling is performed based on the fused features extracted by the intermediate feature processing module to gradually predict noise; The noise introduced during the forward diffusion process at time step k is optimized by designing a loss function so that the predicted noise gradually approaches the introduced noise during the training process. The reverse denoising loss function for each time step k is defined as: loss=||ξ k -ξ(C k ,k)|| 2 (2) Among them, ξ k represents the real noise, ξ(C k ,k) represents the prediction noise; According to the noise ξ(C k ,k) Calculate the variance of the random generation module And the mean μc, as shown in formula (3): β k-1 represents the weight of adding noise at time step k-1; According to formula (4), the image distribution from time step k to time step k-1 is obtained by the mean and variance. When k=1, the randomly generated X-ray contraband image is obtained:
6. The image data enhancement method for target detection tasks according to claim 1, characterized in that: The X-ray pseudo-color image is obtained by security inspection equipment; the dual-energy X-ray RAW image acquisition process is as follows: when X-rays pass through a package containing contraband, the energy of the X-rays attenuates according to the Bell-Lambert law, and the process is expressed by formula (5): I=I0exp(-σT)(5) Where I0 represents the initial intensity of the X-ray, I is the intensity of the X-ray received by the detector after attenuation, σ is the attenuation coefficient, and T is the thickness of the object inside the package that the X-ray passes through; When the detector receives the transmitted X-ray information, a dual-energy X-ray RAW image is formed. The dual-energy X-ray RAW image is a 16-bit high- and low-energy dual-energy X-ray image. The high-energy X-ray intensity I high and low-energy X-ray intensity I low It is expressed as follows: In formula (6), σ high is the attenuation coefficient of high-energy X-rays, σ low is the attenuation coefficient of low-energy X-rays.
7. The image data enhancement method for target detection tasks according to claim 1, characterized in that: An image conversion network is designed to convert an X-ray pseudo-color image into a dual-energy X-ray RAW image; an X-ray contraband image is converted into a dual-energy X-ray contraband RAW image by the image conversion network, and an X-ray wrapping image is converted into a dual-energy X-ray wrapping RAW image by the image conversion network; The image conversion network mainly consists of a downsampling feature extractor and an upsampling generator; The downsampling feature extractor includes a convolution layer and a maximum pooling layer; the convolution layer consists of a two-dimensional convolution module with a size of 3×3 and a stride of 1, a batch normalization module and a ReLU activation function; the upsampling generator includes a deconvolution layer and a convolution layer, and uses jump connections to channel-concatenate the features of each downsampling feature extractor with the deconvolution layer of the corresponding level upsampling generator, ultimately obtaining a dual-energy X-ray RAW image.
8. The image data enhancement method for target detection tasks according to claim 1, characterized in that: The X-ray contraband package image is generated by performing pseudo-color imaging on the X-ray contraband package RAW image; The dual energy value R reflects the recognition ability at low energy and high atomic number, and its calculation formula is shown in formula (11): When X-rays penetrate an object, assuming that the material of the object is uniformly distributed, the material of the object is represented by the effective relative atomic number Z. eff It is expressed by formula (7) and the known RZ eff Curve fitting, Z eff The calculation is shown in formula (12): Z eff =λ1exp(ξ1R)-λ2exp(-ζ2R) (12) Among them, λ i and ζ i is a constant, i=1,2; the RAW image of the X-ray contraband package is based on the effective relative atomic number Z eff The distribution of the artificial lookup table is used for coloring, and finally the X-ray contraband package image is obtained.
Citation Information
Patent Citations
Security check contraband detection method based on multi-scale attention and data enhancement
CN116883933A
Cross-spectrum image semantic segmentation method based on texture irrelevant features
CN117315240A