Contraband X-ray security check image multi-angle generation method based on generative model

Through the multi-angle two-dimensional mapping and feature reconstruction module of the generative model, the quality and attitude control problems of contraband image generation in the intelligent security inspection system are solved, and high-quality contraband X-ray security inspection image generation is realized, which improves the stability of the model and data application capabilities.

CN120355828APending Publication Date: 2025-07-22THE THIRD RES INST OF MIN OF PUBLIC SECURITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510426522.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing intelligent security inspection system faces difficulties in obtaining data, insufficient generalization performance of model and privacy leakage risks in the generation of contraband X-ray images. Generative AI technology has the problem of difficulty in ensuring image quality and insufficient attitude control accuracy of target items.

Method used

The generative model is adopted to build a multi-angle two-dimensional mapping of the 3D model of the contraband, combine the feature reconstruction module and the discriminator, and use the generator and diffusion process to generate high-quality contraband X-ray security images to control the category and posture of the images.

Benefits of technology

It improves the generation quality and attitude control accuracy of contraband X-ray security images, enhances the training stability of the model, is suitable for multi-source data collaborative applications, and reduces the risk of privacy leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355828A_ABST
    Figure CN120355828A_ABST
Patent Text Reader

Abstract

The invention relates to a prohibited goods X-ray security check image multi-angle generation method based on a generative model, and the method comprises the steps: (1) importing a prohibited goods 3D model I3D into a 3D modeling system, carrying out the multi-angle two-dimensional mapping processing according to a set rotation mapping interval angle theta, and constructing a target prohibited goods semantic graph Ix; (2) sampling a nexhexwe probability graph E from the normal distribution N (mu, sigma2), and inputting the nexhexwe probability graph E into a feature reconstruction module to obtain a reconstructed feature graph Fre; and (3) taking the target contraband semantic graph Ix and the reconstructed feature graph Fre as inputs of a generator, outputting a contraband X-ray security check image IG according to the inputs, and completing training for a generative model through continuous optimization iteration processing. The invention further relates to a corresponding device, a processor and a storage medium thereof. By adopting the method, the device, the processor and the storage medium, the quality of the synthesized image is remarkably improved, and accurate control on fine-grained attributes such as types, postures and shielding proportions of prohibited goods is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of X-ray security inspection, and specifically refers to a method, device, processor, and computer-readable storage medium for generating multi-angle X-ray security inspection images of contraband based on a generative model. Background Art

[0002] Under the background of the increasingly severe global security situation, X-ray security inspection equipment has become the core security means in key places such as airports and rail transit. The current intelligent security inspection system is based on deep learning-based automatic detection algorithms, but its performance is highly limited by the quality and scale of training data. Therefore, the intelligent security inspection system faces three major limitations from data: 1. The strict control of contraband makes it difficult to obtain real X-ray images of contraband. 2. Traditional data augmentation methods cannot effectively simulate the poses of contraband in real scenarios, resulting in insufficient generalization performance of the model. 3. The use of real security inspection images poses a risk of privacy leakage, especially in cross-border data sharing scenarios, there are legal compliance obstacles, which seriously restricts the collaborative application of multi-source data.

[0003] In recent years, generative AI technologies such as generative adversarial networks (GANs) and diffusion models have shown significant advantages in data generation, providing a new technical path to alleviate the problem of data scarcity in the security inspection field, but some problems have also emerged, such as: it is difficult to guarantee the quality of the generated contraband images, the accuracy of target object pose control is insufficient, and the lack of annotation information. Therefore, developing an advanced multi-angle generation method for X-ray security inspection images of contraband based on a generative model has important research value and practical significance for promoting the application and development of X-ray security inspection technology. Summary of the Invention

[0004] The object of the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a method, device, processor, and computer-readable storage medium for generating multi-angle X-ray security inspection images of contraband based on a generative model.

[0005] To achieve the above object, the method, device, processor, and computer-readable storage medium for generating multi-angle X-ray security inspection images of contraband based on a generative model of the present invention are as follows:

[0006] The multi-angle generation method for X-ray security inspection images of contraband based on a generative model is mainly characterized in that the method includes the following steps:

[0007] (1) Import the 3D model I of contraband 3D into a 3D modeling system, and perform multi-angle two-dimensional mapping processing according to the set rotation mapping interval angle θ to construct the target contraband semantic map I x ;

[0008] (2) Create a probability map E that follows a normal distribution N(μ,σ 2 ), where the size of the probability map E is n e ×h e ×w e , and input it into the feature reconstruction module to obtain the reconstructed feature map F re ;

[0009] (3) Use the target contraband semantic map I x and the reconstructed feature map F re as the input of the generator, and accordingly output the contraband X-ray security inspection image I G , and complete the training of the generative model through continuous optimization and iterative processing.

[0010] Preferably, the multi-angle two-dimensional mapping process specifically includes:

[0011] (1.1) Construct a 3D model I of a single target item 3D , import it into the Blender software platform, and use the open-source 3D library to collect picture materials and convert them into the.off file format;

[0012] (1.2) Fix the position (x, y, z) of the target item I 3D , initialize the camera position (θ p , θ a ) in the Blender software platform, the image distance r, the rendering mode O, and the rendering resolution h×w, where x, y, and z are the x-axis coordinate, y-axis coordinate, and z-axis coordinate in the three-dimensional space coordinate system respectively, θ p is the pitch angle, θ a is the azimuth angle, r is the straight-line distance from the camera center to the target center, O is the projection mode of the camera to the target, h is the rendering height, and w is the rendering width;

[0013] (1.3) Sample the camera azimuth angle once every θ' degrees from the azimuth angle θ a to θ a ' degrees, and create a camera rotation spherical coordinate system position queue with a total of n = θ' a / θ' positions;

[0014] (1.4) Convert the camera spherical coordinate system position (θ p , θ a ) to the Cartesian coordinate system position (x c , y c , z c );

[0015] (1.5) Move the camera according to the Cartesian coordinate system position queue and project the target object with an output size of n 2d ×h 2d × 2d Two-dimensional Mapping I 2D , where n 2d ,h 2d and w 2d are the channel, height and width of the two-dimensional map respectively;

[0016] (1.6) For the two-dimensional mapping graph I 2D Binarize to construct a size n g ×h g × g The single-channel target semantic graph I′ x , where n g 、h g and w g are the channel, height, and width of the single-channel target semantic map, respectively;

[0017] (1.7) The single-channel target semantic graph I x ' is copied to the three channels of RGB respectively, and the size is n. rgb ×h rgb × rgb Three-channel target semantic graph I x , where n rgb 、h rgb and w rgb are the channel, height and width of the three-channel target semantic map respectively.

[0018] Preferably, the feature reconstruction module specifically performs the following processing:

[0019] The probability map E is upsampled and the feature extraction module composed of a 3×3 convolution layer and an exponential linear rectification layer ELU is used to expand and reorganize the feature space resolution to generate a n-th dimension. up ×h up × up The probability feature map F up , where n up 、h up and w up are the channel, height and width of the probability feature map respectively;

[0020] The probability feature map F up The feature reconstruction process is completed through the reconstruction layer composed of a 3×3 convolutional layer and a hyperbolic tangent activation layer Tanh, and the reconstructed feature map F is obtained. re .、

[0021] Preferably, the generator specifically performs the following processing:

[0022] Input the semantic map I of the target contraband x and the reconstructed feature map F re into the cross layer Cross for feature cross - processing, and then complete the reconstruction of the X - ray image I of the target object through the depth feature extraction and recombination of several 3×3 convolutional layers and the rectified linear unit ReLU G .

[0023] Preferably, the generative model further includes a discriminator, and the discriminator is used to guide the training of the generator and the feature reconstruction module. Specifically:

[0024] Input the image I' generated by the generator G and the corresponding real X - ray image sample I' of the target object R into the discriminator after splicing them in the channel dimension. The input of the discriminator undergoes feature extraction and recombination through several 3×3 convolutional layers and leaky rectified linear units to generate a predicted probability map E D ;

[0025] During the training stage of the generative model, the predicted probability map E D will be used as the input for the next iteration of the feature reconstruction module, and the model parameters of the generator and the feature reconstruction module will be optimized through backpropagation.

[0026] Preferably, the generative model includes a diffusion process, and its processing process is as follows:

[0027] (3.1) Initialize the parameters of the generator, discriminator, and feature reconstruction model in the generative model, and set the input as The target image is

[0028] (3.2) Query the output D of the discriminator at the previous iteration step t - 1 of the current iteration step t, and use it as the input for the feature reconstruction module for feature reconstruction processing; t-1

[0029] (3.3) Input the reconstructed feature map obtained at iteration step t into the generator to obtain a reconstructed image

[0030] (3.4) Input the reconstructed image output by the generator and the target image into the discriminator to obtain a predicted probability map The predicted probability map will be embedded in a query table for use as the input for the feature reconstruction module in the next iteration t + 1.

[0031] ​Preferably, the processing process further includes:

[0032] (3.5) Use the Adam function to optimize the loss function in the model. The feature reconstruction module and generator loss function of the generative model are defined as:

[0033]

[0034] Where, is the adversarial loss, Θ R is the network parameter of the feature reconstruction module of the generative model, Θ G is the network parameter of the generator of the generative model, λ is the weight control parameter of the reconstruction loss function, and L L1 is the reconstruction loss;

[0035] (3.6) Iterate step t = t + 1, and repeat the above process until t > T, where T is the number of denoising time steps in each training cycle; Take out a new image pair from the dataset Data index i = i + 1 until i > N, where is the semantic segmentation map of the (i + 1)-th real contraband X-ray image sample, is the (i + 1)-th real contraband X-ray image sample, and N is the total number of the dataset. Thus, one cycle of training of the generative model is completed;

[0036] (3.7) After Ep times of training, obtain the optimized generator parameter Θ′ G , the feature reconstruction module parameter Θ′ R and the discriminator parameter Θ′ D . For a target semantic image I x projected from a 3D model of contraband at multiple angles and a probability map E sampled from the normal distribution N(μ, σ 2 ), perform semantic reconstruction through I G = G(R(E), I x ) to obtain the X-ray security inspection image of the contraband with the corresponding category and pose, where G(·) is a hidden function representing the mathematical implicit operation process of the generator of the generative model, and R(·) is a hidden function representing the mathematical implicit operation process of the feature reconstruction module.

[0037] The multi-angle generation device for X-ray security inspection images of contraband based on the generative model is mainly characterized in that the device includes:

[0038] A processor configured to execute computer-executable instructions;

[0039] A memory stores one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned method for generating multi-angle X-ray security inspection images of contraband based on a generative model are implemented.

[0040] The processor for generating multi-angle X-ray security inspection images of contraband based on a generative model is mainly characterized in that the processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned method for generating multi-angle X-ray security inspection images of contraband based on a generative model are implemented.

[0041] The computer-readable storage medium is mainly characterized in that a computer program is stored thereon. The computer program can be executed by a processor to implement the steps of the above-mentioned method for generating multi-angle X-ray security inspection images of contraband based on a generative model.

[0042] By adopting the method, device, processor and computer-readable storage medium for generating multi-angle X-ray security inspection images of contraband based on a generative model of the present invention, the semantic image of the target object is obtained by the Blender system through two-dimensional mapping of a multi-angle 3D model, taking into account the physical characteristics of the mapping; the generative model is divided into a feature reconstruction module, a generator and a discriminator, improving the stability of model training; the probability prediction map generated by the feature reconstruction module can better simulate the uncertainty of the discriminator, enabling the generator to learn more complex image paradigms during training; the training of the generative model introduces a diffusion process, improving the quality of the generated contraband images. Compared with the prior art, the technical solution considers the multi-pose attributes of X-ray security inspection images of contraband in the application scenario and assigns different pixel values to the target semantic map according to categories, effectively controlling the category and pose of the generated images. The technical solution is very conducive to the application research of X-ray security inspection technology and related deep learning technologies, and is expected to be widely applied in the fields of X-ray security inspection image synthesis and the like. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flowchart of the method for generating multi-angle X-ray security inspection images of contraband based on a generative model of the present invention.

[0044] Figure 2 It is a schematic diagram of the process of multi-angle mapping of the target semantic map by the method for generating multi-angle X-ray security inspection images of contraband based on a generative model of the present invention.

[0045] Figure 3 It is a schematic diagram of the network structure of the feature reconstruction module in the generative model of the method for generating multi-angle X-ray security inspection images of contraband based on a generative model of the present invention.

[0046] Figure 4 Schematic diagram of the generator network structure in the generative model of the method for generating multi-angle X-ray security inspection images of contraband based on the generative model of the present invention.

[0047] Figure 5 Schematic diagram of the multi-angle projection process of the 3D model in the method for generating multi-angle X-ray security inspection images of contraband based on the generative model of the present invention.

[0048] Figure 6 Schematic diagram of the structure of the generative model applied in the present invention.

[0049] Figure 7 Schematic diagram of the discriminator network structure in the generative model applied in the present invention. Detailed implementation manners

[0050] In order to more clearly describe the technical content of the present invention, the following further description is made in conjunction with specific embodiments.

[0051] Before detailing the embodiments according to the present invention, it should be noted that, hereinafter, the terms "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.

[0052] Please refer to Figure 1 As shown, the method for generating multi-angle X-ray security inspection images of contraband based on the generative model of the present technical solution uses a 3D modeling system to construct a multi-angle planar mapping of the 3D model of the contraband, and obtains a multi-angle target contraband semantic map; uses a feature reconstruction module to generate a probability feature map for feature cross of the target contraband semantic map in the generator; the generator of the generative model performs domain transfer on the preprocessed target contraband semantic map, thereby generating the corresponding X-ray security inspection image of the contraband. The whole process of generating the contraband image and synthesizing the background is divided into the following steps:

[0053] Step 1: Import the 3D model I of the target item 3D into the 3D modeling system, collect multi-angle two-dimensional mappings at intervals of θ, and construct the target semantic map I x . Among them, θ is the rotation mapping interval angle, and the process of the multi-angle mapping semantic map of the 3D model is as Figure 2 shown.

[0054] Step 2: Sample n 2 from the normal distribution N(μ,σ e ×h e ×w e probability maps E, asFigure 3 Reconstructed by the feature reconstruction module as shown to form a reconstructed feature map F of size n R ×h R ×w R , where n re and n R are the image channels, h e and h R are the image height, w e and w R are the image width. Its mathematical model can be expressed as follows: e F

[0055] F re = R(E)

[0056] In the above expression, F re represents the cross feature map, and R(·) is an implicit function used to characterize the mathematical implicit operation process of the reconstruction module.

[0057] Step 3: Use the reconstructed feature map F re and the target semantic map I x as Figure 4 the input of the generator as shown, and finally output the contraband X-ray security inspection image. Its mathematical model can be expressed as:

[0058] I G = G(F re , I x )

[0059] where I G represents the generated contraband X-ray security inspection image, and G(·) is an implicit function used to represent the mathematical implicit operation process of the generator of the generation model.

[0060] In the above solution, the 3D modeling system is the Blender software platform.

[0061] Design of multi-angle two-dimensional mapping method

[0062] (1) Construct a 3D model I of a single target object 3D , import it into Blender. Use the open-source 3D library to collect materials and convert them into the.off file format.

[0063] (2) Fix the position (x, y, z) of the target object I 3D , initialize the camera position (θ p , θ a ) in Blender, the image distance r, the rendering mode O, and the rendering resolution h×w. Among them, x, y, and z are the x-axis coordinate, y-axis coordinate, and z-axis coordinate in the three-dimensional space coordinate system respectively, θ p is the pitch angle, θ aLet θ be the azimuth angle, r be the straight-line distance from the camera center to the target center, O be the projection mode from the camera to the target, h be the rendering height, and w be the rendering width.

[0064] (3) Sample the camera azimuth angle, sampling once every θ' degrees from θ a to θ' a degrees, and create a queue of camera rotation spherical coordinate system positions A total of n = θ' a / θ' positions are sampled.

[0065] (4) Convert the camera spherical coordinate system position (θ p , θ a ) to the Cartesian coordinate system position (x c , y c , z c ), and its mathematical expression is as follows:

[0066] (R1, R2) = (θ p , θ a ) × π / 180

[0067]

[0068] In the above expression, R1 and R2 are the radian representations of the pitch angle and azimuth angle respectively, x c , y c and z c are the x-axis coordinate, y-axis coordinate, and z-axis coordinate of the camera in the Cartesian coordinate system respectively. According to the above expression, convert the queue of camera rotation spherical coordinate system positions into a queue of camera rotation Cartesian coordinate system positions

[0069] (5) Move the camera according to the queue of Cartesian coordinate system positions, and project the target to output a two-dimensional mapping graph I 2d × h 2d × w 2d , as 2D shown, where n Figure 5 , h 2d and w 2d are the channel, height, and width of the two-dimensional mapping graph respectively. 2d

[0070] (6) Binarize the mapping graph I 2D to construct a single-channel target semantic graph I' g × h g × w g , where n x is the channel, h is the height, and w is the width. The binarization process is as follows: ch

[0071]

[0072] In the above expression, I′ x (i, j) is the pixel value of the target semantic map I′ x at the position (i, j), and I 2D (i, j) is the pixel value of the two-dimensional mapping map at the position (i, j). λ is a user-defined constant (determined by the target object category), and Gray(·) is an implicit function of the grayscale calculation process. Its mathematical expression is as follows:

[0073] Gray = α × R + β × G + γ × B

[0074] where R, G, and B represent the pixel values of the red channel, the green channel, and the blue channel, respectively. α, β, and γ are weight coefficients. After generating the single-channel target semantic map, the grayscale values in I′ x are respectively copied to the three RGB channels to construct a three-channel target semantic map I rgb × h rgb × w rgb with the size of x , that is:

[0075]

[0076] where c represents the image channel, and I x (c, i, j) represents the pixel value at the position (i, j) in channel c, and I x ′(1, i, j) represents the pixel value at the position (i, j) in channel 1.

[0077] Internal Structure Design of the Generative Model

[0078] The generative model structure in this solution is as Figure 6 shown. From left to right, this network sequentially includes: a feature reconstruction module, a generator, and a discriminator. During the training phase of the model, the discriminator will be used to guide the training of the generator and the feature reconstruction module respectively.

[0079] The introduction of each module of the generative model is as follows:

[0080] 1. Feature Reconstruction Module

[0081] The reconstruction module structure in this solution is as Figure 3 shown, including 3 upsampling layers Upsample, 4 3×3 convolutional layers, 3 exponential linear unit layers ELU, and 1 hyperbolic tangent activation layer Tanh. In this solution, with the size of n e × h e × w eThe probability map E is used as the input of the reconstruction module. First, through a series of upsampling and a feature extraction module composed of a 3×3 convolutional layer and an exponential linear unit (ELU), the feature space resolution is expanded and reorganized to generate a probability feature map F of size n up ×h up ×w up . In subsequent processing, the probability feature map completes the feature reconstruction process through a reconstruction layer composed of a 3×3 convolutional layer and a hyperbolic tangent activation layer (Tanh) to obtain the feature map F up . re .

[0082] 2. Generator

[0083] In this solution, the input of the generator module is the target semantic map I x and the output feature map F of the feature reconstruction module re . Its internal structure is as shown in Figure 4 , including 1 layer of cross layer (Cross), 4 consecutive stacks and skip connections of 3×3 convolutional layers and rectified linear unit (ReLU), and 1 layer of 3×3 convolutional layer. The target semantic map I x and the reconstructed feature map F re are first passed into the cross layer for feature crossing, and its mathematical model can be expressed as follows:

[0084] F c =(I x ×F re ) + I x

[0085] In the above expression, F c represents the cross feature map. The cross feature map undergoes a series of deep feature extraction and reorganization by 3×3 convolutional layers and rectified linear unit (ReLU) to complete the reconstruction of the X-ray image I of the target object G .

[0086] 3. Discriminator

[0087] The discriminator structure in this solution is as shown in Figure 7 , which consists of 3 layers of 3×3 convolutional layers and 4 leaky rectified linear units. During the training process, the image I' generated by the generator G and the corresponding real X-ray image sample I' of the target object R are concatenated in the channel dimension and input into the discriminator, and its data expression is as follows:

[0088]

[0089] where I D is the input of the discriminator, i represents the index of the height dimension, j represents the index of the width dimension, k is the index of the channel dimension, and n is the generated image I'G The number of channels. I D Feature extraction and recombination via a series of 3×3 convolutional layers and leaky rectified linear units to generate a predicted probability map E D . During the training phase of the generative model, the predicted probability map E D will be used as the input for the next iteration of the feature reconstruction module, and the model parameters of the generator and the feature reconstruction module will be optimized through backpropagation.

[0090] In this solution, the training of the generative model introduces a diffusion process, and its training dataset is is the i-th real contraband X-ray image sample, is the semantic segmentation map of the i-th real contraband X-ray image sample, N is the total number of the dataset, and i is the index. The training period of the generative model in this solution is Ep, and the number of denoising time steps in each training period is T. In the K-th period of training, from t = 1 to t = T, the process is introduced as follows:

[0091] (1) Initialize the parameters of the generator, discriminator, and feature reconstruction model of the generative model, and set the input as The target image is

[0092] (2) Query the output D t-1 of the discriminator at the previous iteration step t - 1 of the current iteration step t, and use it as the input of the feature reconstruction module for feature reconstruction. Its mathematical expression is as follows:

[0093]

[0094] where is the reconstructed feature map at iteration step t, D t-1 is the discriminator output at iteration step t - 1, D 0 is the discriminator output initialized at t = 1, U(μ1, μ2) is a uniform distribution within [μ1, μ2), and R(·) is a hidden function used to represent the mathematical implicit operation process of the reconstruction module.

[0095] (3) Input into the generator, and output the reconstructed image That is:

[0096]

[0097] In the above expression, is the generator output image at iteration step t, and G(·) is a hidden function used to represent the mathematical implicit operation process of the generator of the generative model.

[0098] (4) Input the generator output image Input into the discriminator to obtain the predicted probability map The predicted probability map will be embedded in the query table for use as the input to the feature reconstruction module in the next iteration t+1, i.e., D in step (2) t-1 .

[0099] (5) Calculate the loss function and optimize the loss function using the Adam function. The loss functions of the model feature reconstruction module and the generator are defined as follows:

[0100]

[0101] In the above formula is the adversarial loss, and its mathematical expression is:

[0102]

[0103] where represents the expected value of all possible and , and D(·) is the mathematical implicit operation process of the discriminator. L L1 is the reconstruction loss, and its mathematical expression is:

[0104]

[0105] where represents the expected value of all possible and , and ‖·‖1 is the L1 norm. Θ R is the network parameter of the model reconstruction module, and Θ G is the network parameter of the model generator. The reconstruction module and the generator are optimized simultaneously through the backpropagation of L 1 ({Θ G , Θ R}), and λ is the weight control parameter of the reconstruction loss function. The loss function of the model discriminator is defined as:

[0106]

[0107] In the above expressions represents the expected value of all possible and , represents the expected value of all possible and , and Θ D is the network parameter of the model discriminator.

[0108] ​(6) The iteration step \(t = t + 1\), repeat steps (2) - (5) until \(t>T\), and retrieve a new image pair from the dataset. The data index \(i = i + 1\).

[0109] (7) Repeat steps (2) - (6) until \(i>N\), and the training of one cycle of the model is completed.

[0110] After \(Ep\) times of training, the optimized generator parameters \(\Theta'\) G , the feature reconstruction module parameters \(\Theta'\) R and the discriminator parameters \(\Theta'\) D can be obtained.

[0111] For a target semantic image \(I\) projected from a 3D model of contraband from multiple angles x and a probability map \(E\) sampled from a normal distribution \(N(\mu,\sigma\) 2 ), the semantic reconstruction can be performed through \(I\) G =G(R(E),I x ) to obtain the X-ray security inspection image of contraband with the corresponding category and pose.

[0112] The following gives a specific implementation manner to further illustrate the multi-angle generation method of the X-ray security inspection image of contraband based on the generative model of the present technical solution:

[0113] Step 1: Import the 3D model \(I\) of the target object 3D into the 3D modeling system, and collect multi-angle two-dimensional mappings at intervals of to construct the target semantic map \(I\) x . Among them, the process of the multi-angle mapping semantic map of the 3D model is as shown in Figure 2 .

[0114] Step 2: Sample a \(1\times32\times32\) probability map \(E\) from the normal distribution \(N(0.5,0.1\) 2 ), which is reconstructed by the feature reconstruction module as shown in Figure 3 to form a reconstructed feature map \(F\) of size \(1\times256\times256\) re , and its mathematical model can be expressed as follows:

[0115] F re =R(E)

[0116] In the above expression, \(F\) re represents the cross feature map, and \(R(\cdot)\) is a hidden function used to characterize the mathematical implicit operation process of the reconstruction module.

[0117] Step 3: Use the reconstructed feature map \(F\) re and the target semantic map \(I\) x as Figure 4The input of the generator shown ultimately outputs the X-ray security inspection image of contraband, and its mathematical model can be expressed as:

[0118] I G = G(F re , I x )

[0119] Among them, I G represents the generated X-ray security inspection image of contraband, and G(·) is an implicit function used to represent the mathematical implicit operation process of the generator of the generation model.

[0120] In the above solution, the 3D modeling system is the Blender software platform.

[0121] Design of multi-angle two-dimensional mapping method

[0122] (1) Construct a 3D model I of a single target item 3D , and import it into Blender. Use the open-source 3D library to collect materials and convert them into the.off file format.

[0123] (2) Fix the position of the target object I 3D at (0, 0, 0), and initialize the camera position in Blender with the image distance r = 3, the rendering mode is orthographic projection, and the rendering resolution is 256×256.

[0124] (3) Sample the camera azimuth angle, sampling once every interval from 0 to 2π , and create a queue of camera rotation spherical coordinate system positions A total of n = 12 positions are sampled.

[0125] (4) Convert the camera spherical coordinate system position (θ p , θ a ) to the Cartesian coordinate system position (x c , y c , z c ), and its mathematical expression is as follows:

[0126] (R1, R2) = (θ p , θ a ) × π / 180

[0127]

[0128] In the above expression, R1 and R2 are the radian representations of the pitch angle and the azimuth angle respectively, x c , y c and z cThey are the x-axis coordinate, y-axis coordinate, and z-axis coordinate corresponding to the camera in the Cartesian coordinate system respectively. According to the above expressions, the camera rotation position queue in the spherical coordinate system is converted into the camera rotation position queue in the Cartesian coordinate system

[0129] (5) Move the camera according to the position queue in the Cartesian coordinate system, and project the target object to output a two-dimensional mapping graph I with a size of 3×256×256 2D , such as Figure 5 shown

[0130] (6) Binarize the mapping graph I 2D to construct a single-channel target semantic graph I' with a size of 1×256×256 x . The binarization process is as follows:

[0131]

[0132] In the above expressions, I' x (i,j) is the pixel value of the target semantic graph I' x at the position (i,j), I 2D (i,j) is the pixel value of the two-dimensional mapping graph at the position (i,j), and λ is a custom constant determined by the target object category (in this scheme, there are 9 types of contraband including dagger, fork, lighter, pistol, explosive, detonator, scissors, and pliers, and the corresponding λ lookup table is [1,2,3,4,5,6,7,8,9]). Gray(·) is an implicit function in the grayscale calculation process, and its mathematical expression is as follows:

[0133] Gray = α×R + β×G + γ×B

[0134] where R, G, and B represent the pixel values of the red channel, green channel, and blue channel respectively. α, β, and γ are weight coefficients, which are set to 0.299, 0.587, and 0.114 respectively. After generating the single-channel target semantic graph, the grayscale values in I' x are copied to the three RGB channels respectively to construct a three-channel target semantic graph I with a size of 3×256×256 x , that is:

[0135]

[0136] where c represents the image channel, and I x (c,i,j) represents the pixel value at the position (i,j) in the channel c, and I' x (1,i,j) represents the pixel value at the position (i,j) in the 1 channel

[0137] Internal structure design of the generative model

[0138] The generative model structure in this solution is as follows Figure 6 shown. From left to right, this network sequentially includes: a feature reconstruction module, a generator, and a discriminator. During the training phase of the model, the discriminator will be used to guide the training of the generator and the feature reconstruction module respectively.

[0139] The introduction of each module of the generative model is as follows:

[0140] 1. Feature reconstruction module

[0141] The structure of the reconstruction module in this solution is as follows Figure 3 shown, including 3 layers of upsampling layer Upsample, 4 layers of 3×3 convolutional layers, 3 exponential linear unit (ELU) layers, and 1 hyperbolic tangent activation layer Tanh. In this solution, the probability map E with a size of 1×32×32 is used as the input of the reconstruction module. First, through a series of upsampling and a feature extraction module composed of 3×3 convolutional layers and ELU layers, the feature space resolution is expanded and reorganized to generate a probability feature map F with a size of 1×256×256 up . In subsequent processing, the probability feature map completes the feature reconstruction process through a reconstruction layer composed of a 3×3 convolutional layer and a hyperbolic tangent activation layer Tanh, obtaining the feature map F re .

[0142] 2. Generator

[0143] In this solution, the input of the generator module is the target semantic map I x and the output feature map F of the feature reconstruction module re , and its internal structure is as follows Figure 4 shown, including 1 layer of cross layer Cross, 4 layers of continuous stacking and skip connections of 3×3 convolutional layers and rectified linear unit (ReLU) layers, and 1 layer of 3×3 convolutional layer. The target semantic map I x and the reconstructed feature map F re are first passed into the cross layer for feature crossing, and its mathematical model can be expressed as follows:

[0144] F c =(I x ×F re )+I x

[0145] In the above expression, F c represents the cross feature map. The cross feature map undergoes a series of depth feature extraction and reorganization of 3×3 convolutional layers and ReLU layers to complete the reconstruction of the target object X-ray image I G .

[0146] 3. Discriminator

[0147] The structure of the discriminator in this solution is as followsFigure 7 As shown in the figure, it consists of 3 layers of 3×3 convolutional layers and 4 leaky rectified linear units. During the training process, the generator generates the image I′ G which is concatenated with the corresponding real target X-ray image sample I′ in the channel dimension and input into the discriminator. Its data expression is as follows:

[0148]

[0149] where I D is the input of the discriminator, i represents the index of the height dimension, j represents the index of the width dimension, k is the channel dimension index, and n = 3 is the number of channels of the generated image I′ G . I D is reorganized through feature extraction of a series of 3×3 convolutional layers and leaky rectified linear units to generate the predicted probability map E D . During the training stage of the generative model, the predicted probability map E D will be used as the input for the next iteration of the feature reconstruction module, and the model parameters of the generator and the feature reconstruction module will be optimized through backpropagation.

[0150] In this solution, the training of the generative model introduces a diffusion process, and its training dataset is is the i-th real contraband X-ray image sample, is the semantic segmentation map of the i-th real contraband X-ray image sample, N = 2000 is the total number of the dataset, and i is the index. The training cycle of the generative model in this solution is Ep = 500, and the number of denoising time steps in each training cycle is T = 100. In the K-th cycle of training, from t = 1 to t = 100, the process is introduced as follows:

[0151] (1) Initialize the parameters of the generator, discriminator, and feature reconstruction model of the generative model, and set the input as the target image as

[0152] (2) Query the output D t-1 of the discriminator at the previous iteration step t - 1 of the current iteration step t, and use it as the input of the feature reconstruction module for feature reconstruction. Its mathematical expression is as follows:

[0153]

[0154] where is the reconstructed feature map at iteration step t, D t-1 is the output of the discriminator at iteration step t - 1, D 0The discriminator output initialized at t = 1, U(0,1) is a uniform distribution within [0,1), and R(·) is a hidden function used to represent the mathematical implicit operation process of the reconstruction module.

[0155] (3) Input into the generator to output the reconstructed image That is:

[0156]

[0157] In the above expression, is the output image of the generator at iteration step t, and G(·) is a hidden function used to represent the mathematical implicit operation process of the generator of the generation model.

[0158] (4) Input the generator output image and the target image into the discriminator to obtain the predicted probability map The predicted probability map will be embedded in the query table for use as the input to the feature reconstruction module in the next iteration t + 1, i.e., D in step (2) t-1 .

[0159] (5) Calculate the loss function and use the Adam function to optimize the loss function. The loss functions of the generation model feature reconstruction module and the generator are defined as:

[0160]

[0161] In the above formula is the adversarial loss, and its mathematical expression is:

[0162]

[0163] Among them, represents the expected value of all possible and , and D(·) is the mathematical implicit operation process of the discriminator. L L1 is the reconstruction loss, and its mathematical expression is:

[0164]

[0165] Among them, represents the expected value of all possible and , and ‖·‖1 is the L1 norm. Θ R is the network parameter of the generation model reconstruction module, Θ G is the network parameter of the generator of the generation model. Through L 1 ({Θ G ,Θ RThe backpropagation of λ = 0.5 simultaneously optimizes the reconstruction module and the generator, where λ = 0.5 is the weight control parameter of the reconstruction loss function. The loss function of the discriminator of the generative model is defined as:

[0166]

[0167] In the above expression, represents the expected value of all possible and , represents the expected value of all possible and , and Θ D are the network parameters of the discriminator of the generative model.

[0168] (6) Set the iteration step t = t + 1, and repeat steps (2) - (5) until t > 100. Then, take a new image pair from the dataset and increment the data index i = i + 1.

[0169] (7) Repeat steps (2) - (6) until i > 2000, and the training of one cycle of the generative model is completed.

[0170] After Ep times of training, the optimized generator parameters Θ′ G , the feature reconstruction module parameters Θ′ R and the discriminator parameters Θ′ D can be obtained.

[0171] For a target semantic image I x projected from multiple angles of a 3D model of contraband and a probability map E sampled from the normal distribution N(0.5, 0.1 2 ), the corresponding X-ray security inspection image of the contraband with the corresponding category and pose can be obtained through semantic reconstruction by I G = G(R(E), I x ).

[0172] The multi-angle generation device for X-ray security inspection images of contraband based on the generative model, wherein the device includes:

[0173] A processor configured to execute computer-executable instructions;

[0174] A memory storing one or more computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above-mentioned multi-angle generation method for X-ray security inspection images of contraband based on the generative model are implemented.

[0175] The multi-angle generation processor for contraband X-ray security inspection images based on a generative model, wherein the processor is configured to execute computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above-mentioned multi-angle generation method for contraband X-ray security inspection images based on a generative model are implemented.

[0176] The computer-readable storage medium, wherein a computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the above-mentioned multi-angle generation method for contraband X-ray security inspection images based on a generative model.

[0177] Any process or method description shown in a flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present invention includes additional implementations, where functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field of the embodiments of the present invention.

[0178] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution device.

[0179] Those of ordinary skill in the art of the present technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0180] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.

[0181] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "embodiment" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0182] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0183] The method, device, processor, and computer-readable storage medium for generating multi-angle contraband X-ray security inspection images based on a generative model of the present invention are adopted. The semantic image of the target object is obtained by the Blender system through two-dimensional mapping of a multi-angle 3D model, taking into account the physical characteristics of the mapping. The generative model is divided into a feature reconstruction module, a generator, and a discriminator, improving the stability of model training. The probability prediction map generated by the feature reconstruction module can better simulate the uncertainty of the discriminator, enabling the generator to learn more complex image paradigms during the training process. The training of the generative model introduces a diffusion process, improving the quality of the generated contraband images. Compared with the prior art, the present technical solution takes into account the multi-pose attributes of contraband X-ray security inspection images in the application scenario and assigns different pixel values to the target semantic map according to the category, effectively controlling the category and pose of the generated images. The present technical solution is very conducive to the application research of X-ray security inspection technology and related deep learning technologies, and is expected to be widely applied in fields such as X-ray security inspection image synthesis.

[0184] In this specification, the present invention has been described with reference to its specific embodiments. However, it is obvious that various modifications and transformations can still be made without departing from the spirit and scope of the present invention. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive.

Claims

1. A multi-angle generation method for contraband X-ray security inspection images based on a generative model, characterized in that, The method described above includes the following steps: (1) The contraband 3D model I 3D Import the 3D modeling system and perform multi-angle two-dimensional mapping processing according to the set rotation mapping interval angle θ to construct the semantic map of the target contraband I x ; (2) Randomly generate a normal distribution N(μ,σ 2 ) and randomly sample n e ×h e ×w e probability maps E, and input them into the feature reconstruction module to obtain the reconstructed feature map F re , where n e , h e and w e are the number of channels, height, and width of the probability map E, respectively; (3) Use the target contraband semantic map I x and the reconstructed feature map F re as the input of the generator, and accordingly output the contraband X-ray security inspection image I G , and complete the training of the generative model through continuous optimization and iterative processing.

2. The method for generating multi-angle X-ray security inspection images of contraband based on a generative model according to claim 1, wherein The multi-angle two-dimensional mapping processing specifically includes: (1.1)Construct the 3D model I of a single target object 3D , import it into the Blender software platform, and use the open-source 3D library to collect picture materials and convert them into the.off file format; (1.2) Fix the target object I 3D Position (x, y, z), initialize the camera position in the Blender software platform (θ p , θ a ), image distance r, rendering mode O, rendering resolution h × w, where x, y, and z are the x-axis coordinate, y-axis coordinate, and z-axis coordinate in the three-dimensional space coordinate system, respectively, θ p is the pitch angle, θ a is the azimuth angle, r is the straight-line distance from the camera center to the target object center, O is the projection mode from the camera to the target object, h is the rendering height, and w is the rendering width; (1.3) From the azimuth angle θ a to θ′ a Sample the camera azimuth angle every θ′ degrees from θ degrees to θ′ degrees, and create a camera rotation spherical coordinate system position queue A total of n = θ′ a / θ′ positions are sampled; (1.4) Convert the camera spherical coordinate position (θ p , θ a ) to the Cartesian coordinate position (x c , y c , z c ); (1.5) Move the camera according to the position queue in the Cartesian coordinate system, and project the target item to output a two-dimensional mapping diagram I with a size of n 2d ×h 2d ×w 2d Two-dimensional mapping diagram I 2D , where n 2d , h 2d and w 2d are the channel, height, and width of the two-dimensional mapping diagram, respectively; (1.6) Binarize the two-dimensional mapping diagram I 2D to construct a single-channel target semantic diagram I′ g with a size of n g ×h g ×w x , where n g , h g and w g are the channel, height, and width of the single-channel target semantic diagram, respectively; (1.7) Copy the gray values in the single-channel target semantic map I′ x into the three RGB channels respectively to construct a three-channel target semantic map I rgb of size n rgb × h rgb × w x , where n rgb , h rgb and w rgb are the number of channels, height, and width of the three-channel target semantic map respectively.

3. The multi-angle generation method of contraband X-ray security inspection images based on a generative model according to claim 1, wherein The feature reconstruction module specifically performs the following processing: The probability map E is upsampled and the feature extraction module composed of a 3×3 convolution layer and an exponential linear rectification layer ELU is used to expand and reorganize the feature space resolution to generate a n-th dimension. up ×h up × up The probability feature map F up , where n up 、h up and w up are the channel, height and width of the probability feature map respectively; The probability feature map F up The feature reconstruction process is completed through a reconstruction layer composed of a 3×3 convolutional layer and a hyperbolic tangent activation layer Tanh to obtain a reconstructed feature map F re .

4. The method for generating multi-angle X-ray security inspection images of contraband based on a generative model according to claim 3, wherein The generator specifically performs the following processing: The target contraband semantic map I x and the reconstructed feature map F re are input into the cross layer Cross for feature cross-processing, and then through the depth feature extraction and recombination of several 3×3 convolutional layers and the rectified linear unit ReLU to complete the reconstruction of the X-ray image I of the target object G .

5. The method for generating multi-angle X-ray security inspection images of contraband based on a generative model according to claim 4, wherein The generative model further includes a discriminator, and the discriminator will be used to guide the training of the generator and the feature reconstruction module. Specifically: The generated image I' generated by the generator G is concatenated with the corresponding real target object X-ray image sample I' in the channel dimension and input into the discriminator. The input of the discriminator is reorganized by feature extraction through several 3×3 convolutional layers and leaky rectified linear units to generate a predicted probability map E D ; During the training phase of the generative model, the predicted probability map E D will be used as the input for the next iteration of the feature reconstruction module, and the model parameters of the generator and the feature reconstruction module will be optimized through backpropagation.

6. The method for generating multi-angle X-ray security inspection images of contraband based on a generative model according to claim 5, wherein The generative model includes a diffusion process, and its processing procedure is as follows: (3.1) Initialize the parameters of the generator, discriminator, and feature reconstruction model in the generative model, and set the input as The target image is (3.2) Query the output D of the discriminator at the previous iteration step t - 1 of the current iteration step t, and use it as the input of the feature reconstruction module for feature reconstruction processing; t-1 , and use it as the input of the feature reconstruction module for feature reconstruction processing; (3.3) Input the reconstructed feature map obtained at iteration step t into the aforementioned generator to obtain a reconstructed image (3.4) The reconstructed image output by the generator and the target image are input into the discriminator described above to obtain a predicted probability map The predicted probability map described above will be embedded in a query table for use as the input to the feature reconstruction module in the next iteration t + 1.

7. The method for generating multi-angle X-ray security inspection images of contraband based on a generative model according to claim 6, wherein The processing procedure further includes: (3.5) The Adam function is used to optimize the loss function in the model. The loss function definitions of the feature reconstruction module and the generator of the generative model are: Among them, is the adversarial loss, Θ R is the network parameter of the generative model feature reconstruction module, Θ G is the network parameter of the generative model generator, λ is the weight control parameter of the reconstruction loss function, L L1 is the reconstruction loss; (3.6) Iterative step: t = t + 1, repeat the above process until t > T, where T is the number of denoising time steps in each training cycle; Take out a new image pair from the dataset Data index: i = i + 1 until i > N, where is the semantic segmentation map of the (i + 1)-th real contraband X-ray image sample, is the (i + 1)-th real contraband X-ray image sample, and N is the total number of the dataset. Thus, one cycle of training of the described generative model is completed; After Ep times of training, the optimized generator parameters Θ′ are obtained. G , the feature reconstruction module parameters Θ′ R and the discriminator parameters Θ′ D . For a target semantic image I projected from a 3D model of contraband at multiple angles x and a probability map E sampled from the normal distribution N(μ, σ 2 ), the X-ray security inspection image of contraband with the corresponding category and pose is obtained through semantic reconstruction by I G = G(R(E), I x ), where G(·) is an implicit function representing the mathematical implicit operation process of the generator of the generative model, and R(·) is an implicit function representing the mathematical implicit operation process of the feature reconstruction module.

8. A multi-angle generation device for contraband X-ray security inspection images based on a generative model, characterized in that, The device includes: A processor configured to execute computer-executable instructions; A memory storing one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method for multi-angle generation of contraband X-ray security inspection images based on a generative model according to any one of claims 1 to 7 are implemented.

9. A multi-angle generation processor for contraband X-ray security inspection images based on a generative model, characterized in that, The processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method for multi-angle generation of contraband X-ray security inspection images based on a generative model according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed by the processor to implement the steps of the method for multi-angle generation of contraband X-ray security inspection images based on a generative model according to any one of claims 1 to 7.

Citation Information

Cited By

  • Image decomposition method, device and computer program product

    CN121053146A