Multi-view multi-label X-ray security check image contraband classification method based on information bottleneck

Through the multi-view multi-label X-ray security inspection method, the multi-modal feature representation and information bottleneck algorithm are used to solve the overlapping and hidden problems of contraband identification in X-ray security inspection images, and the accuracy and efficiency of contraband classification are improved.

CN120339727AActive Publication Date: 2025-07-18HANGZHOU DIANZI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510804865.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-18
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

There are problems of overlap, hiddenness and labeling in existing X-ray security images, resulting in low recognition efficiency and high cost, and poor detection effect from a single perspective.

Method used

A multi-view multi-label X-ray security image contraband classification method based on information bottlenecks is adopted. Multi-modal feature representation is generated through hierarchical feature extraction from top view and side view angles, and a tag-aware information bottleneck module and an adaptive lightweight semantic image interaction module are used for comprehensive prediction, and combined with a consistency and specialized information processing algorithm, the shared information is maximized and shared information is minimized.

Benefits of technology

It improves the accuracy and efficiency of the classification of contraband images in X-ray security inspection, reduces interference with identification of accompanying biological products, and enhances the detection ability of overlapping contrabands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339727A_ABST
    Figure CN120339727A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view multi-label X-ray security check image contraband classification method based on information bottleneck. Firstly, a hierarchical feature extractor is used for feature extraction; generating multi-modal feature representation based on the extracted features; then, based on the obtained multi-modal feature representation, a predicted value of each view is generated; and finally, performing comprehensive prediction according to the prediction values of the plurality of views. Through a multi-view multi-label information bottleneck algorithm composed of a label perception information bottleneck module and a self-adaptive lightweight semantic image interaction module, under the condition of label independence, shared information between different views is maximized, and meanwhile, the influence of noise information in the views is minimized. In addition, according to the method, a consistency particularity information processing algorithm is utilized, and a correlation function is calculated by redesigning information, so that the shared information is maximized, and meanwhile, the correlation between specific information in a single view and other views and the difference between the specific information and the shared information are minimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of multi-view multi-label classification learning and X-ray security inspection, and particularly relates to a method for classifying contraband in X-ray security inspection images of multi-view multi-labels based on information bottleneck. Background Art

[0002] With the continuous growth of the flow of people at public transportation stations, ensuring public safety has become particularly crucial. X-ray security inspection technology has become an important means to ensure the safety of public places. After items such as suitcases and backpacks are irradiated by X-rays, some items inside them will be displayed, helping security inspectors discover potential dangerous goods and ensuring the safety of public places. Due to the penetrability of X-rays, compared with natural light images, X-ray images have the following characteristics: (1) Overlap: When stacked items are irradiated by X-rays, an overlap phenomenon will occur in the image, resulting in blurred surface textures of the items in the image; (2) Concealment: Contraband in X-ray security inspection images is small in volume and overlaps with other items, making it difficult to be discovered by the system from a single-view angle; (3) Difficult annotation: Due to characteristics (1) and (2), it takes a huge amount of time and economic cost to perform object-level annotation on X-ray security inspection images.

[0003] Identifying different contraband in X-ray security inspection images can be regarded as a multi-label classification task. Compared with the object detection method, the object classification method does not require object-level annotation of the X-ray security inspection data set, but only requires simpler and more efficient image-level annotation. That is, through the object classification method, it can be judged whether there is contraband in the suitcase from the overall features of the image. In the suitcase, in addition to contraband, there are also various ordinary items, which are called accompanying items. These accompanying items are usually irrelevant to the classification task, but in the classification process, the feature information of these accompanying items will interfere with the normal identification of contraband. Therefore, in the classification process, it is necessary to filter and suppress this interfering feature information. In addition, due to the overlap phenomenon formed by stacked contraband in X-ray security inspection images, single-level visual features may not be able to fully capture the key information of these overlapping contraband. At the same time, in the process of detecting contraband in X-ray security inspection images from a single view, some contraband may be difficult to be identified due to the limitation of the observation angle. Summary of the Invention

[0004] Aiming at the deficiencies in the prior art, the present invention provides a method for classifying contraband in X-ray security inspection images of multi-view multi-labels based on information bottleneck.

[0005] The method for classifying contraband in X-ray security inspection images of multi-view multi-labels based on information bottleneck includes the following steps:

[0006] Obtain the X-ray security inspection images from the top view and side view perspectives as the input of the classification model, and use the corresponding hierarchical feature extractors to extract features respectively.

[0007] Generate multi-modal feature representations based on the extracted features, and generate prediction values for each view respectively; perform comprehensive prediction according to the prediction values of multiple views.

[0008] Set the total loss function, perform iterative training on the classification model; and verify on the validation set to obtain the optimal parameter model, and output the classification results of prohibited items.

[0009] In a possible implementation, divide the set of dual-view images including X-ray security inspection images from the top view and side view perspectives into a training set and a validation set;

[0010] Adjust the size of the images and perform normalization processing on the pixel values;

[0011] For the input images, generate the features under the shared view and individual views, namely the image feature representations, through the hierarchical feature extractor from the top view perspective and the hierarchical feature extractor from the side view perspective.

[0012] In a possible implementation, use the NLP language model BERT to generate the embedding representation of the label space; based on the extracted features and the embedding representation of the label space, use the Adaptive Lightweight Semantic Image Interaction module ALSI to generate multi-modal feature representations. Specifically, first the image feature representation passes through the image feature mapping matrix for mapping. At the same time, the label space embedding representation uses the label embedding mapping matrix for mapping. Perform the Hadamard product on the two mapped features, and apply a non-linear activation function to the product result; finally, perform weighted processing through the adaptive weight matrix to obtain the multi-modal feature representation.

[0013] In a possible implementation, the ALSI module maps the multi-modal feature representation into the prediction values of all views of each input image sample through MLP mapping.

[0014] In a possible implementation, according to the prediction values of all views of the th input image sample, adopt a multi-view fusion strategy for comprehensive prediction, and finally obtain the multi-view prediction output value, that is, perform calculation on all views of the k-th input image sample, then perform the Hadamard product calculation on the calculation result and all views of the original k input image samples, and perform a summation operation on the calculation result to obtain the final prediction value.

[0015] In a possible implementation manner, the total loss function includes a common feature information calculation loss function, a correlation loss function between the shared view and the specific features of each single view, a correlation loss function between each different single view, and a prediction loss function.

[0016] According to Inference 1 and Inference 2 in the label-aware information bottleneck module LAIB, the common feature information calculation loss function is calculated.

[0017] In the consistency and particularity information processing module CSIP, according to the consensus principle and the complementary principle, the correlation loss function between the shared view and the specific features of each single view, and the correlation loss function between each different single view are calculated.

[0018] For the prediction loss function, binary cross-entropy loss is adopted.

[0019] In a possible implementation manner, the prepared training set is input into the classification model, and multiple rounds of training are performed until convergence; the total loss function is used to measure the difference between the predicted output of the model and the true label, and the loss function is minimized to improve the prediction accuracy of the model; during the training process, the mAP metric is used as a measurement metric to continuously monitor the performance of the model on the validation set; whenever the best mAP value is obtained during the training process, the corresponding model configuration is recorded and these configurations are saved as the optimal parameters.

[0020] Compared with the prior art, the beneficial effects of the present invention are:

[0021] The multi-view multi-label X-ray security inspection image contraband classification method based on information bottleneck proposed by the present invention improves the performance of X-ray security inspection image contraband classification. Specifically, the present invention uses the multi-view multi-label information bottleneck MMIB algorithm composed of the label-aware information bottleneck module LAIB and the adaptive lightweight semantic image interaction module ALSI. Under the condition of label independence, while maximizing the shared information between different views, the influence of the noise information within the view is minimized. Among them, the LAIB module applies the information bottleneck technology and its label-related latent variables in the latent space jointly defined by the label and the image sample , while maximizing the task-related shared information, minimizing the redundancy within the view. The ALSI module dynamically adjusts the weight information between the image features and the label embedding representation, enhancing the sensitivity of the model to key image features during the multi-modal feature learning process. In addition, the present invention uses the consistency and particularity information processing CSIP algorithm. Through the multi-view consensus and complementary principle, while maximizing the shared information, minimizing the correlation between the specific information in a single view and other views, and the difference between this specific information and the shared information, improving the classification performance of the algorithm. Description of the Drawings

[0022] Figure 1 This is the flowchart of the method according to an embodiment of the present invention.

[0023] Figure 2 This is the schematic diagram of the overall structure of the model according to an embodiment of the present invention.

[0024] Figure 3 This is the schematic diagram of the structure of the hierarchical feature extractor according to an embodiment of the present invention.

[0025] Figure 4 This is the schematic diagram of the structure of the ALSI module according to an embodiment of the present invention.

[0026] Figure 5 This is the comparison result of the classification mAP. Detailed implementation manners

[0027] To better understand the purpose, structure and function of the present invention, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. The present invention proposes a multi-view multi-label X-ray security inspection image contraband classification method based on information bottleneck, and the overall process is as Figure 1 shown. The overall structure of the model is as Figure 2 shown. The present invention uses a multi-view multi-label information bottleneck MMIB algorithm composed of a label-aware information bottleneck module LAIB and an adaptive lightweight semantic image interaction module ALSI to maximize the shared information between different views under the condition of label independence and minimize the influence of the noise information within the view. Among them, LAIB applies the information bottleneck technology and its label-related latent variables within the latent space jointly defined by the label and the image sample, aiming to maximize the task-related shared information and minimize the redundancy within the view. ALSI dynamically adjusts the weight information between the image feature and the label embedding representation, and improves the sensitivity of the model to the key image features in multi-modal feature learning, as shown. In addition, the present invention uses a consistency special information processing CSIP algorithm. Through the multi-view consensus and complementarity principle, while maximizing the shared information, it minimizes the correlation between the specific information in a single view and other views, as well as the difference between these specific information and the shared information, thereby improving the classification performance of the algorithm. Figure 4 shown.

[0028] In a possible implementation manner, a hierarchical feature extractor is used for feature extraction, and the specific operations are as follows:

[0029] (1-1) Divide the obtained dual-view image set into a training set and a validation set: the training set is used for training the model, that is ; the validation set is used to verify the accuracy of the model, that is . Among them, R is the real number field, denotes the number of image samples in the training set, denotes the i-th training image sample, denotes the number of image samples in the validation set, denotes the j-th validation image sample, H denotes the image height, W denotes the image width, and 3 denotes the number of RGB channels.

[0030] (1 - 2) Each training sample corresponding label ; where denotes the image sample the number of objects contained in it, denotes in the the true category of the k-th object, where C denotes the total number of categories in the dataset. The image is resized and uniformly scaled to 224×224 pixels and the pixel values are normalized to ensure the consistency and effectiveness of the model input.

[0031] (1 - 3) For the input image, through the hierarchical feature extractor (specifically including the top - view perspective hierarchical feature extractor and the side - view perspective hierarchical feature extractor ), generate the feature representations under the shared view and the individual views, i.e., the image feature representations: , , and , where the superscript denotes the common feature information under the shared view, and the superscript denotes the specific feature information within the individual view.

[0032] Among them, in the hierarchical feature extractor , after extracting features from each feature extraction layer, the obtained extracted features are globally average - pooled and then feature concatenation operations are performed to obtain the image feature representation.

[0033] In a possible implementation manner, generate a multi - modal feature representation based on the extracted features, and the specific operations are as follows:

[0034] (2 - 1) Use the pre - trained BERT language model to generate each contraband label embedding representation , where denotes the dimension of each label embedding representation.

[0035] (2 - 2) Use the image feature representations , and the label embedding representation as the input of the adaptive lightweight semantic image interaction module ALSI to generate the multi - modal feature representation 。

[0036] In a possible implementation, based on obtaining the multi-modal feature representation , the prediction values of each view are generated separately in the ALSI module. That is, the MLP mapping is used to map the th view of the th input image sample into the prediction value 。

[0037] In a possible implementation, the comprehensive prediction is performed according to the prediction values of multiple views. All views of the kth input image sample are calculated, and then the calculation result is subjected to the Hadamard product calculation with all views of the original k input image samples, and the sum operation is performed on the calculation result to obtain the final prediction value.

[0038] In a possible implementation, according to Inference 1 and Inference 2 in the Label-Aware Information Bottleneck module LAIB, the loss function for calculating the common feature information is calculated . Specifically, Inference 1:

[0039]

[0040] Among them, represents the joint input of image X and label Y. is the mutual information function. represents entropy. is a variational approximation of the prior probability . In practical applications, is usually assumed to be a spherical Gaussian distribution. It uses the empirical data distribution , where is the Dirac function, used to represent the distribution of data points, is the total number of data points, is the th data point.

[0041] Inference 2:

[0042]

[0043] According to Inference 1 and Inference 2, the loss function for calculating the common feature information is defined as follows:

[0044]

[0045] Among them, E represents the expectation operation.

[0046] In a possible implementation, in the Consistency Specificity Information Processing Module (CSIP), according to the consensus principle and the complementary principle, the correlation loss function between the shared view and the specific features of a single view is calculated , and the correlation loss function between each different single view . The specific operations in step (6) are as follows:

[0047] (6-1) In the Consistency Specificity Information Processing Module (CSIP) is the loss function that measures the correlation between the shared information among views and the specific information within a single view. For the common feature and the specific feature , = , = , v , ; its correlation loss function is obtained by calculating the product between and taking the logarithm, and then summing up.

[0048] (6-2): In CSIP is the loss function that measures the correlation between the specific information within each view and the specific information of other views. For representing the specific features of each single view .

[0049] Using the Consistency Specificity Information Processing Module (CSIP), under the multi-view consensus and complementary principle, while maximizing the shared information, it minimizes the correlation between the specific information in a single view and other views, as well as the difference between this specific information and the shared information, improving the classification performance of the algorithm.

[0050] In a possible implementation, the model is iteratively trained based on the designed total loss function; and it is verified on the validation set to obtain the optimal parameter model and output the contraband classification result. The specific operations are as follows:

[0051] (7-1) The total loss function is designed as follows:

[0052]

[0053] Among them, is the function calculation parameter, that is is 's calculation parameter, is 's calculation parameter, is 's calculation parameter, is 's calculation parameter. is the prediction loss function, which is calculated using binary cross-entropy loss.

[0054] (7-2): Input the prepared training set into the model framework (prediction model) constructed through steps 1-4, and conduct multiple rounds of training until convergence. Use the loss function to measure the difference between the predicted output of the model and the true labels, and minimize the loss function to improve the prediction accuracy of the model. During the training process, use the mAP metric as the measurement index to continuously monitor the performance of the model on the validation set. Whenever the best mAP value is obtained during the training process, the corresponding model configuration will be recorded and these configurations will be saved as the optimal parameters. This approach helps to find the best combination of parameters during the training process, thereby improving the prediction accuracy of the model.

[0055] (7-3): After the training is completed, use the validation set to verify the optimal model obtained in step (7-2), obtain the metric parameters for the dataset detection by the final model, and mark the categories and confidence levels of the detected contraband on the detection results.

[0056] Experimental comparison data:

[0057] The datasets used are DvXray and OPIXray.

[0058] (1) DvXray is the first dual-view X-ray security inspection image dataset, with a total of 15 types of contraband, namely: Gun, Knife, Wrench, Pliers, Scissors, Lighter, Battery, Bat, Razor blade, Saw blade, Fireworks, Hammer, Screwdriver, Dart, and Pressure vessel. Among them, the following types of contraband use abbreviations: Wrench (Wren), Scissors (Scis), Battery (Batt), Razor blade (Razor), Saw blade (Saw), Fireworks (Fires), Hammer (Hamm), Screwdriver (Screw), Pressurevessel (PV). The DvXray dataset is shown in Table 1.

[0059] Table 1: DvXray dataset

[0060]

[0061] (2) The OPIXray dataset contains a total of 8,885 images, and there are five categories of contraband knives, namely: Folding knife, Straight knife, Scissor, Utility knife, and Multi-tool knife. OPIXray focuses on occluded contraband knives, and these knives are usually small in size and easily occluded by other objects, posing a new challenge to X-ray security inspection. The OPIXray dataset is shown in Table 2.

[0062] Table 2: OPIXray Dataset

[0063]

[0064] Experimental comparison description:

[0065] Figure 5 shows the comparison results of the classification mAP (cmAP) of the multi-view multi-label X-ray security inspection image contraband classification method X-MP based on information bottleneck of the present invention and other algorithms on the DvXray dataset. These algorithms include: CHR, DOAM, SXMNet, LIM, AHCR, MVCNN-new, MHBN, and GVCNN. From Figure 5 it can be seen that X-MP achieved the best result of 85.6%. It is 12.5% higher than AHCR, 16.3% higher than CHR, and 21% higher than GVCNN. This shows that X-MP has good performance in identifying contraband.

[0066] Table 3: Detection Comparison Results of X-MP on OPIXray

[0067]

[0068] Table 3 shows the detection comparison results of X-MP and AHCR on the OPIXray dataset. It can be seen from Table 1 that when X-MP uses ResNet-50 as the backbone network for feature extraction, i.e., the hierarchical feature extractor, the cmAP is 8.33% higher than that of AHCR. When X-MP uses ResNet-50 as the hierarchical feature extractor, the cmAP is 5.42% higher than when using Swin-T. In addition, in terms of identifying folding knives, ResNet-50 is 8.36% higher than Swin-T, which indicates that some limitations of the self-attention mechanism may cause Swin-T to miss some key features when detecting small and overlapping contraband items, resulting in performance degradation.

[0069] The above embodiments have elaborated in detail the objectives, technical solutions and advantages of the present invention. It should be clear that these embodiments only represent some preferred implementation ways of the present invention and do not limit its scope. Any form of modification, equivalent replacement or improvement made under the premise of following the core idea and principles of the present invention shall be regarded as being included within the protection scope of the present invention.

Claims

1. A multi-view and multi-label X-ray security inspection image contraband classification method based on information bottleneck, characterized in that, It includes the following steps: Obtain the X-ray security inspection images from the top view and side view perspectives as the input of the classification model, and use the corresponding hierarchical feature extractors to extract features respectively; Generate multi-modal feature representations based on the extracted features, and generate prediction values for each view respectively; make a comprehensive prediction based on the prediction values of multiple views; Set the total loss function and perform iterative training on the classification model; And verify on the validation set to obtain the optimal parameter model and output the contraband classification result.

2. The method for classifying contraband in multi-view and multi-label X-ray security inspection images based on information bottleneck according to claim 1, wherein, Divide the dual-view image set including X-ray security inspection images from the top view and side view perspectives into a training set and a validation set; Adjust the size of the images and perform normalization processing on the pixel values; For the input images, generate the features under the shared view and individual views, that is, the image feature representations, through the hierarchical feature extractor from the top view perspective and the hierarchical feature extractor from the side view perspective.

3. The multi-view multi-label X-ray security inspection image contraband classification method based on information bottleneck according to claim 2, characterized in that, Using the NLP language model BERT, generate the embedding representation of the label space; based on the extracted features and the embedding representation of the label space, use the adaptive lightweight semantic image interaction module to generate the multimodal feature representation; specifically, first the image feature representation is mapped through the image feature mapping matrix Meanwhile, the embedding representation of the label space is mapped using the label embedding mapping matrix The two mapped features are subjected to the Hadamard product, and the nonlinear activation function is applied to the product result; finally, weighted processing is performed through the adaptive weight matrix to obtain the multimodal feature representation.

4. The multi-view and multi-label X-ray security inspection image contraband classification method based on information bottleneck according to claim 3, characterized in that Establish an MLP mapping in the adaptive lightweight semantic image interaction module, use the obtained multi-modal feature representation as the input, and generate the prediction values of all views for each input image sample.

5. The multi-view multi-label X-ray security inspection image contraband classification method based on information bottleneck according to claim 4, characterized in that, According to the predicted values of all views of the th input image sample, a multi-view fusion strategy is adopted for comprehensive prediction, and finally a multi-view prediction output value is obtained, that is, all views of the kth input image sample are calculated, and then the calculation result is subjected to a Hadamard product calculation with all views of the original k input image samples, and the sum operation is performed on the calculation result to obtain the final predicted value.

6. The multi-view multi-label X-ray security inspection image contraband classification method based on information bottleneck according to claim 5, wherein The total loss function includes the loss function for calculating common feature information, the correlation loss function between the shared view and the specific features of the individual views, the correlation loss function between each different individual view, and the prediction loss function.

7. The method for classifying contraband in multi-view and multi-label X-ray security inspection images based on information bottleneck according to claim 6, characterized in that, Calculate the loss function for calculating common feature information according to Inference 1 and Inference 2 in the label-aware information bottleneck module.

8. The method for classifying contraband in multi-view and multi-label X-ray security inspection images based on information bottleneck according to claim 7, characterized in that, In the consistency and particularity information processing module, calculate the correlation loss function between the shared view and the specific features of the individual views, and the correlation loss function between each different individual view according to the consensus principle and the complementary principle.

9. The method for classifying contraband in multi-view and multi-label X-ray security inspection images based on information bottleneck according to claim 8, characterized in that, Input the prepared training set into the classification model and perform multiple rounds of training until convergence; use the total loss function to measure the difference between the predicted output of the model and the true label, and minimize the loss function to improve the prediction accuracy of the model; During the training process, use the mAP metric as the measurement metric and continuously monitor the performance of the model on the validation set; Whenever the best mAP value is obtained during the training process, record the corresponding model configuration and save these configurations as the optimal parameters.

Citation Information

Patent Citations

  • Multi-view multi-label classification method based on consensus-specific semantic fusion

    CN118097234A

  • Multi-label image classification method based on semantic guidance fusion

    CN118154938A

  • X-ray image contraband detection method based on de-overlapping and associated attention mechanism

    CN118261853A

  • Multi-view image contraband detection method based on NeRF and deep learning

    CN118486015A

  • Incomplete multi-view multi-label classification method based on task enhancement and missing view completion

    CN119762843A