In-domain and cross-domain objectness learning method for point supervised x-ray contraband detection

CN119090818BActive Publication Date: 2026-09-25XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411089623.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-09-25
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

[0007]这些方法通常用于具有小的差距和相似类别的领域中,但由于自然图片和X射线图片之间存在很大差异,这些方法在X射线图片中的性能并不令人满意

Benefits of technology

[0080]采用上述的技术方案,本发明与现有技术相比,其具有的有益效果为:本方案提出了一种点监督X射线违禁品检测的域内-域间对象性学习方法,具体来说,为了解决现有方法严重依赖于费时费力的框标注这一问题,本方案能够通过域内-域间对象性模块在点监督下进行X射线违禁物品检测。该发明由两个关键模块组成:域内对象性学习(intra-OL)模块和域间对象性性学习(inter-OL)模块。域内对象性学习模块设计了局部焦点高斯掩蔽块和全局随机高斯掩蔽块,共同学习X射线图片中的对象性。同时,域间对象性学习模块引入了基于小波分解的对抗学习块和对象性块,有效减少了模态差异,并将从带有实例级标注的自然图片中学到的对象性知识迁移到了X射线图片中。基于以上内容,本发明大大缓解了X射线图片中由严重类内变化引起的局部主导的问题。在四个X射线数据集上的实验结果显示,本发明在显著降低注释成本的同时实现了卓越性能,提高了其实用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119090818B_ABST
    Figure CN119090818B_ABST
Patent Text Reader

Abstract

The application discloses an intra-domain-inter-domain object learning method for point supervision X-ray contraband detection, and the method is composed of an intra-domain object learning module and an inter-domain object learning module; the intra-domain object learning module designs a local focal Gaussian mask block and a global random Gaussian mask block to jointly learn the objectiveness in the X-ray picture; the inter-domain object learning module introduces an adversarial learning block and an objectiveness block based on wavelet decomposition, effectively reduces the modal difference, and migrates the objectiveness knowledge learned from natural pictures with instance-level labels to X-ray pictures; and the method alleviates the local dominant problem caused by severe intra-class variation in the X-ray picture. Experimental results on four X-ray data sets show that the application realizes excellent performance while significantly reducing annotation cost, and improves practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to an intra-domain and inter-domain object-oriented learning method for point-supervised X-ray contraband detection. Background Technology

[0002] X-ray images are characterized by complex backgrounds, severe occlusion, limited color and detail information, and a lack of contour detail, which can affect the extraction of overall object information. However, fully supervised object detection methods can obtain holistic information from instance-level annotations and the X-ray images themselves, thus achieving good performance. When it comes to improving the performance of fully supervised object detection, since instance-level annotations are fixed for specific datasets, the holistic information they provide is relatively fixed. Therefore, most existing methods focus on mining the inherent holistic information of X-ray images. It is safe to say that the good performance of these methods is also based on a large amount of expensive instance-level supervision.

[0003] As a special type of object detection task, X-ray prohibited object detection aims to detect the location and category of prohibited objects in X-ray images. Existing prohibited object detection methods are generally based on popular object detection paradigms (including one-stage and two-stage detectors). Unlike object detection in natural images, prohibited object detection in X-ray images is often affected by severe occlusion and object overlap. Several methods have been developed to address the object occlusion problem in X-ray images. For example, Wei et al. developed a Deocclusion Attention Module (DOAM) containing an edge attention module and a region attention module to detect occluded objects. Miao et al. assumed that each input image was sampled from a mixed distribution and introduced a class-balanced hierarchical refinement method (CHR). Tao et al. employed a Local Suppression Module (LIM) to identify discriminative features while ignoring irrelevant information in occluded objects. Wang et al. proposed a Selective Dense Attention Network (SDANet).

[0004] The methods described above all heavily rely on bounding box annotations. While these methods offer good performance, they typically require significant manpower to collect accurate annotations. To balance annotation costs and detection performance, this paper investigates prohibited item detection under point supervision. This method provides only one anchor point to locate the area near the center of the item, thus significantly reducing the total annotation time.

[0005] Weakly supervised object detection methods do not use bounding box annotations, but instead utilize weaker forms of supervision (such as image supervision and point supervision) to reduce annotation costs. Many methods have been proposed; for example, WSDDN was the first to introduce multi-instance learning (MIL) into weakly supervised object detection. OICR alleviates the partial dominance problem through iterative optimization. PCL expands the detection region through proposal clustering. ICMWSD addresses the partial dominance problem by employing a multi-instance self-training method and masks, while reducing memory consumption through batch backpropagation. OD-WSCL utilizes contrastive learning to achieve better pseudo-label accuracy. Essentially, these methods rely on prior bounding box generation to propose regions based on the overall information of the image, and then use MIL to filter these proposals. However, these prior bounding box generation methods are ineffective on X-ray images. Point-supervised object detection methods effectively address the problems of too many or too few instances and high memory consumption. This paper investigates a method based on simple yet effective wavelet decomposition to transfer target information learned from natural images to X-ray images, thus demonstrating the significant potential of utilizing natural images to solve the problem of prohibited object detection in X-ray images.

[0006] Domain-adaptive object detection aims to transfer models trained in the source domain to the target domain for testing. Existing methods can be broadly categorized into three types. The first type focuses on adjusting the feature distributions of different domains by applying adversarial learning or minimizing the maximum mean difference between the source and target domains. The second type employs a self-training strategy to generate high-quality pseudo-labels in the target domain. The third type utilizes a teacher-student framework to achieve domain-adaptive detection through consistency constraints based on detector predictions. Within the first type, Saito et al. proposed an adaptive detector method based on strong local alignment and weak global alignment. The second type employs a self-training strategy to generate high-quality pseudo-labels in the target domain. The third type utilizes a teacher-student framework to achieve domain-adaptive detection through consistency constraints based on detector predictions.

[0007] These methods are typically used in domains with small gaps and similar categories, but their performance in X-ray images is unsatisfactory due to the significant differences between natural images and X-ray images. Therefore, this application investigates a method for transferring target information learned from natural images to X-ray images, based on a simple yet effective wavelet decomposition. This clearly demonstrates the great potential of utilizing natural images to solve the problem of prohibited object detection in X-ray images. Summary of the Invention

[0008] In view of this, the purpose of this invention is to propose a point-supervised intra-domain and inter-domain object-oriented learning method for X-ray contraband detection. This method is reliable in implementation, flexible in application, and can alleviate the problem of local dominance caused by severe intra-class variations in X-ray image detection.

[0009] To achieve the above-mentioned technical objectives, the technical solution adopted by this invention is as follows:

[0010] A point-supervised intra-domain and inter-domain object-oriented learning method for X-ray contraband detection, comprising:

[0011] A. Prepare an X-ray contraband dataset and a natural image dataset, where the former contains only labeled points and the latter contains instance-level annotations;

[0012] B. Input the X-ray contraband dataset and the natural image dataset into the feature extractor to obtain the feature maps corresponding to the two datasets respectively;

[0013] C. Design an in-domain object learning network that includes local focal Gaussian occlusion and global random Gaussian occlusion to mine in-domain object information;

[0014] D. Input the feature map of the X-ray contraband dataset obtained in step B into the domain object learning network to obtain the enhanced feature map, and then feed it into the multi-instance branch learning network to predict the instance score of each instance generated around its labeled point.

[0015] E. Design an inter-domain objectivity learning network, which includes a wavelet decoupling adversarial learning network and an objectivity module;

[0016] F. Decouple the feature maps of the X-ray contraband dataset and the natural image dataset obtained in step B using a wavelet decoupling adversarial learning network, and perform adversarial learning on the low-frequency style features.

[0017] G. Utilize the objectivity module to predict the objectivity score of each target within the feature map of the X-ray contraband dataset and the natural image dataset;

[0018] H. The object score of the feature map of the X-ray contraband dataset and the instance score obtained in step D are weighted to obtain the final instance score of the multi-instance branch learning network; for the natural image dataset, its object score is supervised by instance-level annotation.

[0019] 1. Select instances from the X-ray contraband dataset whose scores are higher than a preset value, use them as pseudo-labels, and then train a fully supervised contraband detection model.

[0020] As a possible implementation, further, in step A of this solution, the X-ray contraband dataset includes the OPIXray dataset, HIXray dataset, SIXray dataset, and PIDray dataset; the OPIXray dataset contains 5 categories, totaling 8,885 images; the SIXray dataset contains 6 categories, totaling 1,059,231 images; the HIXray dataset contains 8 categories, totaling 45,364 images; and the PIDray dataset contains 12 categories, totaling 124,486 images.

[0021] Among them, the OPIXray, HIXray, SIXray, and PIDray datasets all include instance-level annotations. This approach uses 47,677 images containing contraband for training and testing. Points are randomly generated near the center point of the instance-level annotations according to a Gaussian distribution and used as annotation points during training.

[0022] The natural image dataset is the commonly used COCO14 dataset. In the COCO14 dataset, the training set has 82,783 images, the validation set has 40,504 images, and the test set has 40,775 images. This scheme randomly selects the same number of images from the COCO14 training set as the X-ray dataset to alleviate the imbalance in the number of images between the X-ray contraband dataset and the natural image dataset.

[0023] As a preferred implementation option, in step B of this scheme, the feature extractor is constructed by combining ResNet50 and FPN (Feature Pyramid Network); the combination of ResNet50 and FPN forms a multi-layer feature extraction network; ResNet50 is responsible for extracting more detailed and local features at the lower level, while FPN performs feature fusion and introduces contextual information at the higher level.

[0024] Among them, ResNet50 is a classic convolutional neural network architecture with strong feature extraction capabilities. FPN (Feature Pyramid Network), on the other hand, can extract and fuse features at different scales, enabling the network to better handle objects of varying sizes.

[0025] By combining ResNet50 and FPN to construct a feature extractor, features related to prohibited items are extracted from the X-ray prohibited item dataset. Simultaneously, features of general objects are extracted from the natural image dataset, resulting in feature maps for the two datasets, designated as natural image and X-ray image respectively, to obtain feature representations for different datasets. This combined approach allows the network to have a better receptive field and semantic information, thereby improving the performance of prohibited item detection. Furthermore, obtaining feature representations for different datasets provides richer and more accurate information for subsequent prohibited item detection tasks.

[0026] As a preferred implementation option, step C of this solution preferably includes:

[0027] An intradomain object-oriented learning network is designed, incorporating Local Focus Gaussian Masking (LFGM) and Global Random Gaussian Masking (GRGM) to mine intradomain object-oriented information. The purpose of this design is to better mine and utilize intradomain object-oriented information. By combining LFGM and GGRM, this network improves the model's robustness, generalization ability, and recognition accuracy, thereby achieving more accurate and reliable detection and recognition of contraband items in complex scenes.

[0028] The LFGM and GRGM blocks are designed to better utilize object-specific information within the domain. Local Focus Gaussian Occlusion (LFGML) guides the model to focus on non-central regions of the prohibited item by generating a local Gaussian mask. It explicitly forces the model to learn these non-central regions by finding points with local maximum activation values ​​and generating Gaussian masks around them. This approach prevents the model from overemphasizing areas near labeled points, allowing for a more comprehensive understanding of the entire prohibited item's shape and structure. Through the LFGM block, the model can learn and utilize various parts of the prohibited item, improving detection and recognition accuracy.

[0029] Specifically, considering that proposal boxes are generated based on point annotations around the center of the prohibited item, almost all proposal boxes include the annotated center region of the prohibited item. For example, when the model is trained to detect hammers, proposal boxes are generated around the annotation points. This approach may cause the model to focus on the handle around the annotation points while ignoring other non-central, discriminative regions (such as the head). To address this issue, we search for points with local maxima of activation and generate local Gaussian occlusions, explicitly forcing the model to learn the non-central regions of the prohibited item.

[0030] Technically, it includes: for a given X-ray image In and its corresponding dot annotations First, take image I n Input the backbone network and FPN to obtain feature maps at different scales. in L represents the number of feature maps, H l W l and C l Representing feature maps F respectively n,l Height, width, and number of channels; for feature map F n,l Each channel, first in a size of (βW) l )×(βH l Find the point with the highest activation within the window, where β represents the scaling factor; this point is used as the average position for generating the local Gaussian mask. The above process can be expressed as:

[0031]

[0032] in, Represents the feature map of the c-th channel. Average position on;

[0033] At the same time, for each F n,l A covariance matrix ∑ is defined l =αR l ,in It is a two-dimensional diagonal matrix, where α represents the scaling factor;

[0034] Finally, a local Gaussian mask is applied to each channel of the feature map.

[0035]

[0036] in, This represents the feature map after masking. Represents a local Gaussian mask;

[0037] Globally Random Gaussian Occlusion (GRGM) enhances the robustness and generalization ability of a model by introducing globally random Gaussian occlusion. The GRGM block randomly occludes certain regions of an image, forcing the model to learn and rely on other discriminative features. Since the occlusion is randomly generated, the model needs to learn features outside the occluded region to predict the target, thereby enhancing the model's ability to recognize different shapes and locations of contraband. Through the GRGM block, the model can better adapt to various disturbances and variations, improving performance and generalization ability in complex environments.

[0038] By applying the LFGM block, the model is forced to learn non-central regions. However, the model may still focus on local areas; by randomly applying multiple Gaussian masks on the feature map using the GRGM block, the model is prevented from over-focusing on locally discriminative regions and learns the entire object globally, including both central and non-central regions; this approach is beneficial for utilizing object-oriented knowledge of intra-object modalities. Unlike the LFGM block, the GRGM block randomly generates points in each feature map. Specifically, for each feature map F... n,l Randomly generate a set of points in Representing feature map F n,l Let m be a randomly generated point (x-coordinate and y-coordinate), and M represent the total number of points generated; based on this, for each point... Generate a Gaussian mask

[0039] Applying these Gaussian masks to the feature map F n,l And subtract them from the original feature map.

[0040]

[0041] Where F″ n,l This represents the feature map after masking.

[0042] In obtaining F′ n,l and F″ n,l Then, the two feature maps are concatenated, and a 1×1 convolutional layer is used to reduce their channel count to obtain feature E. n,l This feature is used as the input feature for the MIL branch.

[0043] As a preferred implementation option, step E of this solution preferably includes:

[0044] Design an inter-domain objectness learning network to effectively transfer objectness information extracted from natural images to X-ray images; the network consists of two key components, including a wavelet decomposition based adversarial learning network (WDAL) and an objectness block (OB);

[0045] Among them, the wavelet decoupled adversarial learning network uses wavelet decomposition technology to decouple the feature maps of the X-ray contraband dataset and the natural image dataset, dividing them into style and content parts; then, through adversarial learning, it reduces the modal differences caused by style differences, thereby enabling the model to better adapt to the characteristics of X-ray images;

[0046] The object-oriented module leverages transferred object-oriented knowledge and integrates it into the multi-instance learning branch. By calculating the object-oriented score of each Region of Interest (RoI) feature, this module enhances the probability that the entire contraband item is included in the proposal box. These two components work together to enable the model to accurately identify and locate contraband items in X-ray images, improving the efficiency and accuracy of security detection.

[0047] As a preferred implementation option, step F of this solution preferably includes:

[0048] In wavelet decomposition adversarial learning networks, due to the significant modal differences between natural images and X-ray images, directly transferring object-related knowledge extracted from natural images to X-ray images is not ideal. To address this issue, the WDAL module uses wavelet decomposition to decouple the feature maps of the X-ray contraband dataset and the natural image dataset into style and content components. Then, adversarial learning is used to reduce the modal differences caused by style differences. The WDAL module is inspired by the effectiveness of wavelet decomposition in preserving content information.

[0049] Technically, wavelet decomposition is performed using Haar wavelets. Haar wavelets consist of four kernels, namely {LL} T LH T HL T HH T}, where L and H are the low-pass and high-pass filters, respectively, defined as follows:

[0050]

[0051] Wavelet decomposition transforms the feature map M n,l From X-ray image I n Or natural image J n Extracting from it, we break it down into four parts, namely...

[0052] A n,l =M n,l *(LL T )

[0053] H n,l =M n,l *(LH T )

[0054] V n,l =M n,l *(HL T )

[0055] D n,l =M n,l *(HH T )

[0056] Where * denotes the convolution operation; A n,l Indicates the low-frequency range; H n,l V n,l and D n,l Indicates the high-frequency portion;

[0057] The low-frequency component mainly corresponds to style features, capturing detailed style information of the image; the high-frequency component corresponds to content features, focusing on local details and edges; then, A... n,l The input is fed into a convolutional layer with learnable parameters to obtain the transformed style features A′. n,l To reduce modal discrepancies, adversarial learning is applied to the transformed style features, treating the natural image dataset and the X-ray contraband dataset as the source and target domains, respectively. A domain classifier is trained on the activation values ​​of each transformed style feature to predict the modality of each style feature. Mathematically, let... It is the domain label of the nth image, where if the nth image belongs to the source domain, then Otherwise, it is 1; the activation value of the transformed style feature at coordinates (x, y) is represented as Utilizing corresponding style features Predicted value at each position Calculate the adversarial loss L based on cross-entropy adv ,Right now

[0058]

[0059] To align modal distributions and optimize the parameters of the modality classifier by minimizing the adversarial loss, and to optimize the parameters of the base network by maximizing this loss, a gradient inversion layer (GRL) is also used to train the modality classifier.

[0060] Then, the content feature H n,l V n,l and D n,l and the transformed style features A′ n,l By performing a connection, we obtain the mode-free characteristic H. n,l ,Right now

[0061]

[0062] in This indicates a splicing operation.

[0063] As a preferred implementation option, step G of this solution preferably includes:

[0064] The objectivity module is used to predict the objectivity score of each target in the feature map of the X-ray contraband dataset and the natural image dataset. The objectivity information extracted from the natural image is transferred to the X-ray image using the WDAL module. This information is then incorporated into the MIL branch of the objectivity module, which enhances the probability that the proposal box contains the entire contraband.

[0065] The object score reflects the probability that each region of interest (RoI) contains the entire item; it includes: assuming... R represents the RoI feature of a proposal generated from the k-th annotation point of the n-th X-ray contraband image or natural image, where R represents the total number of proposals;

[0066] RoI pooling is performed on each feature map, and the final objectivity score is calculated through a fully connected layer. The objectivity module can be represented as:

[0067]

[0068] Wherein, RoI_Pooling() and FC() represent RoI pooling and fully connected layers, respectively; This represents the objectivity score corresponding to the r-th proposal box generated from the k-th point;

[0069] For natural images, rich bounding box annotations can be used to train an object-oriented module. Specifically, for proposed bounding boxes generated from point annotations, the intersection-union ratio (IUR) between them and the annotations of the standard bounding boxes is calculated. In this model, proposals with an intersection-union ratio (IU) greater than or equal to 0.5 are considered positive samples and labeled as 1, while proposals with an IU less than this value are considered negative samples and labeled as 0. Based on this, the cross-entropy loss function can be used to adjust the model parameters.

[0070]

[0071] in, and These represent the regions annotated by the r-th proposal box and its corresponding annotation box, respectively.

[0072] Finally, I 2 The joint loss of OL-Net is

[0073] L = L mil +λ1L adv +λ2L obj

[0074] Among them, L mil λ represents the total MIL loss defined in P2BNet; λ1 and λ2 are the balancing weights.

[0075] In step H, for the X-ray contraband dataset, when processing X-ray images, the objectivity score and instance score are weighted to obtain the final instance score of the multi-instance branch learning network. This weighting operation aims to ensure that the network can accurately locate and identify contraband items. Technically, this means... Multiply by the instance score of the annotation at the k-th point in the MIL branch, i.e., the k-th bag of MIL; for proposal boxes with higher objectivity scores, larger weights are assigned; the classification score and subsequent MIL loss calculation remain unchanged. This process can be represented as follows:

[0076]

[0077] in, Represents the weighted instance score; Softmax() and FC ins () represent the Softmax function and the instance sub-branch, respectively.

[0078] For natural image datasets, instance-level annotations are used for object-oriented scoring supervision. This means providing detailed annotations for each instance in the natural image, enabling the network to accurately learn and understand target information within the image. In this way, the rich object-oriented information from the instance-level annotations in the natural image dataset can be effectively utilized and applied to the X-ray contraband detection task, thereby improving the network's performance on X-ray images.

[0079] In step I, the highest-scoring instances are first selected from the X-ray contraband dataset. These instances are considered positive samples with high confidence and are used as pseudo-labels. Next, these pseudo-labels are used to train a fully supervised contraband detection model. This method fully utilizes information already present in the X-ray contraband dataset and natural images, further improving the model's performance in detecting contraband. This training strategy not only strengthens the model's ability to identify contraband but also improves its generalization ability to various contraband items, making the model more reliable and robust in practical applications.

[0080] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: This solution proposes a point-supervised intra-domain and inter-domain objectivity learning method for X-ray contraband detection. Specifically, to address the problem that existing methods heavily rely on time-consuming and laborious bounding box annotation, this solution enables X-ray contraband detection under point supervision through an intra-domain and inter-domain objectivity module. The invention consists of two key modules: an intra-domain objectivity learning (intra-OL) module and an inter-domain objectivity learning (inter-OL) module. The intra-domain objectivity learning module designs local focal Gaussian masking blocks and global random Gaussian masking blocks to jointly learn objectivity in X-ray images. Simultaneously, the inter-domain objectivity learning module introduces wavelet decomposition-based adversarial learning blocks and objectivity blocks, effectively reducing modal differences and transferring objectivity knowledge learned from natural images with instance-level annotations to X-ray images. Based on the above, the present invention significantly alleviates the problem of local dominance caused by severe intra-class variations in X-ray images. Experimental results on four X-ray datasets show that the present invention achieves superior performance while significantly reducing annotation costs, thus improving its practicality. Attached Figure Description

[0081] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0082] Figure 1 This is a flowchart illustrating the entire implementation process of the method in this embodiment of the invention.

[0083] Figure 2 This is a diagram of the entire network structure of the detection model in an embodiment of the present invention. Detailed Implementation

[0084] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0085] This embodiment presents a point-supervised X-ray contraband detection method with intra-domain and inter-domain object-oriented learning capabilities, comprising:

[0086] Step A: Prepare an X-ray contraband dataset and a natural image dataset. The former contains only point annotations, while the latter contains instance-level annotations.

[0087] This approach uses four commonly used X-ray contraband detection datasets: OPIXray, HIXray, SIXray, and PIDray. The OPIXray dataset contains 8,885 images across 5 categories. The SIXray dataset contains 1,059,231 images across 6 categories. This approach uses 8,929 images containing contraband for training and testing. The HIXray dataset contains 45,364 images across 8 categories. The PIDray dataset contains 124,486 images across 12 categories, of which this approach uses 47,677 images containing contraband for training and testing. All these datasets include instance-level annotations; however, in our experimental setup, points are randomly generated near the center of each instance-level annotation using a Gaussian distribution as annotation points for training.

[0088] The natural image dataset used is the commonly used COCO14 dataset. In the COCO14 dataset, the training set has 82,783 images, the validation set has 40,504 images, and the test set has 40,775 images. To alleviate the image imbalance between the X-ray contraband dataset and the natural image dataset, this approach randomly selected the same number of images from the COCO14 training set as the X-ray dataset.

[0089] The X-ray contraband dataset and the natural image dataset from Step B and Step A are input into the feature extractor to obtain feature maps for the two datasets respectively.

[0090] For the feature extractor, this scheme uses a combination of ResNet50 and FPN. ResNet50 is a classic convolutional neural network architecture with strong feature extraction capabilities. FPN (Feature Pyramid Network) can extract and fuse features at different scales, enabling the network to better handle objects of varying sizes.

[0091] This approach combines ResNet50 with FPN to form a multi-layered feature extraction network. ResNet50 is responsible for extracting more detailed and local features at the lower levels, while FPN performs feature fusion and incorporates contextual information at higher levels. This combination allows the network to have a better receptive field and semantic information, thereby improving the performance of contraband detection.

[0092] By using a feature extractor built with ResNet50 and FPN, we can extract features related to prohibited items from X-ray prohibited item datasets, and also extract features of general objects from natural image datasets. This allows us to obtain feature representations for different datasets, providing richer and more accurate information for subsequent prohibited item detection tasks.

[0093] Step C: Design an in-domain object-oriented learning network. This network includes local focal Gaussian occlusion and global random Gaussian occlusion to mine in-domain object-oriented information.

[0094] Combination Figure 2 As shown, this scheme designs an intra-domain object-oriented learning network. This network includes Local Focus Gaussian Masking (LFGM) and Global Random Gaussian Masking (GRGM) to mine intra-domain object-oriented information. The purpose of designing the intra-domain object-oriented learning network is to better mine and utilize intra-domain object-oriented information. By combining LFGM and GGRM, this network can improve the model's robustness, generalization ability, and recognition accuracy, thereby achieving more accurate and reliable detection and recognition of contraband items in complex scenes.

[0095] The LFGM and GRGM blocks in this scheme are designed to better utilize the object-specific information within the domain. The LFGM block guides the model to focus on the non-central regions of the prohibited item by generating a local Gaussian mask. This method explicitly forces the model to learn the non-central regions of the prohibited item by finding points with local maximum activation values ​​and generating Gaussian masks around those points. The benefit of this is that the model no longer overemphasizes the area near the labeled point, but can more comprehensively understand the shape and structure of the entire prohibited item. Through the LFGM block, the model can learn and utilize various parts of the prohibited item, improving the accuracy of detection and recognition.

[0096] Specifically, considering that proposal boxes are generated based on point annotations around the center of the contraband, almost all proposal boxes include the annotated center region of the contraband. For example, when the model is trained to detect hammers, proposal boxes are generated around the annotation points. This approach may cause the model to focus on the handle around the annotation points, while ignoring other non-central, discriminative regions (such as the head). To address this issue, we search for points with local maxima of activation and generate local Gaussian occlusion, explicitly forcing the model to learn the non-central regions of the contraband. Technically, given an X-ray image I... n and its corresponding dot annotations First, take image I n Input the backbone network and FPN to obtain feature maps at different scales. in L represents the number of feature maps, H l W l and C l Representing feature maps F respectively n,l The height, width, and number of channels of the feature map F. n,l For each channel, we first in a space of size (βW) l Find the point with the highest activation within a window of 1 / (βH1)×(βH1) (β represents the scaling factor), and use this point as the average position for generating the local Gaussian mask. The above process can be expressed as:

[0097]

[0098] in Represents the feature map of the Cth channel. The average position on.

[0099] At the same time, for each F n,l We also defined a covariance matrix ∑ l =αR l ,in It is a two-dimensional diagonal matrix, and α represents the scaling factor.

[0100] Finally, we apply a local Gaussian mask to each channel of the feature map.

[0101]

[0102] in This represents the feature map after masking. This represents a local Gaussian mask.

[0103] On the other hand, the GRGM block enhances the model's robustness and generalization ability by introducing globally random Gaussian occlusion. This block forces the model to learn and rely on other discernible features by randomly occluding certain regions of the image. Since the occlusion is randomly generated, the model needs to learn features outside the occluded areas to predict targets, thereby enhancing the model's ability to recognize different shapes and locations of contraband. Through the GRGM block, the model can better adapt to various disturbances and variations, improving performance and generalization ability in complex environments.

[0104] By applying the LFGM block, the model is forced to learn non-central regions. However, the model may still focus on localized areas. Therefore, we further designed the GRGM block, which randomly applies multiple Gaussian masks to the feature map to prevent the model from over-focusing on locally discriminative regions and to globally learn the entire object (including both central and non-central regions). This approach is beneficial for leveraging object-oriented knowledge of intra-object modalities. Unlike the LFGM block, the GRGM block randomly generates points in each feature map. Specifically, for each feature map F... n,l Randomly generate a set of points in Representing feature map F n,l Let M be the m-th randomly generated point (x-coordinate and y-coordinate), and M represent the total number of generated points. Therefore, for each point... Generate a Gaussian mask

[0105] Finally, this scheme applies these Gaussian masks to the feature map F. n,l And subtract them from the original feature map.

[0106]

[0107] Where F″ n,l This represents the feature map after masking.

[0108] In obtaining F′ n,l and F″ n,l Next, we concatenate these two feature maps and use a 1×1 convolutional layer to reduce their channel count, obtaining feature E. n,l This feature is then used as the input feature for the MIL branch.

[0109] Step D: Input the X-ray contraband dataset from Step B into the domain object learning network to obtain the enhanced feature map, and then feed it into the multi-instance branch learning network to predict the instance score of each instance generated around the labeled point.

[0110] The X-ray contraband dataset from step B is input into the domain object-oriented learning network, and after processing, enhanced feature maps are obtained. These feature maps are then fed into a multi-instance branch learning network, whose task is to predict the instance score for each instance generated around the labeled point.

[0111] E. Design an inter-domain object-oriented learning network. This network includes a wavelet decoupling adversarial learning network and an object-oriented module.

[0112] An inter-domain objectivity learning network was designed to effectively transfer objectivity information extracted from natural images to X-ray images. This network consists of two key components: a Wavelet Decomposition Based Adversarial Learning (WDAL) network and an Objectness Block (OB). First, the WDAL network uses wavelet decomposition to decouple the feature maps of natural and X-ray images, dividing them into style and content components. Then, adversarial learning reduces modal differences caused by style differences, allowing the model to better adapt to the characteristics of X-ray images. Second, the Objectness Block utilizes the transferred objectivity knowledge, integrating it into a multi-instance learning branch. By calculating the objectivity score for each region of interest (RoI) feature, this module enhances the probability that the entire contraband item is contained within the proposal box. These two components work together to enable the model to accurately identify and locate contraband items in X-ray images, improving the efficiency and accuracy of security detection.

[0113] F. Decouple the feature maps of the X-ray contraband dataset and the natural image dataset from step B using a wavelet decoupling adversarial learning network, and perform adversarial learning on the low-frequency style features.

[0114] In wavelet decomposition adversarial learning networks, directly transferring object-related knowledge extracted from natural images to X-ray images is not ideal due to the significant modal differences between natural and X-ray images. To address this issue, our proposed WDAL module first uses wavelet decomposition to decouple the feature maps of natural and X-ray images into style and content components. Then, adversarial learning is used to reduce the modal differences caused by style variations. Our WDAL module is inspired by the effectiveness of wavelet decomposition in preserving content information.

[0115] Technically, this scheme uses Haar wavelets for wavelet decomposition. Haar wavelets consist of four kernels, namely {LL... T LH T HL T HH T}, where L and H are the low-pass and high-pass filters, respectively, defined as follows:

[0116]

[0117] Wavelet decomposition transforms the feature map M n,l (from X-ray image I) n Or natural image J n (Extracted) and decomposed into four parts, namely

[0118] A n,l =M n,l *(LL T )

[0119] H n,l =M n,l *(LH T )

[0120] V n,l =M n,l *(HL T )

[0121] D n,l =M n,l *(HH T )

[0122] Where * denotes the convolution operation; A n,l Indicates the low-frequency range; H n,l V n,l and D n,l This indicates the high-frequency component.

[0123] Typically, low-frequency components primarily correspond to stylistic features, capturing detailed stylistic information of the image; high-frequency components correspond to content features, focusing on local details and edges. Then, A... n,l The input is fed into a convolutional layer with learnable parameters to obtain the transformed style features A′. n,l To reduce modal variance, we employ adversarial learning on the transformed style features. We treat the natural image dataset and the X-ray image dataset as the source and target domains, respectively. We train a domain classifier on the activation values ​​of each transformed style feature to predict the modality of each style feature. This approach helps to increase the number of training samples for adversarial learning and reduce global image variance. Mathematically, let... It is the domain label of the nth image, where if the nth image belongs to the source domain, then... Otherwise, it is 1. This scheme represents the activation value of the transformed style feature at coordinates (x, y) as... Utilizing corresponding style features Predicted value at each position Calculate the adversarial loss L based on cross-entropy adv ,Right now

[0124]

[0125] To align the modality distributions, we simultaneously optimize the parameters of the modality classifier by minimizing the adversarial loss and optimize the parameters of the base network by maximizing this loss. We also use a gradient inversion layer (GRL) to train the modality classifier.

[0126] Then, the content feature Hn,l V n,l and D n,l and the transformed style features A′ n,l By performing a connection, we obtain the mode-free characteristic H. n,l ,Right now

[0127]

[0128] in This indicates a splicing operation.

[0129] G. Using the objectivity module, predict the objectivity score of each target in the X-ray security inspection image dataset and the natural image dataset.

[0130] We utilize an objectivity module to predict the objectivity score of each target in both the X-ray security image dataset and the natural image dataset. After transferring objectivity information extracted from natural images to X-ray images using the WDAL module, we designed an objectivity module that incorporates this information into the MIL branch, enhancing the probability that the proposal box contains the entire contraband.

[0131] This scheme calculates an objectivity score, reflecting the probability that each region of interest (RoI) feature contains the entire item. Let... Let R represent the RoI features generated from the k-th annotation point of the n-th image (X-ray or natural image). Here, R represents the total number of proposals. We then perform RoI pooling on each feature map and compute the final objectivity score through a fully connected layer. The objectivity module can be represented as:

[0132]

[0133] RoI_Pooling() and FC() represent RoI pooling and fully connected layers, respectively. This represents the objectivity score corresponding to the r-th proposal box generated from the k-th point.

[0134] For natural images, we can leverage rich bounding box annotations to train an object-oriented module. Specifically, for proposed bounding boxes generated from point annotations, we calculate their intersection-union ratio (IU) with the annotations of the standard bounding boxes. We consider proposals with an intersection-union ratio (IU) greater than or equal to 0.5 as positive samples and label them as 1, while proposals with an IU less than this value are considered negative samples and labeled as 0. Based on this, we can use the cross-entropy loss function to adjust the model parameters, i.e.

[0135]

[0136] here, and These represent the areas marked by the r-th proposal box and its corresponding annotation box, respectively.

[0137] Finally, I 2 OL-Net joint loss is

[0138] L = L mil +λ1L adv +λ2L obj

[0139] Where L mil λ represents the total MIL loss defined in P2BNet; λ1 and λ2 are the balancing weights.

[0140] H. The object score from the X-ray contraband dataset and the instance score from D are weighted to obtain the final instance score for the multi-instance branch learning network. For the natural image dataset, the object score is supervised by instance-level annotations.

[0141] For X-ray contraband datasets, this scheme weights objectivity and instance scores when processing X-ray images to obtain the final instance score for the multi-instance branch learning network. This weighting operation aims to ensure that the network can accurately locate and identify contraband items. Technically, this involves... Multiply by the instance score of the annotation at the k-th point in the MIL branch, i.e., the k-th bag of MIL; for proposal boxes with higher objectivity scores, larger weights are assigned; the classification score and subsequent MIL loss calculation remain unchanged. This process can be represented as follows:

[0142]

[0143] in, Represents the weighted instance score; Soffmax() and FC ins () represent the Softmax function and the instance sub-branch, respectively.

[0144] For natural image datasets, we employ instance-level annotation for object-oriented scoring supervision. This means we provide detailed annotations for each instance in the natural images, enabling the network to accurately learn and understand target information within them. In this way, we can effectively utilize the rich object-oriented information in the instance-level annotations of natural image datasets and apply it to the X-ray contraband detection task, thereby improving the network's performance on X-ray images.

[0145] I. Select the instances with the highest scores from the X-ray contraband dataset and use them as pseudo-labels to train a fully supervised contraband detection model.

[0146] This approach first selects the instances with the highest scores from the X-ray contraband dataset. These instances are considered positive samples with high confidence and are used as pseudo-labels.

[0147] Next, we utilize these pseudo-labels to train a fully supervised contraband detection model. This approach allows us to fully leverage information from both the X-ray contraband dataset and existing natural images, further enhancing the model's performance in detecting contraband. This training strategy not only strengthens the model's ability to identify contraband but also improves its generalization ability across various contraband categories, making the model more reliable and robust in practical applications.

[0148] To facilitate performance comparison testing of our proposed method, we evaluated it on four commonly used X-ray datasets: OPIXray, HIXray, SIXray, and PIDray. The OPIXray dataset contains 8,885 images across 5 categories. The SIXray dataset contains 1,059,231 images across 6 categories. We used 8,929 images containing contraband for training and testing. The HIXray dataset contains 45,364 images across 8 categories. The PIDray dataset contains 124,486 images across 12 categories; our proposed method used 47,677 images containing contraband for training and testing.

[0149] During training, this approach also uses the COCO14 training set as an additional dataset. To address the image imbalance between the natural image dataset and the X-ray dataset, this approach randomly selects the same number of images from the COCO14 training set as the X-ray dataset, and merges the randomly selected subset with the X-ray training set to construct the final training set.

[0150] This approach is based on P2BNet and implemented using MMDetection. The model is trained using two NVIDIA RTX 3090 GPUs with an SGD algorithm at a learning rate of 0.005. The total number of training epochs is set to 12, and the batch size is set to 4. Data for each batch is randomly selected from the X-ray training set and the additional natural image dataset to fully utilize the parallelism of the GPU. In the loss function, the balancing weights λ_1 and λ_2 are set to 0.1 and 1, respectively. The scaling factors α and β are set to 0.1 and 0.15, respectively. The number of random points M in GRGM is set to 2. All other settings are identical to those in P2BNet except for the parameters mentioned above.

[0151] Correspondingly, the control network model is trained using a similar training method or parameters as this scheme. The performance comparison of this scheme with other weakly supervised object detectors, point-supervised object detectors, and fully supervised object detectors on the OPIXray+HIXray+SIXray+PIDray dataset is shown in Table 1 below:

[0152] Table 1

[0153]

[0154] Note: Faster R-CNN corresponds to the method proposed by Ren, S. et al. (Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems 28 (2015));

[0155] RetinaNet corresponds to the method proposed by Lin, TY, et al. (Lin, TY, Goyal, P., Girshick, R., He, K., Dollár, P.: Focal loss for dense object detection. In: Proceedings of the IEEE / CVF International Conference on Computer Vision.pp.2980–2988(2017));

[0156] YOLOv8-s corresponds to the method proposed by Jocher, G. et al. (Jocher, G., Chaurasia, A., Qiu, J.: Ultralytics YOLO (Jan 2023), https: / / github.com / ultralytics / ultralytics);

[0157] FCOS corresponds to the method proposed by Tian, ​​Z. et al. (Tian, ​​Z., Shen, C., Chen, H., He, T.: FCOS: Fully convolutional one-stage object detection. In: Proceedings of the IEEE / CVF International Conference on Computer Vision. pp. 9627–9636 (2019));

[0158] Deformable DETR corresponds to the method proposed by Zhu, X. et al. (Zhu, X., Su, W., Lu, L., Li, B., Wang, X., Dai, J.: Deformable DETR: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159(2020));

[0159] PCL corresponds to the method proposed by Tang, P. et al. (Tang, P., Wang, X., Bai, S., Shen, W., Bai, X., Liu, W., Yuille, A.: PCL: Proposal cluster learning for weakly supervised object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 42(1), 176–191(2018));

[0160] WSOD2 corresponds to the method proposed by Zeng, Z. et al. (Zeng, Z., Liu, B., Fu, J., Chao, H., Zhang, L.: WSOD2: Learning bottom-up and top-down objectness distillation for weakly-supervised object detection. In: Proceedings of the IEEE / CVF International Conference on Computer Vision. pp. 8292–8300 (2019));

[0161] OD-WSCL corresponds to the method proposed by Seo, J. et al. (Seo, J., Bae, W., Sutherland, DJ, Noh, J., Kim, D.: Object discovery via contrastive learning for weakly supervised object detection. In: Proceedings of the European Conference on Computer Vision. pp. 312–329 (2022));

[0162] P2BNet corresponds to the method proposed by Chen, P. et al. (Chen, P., Yu, X., Han, X., Hassan, N., Wang, K., Li, J., Zhao, J., Shi, H., Han, Z., Ye, Q.: Point-to-box network for accurate object detection via single point supervision. In: Proceedings of the European Conference on Computer Vision. pp. 51–67 (2022)).

[0163] The comparison results above demonstrate that the proposed method and its corresponding fully supervised contraband detection model exhibit significant improvements and advantages compared to existing methods. This approach effectively mines intra-domain object information and can transfer the rich object information contained in instance-level annotations of natural image datasets to X-ray contraband datasets, greatly alleviating the local dominance problem caused by severe intra-domain variations in X-ray images. This invention achieves excellent performance on multiple X-ray contraband datasets, showing a substantial performance improvement compared to traditional weakly supervised and point-supervised object detection algorithms.

[0164] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A point-supervised X-ray contraband detection method with intra-domain and inter-domain object-oriented learning characteristics, It includes: A. Prepare an X-ray contraband dataset and a natural image dataset, where the former contains only labeled points and the latter contains instance-level annotations; B. Input the X-ray contraband dataset and the natural image dataset into the feature extractor to obtain the feature maps corresponding to the two datasets respectively; C. Design an in-domain object learning network that includes local focal Gaussian occlusion and global random Gaussian occlusion to mine in-domain object information; D. Input the feature map of the X-ray contraband dataset obtained in step B into the domain object learning network to obtain the enhanced feature map, and then feed it into the multi-instance branch learning network to predict the instance score of each instance generated around its labeled point. E. Design an inter-domain objectivity learning network, which includes a wavelet decoupling adversarial learning network and an objectivity module; F. Decouple the feature maps of the X-ray contraband dataset and the natural image dataset obtained in step B using a wavelet decoupling adversarial learning network, and perform adversarial learning on the low-frequency style features. G. Utilize the objectivity module to predict the objectivity score of each target within the feature map in the X-ray contraband dataset and the natural image dataset; H. Weight the object score of the feature map of the X-ray contraband dataset and the instance score obtained in step D to obtain the final instance score of the multi-instance branch learning network. For natural image datasets, objectivity scores are supervised using instance-level annotations; I. Select instances from the X-ray contraband dataset whose scores are higher than a preset value, use them as pseudo-labels, and then train a fully supervised contraband detection model; Step C includes: Design an in-domain object learning network that includes Locally Focused Gaussian Occlusion (LFGM) and Globally Randomized Gaussian Occlusion (GRGM) to mine in-domain object information; Among them, Local Focus Gaussian Occlusion (LFGML) guides the model to focus on the non-central region of the prohibited item by generating a local Gaussian mask; it explicitly forces the model to learn the non-central region of the prohibited item by finding the point with the local maximum activation value and generating a Gaussian mask around that point; its principle includes: Given an X-ray image and its corresponding dot annotations First, take the picture Input the backbone network and FPN to obtain feature maps at different scales. ,in , Indicates the number of feature maps. , and Representing feature maps respectively Height, width, and number of channels; for feature maps Each channel, first in size of Find the point with the highest activation within the window. This represents the scaling factor, and the point serves as the average location for generating the local Gaussian mask; the above process can be expressed as: in, Indicates the first Channel feature map Average position on; At the same time, for each A covariance matrix is ​​defined. ,in It is a two-dimensional diagonal matrix. Indicates the scaling factor; Finally, a local Gaussian mask is applied to each channel of the feature map. in, This represents the feature map after masking. Represents a local Gaussian mask; Globally Stochastic Gaussian Occlusion (GRGM) enhances the robustness and generalization ability of a model by introducing globally stochastic Gaussian occlusion. It randomly applies multiple Gaussian occlusions to the feature map to prevent the model from over-focusing on locally discriminative regions and to learn the entire item globally, including both central and non-central regions. Its principles include: For each feature map Randomly generate a set of points ,in Representation of feature map Upper Randomly generated points ( coordinates and coordinate), This represents the total number of points generated; based on this, for each point... Generate a Gaussian mask ; Applying these Gaussian masks to the feature map And subtract them from the original feature map. in This represents the feature map after masking. In obtaining and Then, the two feature maps are concatenated and a single... Convolutional layers reduce their channel count to obtain features. This feature is used as the input feature for the MIL branch; Step E includes: Design an inter-domain objectivity learning network to effectively transfer objectivity information extracted from natural images to X-ray images; the network consists of two key components, including a wavelet decoupling adversarial learning network and an objectivity module; Among them, the wavelet decoupled adversarial learning network uses wavelet decomposition technology to decouple the feature maps of the X-ray contraband dataset and the natural image dataset, dividing them into style and content parts; then, through adversarial learning, it reduces the modal differences caused by style differences, thereby enabling the model to better adapt to the characteristics of X-ray images; The objectivity module leverages the transferred objectivity knowledge and integrates it into the multi-instance learning branch. By calculating the objectivity score of each region of interest (RoI) feature, this module can enhance the probability that the entire prohibited item is included in the proposal box.

2. The intra-domain and inter-domain object-oriented learning method for point-supervised X-ray contraband detection as described in claim 1, characterized in that, In step A, the X-ray contraband dataset includes the OPIXray dataset, HIXray dataset, SIXray dataset, and PIDray dataset; the OPIXray dataset contains 5 categories and a total of 8,885 images; the SIXray dataset contains 6 categories and a total of 1,059,231 images; the HIXray dataset contains 8 categories and a total of 45,364 images; and the PIDray dataset contains 12 categories and a total of 124,486 images. Among them, the OPIXray, HIXray, SIXray, and PIDray datasets all include instance-level annotations. Points are randomly generated near the center point of the instance-level annotations according to a Gaussian distribution and used as annotation points during training. The natural image dataset is the commonly used COCO14 dataset; the same number of images as the X-ray dataset are randomly selected from the COCO14 training set to alleviate the imbalance in the number of images between the X-ray contraband dataset and the natural image dataset.

3. The intra-domain and inter-domain object-oriented learning method for point-supervised X-ray contraband detection as described in claim 1 or 2, characterized in that, In step B, the feature extractor is constructed using a combination of ResNet50 and FPN. The combination of ResNet50 and FPN forms a multi-layered feature extraction network. ResNet50 is responsible for extracting more detailed and local features at the lower level, while FPN performs feature fusion and introduces contextual information at the higher level. By combining ResNet50 and FPN to construct a feature extractor, features related to contraband items are extracted from the X-ray contraband dataset. At the same time, features of general objects are extracted from the natural image dataset. Feature maps corresponding to the two datasets are obtained and labeled as natural images and X-ray images, respectively, to obtain feature representations for different datasets.

4. The intra-domain and inter-domain object-oriented learning method for point-supervised X-ray contraband detection as described in claim 3, characterized in that, Step F includes: In the wavelet decomposition adversarial learning network, the feature maps of the X-ray contraband dataset and the natural image dataset are decoupled into style and content parts by wavelet decomposition through the WDAL module. Then, adversarial learning is used to reduce the modal differences caused by style differences. The WDAL module is a wavelet decoupling adversarial learning network module. Wavelet decomposition was performed using Haar wavelets, which consist of four kernels: ,in and These are low-pass and high-pass filters, defined as follows: Wavelet decomposition transforms the feature map From X-ray images Or natural pictures Extracting from it, we break it down into four parts, namely... in, Indicates the convolution operation; Indicates the low-frequency component; , and Indicates the high-frequency portion; The low-frequency component corresponds to style features, capturing detailed stylistic information of the image; the high-frequency component corresponds to content features, focusing on local details and edges; then, The input is fed into a convolutional layer with learnable parameters to obtain the transformed style features. To reduce modal discrepancies, adversarial learning is applied to the transformed style features, treating the natural image dataset and the X-ray contraband dataset as the source and target domains, respectively. A domain classifier is trained on the activation values ​​of each transformed style feature to predict the modality of each style feature. Mathematically, let... It is the first The domain label of the image, where if the first image... If an image belongs to the source domain, then Otherwise, it is 1; the transformed style features are plotted on coordinates. The activation value at the location is represented as Utilizing corresponding style features Predicted value at each position Calculate adversarial loss based on cross-entropy ,Right now To align modal distributions and optimize the parameters of the modality classifier by minimizing the adversarial loss, and to optimize the parameters of the base network by maximizing this loss, a gradient inversion layer (GRL) is also used to train the modality classifier. Then, content features , and and the transformed style characteristics By performing a connection, non-modality-specific features are obtained. ,Right now in This indicates a splicing operation.

5. The intra-domain and inter-domain object-oriented learning method for point-supervised X-ray contraband detection as described in claim 4, characterized in that, Step G includes: The objectivity module is used to predict the objectivity score of each target in the feature map of the X-ray contraband dataset and the natural image dataset. The objectivity information extracted from the natural image is transferred to the X-ray image using the WDAL module. This information is then incorporated into the MIL branch of the objectivity module, which enhances the probability that the proposal box contains the entire contraband. The object score reflects the probability that each region of interest (RoI) contains the entire item; it includes: assuming... Indicates by the first The first X-ray image of prohibited items or a natural image. The proposed RoI features generated from annotation points, where... Indicates the total number of proposals; RoI pooling is performed on each feature map, and the final objectivity score is calculated through a fully connected layer; the objectivity module can be represented as: in, and These represent RoI pooling and fully connected layers, respectively. Indicates the first The nth point generated The objectivity score corresponding to each suggestion box; For natural images, rich bounding box annotations can be used to train the object-oriented module. Specifically, for proposed bounding boxes generated from point annotations, the intersection-union ratio (IU) between them and the labeled bounding boxes is calculated. In this model, proposals with an intersection-union ratio (IU) greater than or equal to 0.5 are considered positive samples and labeled as 1, while proposals with an IU less than this value are considered negative samples and labeled as 0. Based on this, the cross-entropy loss function can be used to adjust the model parameters. in, and They represent the first The area of ​​each suggestion box and its corresponding label box; Finally, the joint loss of the model is in, This represents the total MIL loss defined in P2BNet; and It is a balancing weight.

6. The intra-domain and inter-domain object-oriented learning method for point-supervised X-ray contraband detection as described in claim 5, characterized in that, In step H, the object score of the feature map of the X-ray contraband dataset and the instance score obtained in step D are weighted to obtain the final instance score of the multi-instance branch learning network, which includes: For X-ray images, Multiply by the first branch in MIL The instance score of the point annotation, i.e., the score of the MIL. Each box contains a subset of boxes; for proposal boxes with higher objectivity scores, larger weights are assigned; the classification score and subsequent MIL loss calculation remain unchanged, and this process can be expressed as: in, Indicates the weighted instance score; and These represent the Softmax function and the instance sub-branch, respectively.