X-ray security check image contraband detection method based on generative sparsity
By optimizing the dynamic instance interaction module and label allocation strategy of DiffusionDET, combined with the IoU perceived classification loss function, the detection accuracy and efficiency of contraband in X-ray security images are improved, and the problem of insufficient detection accuracy in the prior art is solved.
Patent Information
- Application Number
- CN202510772195.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The existing sparse object detection algorithm DiffusionDET is generally performed in the detection of contraband in X-ray security images, especially when dealing with overlap, background similarity and hidden problems.
By introducing dynamic filtering fusion module DFF, Hungarian Top-K matching strategy HMK and IoU-aware classification loss function ACR-Loss, the dynamic instance interaction module and label allocation strategy of DiffusionDET are optimized to improve the accuracy of feature extraction and target positioning.
The detection accuracy and efficiency of contraband in X-ray security images are significantly improved, the robustness and recall of the model are enhanced, and the impact of hyperparameters on the detection results are reduced.
Smart Images

Figure CN120339264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and object detection. Specifically, the present invention proposes a method for detecting contraband in X-ray security inspection images based on generative sparsity Background Art
[0002] With the increase in passenger flow, X-ray security inspection has become a key technology to ensure people's safety in public places. X-rays have penetrability, which can help security personnel detect potential contraband in luggage and backpacks. However, affected by penetrability, X-ray security inspection images exhibit characteristics different from natural light images, including: (1) Overlap: When items stacked together in a luggage or backpack are irradiated by X-rays, an overlap phenomenon will occur in the image, resulting in the surface texture of the items in the image becoming blurred; (2) Background similarity: Since the materials of some contraband items in X-ray security inspection images are extremely similar to those of background items, the contraband in X-ray imaging is very similar to the background, and it is very easy to be treated as the background during the detection process; (3) Hiddenness: Some contraband items in X-ray security inspection images are small in size and overlap with other items, making them difficult to be detected by the system
[0003] With the development of deep learning object detection technology, the detection of contraband in X-ray security inspection images has achieved automated detection. Common detection frameworks are divided into single-stage object detection, two-stage object detection, and sparsity object detection. Among the sparsity object detection algorithms, the object detection algorithm DiffusionDET based on the diffusion model performs well in terms of scalability, real-time performance, and accuracy. However, during the process of using DiffusionDET for X-ray contraband detection, its performance is average Summary of the Invention
[0004] To address the above problems, the present invention proposes an X-ray security inspection image contraband detection method based on generative sparsity, which optimizes and improves DiffusionDET. By optimizing and improving the dynamic instance interaction module, label assignment algorithm, and loss function of DiffusionDET, the detection accuracy of contraband in X-ray security inspection images is improved. Specifically, first, a dynamic filtering and fusion module DFF is designed to iteratively filter the pre-scheme feature map in the vertical and horizontal directions to extract the spatial coordinate information in the feature map. During the training process, the RoI features generated in the previous iteration are extracted and fused with the pre-scheme feature map to generate a refined feature map for object localization and classification. Then, a label assignment strategy HMK between one-to-one and one-to-many is designed. Through global matching, K equally important optimal positive samples are found for each ground truth bounding box GT. Finally, a new IoU-aware classification loss function ACR-Loss that emphasizes the importance of spatial information is designed. By adding the ratio of IoU to the classification loss in the classification loss function, it helps the model to pay more attention to the samples where classification and localization are inconsistent, thus improving the detection performance. The method of the present invention adopts a dynamic filtering and fusion module, an advanced label assignment strategy, and an optimized training loss function design. The method aims to improve the accuracy and efficiency of contraband detection in the security inspection field.
[0005] The X-ray security inspection image contraband detection method improved based on generative sparsity includes the following steps:
[0006] Step (1): First, obtain an X-ray security inspection image dataset containing object bounding boxes and class annotations, divide the dataset into a training set and a validation set, and preprocess the images. Successively use the ResNet backbone and the feature pyramid structure to extract multi-level features in the security inspection images.
[0007] Step (2): Use the diffusion model to generate pre-scheme boxes with noise . And use to generate region of interest features and context-aware feature representations .
[0008] Step (3): Replace the dynamic instance interaction part of the diffusion model DiffusioDET with the dynamic filtering and fusion module DFF. DFF further refines the features during the iterative filtering process, and through dynamic parameter screening operations, selects appropriate and . Then, the features are concatenated to generate a refined feature map , and obtain the target prediction box and the corresponding class scores.
[0009] In step (4) during the label assignment phase, the Hungarian Top-K matching strategy (HMK strategy) is used to replace the original assignment strategy. Through global matching, K equally important optimal positive samples are found for each ground truth (GT). This not only improves the model recall rate but also reduces the impact of hyperparameters on the final detection results, thereby enhancing the robustness of the model.
[0010] In step (5), set the training parameters and replace the DiffusionDET loss function with a new loss function, ACR-Loss. Input the multi-level features generated in step (1) into the model constructed in steps (2)-(4) for iterative training. Then validate on the validation set to obtain the optimal parameter model and output the prohibited item detection effect diagram.
[0011] Furthermore, step (2) is specifically as follows:
[0012] (2-1) Use the diffusion model to generate anchor box samples.
[0013] (2-1-1) During the forward noise addition phase, the model adds Gaussian noise to the ground truth (GT) to generate a noisy preliminary box, denoted as . Specifically, set , to represent the center point coordinates of the preliminary box and the width and height of the preliminary box respectively, and calculate the corresponding , where .
[0014] (2-1-2) During the inference phase, first create an inverse time pair , where represents the final time. Then for each time pair , the diffusion model uses the at the current time point to predict the at the time point. Then use the denoising diffusion implicit model (DDIM) to perform denoising sampling on to calculate the corresponding bounding box, i.e., the anchor box:
[0015] (2-2) Apply the region of interest alignment (RoIAlign) operation on the multi-level feature map to extract the region of interest features corresponding to the anchor box from the feature map .
[0016] (2-3) Adopt the self-attention mechanism to perform cross-attention mechanism feature aggregation on all features to achieve information interaction between features, thereby obtaining a context-aware feature representation .
[0017] Furthermore, step (3) is specifically as follows:
[0018] (3-1) First, a fully connected network is adopted to map the feature spf into the parameter space, thereby obtaining the dynamic filtering parameter dfp. Then, through the kernel splitting technique, dfp is split into two independent parameters: dfv responsible for the vertical direction and dfh responsible for the horizontal direction.
[0019] (3-2) Use the CA attention mechanism to extract features from roif, generating the feature map ROICAf. Then, use the vertical filtering parameter dfv and the horizontal filtering parameter dfh to filter the region of interest in features using the spatial attention mechanism CA, obtaining the filtered features xfv and xfh.
[0020] (3-3) Concatenate the filtered features xfh and xfv, and process them through a non-linear activation module. Finally, map the result back to the feature space to obtain the refined feature objf after filtering and fusion.
[0021] (3-4) Use MLP to perform corresponding scaling and translation operations, and then calculate the prediction score and prediction box.
[0022] Furthermore, step (4) is specifically as follows:
[0023] (4-1) HMK calculates the cost loss based on the prediction score, prediction box, GT ground truth box, and GT label, including the cost classification loss , the cost regression loss , and the cost bounding box loss . During the calculation, it will be judged whether the center point of the prediction box is inside the GT ground truth box and whether it is inside the circle with the center of the GT ground truth box as the center and a radius of 2.5, that is, check whether the center point of the prediction box is within 5 unit squares around the center of the GT ground truth box. Penalties are imposed on those prediction boxes whose center points are not inside the GT ground truth box or not within its center radius range.
[0024] (4-2) Calculate the IoU value between the prediction box and GT. Through the summation of these IoU values, HMK determines the number of positive samples dk required for each GT. Then, according to dk, the column corresponding to each GT in the total loss is expanded dk times to construct the final loss matrix.
[0025] (4-3) Use the Hungarian algorithm to find the dk positive sample boxes that each GT needs to match according to the loss matrix. This process ensures that each GT can be effectively matched with a certain number of prediction boxes, thereby improving the accuracy and recall rate of object detection.
[0026] Compared with the prior art, the beneficial effects of the present invention are:
[0027] The present invention improves the performance of the generative sparsity object detection network in X-ray security inspection images by making improvements. Specifically, first, the present invention introduces a new dynamic filtering fusion module DFF in the iterative filtering process, which filters out useless background features and impurities in the horizontal and vertical directions and retains the refined features for detection. Then, a Hungarian Top-K matching strategy HMK is adopted to perform one-to-K matching in the global label matching process, which improves the detection recall rate and reduces the influence of hyperparameters on the detection results. Finally, a new IoU-aware classification loss function ACR-Loss that emphasizes the importance of spatial information is used. By introducing spatial information consistency, the dependence of the model on the prediction confidence is reduced. Compared with the existing sparsity detection algorithms, the present invention performs well both in terms of the accuracy of detecting X-ray contraband and in terms of timeliness. Brief Description of the Drawings
[0028] Figure 1 is the processing flow chart of the present invention;
[0029] Figure 2 is the overall structure diagram of the model of the present invention;
[0030] Figure 3 is the schematic diagram of HMK label assignment. Detailed Description of the Invention
[0031] To better understand the purpose, structure and function of the present invention, the technical solution of the present invention will be further described in detail below with reference to the drawings. The present invention proposes a method for detecting contraband in X-ray security inspection images based on generative sparsity, and the overall process is as Figure 1 shown. The overall structure of the model is as Figure 2 shown. First, the present invention designs a dynamic filtering fusion module DFF. By iteratively filtering the pre-scheme feature map in the vertical and horizontal directions, the spatial coordinate information in the feature map is extracted. And iteratively using multiple fused features, a refined feature map for object localization and classification is generated. Then, a label assignment strategy HMK that is between one-to-one and one-to-many is adopted, Figure 3 shown. This strategy finds K equally important optimal positive samples for each GT through global matching. This method not only improves the recall rate but also reduces the influence of hyperparameters on the final detection results, thus enhancing the robustness of the model. Finally, a new IoU-aware classification loss function ACR-Loss that emphasizes the importance of spatial information is used. ACR-Loss helps the model to pay more attention to the samples with inconsistent classification and localization by adding the ratio of IoU to the classification loss in the classification loss function, thereby improving the detection performance. The present invention has a great improvement in the accuracy of detecting contraband compared with the previous sparsity object detection algorithms.
[0032] The method of the present invention includes the following steps:
[0033] Step (1) First, obtain an X-ray security inspection image dataset containing target bounding boxes and class annotations, divide the dataset into a training set and a validation set, and preprocess the images; use the ResNet backbone to extract multi-level features from the images. Specifically:
[0034] (1-1) Divide the image set into two main parts: one part is used for training the model, that is ; the other part is used to verify the accuracy of the model, that is . Among them, R is the real number field, represents the number of image samples in the training set, represents the i-th training image sample, represents the number of image samples in the validation set, represents the j-th validation image sample, H represents the image height, W represents the image width, and 3 represents the number of RGB channels;
[0035] (1-2) The label corresponding to each training sample ; where represents the number of targets contained in the image sample represents the image sample contains, represents in the th target's true class, where C represents the total number of classes in the dataset, represents in the th target's bounding box, which consists of the center point abscissa x, the center point ordinate y, the target width w, and the height h.
[0036] (1-3) Before formal training, perform a series of data augmentation operations on the images. This includes randomly selecting four images from the training set, performing a series of transformation operations on them, such as flipping, scaling, and adjusting the hue, and then stitching these images into a new sample while retaining their original label information. In addition, the images are also resized, uniformly scaled to 640×640 pixels, and pixel value normalization is performed to ensure the consistency and effectiveness of the model input.
[0037] (1-4) Use ResNet to extract features from the image , and then use the feature pyramid structure to generate multi-level feature maps from P2 to P5, that is .
[0038] Step (2) Use the diffusion model to generate noisy pre-determined boxes . And use Generate region of interest features and context-aware feature representations . Specifically:
[0039] (2-1) Use a diffusion model to generate anchor box samples.
[0040] (2-1-1) In the forward noise addition stage, the model adds Gaussian noise to the ground truth GT to generate a noisy proposal box, denoted as . Specifically, set , and calculate the corresponding using the following formula, where .
[0041] (1)
[0042] where is a hyperparameter, set the variance scheduling variable and , then and .
[0043] where , represents the merging of two Gaussian functions with different variances. That is, for two Gaussian functions, and , the new distribution after addition and merging is , and its variance is expressed as: .
[0044] (2-1-2) In the inference stage, first create the reverse time pair , where represents the final time. Then for each time pair , the diffusion model uses the at the current time point to predict the at the time point. Then use the DDIM algorithm to perform noise reduction sampling on and calculate the corresponding anchor box according to the following formula:
[0045] (2)
[0046] (2-2) Apply the region of interest alignment RoIAlign operation on the image feature map to extract the region of interest features corresponding to the anchor box from the feature map.
[0047] (2-3) Adopt a self-attention mechanism for all features Perform feature aggregation to achieve information interaction between features, thereby obtaining a context-aware feature representation .
[0048] In step (3), the dynamic instance interaction part of the diffusion model DiffusioDET is replaced with the dynamic filtering and fusion module DFF. DFF further refines the features during the iterative filtering process. Through dynamic parameter screening operations, appropriate and are selected. After that, the features are concatenated to generate a refined feature map . Specifically:
[0049] (3-1) First, use a fully connected network to map the feature representation into the parameter space to obtain the dynamic filtering parameter dfp. Then, through kernel segmentation technology, dfp is segmented into two independent parameters: dfv responsible for the vertical direction and dfh responsible for the horizontal direction.
[0050] (3-2) Use the CA attention mechanism to extract features to generate a feature map. Then, use the vertical filtering parameter dfv and the horizontal filtering parameter dfh to filter the region of interest of the features using the spatial attention mechanism CA to obtain the filtered features xfv and xfh.
[0051] (3-3) Concatenate the filtered features xfh and xfv, and process them through a non-linear activation module. Finally, map the result back to the feature space to obtain the refined feature after filtering and fusion .
[0052] (3-4) Use MLP to perform corresponding scaling and translation operations, and then calculate the prediction score and prediction box.
[0053] In step (4), during the label assignment stage, use the HMK strategy to replace the original assignment strategy. Through global matching, K equally important optimal positive samples are found for each GT. While improving the model recall rate, it also reduces the impact of hyperparameters on the final detection result, thereby enhancing the robustness of the model. Specifically:
[0054] (4-1) HMK calculates the cost loss based on the prediction score predLogits, the prediction box predBoxes, the GT ground truth box gtBoxes, and the GT label gtLabels, including the cost classification loss , the cost regression loss , and the cost bounding box loss During the calculation process, it is determined whether the center point of the predicted box predBoxes is located within the GT ground truth box gtBoxes and whether it is within a circle with the center of the GT ground truth box as the center and a radius of 2.5, that is, it is checked whether the center point of the predicted box is within 5 unit squares around the center of the GT ground truth box. Penalties are imposed on those predicted boxes predBoxes whose center points are not within the GT ground truth box or not within its center radius range.
[0055] (4-2) Cost Classification Loss Adopt Loss function calculation:
[0056] (3)
[0057] Among them, is the weight factor used to adjust the weight balance of easy and difficult samples, and more difficult samples are given higher weights. represents the probability that the current predicted target is a positive sample. represents the probability of predicting a negative sample. is a hyperparameter used to adjust the degree of attention of the model to easy and difficult samples. When increases, the model pays attention to difficult-to-classify samples. Cost regression loss Adopt Loss function calculation. Cost bounding box loss Adopt Loss function calculation:
[0058] (4)
[0059] Among them, represents calculating the IoU between the anchor box Ar and the GT. represents the area of the minimum enclosed region of Ar and the GT. Summing up all these losses gives the total loss.
[0060] (4-3) Calculate the IoU value between the predicted box predBoxes and the GT. By summing up these IoU values, HMK determines the number of positive samples dk required for each GT. Then, according to dk, each column corresponding to each GT in the total loss is expanded dk times to construct the final loss matrix.
[0061] (4-4) Use the Hungarian algorithm to find the dk positive sample boxes that each GT needs to match according to the loss matrix. This process ensures that each GT can be effectively matched with a certain number of predicted boxes, thereby improving the accuracy and recall rate of object detection.
[0062] Step (5) Set the training parameters and replace the DiffusionDET loss function with the new loss function ACR-Loss. Input the features generated in step (1) into the model constructed by steps (2)-(4) for iterative training. And verify on the validation set to obtain the optimal parameter model and output the prohibited item detection effect diagram. Specifically:
[0063] (5-1) At the initial stage of model training, a set of key hyperparameters need to be set first, such as the initial learning rate, momentum factor, learning rate decay rate, batch size, the number of GPUs used, the selected optimization algorithm, and the total number of training epochs.
[0064] (5-2) Set the new IoU-aware loss function ACR-Loss for model iterative training. The ACR-Loss function is:
[0065] (5)
[0066] Among them, represents the classification loss function, adopts the GIoU loss calculation function, adopts the bounding box regression loss. Specifically, the classification loss function is:
[0067] (6)
[0068] Among them, , is the value obtained from object detection, is the binary cross-entropy function, is the classification value. The regression loss is:
[0069] (7)
[0070] Among them, represents the predicted bounding box, represents the ground truth bounding box.
[0071] (5-3) Input the prepared dataset into the model framework constructed by the previous steps and conduct multiple rounds of training. During the training process, use the mAP50 metric as the measurement metric to continuously monitor the performance of the model on the validation set. Whenever the best mAP50 value is obtained during the training process, record the corresponding model configuration and save these configurations as the optimal parameters. This approach helps to find the best parameter combination during the training process, thereby improving the prediction accuracy of the model.
[0072] (5-4) After the training is completed, use the validation set to verify the optimal model obtained in step (5-1), obtain the metric parameters for the dataset detection by the final model, and mark the categories and confidence levels of the detected contraband on the detection results.
[0073] The above embodiments have elaborated in detail the objectives, technical solutions, and their advantages of the present invention. It should be clear that these embodiments only represent some preferred implementation approaches of the present invention and do not limit its scope. Any form of modification, equivalent replacement, or improvement made under the premise of following the core ideas and principles of the present invention should be considered as being included within the protection scope of the present invention.
[0074] Experimental comparison description:
[0075] SparseGenX uses ResNet-50 as the feature extraction backbone, denoted as SparseGenX-R; uses Swin-T as the feature extraction backbone, denoted as SparseGenX-S. Unless otherwise specified, SparseGen is default to represent SparseGenX-R.
[0076] Table 1 Comparison results of SparseGenX and other SOTA algorithms on CLCXray
[0077]
[0078] As can be seen from Table 1, the mAP50 accuracy of SparseGenX-S on the CLCXray test set is 89.7%, achieving the best detection performance. SparseGenX-S has the highest detection rate in detecting contraband such as Dagger, PlasticBottle, Cans, GlassBottle, and Tin. When detecting other contraband, the overall performance of SparseGenX is also better than other similar algorithms. In addition, by comparing SparseGenX-R and SparseGenX-S, it can be seen that the detection rate of SparseGenX-S on GlassBottle reaches 33.7%, and the detection rate on PlasticBottle reaches 94.1%. This indicates that when identifying contraband similar to the background, using shape features for detection is more effective than relying on texture features. Compared with the detection results of Sparse R-CNN at 87.6% and YOLOv7u6 at 86.2%, SparseGenX-S is respectively 2.1% and 3.5% higher.
[0079] Table 2 Comparison results of SparseGenX and other algorithms on the WIXray dataset
[0080]
[0081] Table 2 shows the detection results of SparseGenX, Faster R-CNN, Sparse R-CNN, and DiffusionDET on the WIXray dataset. The results show that when the SparseGenX algorithm uses ResNet-50 as its feature extraction backbone, its mAP reaches 46.8% and its mAP50 reaches 64.0%. Compared with DiffusionDET, its mAP and mAP50 are improved by 2.2% and 0.2% respectively. When SparseGenX uses the larger ResNet-101, its mAP and mAP50 are improved by 2.7% and 1.4% respectively compared with using ResNet-50, 0.6% and 0.5% higher than Sparse R-CNN, and 6% and 3% higher than Faster R-CNN. This result highlights the certain advantages of the SparseGenX algorithm when dealing with other X-ray images with similar backgrounds.
Claims
1. A method for detecting contraband in X-ray security inspection images based on generative sparsity, characterized in that, It includes the following steps: Step 1: Obtain an X-ray security inspection image dataset containing target bounding boxes and class annotations, and successively use the ResNet backbone and feature pyramid structure to extract multi-level features in the security inspection images; Step 2: Use a diffusion model to generate a noisy pre-plan box; and generate region-of-interest features using the pre-plan box and a context-aware feature representation ; Step 3: Replace the dynamic instance interaction of the diffusion model DiffusioDET with the dynamic filtering and fusion module DFF to generate a refined feature map , and obtain the target prediction box and the corresponding class scores; Step 4: In the label assignment stage, use the Hungarian Top-K matching strategy HMK to find K equally important optimal positive samples for each ground truth GT through global matching; Step 5: Construct a loss function, input the multi-level features in Step 1 into the model constructed by Steps 2 - 4 for iterative training, and verify on the validation set to output the detection effect diagram of prohibited items.
2. The method for detecting contraband in X-ray security inspection images based on generative sparsity according to claim 1, wherein In the said step 1, specifically, an X-ray security inspection image dataset containing target bounding boxes and class annotations is obtained, the dataset is divided into a training set and a validation set, and data augmentation is performed on the images through preprocessing; features in the images are extracted using a ResNet backbone, and then multi-level features are obtained using a feature pyramid structure .
3. The method for detecting contraband in X-ray security inspection images based on generative sparsity according to claim 2, wherein The specific implementation process of Step 2 is as follows: Step 2-1: Generate anchor box samples using a diffusion model; Step 2-2. Apply the Region of Interest (RoI) Align operation on the image feature map to extract the region of interest features corresponding to the anchor boxes from the feature map ; Step 2-3: Using the self-attention mechanism, for all features perform cross-attention mechanism feature aggregation Implement information interaction between features to obtain context-aware feature representations .
4. The method for detecting contraband in X-ray security inspection images based on generative sparsity according to claim 3, wherein, The specific process of generating anchor box samples using a diffusion model is as follows: Step 2-1-1: In the forward noise addition stage, the diffusion model adds Gaussian noise to the ground truth GT to generate a noisy proposed box , specifically: set , representing the center point coordinates of the proposed box and the width and height of the proposed box respectively, and calculate the corresponding ; Step 2-1-2: In the inference stage, create an inverse time pair , where represents the final moment; then for each time pair , the diffusion model uses the current time point on to predict at the time point ; then use the denoising diffusion implicit model DDIM to perform denoising sampling to calculate the corresponding anchor box.
5. The method for detecting contraband in X-ray security inspection images based on generative sparsity according to claim 4, wherein The specific implementation process of Step 3 is as follows: Step 3-1: First, use a fully connected network to map the feature representation into the parameter space to obtain the dynamic filtering parameter dfp; then, through the kernel splitting technique, split dfp into two independent parameters: dfv responsible for the vertical direction and dfh responsible for the horizontal direction; Step 3-2: Use the CA attention mechanism to extract features to generate a feature map, and then use the vertical filtering parameter dfv and the horizontal filtering parameter dfh respectively to filter the region of interest in features using the spatial attention mechanism CA to obtain the filtered features xfv and xfh; Step 3-3: Concatenate the filtered features xfv and xfh, process them through a non-linear activation module, and finally map the result back to the feature space to obtain the refined features after filtering and fusion. ; Step 3-4: Use the MLP to perform corresponding scaling and translation operations, and then calculate the prediction scores and prediction bounding boxes.
6. The method for detecting contraband in X-ray security inspection images based on generative sparsity according to claim 5, wherein The specific implementation process of Step 4 is as follows: Step 4-1: HMK calculates the cost loss according to the predicted score, predicted box, GT ground truth box, and GT label, including cost classification loss, cost regression loss, and cost bounding box loss; during the calculation, it is judged whether the center point of the predicted box is inside the GT ground truth box and whether it is inside the circle with the center of the GT ground truth box as the center and a radius of 2.5 units; Penalize the predicted boxes whose center points are not inside the GT ground truth box or not within its center radius range; Step 4-2: Calculate the IoU value between the predicted box and the GT. Through the summation of the IoU values, HMK determines the number of positive samples dk required for each GT; then, according to dk, expand each column corresponding to each GT in the total loss dk times to construct the final loss matrix; Step 4-3: Use the Hungarian algorithm to find the dk positive sample boxes to be matched for each GT according to the loss matrix. This process ensures that each GT can be effectively matched with the predicted box.
Citation Information
Patent Citations
Monocular image-oriented three-dimensional object detection method based on three-dimensional reconstruction
CN110689008A
Cloth length online accurate metering method
CN111504203A
X-ray contraband package detection method and device based on feature map re-empowerment
CN112070079A
Stereoscopic scene target detection method and system based on key point multi-scale fusion
CN114332792A
X-ray image contraband detection method based on de-overlapping and associated attention mechanism
CN118261853A