X-ray security check image contraband detection method based on hybrid self-supervised learning
Through the hybrid self-supervised learning pre-training framework ESLA and head-tail feature pyramid module HTFP, the problem of low detection accuracy of contraband in X-ray security images is solved, and high-precision detection on small-scale data sets is achieved.
Patent Information
- Application Number
- CN202510670224.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The contraband in X-ray security images have low detection accuracy due to the low object-level label data or the domain differences between pre-trained data sets.
Using the pre-trained framework ESLA and head-tail feature pyramid module HTFP based on hybrid self-supervised learning, the self-supervised signal and feature representation are integrated, and the low-level image features in the multi-level feature map are strengthened through the horizontal connection structure.
The detection accuracy of the model on the small-scale object-level annotation data set is improved, the richness of features and the accuracy of detection is improved, and the problem of low detection accuracy caused by domain differences is solved.
Smart Images

Figure CN120198425A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of self-supervised learning and object detection, and specifically relates to an X-ray security inspection image contraband detection method based on a hybrid self-supervised pre-training module and a head-tail feature pyramid. This method is designed specifically for the security inspection field and aims to improve the accuracy of contraband detection. Background Art
[0002] With the convenience of transportation, X-ray security inspection has become an important means to ensure the safety of public places. After being irradiated by X-rays, the internal items in luggage, backpacks, etc. will be displayed in the X-ray security inspection image, helping security inspectors discover potential contraband therein and ensuring the safety of public places. Due to the penetrability of X-rays, compared with natural light images, X-ray security inspection images have the following characteristics: (1) Overlap: When stacked items are irradiated by X-rays, an overlap phenomenon will occur in the image, resulting in blurred surface textures of the items in the image; (2) Hiddenness: Contraband in X-ray security inspection images is small in volume and overlaps with other items, making it difficult to be detected by the system; (3) Difficult annotation: Due to characteristics (1) and (2), object-level annotation on X-ray security inspection images requires huge time and economic costs.
[0003] With the development of deep learning object detection technology, great progress has been made in the automated detection of X-ray contraband. There are two common object detection training methods. One is to directly train a detection model on the target dataset, and the other is to initiate detection on this dataset by fine-tuning a model pre-trained on other datasets. The first method requires training on a dataset with large-scale object-level annotations, but it is very difficult to obtain such a dataset; the second method has certain domain differences, that is, there are certain domain differences between X-ray security inspection images and natural light images, and there are also certain domain differences between X-ray security inspection images generated by different X-ray machines, resulting in poor detection effects. Therefore, the present invention provides an X-ray security inspection image contraband detection method based on hybrid self-supervised learning, which improves the detection accuracy of the model on a small-scale object-level annotation dataset without the need for large-scale object-level label annotation. The present invention is based on a hybrid self-supervised method and a head-tail feature pyramid module, and solves the problem of low detection accuracy of X-ray security inspection image contraband due to less object-level label data or domain differences between pre-trained datasets. Summary of the Invention
[0004] In view of the deficiencies in the existing technology, the present invention provides an X-ray security inspection image contraband detection method based on hybrid self-supervised learning. By improving the hybrid self-supervised learning algorithm and optimizing the process of converting flat-level features into multi-level features, the detection accuracy of the model on a small-scale object-level labeled dataset is improved. Specifically, the present invention integrates the self-supervised signal and feature representation in the image by truncating and fusing the dual-augmented view branch features in the hybrid framework, and designs a hybrid self-supervised pre-training framework ESLA. It does not require additional encoders and decoders. On the one hand, ESLA uses the augmented view features for contrastive learning to deepen the model's understanding of the image content; on the other hand, it performs decoding operations through the truncated and fused dual-augmented view features to achieve contrastive reconstruction, enabling the model to learn more complete self-supervised signals. In addition, the present invention also designs a head and tail feature pyramid module HTFP to convert the single-level flat features generated by ESLA into multi-level features. During the conversion process, HTFP strengthens the low-level image features in the multi-level feature maps, such as shapes and edges, by introducing a lateral connection structure, thereby improving the detection performance.
[0005] An X-ray security inspection image contraband detection method based on hybrid self-supervised pre-training includes the following steps:
[0006] Extract features from the X-ray security inspection images to generate flat-level feature maps.
[0007] Divide the obtained set of X-ray security inspection images into two parts: one part is used for training the model, and the other part is used to verify the accuracy of the model.
[0008] Adjust the size of the images and perform normalization processing on the pixel values.
[0009] Use the hybrid self-supervised pre-training framework ESLA for pre-training the feature extraction backbone.
[0010] The feature extraction process of the feature extraction backbone is responsible for two ViT encoders: the masked online encoder and the masked target encoder .
[0011] Respectively use and the corresponding projectors and to project and map the extracted features to obtain latent representations.
[0012] Through feature fusion technology, the feature representations of two data-augmented views are combined in the feature space to supplement the comprehensive feature information that may not be captured by a single view. The fused features are fed into the decoder Decode in it to reconstruct the original image.
[0013] After being extracted by the pre-trained ViT feature extractor, a flat-level feature map is generated, that is .
[0014] Convert the single-level feature representation into a multi-level feature representation.
[0015] For perform downsampling operations to generate four feature maps of different scales; for the feature map with the initial scale of of feature map, respectively through the convolution operation with a stride of to generate } scale feature map .
[0016] Add a lateral connection structure; on the feature map , through convolution to and perform lateral connection, and use the underlying graphic information in to enhance the model's ability to identify dangerous goods; perform downsampling operations on to obtain a feature map with halved size .
[0017] Input the multi-scale feature maps into the multi-level detection head for processing to obtain the ROI feature map and generate a feature map of a unified size.
[0018] The processing process of the multi-level detection head is as follows:
[0019] In the region candidate layer of the multi-level detection head, for each layer of the multi-scale feature map , a series of prediction scores and bounding box regression coefficients are generated in the form of a sliding window.
[0020] Create multiple anchor boxes using the anchor box generator in the pre-plan layer and the anchor box target layer of the multi-level detection head, and screen out high-quality candidate regions through the class prediction score and the NMS algorithm; then, finely adjust the candidate regions through the bounding box regression coefficients regression coefficients.
[0021] Screen and optimize the ROI list through the pre-plan target layer of the multi-level detection head.
[0022] Perform a spatial transformation on the ROI feature map through the region of interest alignment layer of the multi-level detection head to generate a feature map of a unified size .
[0023] The feature map is detected using a fully connected layer network, and the specific operations are as follows:
[0024] The feature map is further processed using two fully connected layer networks, which respectively perform the tasks of classification and bounding box regression; the first fully connected layer is responsible for classifying each feature map to determine the target class it belongs to; the second fully connected layer focuses on the precise regression of the bounding box, adjusting the position and size of each ROI to more accurately locate the target object.
[0025] The overall model is iteratively trained to obtain an optimal parameter model, and the detection effect diagram of contraband is output.
[0026] The prepared dataset is input into the model framework constructed by the ViT feature extractor, the head and tail feature pyramid module HTFP, the multi-level detection head, and the fully connected layer network, and is trained for multiple rounds until convergence; during the training process, the mAP50 metric is used as a measurement metric to continuously monitor the performance of the model on the validation set; whenever the best mAP50 value is obtained during the training process, the corresponding model configuration is recorded and these configurations are saved as the optimal parameters.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] The present invention innovates in the training of X-ray contraband detection by using a hybrid self-supervised model, improving its performance in X-ray security inspection images. Specifically, the present invention designs a hybrid self-supervised pre-training framework ESLA. It integrates the self-supervised signal and feature representation of the image by truncating and fusing the dual-augmented view branch features in the hybrid framework, without adding additional encoders and decoders. On the one hand, ESLA uses the augmented view features for contrastive learning to deepen the model's understanding of the image content; on the other hand, it performs decoding operations through the truncated and fused dual-augmented view features to achieve contrastive reconstruction, enabling the model to learn a more complete self-supervised signal. In addition, the invention also designs a head and tail feature pyramid HTFP to convert the single-level flat features extracted by ESLA into multi-level features. HTFP enhances the low-level image features in the multi-level feature map, such as key information like shape and edge, by introducing a lateral connection structure, thereby improving the richness of the features and the accuracy of detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a processing flow chart of an embodiment of the present invention.
[0030] Figure 2 It is a schematic diagram of the overall structure of the model of an embodiment of the present invention.
[0031] Figure 3 It is a schematic diagram of the structure of the hybrid self-supervised pre-training framework ESLA of an embodiment of the present invention.
[0032] Figure 4 This is a schematic diagram of the head and tail feature pyramid HTFP structure of the embodiment of the present invention. Detailed implementation manner
[0033] To better understand the purpose, structure and function of the present invention, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. The present invention proposes an X-ray security inspection image contraband detection method based on hybrid self-supervised learning, and the overall process is as Figure 1 shown. The overall structure of the designed model is as Figure 2 shown. The present invention first designs a hybrid self-supervised pre-training framework ESLA. It integrates the self-supervised signal and feature representation of the image by truncating and fusing the dual-augmented view branch features in the hybrid framework, and no additional encoders and decoders are required. On the one hand, ESLA uses the augmented view features for contrastive learning to deepen the model's understanding of the image content; on the other hand, through the truncated and fused dual-augmented view features for decoding operations, contrastive reconstruction is realized, enabling the model to learn a more complete self-supervised signal. Then, the present invention also designs a head and tail feature pyramid HTFP to convert the single-level flat features extracted by ESLA into multi-level features. HTFP enhances the low-level image features in the multi-level feature maps, such as key information such as shapes and edges, by introducing a lateral connection structure, thereby improving the richness of features and the accuracy of detection.
[0034] The method of the present invention includes the following steps:
[0035] An X-ray security inspection image contraband detection method based on hybrid self-supervised pre-training includes the following steps:
[0036] Step (1): Extract features from the X-ray security inspection images.
[0037] Use the ViT feature extractor pre-trained by the hybrid self-supervised pre-training framework ESLA to extract features from the X-ray security inspection images to generate a flat feature map, that is . The specific operations are as follows:
[0038] (1-1): Divide the obtained set of X-ray security inspection images into two parts: one part is used for training the model, that is ; the other part is used to verify the accuracy of the model, that is . Among them, R is the real number field, represents the number of image samples in the training set, represents the i-th training image sample, represents the number of image samples in the validation set, represents the j-th validation image sample, H represents the image height, W represents the image width, and 3 represents the number of RGB channels.
[0039] (1-2): Each training sample The corresponding label ; where Indicates the number of objects contained in the image sample in Indicates in the th object's true class, where C represents the total number of classes in the dataset Indicates in the th object's bounding box, which consists of the abscissa x of the center point, the ordinate y of the center point, the width w of the object, and the height h. The image was resized and uniformly scaled to 224×224 pixels, and pixel value normalization was performed to ensure the consistency and effectiveness of the model input.
[0040] (1-3): Use the hybrid self-supervised pre-training framework ESLA to pre-train the feature extraction backbone.
[0041] (1-3-1): Define two data augmentation functions and , and apply them to the original image to generate two different augmented views. Specifically, applying the first augmentation function gives , applying the second augmentation function gives , where represents the original image represents the pixel space.
[0042] (1-3-2): The data augmentation functions and randomly select data augmentation strategies, including:
[0043] 1) Color perturbation: By changing attribute parameters such as the color, hue, saturation, brightness, and contrast of the image, it helps the hybrid self-supervised pre-training framework ESLA to identify object features under different lighting conditions.
[0044] 2) Random grayscale processing: Converting the image to grayscale helps the model learn color-independent features.
[0045] 3) Blurring processing: By simulating images under different focusing conditions and adding random Gaussian noise to the image, it makes the model pay more attention to the features of blurred pixels.
[0046] 4) Random horizontal flipping: Flipping the image helps the model learn mirror symmetry invariance.
[0047] 5) Random cropping and scaling: Randomly crop and scale the image so that the model can learn the features of the image at different sizes and perspectives.
[0048] 6) Random rotation: Rotate the image by 90 degrees, 180 degrees, or 270 degrees to help the model learn rotational invariance of the image.
[0049] (1-3-3): The feature extraction process of the feature extraction backbone is responsible for two ViT encoders: the masked online encoder and the masked object encoder . These two encoders are applied to two augmented views and , that is To maintain the stability of the model, the parameters in the object encoder are slowly updated by the online encoder through the EMA strategy. The update formula is: , where and represent and parameters of, represents the smoothness of the update.
[0050] (1-3-4): Respectively use and corresponding projectors and to project the extracted features to obtain the latent representation, that is and . These projectors are implemented by a multi-layer perceptron (MLP), including a linear transformation layer, BN, and ReLU activation function, to further refine and transform the feature representation, providing richer and more robust features for subsequent self-supervised learning tasks.
[0051] (1-3-5): Use the predictor to predict the projected features to obtain the predicted features . Then, the hybrid self-supervised pre-training framework ESLA compares the predicted features with and calculates the contrastive loss to enhance the signal strength of self-supervised learning, strengthening the recognition and understanding ability of the hybrid self-supervised pre-training framework ESLA for image features. In addition, a gradient stop strategy is adopted to ensure the consistency of the prediction process and the overall stability of the model by preventing gradient information from backpropagating from the predictor to the online network.
[0052] (1-3-6): Through feature fusion technology, in the feature space The feature representations of two data-augmented views are combined internally to supplement the comprehensive feature information that a single view may fail to capture. This process is achieved through a specific fusion function as follows, and its fusion formula is expressed as:
[0053] (1)
[0054] where and are the feature representations generated by the encoder in the feature space , that is . The function represents a truncation operation, and represents a concatenation operation. The fusion function aligns and combines the two feature representations according to a certain mixing ratio , which is set to 0.5 by default. The mixed feature representation maintains the same length as the feature representation generated by the encoder. The fused feature is fed into the decoder for decoding to reconstruct the original image.
[0055] (1 - 4) Through the extraction of the pre-trained ViT feature extractor, a flat feature map is generated, that is .
[0056] Step (2): Convert the single-level feature representation into a multi-level feature representation.
[0057] Use the head and tail feature pyramid module HTFP to perform upsampling on the final layer feature map of the ViT feature extractor to generate four feature maps of different scales, that is . Perform downsampling on to obtain a feature map with half the size . The specific operations are as follows:
[0058] (2 - 1): Perform downsampling on to generate four feature maps of different scales. For the feature map (stride = 16) with the initial scale of , through convolutional operations with a stride of , more detailed feature maps of the } scale are generated.
[0059] (2 - 2): Add a lateral connection structure. On the feature map , through convolution, is combined with Perform a horizontal connection and enhance the model's ability to identify dangerous goods by leveraging the underlying graphic information in . Downsample to obtain a feature map with halved dimensions .
[0060] Step (3): Process the ROI features to generate a feature map of a unified size.
[0061] Input the multi-scale feature map into a multi-level detection head for calculation to obtain an ROI feature map. Perform a spatial transformation on the ROI feature map through an ROI alignment layer to generate a feature map of a unified size . The specific operations are as follows:
[0062] (3-1): In the region proposal layer of the multi-level detection head, for each layer of the multi-level feature map , generate a series of prediction scores and bounding box regression coefficients in a sliding window manner. Specifically, first, on each layer of the feature map , generate a feature map through a convolution operation. Then, at each pixel point of it, generate a set of anchor boxes of predefined sizes. For each anchor box in , use a Region Proposal Network (RPN) to predict the scores of objects with different sizes and shapes in each anchor box through a binary classifier to determine whether the target to be detected is contained in this anchor box. If the anchor box contains the target, mark it as a positive sample. If not, mark it as a negative sample. Finally, for the positive sample anchor boxes, the RPN further calculates the regression coefficients of the bounding boxes through regression analysis. The regression coefficients include the center coordinates of the anchor box ( ) and the width and height ( ).
[0063] (3-2): In the Proposal Layer and Anchor Target Layer of the multi-level detection head, use an anchor box generator to create multiple anchor boxes, and filter out high-quality candidate regions through the class prediction scores and the NMS algorithm to reduce redundancy and retain the most promising regions. Then, finely adjust these candidate regions through the bounding box regression coefficients to ensure the accuracy of the bounding boxes and their fit with the target object.
[0064] (3-3): Screen and optimize the ROI list through the Proposal Target Layer of the multi-level detection head. Filter the bounding boxes that highly overlap with the GT by setting the IoU threshold, and use NMS to remove the overlapping low-score candidate regions in the fine-tuned candidate regions, finally generating an ROI list ROIs. Select the high-quality ROI feature maps that highly overlap with the real targets through the IoU threshold, and at the same time generate positive and negative samples to balance the class distribution.
[0065] (3-4): Perform spatial transformation on the ROI feature maps through the Region of Interest Alignment Layer of the multi-level detection head to generate feature maps of a unified size. .
[0066] Step (4): Use the fully connected layer network to detect the feature maps. The specific operations are as follows:
[0067] Use two fully connected layer networks FC to further process the and respectively perform the tasks of classification and bounding box regression. The first fully connected layer is responsible for classifying each feature map to determine the target class it belongs to. The second fully connected layer focuses on the precise regression of the bounding boxes, adjusting the position and size of each ROI to more accurately locate the target object.
[0068] Step (5): Iteratively train the overall model to obtain the optimal parameter model and output the detection effect diagram of contraband.
[0069] Set the training parameters and perform iterative training on the model constructed by the ViT feature extractor, the head and tail feature pyramid module HTFP, the multi-level detection head, and the fully connected layer network. And verify on the validation set to obtain the optimal parameter model and output the detection effect diagram of contraband. The specific operations are as follows:
[0070] (5-1): Input the prepared dataset into the model framework constructed by the ViT feature extractor, the head and tail feature pyramid module HTFP, the multi-level detection head, and the fully connected layer network, and perform multiple rounds of training until convergence. During the training process, use the mAP50 metric as the measurement index to continuously monitor the performance of the model on the validation set. Whenever the best mAP50 value is obtained during the training process, the corresponding model configuration will be recorded and these configurations will be saved as the optimal parameters. This approach helps to find the best parameter combination during the training process, thus improving the prediction accuracy of the model.
[0071] (5-2): After the training is completed, use the validation set to validate the optimal model obtained in step (5-1), obtain the metric parameters for the dataset detection by the final model, and mark the detected contraband category and confidence level on the detection results.
[0072] Experimental comparison data:
[0073] The dataset uses the APIDray dataset, which is an extended version of the PIDray dataset. The PIDray dataset covers various situations in real security inspection scenarios, especially including those deliberately hidden contraband items, and contains a total of 12 categories, namely Baton, Bullet, Gun, Powerbank, Sprayer, Lighter, Handcuffs, Pliers, Scissors, Knife, Hammer, Wrench.
[0074] The APIDray dataset adjusts the image size in PIDray to 440×448 pixels and grayscales the images to simulate the effect of CRT (cathode ray tube) phosphors. Then, the APIDray dataset is further enriched through a series of data augmentation techniques. These augmentation means include: (1) Horizontally flip the image with a 50% probability; (2) Vertically flip the image with a 50% probability; (3) Randomly select one of the following 90-degree rotations: no rotation, 90-degree clockwise rotation, 90-degree counterclockwise rotation, or 180-degree rotation (i.e., the image is upside down); (4) Apply a random rotation transformation of -45 degrees to +45 degrees; (5) Perform a random shear of -0 to +0 degrees in the horizontal direction and a random shear of -45 degrees to +45 degrees in the vertical direction. After expansion, APIDray contains a total of 70,619 images, with nearly 60K in the training set.
[0075] Experimental comparison description:
[0076] Table 1 Comparison of detection results of EslaXDET and other models on APIDray
[0077]
[0078] Table 1 shows the result comparison of the EslaXDET, a method for detecting contraband in X-ray security inspection images based on hybrid self-supervised learning, with several other models on the APIDray test set. These models include: BYOL based on HTFP (BYOL+HTFP); MAE based on HTFP (MAE+HTFP), where MAE is pre-trained for 1600 epochs; two-stage detection frameworks Mask R-CNN and Faster R-CNN, and CAE based on HTFP (CAE+HTFP). As can be seen from Table 1, the mAP detection result of EslaXDET is 80.3%, which is the current optimal detection result. It is 1.1% and 2.7% higher than MAE+HTFP and BYOL+HTFP respectively, indicating that the hybrid self-supervised learning algorithm ESLA can learn more feature representations beneficial to downstream tasks than MAE or BYOL. In addition, EslaXDET is 1.9% higher than CAE+HTFP in detection results, which shows that the training strategy of only decoding one enhanced view in the CAE framework may cause the model to only capture partial features of the image, thereby affecting its performance in downstream tasks. In addition, EslaXDET is 5.7% and 8.2% higher than the mAP detection results of Mask R-CNN and Faster R-CNN respectively, indicating that pre-training using self-supervised learning on a homologous dataset performs better than fine-tuning a model pre-trained on a natural light dataset. These results highlight the effectiveness and superiority of EslaXDET in the contraband detection task.
[0079] The above embodiments elaborate in detail the objectives, technical solutions and their advantages of the present invention. It should be clear that these embodiments only represent some preferred implementation approaches of the present invention and do not limit its scope. Any form of modification, equivalent replacement or improvement made on the premise of following the core idea and principles of the present invention shall be regarded as being included within the protection scope of the present invention.
Claims
1. A method for detecting contraband in X-ray security inspection images based on hybrid self-supervised pre-training, characterized in that, It includes the following steps: Extract features from the X-ray security inspection images to generate flat-level feature maps; convert the single-level feature representation into a multi-level feature representation; Input the multi-scale feature maps into the multi-level detection heads for processing to obtain ROI feature maps and generate feature maps of a unified size; use a fully connected layer network to detect the feature maps; Iteratively train the overall model to obtain an optimal parameter model and output the detection effect diagram of prohibited items.
2. The method for detecting contraband in X-ray security inspection images based on hybrid self-supervised pre-training according to claim 1, wherein Extract features from the X-ray security inspection images to generate flat-level feature maps. The specific operations are as follows: Divide the obtained set of X-ray security inspection images into two parts: one part is used for training the model, and the other part is used to verify the accuracy of the model; Adjust the size of the images and perform normalization processing on the pixel values; Use the hybrid self-supervised pre-training framework ESLA to pre-train the feature extraction backbone; After the extraction by the pre-trained ViT feature extractor, a flat feature map is generated, that is .
3. The method for detecting contraband in X-ray security inspection images based on hybrid self-supervised pre-training according to claim 2, wherein Convert the single-level feature representation into a multi-level feature representation. The specific operations are as follows: Perform a downsampling operation to generate four feature maps of different scales; for the feature map with an initial scale of and stride of , generate a feature map of scale through a convolution operation ; Add a horizontal connection structure; on the feature map , through convolution, perform a horizontal connection between and , and utilize the underlying graphic information in to enhance the model's ability to identify dangerous goods; perform a downsampling operation on to obtain a feature map with half the size.
4. The method for detecting contraband in X-ray security inspection images based on hybrid self-supervised pre-training according to claim 3, wherein The processing process of the multi-level detection heads. The specific operations are as follows: In the region candidate layer of the multi-level detection head, for each layer of the multi-level feature map a series of prediction scores and bounding box regression coefficients ; Create multiple anchor boxes using an anchor box generator in the pre-plan layer and the anchor box target layer of the multi-level detection head, and filter out high-quality candidate regions through the class prediction score and the NMS algorithm; then, finely adjust the candidate regions through the regression coefficients of the bounding box Screen and optimize the ROI list through the pre-defined target layer of the multi-level detection heads; Spatially transform the ROI feature map through the region of interest alignment layer of the multi-level detection head to generate a feature map of a unified size .
5. The method for detecting contraband in X-ray security inspection images based on hybrid self-supervised pre-training according to claim 4, wherein Use a fully connected layer network to detect the feature maps. The specific operations are as follows: Use two fully connected layer networks to further process the feature maps, respectively performing the tasks of classification and bounding box regression; the first fully connected layer is responsible for classifying each feature map to determine its target category; the second fully connected layer focuses on the precise regression of the bounding box, adjusting the position and size of each ROI to more accurately locate the target object.
6. The method for detecting contraband in X-ray security inspection images based on hybrid self-supervised pre-training according to claim 5, wherein Iteratively train the overall model. The specific operations are as follows: Input the prepared dataset into the model framework constructed by the ViT feature extractor, the head and tail feature pyramid module HTFP, the multi-level detection heads, and the fully connected layer network, and perform multiple rounds of training until convergence; during the training process, use the mAP50 metric as the measurement index to continuously monitor the performance of the model on the validation set; Whenever the best mAP50 value is obtained during the training process, record the corresponding model configuration and save these configurations as the optimal parameters.
Citation Information
Patent Citations
Lightweight deep neural network rotating target detection method and system
CN111242122A
Security check contraband detection method based on multi-scale attention and data enhancement
CN116883933A
Small sample image defect target detection method based on self-supervised pre-training
CN116994047A
X-ray image contraband detection method based on de-overlapping and associated attention mechanism
CN118261853A