Occluded object detection method in X-ray images based on local-global visibility analysis

By adopting the local-global visibility analysis method in X-ray image occlusion object detection, the visibility features of the image are extracted and utilized, combined with the YOLOv8 model and the visibility attention model, the problem of underutilizing global features and other color space information in the prior art is solved, and more efficient and accurate occlusion object detection is achieved.

CN118968027BActive Publication Date: 2025-05-13HEFEI UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411106661.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-13
Publication Date
2025-05-13
Estimated Expiration
2044-08-13

AI Technical Summary

Technical Problem

The prior art does not fully consider the relationship between global features and fails to effectively utilize information in other color spaces when processing X-ray image occlusion object detection, resulting in limited effect in complex occlusion object detection.

Method used

Using a method based on local-global visibility analysis, the local and global visibility features of X-ray images are extracted, and the prediction box is generated using the visibility attention model, combined with the YOLOv8 model for detection, and the local-global visibility mask loss and occlusion score gradient adjustment loss are additionally calculated to improve the accuracy of the detection.

Benefits of technology

Effectively analyze the visibility mask of the target, alleviate the missed detection of the occlusion target, and improve the accuracy and efficiency of X-ray image occlusion target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968027B_ABST
    Figure CN118968027B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting occluded targets in X-ray images using local-global visibility analysis. The present invention uses YOLOv8 to extract local block features, and calculates a variety of local visibility masks and local visibility features. The present invention calculates the global relationship between local blocks, and calculates a variety of global visibility masks and global visibility features. In order to verify the accuracy of visibility mask estimation in the case of occluded targets, the present method additionally calculates the local-global visibility mask loss and the occlusion score gradient adjustment loss when training the target detection model. The present invention takes into account the local-global visibility mask loss and the gradient adjustment loss, and can effectively analyze the visibility mask of the target and alleviate the missed detection of occluded targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection and artificial intelligence, and in particular to an X-ray image occluded target detection method based on local-global visibility analysis. Background Art

[0002] With the continuous improvement of public safety needs, the traditional method of manually judging whether X-ray images contain prohibited items faces challenges such as low recognition readiness rate, low efficiency, and a large amount of manpower consumption. In recent years, the use of deep learning algorithms to assist manual work in judging prohibited items in X-ray images has gradually begun to be put into use. It has greatly improved the efficiency and accuracy of prohibited item detection, and at the same time, it has become more convenient to manage and analyze detection data.

[0003] China's public patent number CN 117765378 B "Method and device for detecting prohibited items in complex environments with multi-scale feature fusion" mainly enhances the network's ability to extract local features of overlapping objects by designing a multi-scale attention module backbone, and introduces a squeeze-incentive attention mechanism to reduce redundant information in the target area; in order to address the problem of information loss of small targets, an adaptive fusion feature pyramid network is designed to introduce shallow features containing detail information and deep features containing semantic information to prevent information loss of small targets; an adaptive weight fusion strategy and a channel attention mechanism are used to avoid target information loss caused by direct fusion.

[0004] China's public patent number CN 114548230 B "X-ray prohibited items detection method based on RGB color separation and dual-path feature fusion" includes the following steps: obtaining training sample sets and test sample sets; constructing a dual-path feature fusion network model for RGB color separation; iteratively training the dual-path feature fusion network for RGB color separation; and obtaining X-ray prohibited items image recognition results. First, the RGB color separation structure is constructed, then the feature extraction network structure is constructed, and then the feature fusion network structure is constructed, and then training is performed, which solves the problem of the prior art excluding irrelevant information from affecting the detection of prohibited items, thereby improving the detection accuracy of prohibited items in X-ray scenes.

[0005] China's public patent number CN 118053122 A "X-ray security inspection dual-view detection method based on improved FCOS network" is improved by adding an image encoder, an HSV-guided encoder and a multi-scale fusion encoder. The image encoder converts the color space of the input top-view X-ray image and the side-view X-ray image from RGB to HSV for feature extraction and fusion. The HSV-guided encoder fuses the HSV-guided features extracted by the image encoder with the input original X-ray top view and side view respectively. The multi-scale fusion encoder is used to extract the multi-scale features of the top view and the side view, and fuse the multi-scale features of the two views. The present invention reduces the interference of stacked targets under X-ray transmission imaging, makes full use of image information in different color spaces, strengthens the network feature perception capability, and improves the detection accuracy and speed of X-ray objects.

[0006] Existing patents have some shortcomings in dealing with occluded target detection: first, although multi-scale attention modules and squeeze-excited attention mechanisms are introduced, the relationship between global features is not fully considered; second, they only rely on the RGB color separation structure, fail to fully utilize the information in other color spaces, and have limited effects when dealing with occluded objects; third, although multiple encoders are introduced, the effect of dual-view feature fusion may be insufficient when facing complex occluded targets. Summary of the invention

[0007] The purpose of the present invention is to remedy the defects of the prior art and provide an X-ray image occluded target detection method based on local-global visibility analysis.

[0008] The present invention is achieved through the following technical solutions:

[0009] A method for detecting occluded objects in X-ray images based on local-global visibility analysis specifically comprises the following steps:

[0010] S1: Extract local visibility features of X-ray images;

[0011] S2: Extract global visibility features of X-ray images;

[0012] S3: Generate prediction box using local visibility features and global visibility features;

[0013] S4: Training visibility attention model;

[0014] S5: Detecting occluded object categories in X-ray images using visibility attention model.

[0015] The step S1 of extracting the local visibility features of the X-ray image specifically includes the following steps:

[0016] S1-1: Input X-ray image dataset;

[0017] S1-2: Input the X-ray image x into the YOLOv8 model and extract the feature F∈R after the 21st layer processing C×H×W , where C represents the number of channels, H represents height, and W represents width;

[0018] S1-3: Divide the feature F into blocks [k, k] with a height of k and a width of k, and sum the feature values ​​in each block to obtain the local block feature LF = {LF j}, LF∈R C×h×w , j represents the subscript index of the block, and its value is {1, ..., hw}, C represents the number of channels,

[0019] S1-4: Calculate local block score;

[0020] S1-5: Calculate multiple local visibility masks;

[0021] S1-6: Extract local visibility features LVF∈R C×h×w .

[0022] The input X-ray image data set described in step S1-1 specifically includes the following steps:

[0023] S1-1-1: X-ray image dataset Data =<X,Y,B,M> , where X represents the X-ray image set, Y represents the target category label set, B represents the target detection box position label set, and M represents the visibility mask label set;

[0024] S1-1-2: x∈X, x represents an image in the image set;

[0025] S1-1-3: y∈Y, y represents the category labels of all objects in an image, and the value of y is {0,…,5}, where 0 represents plastic bottles, 1 represents cans, 2 represents vacuum cups, 3 represents glass bottles, 4 represents paper cups, and 5 represents spray cans;

[0026] S1-1-4: b∈B, b represents the detection box position label of all targets in an image;

[0027] S1-1-5: m∈M, m represents the visibility mask label of an image. The label is pixel-level, and the element value is {0, 1}, where 0 represents occlusion and 1 represents visibility;

[0028] The calculation of the local block score described in step S1-4 specifically includes the following steps:

[0029] S1-4-1: Given local block feature LF j, calculate the sum of the eigenvalues ​​of all channels of the jth block and get the local block score LS j ;

[0030] S1-4-2: For each local block feature, call S1-4-1 to calculate the local block score, and obtain the local block score LS = {LS j}, LS∈R h×w .

[0031] The calculation of multiple local visibility masks described in step S1-5 specifically includes the following steps:

[0032] S1-5-1: Define a set of local visibility thresholds {lτ s}, where lτ s is a threshold, the value of s is {1, ..., S}, S represents the total number of local visibility thresholds;

[0033] S1-5-2: Using a local visibility threshold lτ s , the local block score LS∈R h×w Binarization to obtain a local visible mask LMask s ∈R h×w The formula is:

[0034]

[0035] S1-5-3: For each local visibility threshold {lτ s}, call S1-5-2 to obtain multiple local visibility masks, LMask = {LMask s};

[0036] The local visibility feature LVF∈R extracted in step S1-6 C×h×w , specifically including the following steps:

[0037] S1-6-1: LMask s As attention, weight the local block feature LF to get LVF′ s ,for:

[0038] LVF′ s =LMask s ⊙LF

[0039] Among them, ⊙ represents the element-by-element product;

[0040] S1-6-2: For each local visibility mask, LMask = {LMask s}, call S1-6-1 to obtain multiple local visibility features LVF = {LVFs};

[0041] S1-6-3: Merge multiple local visibility features to obtain a local visibility feature LVF, which is:

[0042]

[0043] Where S represents the total number of local visibility thresholds.

[0044] The step S2 of extracting the visibility features of the global X-ray image specifically includes the following steps:

[0045] S2-1: Calculate the values ​​of Query, Key, and Value in the attention mechanism;

[0046] S2-2: Calculate the attention weight attention∈R hw×hw ,for:

[0047]

[0048] Among them, the softmax function is used to normalize the similarity, and Trans is the transposition operation;

[0049] S2-3: Calculate global block features;

[0050] S2-4: Calculate the global block score:

[0051] S2-5: Calculate multiple global visibility masks;

[0052] S2-6: Extract global visibility feature GVF∈R C×h×w .

[0053] The values ​​of Query, Key, and Value in the calculation attention mechanism described in step S2-1 are as follows:

[0054] S2-1-1: Input local block features LF∈R extracted from S1-3 C×h×w ;

[0055] S2-1-2: Perform matrix dimension transformation on the local block feature LF to obtain LF′∈R C×hw ;

[0056] S2-1-3: Given W Q ∈R C×C , calculate Query, as:

[0057] Query = W Q LF′, Query∈R C×hw

[0058] S2-1-4: Given W K ∈R C×C , calculate the Key, as:

[0059] Key=W K LF′, Key∈R C×hw

[0060] S2-1-5: Given W V ∈R C×C , calculate Value, as:

[0061] Value = W V LF′, Value∈R C×hw ;

[0062] The calculation of the global block features described in step S2-3 is as follows:

[0063] S2-3-1: Multiply the attention weights attention and Value to calculate the flattened form of the global block feature, GF′∈R C×hw ,for:

[0064] GF′=attention.Value, GF′∈R C×hw

[0065] S2-3-2: Transform the matrix dimension of GF′ to obtain the global block feature GF∈R C×h×w ;

[0066] The calculation of the global block score described in step S2-4 is as follows:

[0067] S2-4-1: Given the global block feature GF j , j represents the subscript index of the block, whose value is {1, ..., hw}, and the sum of all channel feature values ​​of the jth block is calculated to obtain the global block score GS j ;

[0068] S2-4-2: For each global block feature, call S2-4-1 to obtain the global block score GS = {GS j}, GS∈R h×w .

[0069] The calculation of multiple global visibility masks described in step S2-5 is as follows:

[0070] S2-5-1: Define a set of global visibility thresholds {gτ s}, where gτ s is a threshold, the value of s is {1, ..., S}, S represents the total number of global visibility thresholds;

[0071] S2-5-2: Using the global visibility threshold gτ s , the global block score GS∈R h×w Binarization to obtain multiple global visibility masks GMask s ∈R h×w ,for:

[0072]

[0073] S2-5-3: For each global visibility threshold {gτ s}, call S2-5-2 to obtain multiple global visibility masks, GMask = {GMask s};

[0074] Step S2-6 extracts the global visibility feature GVF∈R C×h×w , as follows:

[0075] S2-6-1: GMask s As attention, weight the global block feature GF to obtain GVF′ s ,for:

[0076] GVF′ s =LMask s ⊙GF

[0077] Among them, ⊙ represents the element-by-element product;

[0078] S2-6-2: For each global visibility mask, GMask = {GMask s}, call S2-6-1, and obtain multiple global visibility features GVF′={GVF′ s};

[0079] S2-6-3: Merge multiple global visibility features GVF′ to obtain a global visibility feature GVF, which is:

[0080]

[0081] The step S3 of generating a prediction frame using local visibility features and global visibility features specifically includes the following steps:

[0082] S3-1: Combine the local visibility feature LVF and the global visibility feature GVF to obtain the hybrid visibility feature HVF∈R C×h×w ,for:

[0083]

[0084] S3-2: Input HVF into the prediction head of YOLOv8 to obtain the output of the prediction head, including the predicted category score ps, the predicted bounding box distance pd and the predicted boundary pb, and generate the prediction box of YOLOv8;

[0085] S3-3: Calculating multiple visibility mask losses mask ;

[0086] S3-4: Calculate the classification loss and regression loss of YOLOv8;

[0087] S3-5: Calculate the score gradient adjustment feature RHVF∈R C×h×w ;

[0088] S3-6: Input RHVF into the prediction head of YOLOv8 to obtain the output of the prediction head, including the predicted category score ps′, the predicted bounding box distance pd′ and the predicted bounding box pb′, and generate the gradient-adjusted prediction box;

[0089] S3-7: Calculate the classification loss and regression loss after gradient adjustment.

[0090] Step S3-3 calculates multiple visibility mask losses mask , generate the prediction box as follows:

[0091] S3-3-1: The pixel-level visibility mask label m∈R in S1-1-5 H×W , divided into blocks [k, k] with a height of k and a width of k, and the label values ​​of all pixels in the block are summed to obtain m′∈R h×w , let j be the subscript index of the block, whose value is {1,...,hw}, which is:

[0092] m′={m′ j}

[0093] in,

[0094] S3-3-2: Binarize m′ to obtain the visibility mask label vm∈R h×w ,for:

[0095]

[0096] Among them, when When , it means that more than half of the pixels in the block have a mask label value of 1, and the value of the visibility mask label is set to 1;

[0097] S3-3-3: Calculate visibility mask loss mask , the specific steps are as follows:

[0098] S3-3-3-1: Input the visibility mask label vm and various local visibility masks LMask in S1-5, and calculate the local visibility mask loss loss Lmask ,for:

[0099]

[0100] S3-3-3-2: Input the visibility mask label vm and various global visibility masks GMask in S2-5, and calculate the global visibility mask loss loss Gmask ,for:

[0101]

[0102] S3-3-3-3: Calculate multiple visibility mask losses mask ,for:

[0103] loss mask =loss Lmask +loss Gmask ;

[0104] Calculate YOLO as described in step S3-4 v The classification loss and regression loss of 8 are as follows:

[0105] S3-4-1: Input the predicted category score ps and category label y, and calculate the classification loss loss class ;

[0106] S3-4-2: Input the predicted bounding box distance pd predicted bounding box pb and the detection box position label b, and calculate the regression loss loss reg ;

[0107] The calculated score gradient adjustment feature RHVF∈R described in step S3-5 C×h×w , as follows:

[0108] S3-5-1: Convert the correct category label in the predicted box into a one-hot vector, hy∈R 1×6 ;

[0109] S3-5-2: Input the predicted category score ps and calculate the correct category score weight w1, which is:

[0110]

[0111] in, Represents gradient calculation;

[0112] S3-5-3: Convert the misleading category label with the highest score in the predicted box into a one-hot vector, hy * ∈R 1×6 ;

[0113] S3-5-4: Input the predicted category score ps and calculate the weight w2 of the misleading category with the highest score, which is:

[0114]

[0115] in, Represents gradient calculation;

[0116] S3-5-5: Use the category score gradient to adjust HVF to obtain the score gradient adjusted feature RHVF: RHVF = (1 + w1 - w2) · HVF;

[0117] The classification loss and regression loss after calculating the gradient adjustment described in step S3-7 are as follows:

[0118] S3-7-1: Input the predicted category score ps and category label y, and calculate the classification loss loss′ class ;

[0119] S3-7-2: Calculate the regression loss loss′ based on the predicted bounding box distance pd′, the predicted bounding box pb′ and the detection box position label b reg .

[0120] The training visibility attention model described in step S4 specifically includes the following steps:

[0121] S4-1: input X-ray image training set;

[0122] S4-2: Call S1 to extract local visibility features;

[0123] S4-3: Call S2 to extract global visibility features;

[0124] S4-4: Call S3-1 to merge local / global visibility features;

[0125] S4-5: Call S3-3 to calculate multiple visibility mask losses mask ;

[0126] S4-6: Call S3-4 to calculate the classification loss of YOLOv8 class And regression loss loss reg ;

[0127] S4-7-1: Call S3-7 to calculate the classification loss loss′ after gradient adjustment classand regression loss loss′ reg ;

[0128] S4-7-2: Calculate the total model loss Loss, which is:

[0129] Loss = loss mask +loss class +loss reg +loss′ class +loss′ reg

[0130] S4-8: training to obtain the optimal model parameters;

[0131] The step S5 of using the visibility attention model to detect the category of the X-ray image occluded target specifically includes the following steps:

[0132] S5-1: input X-ray image;

[0133] S5-2: Call S1 according to the optimal parameters of the model to calculate the local visibility features;

[0134] S5-3: Call S2 according to the optimal parameters of the model to calculate the global visibility features;

[0135] S5-4: Call S3-1 according to the optimal parameters of the model to merge local / global visibility features;

[0136] S5-5: According to the optimal parameters of the model, the merged features are input into the YOLOv8 detection head to generate a YOLOv8 prediction box;

[0137] S5-5: Detect the X-ray image occluded target.

[0138] The advantages of the present invention are as follows: the present invention uses YOLOv8 to extract local block features and calculate a variety of local visibility masks and features, and at the same time calculates the global relationship between local blocks, and generates a global visibility mask and features; in order to verify the accuracy of the visibility mask under the occluded target, when training the target detection model, the local-global visibility mask loss and the occlusion score gradient adjustment loss are additionally calculated, and by considering these losses, the visibility mask of the target is effectively analyzed, and the problem of missed detection of the occluded target is alleviated. BRIEF DESCRIPTION OF THE DRAWINGS

[0139] Figure 1 It is a flow chart of the X-ray image occluded target detection method based on local-global visibility analysis of the present invention;

[0140] Figure 2 The process of extracting local visibility features of X-ray images of the present invention;

[0141] Figure 3 The process of extracting the global visibility features of the X-ray image of the present invention;

[0142] Figure 4 The process of generating a prediction frame using visibility features of the present invention;

[0143] Figure 5 The process of training the visibility attention model of the present invention;

[0144] Figure 6 This is the process of detecting the category of occluded objects in X-ray images according to the present invention. DETAILED DESCRIPTION

[0145] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. The present invention is a method for detecting occluded objects in X-ray images based on local-global visibility analysis. The specific process is as follows: Figure 1 As shown, the implementation scheme of the present invention is divided into the following steps:

[0146] S1: Extract local visibility features of X-ray images, such as Figure 2 As shown;

[0147] S1-1: Input X-ray image dataset. The specific steps are as follows:

[0148] S1-1-1: X-ray image dataset Data =<X,Y,B,M> , where X represents the X-ray image set, Y represents the target category label set, B represents the target detection box position label set, and M represents the visibility mask label set;

[0149] S1-1-2: x∈X, x represents an image in the image set;

[0150] S1-1-3: y∈Y, y represents the category labels of all objects in an image, and the value of y is {0,…,5} (0 represents plastic bottles, 1 represents cans, 2 represents vacuum cups, 3 represents glass bottles, 4 represents paper cups, and 5 represents spray cans);

[0151] S1-1-4: b∈B, b represents the detection box position label of all targets in an image;

[0152] S1-1-5: m∈M, m represents the visibility mask label of an image. The label is pixel-level, and the element value is {0, 1} (0 represents occlusion, 1 represents visibility);

[0153] S1-2: Input the X-ray image x into the YOLOv8 model and extract the feature F∈R after the 21st layer processing C ×H×W, where C represents the number of channels, H represents height, and W represents width;

[0154] S1-3: Divide the feature F into blocks [k, k] with a height of k and a width of k, and sum the feature values ​​in each block to obtain the local block feature (Local Features) LF = {LF j}, LF∈R C×h×w , j represents the subscript index of the block, and its value is {1, ..., hw}, C represents the number of channels,

[0155] S1-4: Calculate the local block score. The specific steps are as follows:

[0156] S1-4-1: Given local block feature LF j , calculate the sum of the eigenvalues ​​of all channels of the jth block and get the local block score (Local Score) LS j ;

[0157] S1-4-2: For each local block feature, call S1-4-1 to calculate the local block score, and obtain the local block score LS = {LS j}, LS∈R h×w ;

[0158] S1-5: Calculate multiple local visibility masks. The specific steps are as follows:

[0159] S1-5-1: Define a set of local visibility thresholds {lτ s}, where lτ s is a threshold, the value of s is {1, ..., S}, S represents the total number of local visibility thresholds;

[0160] S1-5-2: Using a local visibility threshold lτ s , the local block score LS∈R h×w Binarization to obtain the local visibility mask (Local mask) LMask s ∈R h×w The formula is:

[0161]

[0162] S1-5-3: For each local visibility threshold {lτ s}, call S1-5-2 to obtain multiple local visibility masks, LMask = {LMask s};

[0163] S1-6: Extracting local visibility features (LVF∈R) C×h×w , the specific steps are as follows:

[0164] S1-6-1: LMask s As attention, weight the local block feature LF to get LVF′ s ,for:

[0165] LVF′ s =LMask s ⊙LF

[0166] Among them, ⊙ represents the element-by-element product;

[0167] S1-6-2: For each local visibility mask, LMask = {LMask s}, call S1-6-1 to obtain multiple local visibility features LVF = {LVF s};

[0168] S1-6-3: Merge multiple local visibility features to obtain a local visibility feature LVF, which is:

[0169]

[0170] Where S represents the total number of local visibility thresholds;

[0171] S2: Extract global visibility features of X-ray images, such as Figure 3 Shown

[0172] S2-1: Calculate Query, Key, and Value. The specific steps are as follows:

[0173] S2-1-1: Input local block features LF∈R extracted from S1-3 C×h×w ;

[0174] S2-1-2: Perform matrix dimension transformation on the local block feature LF to obtain LF′∈R C×hw ;

[0175] S2-1-3: Given W Q ∈R C×C , calculate Query, as:

[0176] Query = W Q LF′, Query∈R C×hw

[0177] S2-1-4: Given W K ∈R C×C , calculate the Key, as:

[0178] Key=W K LF′, Key∈R C×hw

[0179] S2-1-5: Given W V ∈R C×C , calculate Value, as:

[0180] Value = W V LF′, Value∈R C×hw

[0181] S2-2: Calculate the attention weight attention∈R hw×hw ,for:

[0182]

[0183] Among them, the softmax function is used to normalize the similarity, and Trans is the transposition operation;

[0184] S2-3: Calculate the global block features. The specific steps are as follows:

[0185] S2-3-1: Multiply attention and Value to calculate the flattened form of the global block feature, (GlobalFeatures′)GF′∈R C×hw ,for:

[0186] GF′=attention.Value, GF′∈R C×hw

[0187] S2-3-2: Transform the matrix dimension of GF′ to obtain the global block feature GF∈R C×h×w ;

[0188] S2-4: Calculate the global block score. The specific steps are as follows:

[0189] S2-4-1: Given the global block feature GF j , j represents the subscript index of the block, whose value is {1, ..., hw}, and the sum of all channel eigenvalues ​​of the jth block is calculated to obtain the global block score (Global Score) GS j ;

[0190] S2-4-2: For each global block feature, call S2-4-1 to obtain the global block score GS = {GS j}, GS∈R h×w ;

[0191] S2-5: Calculate multiple global visibility masks. The specific steps are as follows:

[0192] S2-5-1: Define a set of global visibility thresholds {gτ s}, where gτ s is a threshold, the value of s is {1, ..., S}, S represents the total number of global visibility thresholds;

[0193] S2-5-2: Using the global visibility threshold gτ s , the global block score GS∈R h×w Binarization to obtain various global visibility masks (Global mask) GMask s ∈R h×w ,for:

[0194]

[0195] S2-5-3: For each global visibility threshold {gτ s}, call S2-5-2 to obtain multiple global visibility masks, GMask = {GMask s};

[0196] S2-6: Extracting global visibility features GVF∈R C×h×w , the specific steps are as follows:

[0197] S2-6-1: GMask s As attention, weight the global block feature GF to obtain GVF′ s ,for:

[0198] GVF′ s =LMask s ⊙GF

[0199] Among them, ⊙ represents the element-by-element product;

[0200] S2-6-2: For each global visibility mask, GMask = {GMask s}, call S2-6-1, and obtain multiple global visibility features GVF′={GVF′ s};

[0201] S2-6-3: Merge multiple global visibility features GVF′ to obtain a global visibility feature GVF, which is:

[0202]

[0203] S3: Generate prediction boxes using visibility features, such as Figure 4As shown;

[0204] S3-1: Combine the local visibility feature LVF and the global visibility feature GVF to obtain the hybrid visibility feature HVF∈R C×h×w ,for:

[0205]

[0206] S3-2: Input HVF into the prediction head of YOLOv8 to obtain the output of the prediction head, including the predicted category score ps, the predicted bounding box distance pd, and the predicted boundary pb, and generate the prediction box of YOLOv8;

[0207] S3-3: Calculating multiple visibility mask losses mask , the specific steps are as follows:

[0208] S3-3-1: The pixel-level visibility mask label m∈R in S1-1-5 H×W , divided into blocks [k, k] with a height of k and a width of k, and the label values ​​of all pixels in the block are summed to obtain m′∈R h×w , let j be the subscript index of the block, whose value is {1,...,hw}, which is:

[0209] m′={m′ j}

[0210] in,

[0211] S3-3-2: Binarize m′ to obtain the visibility mask label vm∈R h×w ,for:

[0212]

[0213] Among them, when When , it means that more than half of the pixels in the block have a mask label value of 1, and the value of the visibility mask label is set to 1;

[0214] S3-3-3: Calculate visibility mask loss mask , the specific steps are as follows:

[0215] S3-3-3-1: Input the visibility mask label vm and various local visibility masks, LMask in S1-5, and calculate the local visibility mask loss loss Lmask ,for:

[0216]

[0217] S3-3-3-2: Input the visibility mask label vm and various global visibility masks, GMask in S2-5, and calculate the global visibility mask loss loss Gmask ,for:

[0218]

[0219] S3-3-3-3: Calculate multiple visibility mask losses mask ,for:

[0220] loss mask =loss Lmask +loss Gmask

[0221] S3-4: Calculate the classification loss and regression loss of YOLOv8. The specific steps are as follows:

[0222] S3-4-1: Input the predicted category score ps and category label y, and calculate the classification loss loss class ;

[0223] S3-4-2: Input the predicted bounding box distance pd, the predicted bounding box pb, and the detection box position label b, and calculate the regression loss loss reg ;

[0224] S3-5: Calculate the score gradient adjustment feature RHVF∈R C×h×w , the specific steps are as follows:

[0225] S3-5-1: Convert the correct category label in the predicted box into a one-hot vector, hy∈R 1×6 ;

[0226] S3-5-2: Input the predicted category score ps and calculate the correct category score weight w1, which is:

[0227]

[0228] in, Represents gradient calculation;

[0229] S3-5-3: Convert the misleading category label with the highest score in the predicted box into a one-hot vector, hy * ∈R 1×6 ;

[0230] S3-5-4: Input the predicted category score ps and calculate the weight w2 of the misleading category with the highest score, which is:

[0231]

[0232] in, Represents gradient calculation;

[0233] S3-5-5: Use the category score gradient to adjust HVF and obtain the score gradient adjusted feature RHVF:

[0234] RHVF=(1+w1-w2)·HVF

[0235] S3-6: Input RHVF into the prediction head of YOLOv8 to obtain the output of the prediction head, including the predicted category score ps′, the predicted bounding box distance pd′, and the predicted bounding box pb′, and generate the gradient-adjusted prediction box;

[0236] S3-7: Calculate the classification loss and regression loss after gradient adjustment. The specific steps are as follows:

[0237] S3-7-1: Input the predicted category score ps and category label y, and calculate the classification loss loss′ class ;

[0238] S3-7-2: Calculate the regression loss loss′ based on the predicted bounding box distance pd′, the predicted bounding box pb′, and the detection box position label b reg ;

[0239] S4: Train the visibility attention model, such as Figure 5 As shown;

[0240] S4-1: input X-ray image training set;

[0241] S4-2: Call S1 to extract local visibility features;

[0242] S4-3: Call S2 to extract global visibility features;

[0243] S4-4: Call S3-1 to merge local / global visibility features;

[0244] S4-5: Call S3-3 to calculate multiple visibility mask losses mask ;

[0245] S4-6: Call S3-4 to calculate the classification loss of YOLOv8 class And regression loss loss reg ;

[0246] S4-7-1: Call S3-7 to calculate the classification loss loss′ after gradient adjustment class and regression loss loss′ reg ;

[0247] S4-7-2: Calculate the total model loss Loss, which is:

[0248] Loss = loss mask +loss class +loss reg +loss′ class +loss′ reg

[0249] S4-8: training to obtain the optimal model parameters;

[0250] S5: Detect the occluded target category in the X-ray image, such as Figure 6 As shown;

[0251] S5-1: input X-ray image;

[0252] S5-2: Call S1 according to the optimal parameters of the model to calculate the local visibility features;

[0253] S5-3: Call S2 according to the optimal parameters of the model to calculate the global visibility features;

[0254] S5-4: Call S3-1 according to the optimal parameters of the model to merge local / global visibility features;

[0255] S5-5: According to the optimal parameters of the model, the merged features are input into the YOLOv8 detection head to generate a YOLOv8 prediction box;

[0256] S5-5: Detect the X-ray image occluded target.

Claims

1. A method for detecting occluded objects in X-ray images based on local-global visibility analysis, characterized in that: The specific steps include: S1: Extract local visibility features of X-ray images; S2: Extract global visibility features of X-ray images; S3: Generate prediction box using local visibility features and global visibility features; S4: Training visibility attention model; S5: Detecting occluded target categories in X-ray images using visibility attention model; The step S1 of extracting the local visibility features of the X-ray image specifically includes the following steps: S1-1: Input X-ray image dataset; S1-2: Input the X-ray image x into the YOLOv8 model and extract the feature F∈R after the 21st layer processing C×H×W , where C represents the number of channels, H represents height, and W represents width; S1-3: Divide the feature F into blocks [k, k] with a height of k and a width of k, and sum the feature values ​​in each block to obtain the local block feature LF = {LF j }, LF∈R C×h×w , j represents the subscript index of the block, and its value is {1, ..., hw}, C represents the number of channels, S1-4: Calculate local block score; S1-5: Calculate multiple local visibility masks; S1-6: Extract local visibility features LVF∈R C×h×w ; The step S2 of extracting the visibility features of the global X-ray image specifically includes the following steps: S2-1: Calculate the values ​​of Query, Key, and Value in the attention mechanism; S2-2: Calculate the attention weight attention∈R hw×hw ,for: Among them, the softmax function is used to normalize the similarity, and Trans is the transposition operation; S2-3: Calculate global block features; S2-4: Calculate the global block score: S2-5: Calculate multiple global visibility masks; S2-6: Extract global visibility feature GVF∈R C×h×w ; The step S3 of generating a prediction frame using local visibility features and global visibility features specifically includes the following steps: S3-1: Combine the local visibility feature LVF and the global visibility feature GVF to obtain the hybrid visibility feature HVF∈R C×h×w ,for: S3-2: Input HVF into the prediction head of YOLOv8 to obtain the output of the prediction head, including the predicted category score ps, the predicted bounding box distance pd and the predicted boundary pb, and generate the prediction box of YOLOv8; S3-3: Calculate multiple visibility mask losses mask ; S3-4: Calculate the classification loss and regression loss of YOLOv8; S3-5: Calculate the score gradient adjustment feature RHVF∈R C×h×w ; S3-6: Input RHVF into the prediction head of YOLOv8 to obtain the output of the prediction head, including the predicted category score ps′, the predicted bounding box distance pd′ and the predicted bounding box pb′, and generate the gradient-adjusted prediction box; S3-7: Calculate the classification loss and regression loss after gradient adjustment.

2. The method for detecting obstructed objects in X-ray images based on local-global visibility analysis according to claim 1, characterized in that: The input X-ray image data set described in step S1-1 specifically includes the following steps: S1-1-1: X-ray image dataset Data =<X,Y,B,M> , where X represents the X-ray image set, Y represents the target category label set, B represents the target detection box position label set, and M represents the visibility mask label set; S1-1-2: x∈X, x represents an image in the image set; S1-1-3: y∈Y, y represents the category labels of all objects in an image, and the value of y is {0,…,5}, where 0 represents plastic bottles, 1 represents cans, 2 represents vacuum cups, 3 represents glass bottles, 4 represents paper cups, and 5 represents spray cans; S1-1-4: b∈B, b represents the detection box position label of all targets in an image; S1-1-5: m∈M, m represents the visibility mask label of an image. The label is pixel-level, and the element value is {0, 1}, where 0 represents occlusion and 1 represents visibility; The calculation of the local block score described in step S1-4 specifically includes the following steps: S1-4-1: Given local block feature LF j , calculate the sum of the eigenvalues ​​of all channels of the jth block and get the local block score LS j ; S1-4-2: For each local block feature, call S1-4-1 to calculate the local block score, and obtain the local block score LS = {LS j }, LS∈R h×w .

3. The method for detecting obstructed objects in X-ray images based on local-global visibility analysis according to claim 2, characterized in that: The calculation of multiple local visibility masks described in step S1-5 specifically includes the following steps: S1-5-1: Define a set of local visibility thresholds {lτ s }, where lτ s is a threshold, the value of s is {1, ..., S}, S represents the total number of local visibility thresholds; S1-5-2: Using a local visibility threshold lτ s , the local block score LS∈R h×w Binarization, get the local visible mask, LMask s ∈R h×w The formula is: S1-5-3: For each local visibility threshold {lτ s }, call S1-5-2 to obtain multiple local visibility masks, LMask = {LMask s }; The local visibility feature LVF∈R extracted in step S1-6 C×h×w , specifically including the following steps: S1-6-1: LMask s As attention, weight the local block feature LF to get LVF′ s ,for: LVF′ s =LMask s ⊙LF Among them, ⊙ represents the element-by-element product; S1-6-2: For each local visibility mask, LMask = {LMask s }, call S1-6-1 to obtain multiple local visibility features LVF = {LVF s }; S1-6-3: Merge multiple local visibility features to obtain a local visibility feature LVF, which is: Where S represents the total number of local visibility thresholds.

4. The method for detecting obstructed objects in X-ray images based on local-global visibility analysis according to claim 3, characterized in that: The values ​​of Query, Key, and Value in the calculation attention mechanism described in step S2-1 are as follows: S2-1-1: Input local block features LF∈R extracted from S1-3 C×h×w ; S2-1-2: Perform matrix dimension transformation on the local block feature LF to obtain LF′∈R C×hw ; S2-1-3: Given W Q ∈R C×C , calculate Query, as: Query=W Q ·LF′,Query∈R C×hw S2-1-4: Given W K ∈R C×C , calculate the Key, as: Key=W K ·LF′,Key∈R C×hw S2-1-5: Given W V ∈R C×C , calculate Value, as: Value=W V ·LF′,Value∈R C×hw ; The calculation of the global block features described in step S2-3 is as follows: S2-3-1: Multiply the attention weights attention and Value to calculate the flattened form of the global block feature, GF′∈R C×hw ,for: GF′=attention·Value,GF′∈R C×hw S2-3-2: Transform the matrix dimension of GF′ to obtain the global block feature GF∈R C×h×w ; The calculation of the global block score described in step S2-4 is as follows: S2-4-1: Given the global block feature GF j , j represents the subscript index of the block, whose value is {1, ..., hw}, and the sum of all channel feature values ​​of the jth block is calculated to obtain the global block score GS j ; S2-4-2: For each global block feature, call S2-4-1 to obtain the global block score GS = {GS j }, GS∈R h×w .

5. The method for detecting X-ray image occluded objects based on local-global visibility analysis according to claim 4, characterized in that: The calculation of multiple global visibility masks described in step S2-5 is as follows: S2-5-1: Define a set of global visibility thresholds {gτ s }, where gτ s is a threshold, the value of s is {1, ..., S}, S represents the total number of global visibility thresholds; S2-5-2: Using the global visibility threshold gτ s , the global block score GS∈R h×w Binarization, to obtain a variety of global visibility masks, GMask s ∈R h×w ,for: S2-5-3: For each global visibility threshold {gτ s }, call S2-5-2 to obtain multiple global visibility masks, GMask = {GMask s }; Step S2-6 extracts the global visibility feature GVF∈R C×h×w , as follows: S2-6-1: GMask s As attention, weight the global block feature GF to obtain GVF′ s ,for: <h2 style=";text-align:left;direction:ltr">GVF′<h2 style=";text-align:left;direction:ltr"> s <h2 style=";text-align:left;direction:ltr"> =LMas k<h2 style=";text-align:left;direction:ltr"> s <h2 style=";text-align:left;direction:ltr"> ⊙GF Among them, ⊙ represents the element-by-element product; S2-6-2: For each global visibility mask, GMask = {GMask s }, call S2-6-1 to obtain multiple global visibility features GVF′={GVF′s}; S2-6-3: Merge multiple global visibility features GVF′ to obtain a global visibility feature GVF, which is:

6. The method for detecting obstructed objects in X-ray images based on local-global visibility analysis according to claim 5, characterized in that: Step S3-3 calculates multiple visibility mask losses mask , generate the prediction box as follows: S3-3-1: The pixel-level visibility mask label m∈R in S1-1-5 H×W , divided into blocks [k, k] with a height of k and a width of k, and the label values ​​of all pixels in the block are summed to obtain m′∈R h×w , let j be the subscript index of the block, whose value is {1,...,hw}, which is: m′={m′ j } in, S3-3-2: Binarize m′ to obtain the visibility mask label vm∈R h×w ,for: Among them, when When , it means that more than half of the pixels in the block have a mask label value of 1, and the value of the visibility mask label is set to 1; S3-3-3: Calculate visibility mask loss mask , the specific steps are as follows: S3-3-3-1: Input the visibility mask label vm and multiple local visibility masks in S1-5, LMask, and calculate the local visibility mask loss loss Lmask ,for: S3-3-3-2: Input the visibility mask label vm and multiple global visibility masks in S2-5, GMask, and calculate the global visibility mask loss loss Gmask ,for: S3-3-3-3: Calculate multiple visibility mask losses mask ,for: loss mask =loss Lmask +loss Gmask; The classification loss and regression loss of YOLOv8 are calculated as described in step S3-4 as follows: S3-4-1: Input the predicted category score ps and category label y, and calculate the classification loss loss class ; S3-4-2: Input the predicted bounding box distance pd predicted bounding box pb and the detection box position label b, and calculate the regression loss loss reg ; The calculated score gradient adjustment feature RHVF∈R described in step S3-5 C×h×w , as follows: S3-5-1: Convert the correct category label in the predicted box into a one-hot vector, hy∈R 1×6 ; S3-5-2: Input the predicted category score ps and calculate the correct category score weight w1, which is: in, Represents gradient calculation; S3-5-3: Convert the misleading category label with the highest score in the predicted box into a one-hot vector, hy*∈R 1×6 ; S3-5-4: Input the predicted category score ps and calculate the weight w2 of the misleading category with the highest score, which is: in, Represents gradient calculation; S3-5-5: Use the category score gradient to adjust HVF to obtain the score gradient adjusted feature RHVF: RHVF = (1 + w1 - w2) · HVF; The classification loss and regression loss after calculating the gradient adjustment described in step S3-7 are as follows: S3-7-1: Input the predicted category score ps and category label y, and calculate the classification loss loss′ class ; S3-7-2: Calculate the regression loss loss′ based on the predicted bounding box distance pd′, the predicted bounding box pb′ and the detection box position label b reg .

7. The method for detecting X-ray image occluded objects based on local-global visibility analysis according to claim 6, characterized in that: The training of the visibility attention model in step S4 specifically includes the following steps: S4-1: input X-ray image training set; S4-2: Call S1 to extract local visibility features; S4-3: Call S2 to extract global visibility features; S4-4: Call S3-1 to merge local / global visibility features; S4-5: Call S3-3 to calculate multiple visibility mask losses mask ; S4-6: Call S3-4 to calculate the classification loss of YOLOv8 class And regression loss loss reg ; S4-7-1: Call S3-7 to calculate the classification loss loss′ after gradient adjustment class and regression loss loss′ reg ; S4-7-2: Calculate the total model loss Loss, which is: Loss=loss mask +loss class +loss reg +loss′ class +loss′ reg S4-8: training to obtain the optimal model parameters; The step S5 of using the visibility attention model to detect the category of the X-ray image occluded target specifically includes the following steps: S5-1: input X-ray image; S5-2: Call S1 according to the optimal parameters of the model to calculate the local visibility features; S5-3: Call S2 according to the optimal parameters of the model to calculate the global visibility features; S5-4: Call S3-1 according to the optimal parameters of the model to merge local / global visibility features; S5-5: According to the optimal parameters of the model, the merged features are input into the YOLOv8 detection head to generate a YOLOv8 prediction box; S5-5: Detect the X-ray image occluded target.

Citation Information

Patent Citations

  • X-ray prohibited items detection method based on RGB color separation and dual-path feature fusion

    CN114548230B

  • Method and device for detecting prohibited items in complex environments based on multi-scale feature fusion

    CN117765378B

  • X-ray security check double-view detection method based on improved FCOS network

    CN118053122A

  • Target detection method based on multi-view feature aggregation

    CN118154854A