X-ray security inspection dual-view detection method based on improved FCOS network

By improving the FCOS network and combining the image encoder with the multi-scale fusion encoder, the problem of low detection accuracy caused by multi-object stacking and complex background in X-ray security inspection is solved, and efficient and accurate contraband detection is achieved.

CN118053122BActive Publication Date: 2025-09-26SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410214986.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-27
Publication Date
2025-09-26
Estimated Expiration
2044-02-27

AI Technical Summary

Technical Problem

Existing X-ray security inspection image detection methods have low detection accuracy when faced with multiple objects stacked and complex backgrounds, and traditional deep learning algorithms find it difficult to effectively utilize multi-view information, resulting in missed detections and false detections.

Method used

An improved FCOS network is adopted, combined with an image encoder, an HSV-guided encoder and a multi-scale fusion encoder. Through color space conversion and feature fusion of top and side views, multi-scale features are extracted and fused to improve detection accuracy and robustness.

Benefits of technology

It achieves the goal of significantly improving detection accuracy and efficiency while maintaining detection speed, reducing missed detection rate, and enhancing the ability to identify contraband and the adaptability of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118053122B_ABST
    Figure CN118053122B_ABST
Patent Text Reader

Abstract

The present invention discloses an X-ray security inspection dual-view detection method based on an improved FCOS network. The improvement of the FCOS network lies in the addition of an image encoder, an HSV-guided encoder and a multi-scale fusion encoder. The image encoder converts the color space of the input top-view X-ray image and the side-view X-ray image from RGB to HSV before feature extraction and fusion. The HSV-guided encoder fuses the HSV-guided features extracted by the image encoder with the input original X-ray top view and side view respectively. The multi-scale fusion encoder is used to extract multi-scale features of the top view and the side view, and fuse the multi-scale features of the two views. The present invention reduces the interference of stacked targets under X-ray transmission imaging, makes full use of image information in different color spaces, strengthens the network feature perception capability, and improves the detection accuracy and speed of X-ray objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of X-ray contraband detection, and in particular to an X-ray security inspection dual-viewing angle detection method based on an improved FCOS network. Background Art

[0002] X-ray security inspection technology plays a vital role in security checks in public transportation, logistics, and express delivery. However, most security inspections currently rely on manual labor, which is susceptible to various unstable factors, resulting in low efficiency and a high risk of missed or misdetected prohibited items. Therefore, deep learning-based X-ray security inspection image detection methods hold great potential. However, X-ray security inspection images suffer from issues such as stacked objects, cluttered backgrounds, and multi-scale representation of prohibited items. These issues severely interfere with detection and represent a bottleneck in the current development of X-ray security inspection technology.

[0003] In this context, dual-view detection has emerged as a new approach to addressing image interference in X-ray security inspections. By utilizing X-ray images acquired from multiple viewpoints, a more comprehensive view of all aspects of the inspected object can be achieved, thereby improving detection accuracy. Dual-view detection not only mitigates the effects of image overlap and background clutter, but also compensates for potential blind spots in single-view detection, further enhancing detection performance.

[0004] Existing X-ray security inspection image detection methods primarily utilize deep learning algorithms, including two-stage or single-stage detection algorithms. These algorithms typically first capture X-ray security inspection images of baggage and parcels and perform image preprocessing, including color space conversion and resizing. They then manually annotate the images with item labels to construct a security inspection dataset. The dataset is then divided into training and test sets. The training set is then fed into a deep learning network for training to produce a detection model. Finally, the test set is used to validate the network's detection performance.

[0005] However, traditional deep learning algorithms are primarily designed for visible spectrum images and cannot be directly applied to X-ray security images. The imaging principles of X-ray security images differ significantly from those of visible spectrum images. Therefore, traditional detection algorithms often face challenges such as high noise levels, blurred edges, complex backgrounds, and stacked objects, resulting in low detection accuracy.

[0006] Therefore, to address the existing challenges in X-ray security image detection, a new deep recognition network for contraband inspection needs to be designed, incorporating dual-view detection technology. This new network can better utilize information from multiple viewpoints, overcome interference from X-ray security images, and improve detection accuracy and efficiency, thereby better meeting the needs of security inspections. Summary of the Invention

[0007] The present invention mainly addresses the problems of low detection accuracy of X-ray images of multi-object stacking in the existing technology and insufficient utilization of image information in single-view detection. It proposes an X-ray security inspection dual-view detection method based on an improved FCOS network, which can realize fast and accurate detection of X-ray images, reduce working costs, improve detection efficiency and detection accuracy of intelligent security inspection systems.

[0008] In order to achieve the above-mentioned purpose, the technical solution provided by the present invention is: an X-ray security inspection dual-view detection method based on an improved FCOS network, which realizes accurate detection of contraband in dual-view X-ray security inspection based on the improved FCOS network. The improvement of the FCOS network lies in the addition of an image encoder, an HSV-guided encoder and a multi-scale fusion encoder. The image encoder converts the color space of the input X-ray top view and X-ray side view from RGB to HSV before feature extraction and fusion, which is used to improve the recognition sensitivity of the FCOS network for the target to be detected; the HSV-guided encoder fuses the HSV-guided features extracted by the image encoder with the input original X-ray top view and side view respectively, thereby reducing the aliasing interference of the background in the X-ray image and improving the recognition ability of the FCOS network for contraband; the multi-scale fusion encoder is used to extract multi-scale features of the top view and side view, and fuse the multi-scale features of the two viewpoints, thereby enhancing the recognition ability of the FCOS network for targets of different sizes to be detected;

[0009] The specific implementation of the X-ray security inspection dual-viewing angle detection method includes:

[0010] Collect X-ray top view and X-ray side view of the object to be inspected on the conveyor belt;

[0011] The trained improved FCOS network is used to perform the following processing on the X-ray top view and X-ray side view:

[0012] The X-ray top view and the X-ray side view are input into the image encoder and first subjected to color space transformation to obtain the top view HSV image T and the side view HSV image S. The top view HSV image T and the side view HSV image S are then subjected to feature extraction and fusion to obtain the HSV guided feature E.

[0013] The HSV guidance encoder fuses and encodes the HSV guidance feature E with the X-ray top view to obtain the top-view guidance feature G; the HSV guidance feature E is fused and encoded with the X-ray side view to obtain the side-view guidance feature B.

[0014] The multi-scale fusion encoder extracts and fuses the top-view guided feature G and the side-view guided feature B to obtain the global multi-scale feature M;

[0015] The global multi-scale feature M is input into the feature extraction backbone network of the improved FCOS network to obtain the backbone feature, which is then input into the prediction module of the improved FCOS network for target positioning and classification, and finally the detection information is output and marked in the X-ray image.

[0016] Furthermore, collecting the X-ray top view and the X-ray side view includes the following operations:

[0017] The object to be inspected is placed on a tray. When the tray is transported to the inspection area by a conveyor belt, the multi-view X-ray instrument emits an X-ray beam from directly above the object and on the right side of the conveyor belt's travel direction to scan the object to be inspected. The X-ray beam passes through the object to be inspected and reaches the receiver, which is rendered by a computer program to obtain an X-ray top view and an X-ray side view of the object to be inspected.

[0018] Furthermore, the improved FCOS network includes:

[0019] An image encoder is divided into a color space conversion module for calculating a top-view HSV image and a side-view HSV image based on the X-ray top view and the X-ray side view, and an image encoding module for calculating HSV guide features based on the top-view HSV image and the side-view HSV image;

[0020] An HSV guidance encoder is divided into a top-view guidance module for fusing HSV guidance features with an X-ray top view and encoding and calculating top-view guidance features, and a side-view guidance module for fusing HSV guidance features with an X-ray side view and encoding and calculating side-view guidance features;

[0021] a multi-scale fusion encoder comprising a top-view multi-scale module for extracting top-view multi-scale features based on top-view guided features, a side-view multi-scale module for extracting side-view multi-scale features based on side-view guided features, and a cross-attention module for calculating global multi-scale features based on the top-view multi-scale features and the side-view multi-scale features;

[0022] Feature extraction backbone network, consisting of 5 computational layers and a residual structure;

[0023] The prediction module consists of four linear convolutional layers, ReLU activation function and Group Normalization. The use of Group Normalization can effectively increase the robustness of the network.

[0024] Furthermore, the color space conversion module transforms the input X-ray top view into a top-view HSV image and the input X-ray side view into a side-view HSV image through color space transformation. The calculation process is as follows:

[0025] T=trans(X)

[0026] S=trans(Y)

[0027] Where X is the input X-ray top view, Y is the input X-ray side view, trans() represents the color space transformation that transforms the RGB channels into HSV channels, T is the top view HSV image, and S is the side view HSV image;

[0028] The image encoding module consists of multiple convolutional layers, gated convolutional layers, pooling layers, batch normalization layers, and nonlinear activation layers. It takes the top-view HSV image T and the side-view HSV image S output by the color space conversion module as input and finally obtains the HSV guided feature E through calculation. The calculation process is as follows:

[0029] f ac =σ(W ac ·S+b ac )

[0030] E=σ(W e ·T+U e ·(f ac ⊙avgpool(T))+b e )

[0031] Where W ac and W e represents the convolution kernel parameters, b ac and b e represents the offset distance, σ represents the nonlinear activation, f ac represents the gated feature map, U e Represents the gated kernel of the gated convolution kernel, ⊙ represents element-wise multiplication, avgpool represents the average pooling operation, and the HSV guided feature E improves the recognition sensitivity of the FCOS network for contraband.

[0032] Furthermore, the top-view guidance module is composed of multiple convolutional layers, gated convolutional layers, batch normalization layers, and nonlinear activation layers. After obtaining the HSV guidance features, the top-view guidance module inputs the HSV guidance features and the X-ray top view, and finally obtains the top-view guidance features G through calculation, which is calculated as follows:

[0033] f gt1 =σ(W gt1 E+b gt1 )

[0034] f x =σ(W x ·X+b x )

[0035] G=σ(W g ·f x +U g ·(fgt1 ⊙f x )+b g )

[0036] Where W gt1 , W x and W g represents the convolution kernel parameters, b gt1 , b x and b g represents the offset distance, σ represents the nonlinear activation, f gt1 represents the gated activation map, X is the X-ray top view of the input, and f x Indicates the middle feature of the X-ray top view, U g represents the gated convolution kernel, ⊙ represents element-wise multiplication, and the top-down view guided feature G improves the recognition sensitivity of the FCOS network for contraband;

[0037] The side view guidance module inputs the HSV guidance feature and the X-ray side view, and finally obtains the side view guidance feature B through calculation, which is calculated as follows:

[0038] f gt2 =σ(W gt2 E+b gt2 )

[0039] f y =σ(W y ·Y+b y )

[0040] B=σ(W o ·f y +U o ·(f gt2 ⊙f y )+b o )

[0041] Where W gt2 、W y and W o represents the convolution kernel parameters, b gt2 、b y and b o represents the offset distance, σ represents the nonlinear activation, f gt2 represents the gated activation map, Y is the input X-ray side view, f y Indicates the middle feature of the X-ray side view, U o Represents the gated kernel of the gated convolution kernel, ⊙ represents element-by-element multiplication, and the side view guided feature B reduces the interference of the transmissive occlusion in the image on the FCOS network.

[0042] Furthermore, the bird's-eye view multi-scale module takes the bird's-eye view guided feature G as input and extracts the bird's-eye view multi-scale feature I through multiple convolutional layers. The calculation is expressed as follows:

[0043] I=σ(W i G+b i )

[0044] Where σ represents the Sigmoid function, W i represents the convolution kernel, b i Represents the offset distance. The multi-scale feature I of the top-down view improves the recognition sensitivity of the FCOS network for contraband of different sizes.

[0045] The side view multi-scale module takes the side view guided feature B as input and extracts the side view multi-scale feature P through multiple convolutional layers. The calculation is expressed as follows:

[0046] P=σ(W p B+b p )

[0047] Where σ represents the Sigmoid function, W p represents the convolution kernel, b p Represents the offset distance. The side view multi-scale feature P reduces the interference of non-contraband items of different sizes on the FCOS network.

[0048] The cross attention module includes multiple convolutional layers, softmax layers, batch normalization layers, and nonlinear activation layers. First, the top-view multi-scale feature I is input into the cross attention module, and the query feature Q is calculated through the convolution layer. The calculation is expressed as follows:

[0049] Q=W q I+b q

[0050] Where W q represents the convolution kernel, b q Indicates the offset distance;

[0051] Then the side view multi-scale feature P is input into the cross attention module, and the key feature K and value feature V are calculated through different convolutional layers. The calculation is expressed as follows:

[0052] K=W k P+b k

[0053] V=W v P+b v

[0054] Where W k and W v represents the convolution kernel, b k and b v Indicates the offset distance;

[0055] After the query feature Q is dot-multiplied with the key feature K, the attention weight A is obtained through the Softmax layer. The attention weight A is dot-multiplied with the value feature V and then added to the top-down multi-scale feature I to obtain the global multi-scale feature M. The calculation is expressed as follows:

[0056]

[0057] M=A·V+I

[0058] Where, d k Represents the scaling factor of the attention weight, which is set to the feature dimension of the query feature Q.

[0059] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0060] 1. Compared with other X-ray target detection methods, the present invention improves detection accuracy while maintaining detection speed. The proposed method sets an image encoder to mine more discriminative features in the HSV color space, and then passes the significant features to the network for fusion, thereby reducing feature interference under X-ray transmission imaging when multiple objects are stacked. The improved FCOS network proposed in the present invention has higher detection accuracy.

[0061] 2. The present invention combines the HSV guided encoder in X-ray detection to extract the image information of the target to be detected in different color spaces, and then fuses the feature information in different color spaces, introducing additional effective information, effectively enhancing the network's perception ability of the object to be detected, reducing the missed detection rate, and improving detection efficiency.

[0062] 3. The present invention sets up a multi-scale fusion encoder, which can simultaneously extract feature information at different scales, including detail information and global information. By performing feature extraction in networks at different levels, it can capture important features of different scales in the image, thereby more comprehensively describing the image content and improving the detection performance and robustness of the model.

[0063] 4. The method of the present invention has wide applicability in computer vision tasks, can achieve end-to-end training, has strong adaptability to various scenarios and tasks, and shows broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 2 is a general framework diagram of the method of the present invention; in the figure, Concatenation represents a concatenation operation.

[0065] Figure 2 This is a top view of the test image used in this embodiment.

[0066] Figure 3This is a side view of the test picture used in this embodiment.

[0067] Figure 4 This is the top view detection result obtained in this embodiment.

[0068] Figure 5 This is the side view detection result obtained in this embodiment. DETAILED DESCRIPTION

[0069] The present invention will be further described below with reference to specific embodiments.

[0070] like Figure 1 As shown, this embodiment discloses an X-ray security inspection dual-view detection method based on an improved FCOS network. The method is based on the improved FCOS network to achieve accurate detection of contraband in X-ray dual-view security inspection. The improvement of the FCOS network lies in the addition of an image encoder, an HSV-guided encoder, and a multi-scale fusion encoder. The image encoder converts the color space of the input X-ray top view and X-ray side view from RGB to HSV before feature extraction and fusion, which is used to improve the recognition sensitivity of the FCOS network for the target to be detected; the HSV-guided encoder fuses the HSV-guided features extracted by the image encoder with the input original X-ray top view and side view respectively, thereby reducing the aliasing interference of the background in the X-ray image and improving the FCOS network's recognition ability for contraband; the multi-scale fusion encoder is used to extract multi-scale features of the top view and side view, and fuse the multi-scale features of the two views, thereby enhancing the FCOS network's recognition ability for targets of different sizes to be detected and improving the detection accuracy of the FCOS network.

[0071] The specific implementation of the X-ray security inspection dual-viewing angle detection method includes the following steps:

[0072] 1) Place the object to be inspected with the dagger on a tray. When the tray is transported to the inspection area by the conveyor belt, the multi-view X-ray instrument emits an X-ray beam from above the object and on the right side of the conveyor belt to scan the object. The X-ray beam passes through the object to be inspected and reaches the receiver. The X-ray top view and X-ray side view of the object to be inspected are rendered by a computer program, as shown in the figure below. Figure 2 and Figure 3 shown.

[0073] 2) Use the trained improved FCOS network to perform the following processing on the X-ray top view and X-ray side view:

[0074] The X-ray top view and the X-ray side view are input into the image encoder and first subjected to color space transformation to obtain the top view HSV image T and the side view HSV image S. The top view HSV image T and the side view HSV image S are then subjected to feature extraction and fusion to obtain the HSV guided feature E.

[0075] The HSV guidance encoder fuses and encodes the HSV guidance feature E with the X-ray top view to obtain the top-view guidance feature G; the HSV guidance feature E is fused and encoded with the X-ray side view to obtain the side-view guidance feature B.

[0076] The multi-scale fusion encoder extracts and fuses the top-view guided feature G and the side-view guided feature B to obtain the global multi-scale feature M;

[0077] The global multi-scale feature M is input into the feature extraction backbone network of the improved FCOS network to obtain the backbone feature, which is then input into the prediction module of the improved FCOS network to locate and classify the target. Finally, the detection information is output and marked in the X-ray top view and X-ray side view, as shown in the figure. Figure 4 and Figure 5 shown.

[0078] Specifically, the improved FCOS network includes:

[0079] An image encoder is divided into a color space conversion module for calculating a top-view HSV image and a side-view HSV image based on the X-ray top view and the X-ray side view, and an image encoding module for calculating HSV guide features based on the top-view HSV image and the side-view HSV image;

[0080] An HSV guidance encoder is divided into a top-view guidance module for fusing HSV guidance features with an X-ray top view and encoding and calculating top-view guidance features, and a side-view guidance module for fusing HSV guidance features with an X-ray side view and encoding and calculating side-view guidance features;

[0081] a multi-scale fusion encoder comprising a top-view multi-scale module for extracting top-view multi-scale features based on top-view guided features, a side-view multi-scale module for extracting side-view multi-scale features based on side-view guided features, and a cross-attention module for calculating global multi-scale features based on the top-view multi-scale features and the side-view multi-scale features;

[0082] Feature extraction backbone network, consisting of 5 computational layers and a residual structure;

[0083] The prediction module consists of four linear convolutional layers, ReLU activation function and Group Normalization. The use of Group Normalization can effectively increase the robustness of the network.

[0084] Specifically, the color space conversion module transforms the input X-ray top view into a top-view HSV image and the input X-ray side view into a side-view HSV image through color space transformation. The calculation process is as follows:

[0085] T=trans(X)

[0086] S=trans(Y)

[0087] Where X is the input X-ray top view, Y is the input X-ray side view, trans() represents the color space transformation that transforms the RGB channels into HSV channels, T is the top view HSV image, and S is the side view HSV image;

[0088] The image encoding module consists of multiple convolutional layers, gated convolutional layers, pooling layers, batch normalization layers, and nonlinear activation layers. It takes the top-view HSV image T and the side-view HSV image S output by the color space conversion module as input and finally obtains the HSV guided feature E through calculation. The calculation process is as follows:

[0089] f ac =σ(W ac ·S+b ac )

[0090] E=σ(W e ·T+U e ·(f ac ⊙avgpool(T))+b e )

[0091] Where W ac and W e represents the convolution kernel parameters, b ac and b e represents the offset distance, σ represents the nonlinear activation, f ac represents the gated feature map, U e Represents the gated kernel of the gated convolution kernel, ⊙ represents element-wise multiplication, avgpool represents the average pooling operation, and the HSV guided feature E improves the recognition sensitivity of the FCOS network for contraband.

[0092] The top-view guidance module consists of multiple convolutional layers, gated convolutional layers, batch normalization layers, and nonlinear activation layers. After obtaining the HSV guidance features, the top-view guidance module inputs the HSV guidance features and the X-ray top view, and finally obtains the top-view guidance features G through calculation, which is calculated as follows:

[0093] f gt1 =σ(W gt1 E+b gt1 )

[0094] f x =σ(W x ·X+b x )

[0095] G=σ(W g ·fx +U g ·(f gt1 ⊙f x )+b g )

[0096] Where W gt1 , W x and W g represents the convolution kernel parameters, b gt1 , b x and b g represents the offset distance, σ represents the nonlinear activation, f gt1 represents the gated activation map, X is the X-ray top view of the input, and f x Indicates the middle feature of the X-ray top view, U g represents the gated convolution kernel, ⊙ represents element-wise multiplication, and the top-down view guided feature G improves the recognition sensitivity of the FCOS network for contraband;

[0097] The side view guidance module inputs the HSV guidance feature and the X-ray side view, and finally obtains the side view guidance feature B through calculation, which is calculated as follows:

[0098] f gt2 =σ(W gt2 E+b gt2 )

[0099] f y =σ(W y ·Y+b y )

[0100] B=σ(W o ·f y +U o ·(f gt2 ⊙f y )+b o )

[0101] Where W gt2 、W y and W o represents the convolution kernel parameters, b gt2 、b y and b o represents the offset distance, σ represents the nonlinear activation, f gt2 represents the gated activation map, Y is the input X-ray side view, f y Indicates the middle feature of the X-ray side view, U o Represents the gated kernel of the gated convolution kernel, ⊙ represents element-by-element multiplication, and the side view guided feature B reduces the interference of the transmissive occlusion in the image on the FCOS network.

[0102] The top-down view multi-scale module takes the top-down view guided feature G as input and extracts the top-down view multi-scale feature I through multiple convolutional layers. The calculation is expressed as follows:

[0103] I=σ(W i G+b i )

[0104] Where σ represents the Sigmoid function, W i represents the convolution kernel, b i Represents the offset distance. The multi-scale feature I of the top-down view improves the recognition sensitivity of the FCOS network for contraband of different sizes.

[0105] The side view multi-scale module takes the side view guided feature B as input and extracts the side view multi-scale feature P through multiple convolutional layers. The calculation is expressed as follows:

[0106] P=σ(W p B+b p )

[0107] Where σ represents the Sigmoid function, W p represents the convolution kernel, b p Represents the offset distance. The side view multi-scale feature P reduces the interference of non-contraband items of different sizes on the FCOS network.

[0108] The cross attention module includes multiple convolutional layers, softmax layers, batch normalization layers, and nonlinear activation layers. First, the top-view multi-scale feature I is input into the cross attention module, and the query feature Q is calculated through the convolution layer. The calculation is expressed as follows:

[0109] Q=W q I+b q

[0110] Where W q represents the convolution kernel, b q Indicates the offset distance;

[0111] Then the side view multi-scale feature P is input into the cross attention module, and the key feature K and value feature V are calculated through different convolutional layers. The calculation is expressed as follows:

[0112] K=W k P+b k

[0113] V=W v P+b v

[0114] Where W k and W v represents the convolution kernel, b kand b v Indicates the offset distance;

[0115] After the query feature Q is dot-multiplied with the key feature K, the attention weight A is obtained through the Softmax layer. The attention weight A is dot-multiplied with the value feature V and then added to the top-down multi-scale feature I to obtain the global multi-scale feature M. The calculation is expressed as follows:

[0116]

[0117] M=A·V+I

[0118] Where, d k Represents the scaling factor of the attention weight, which is set to the feature dimension of the query feature Q.

[0119] The above-described embodiments are only preferred embodiments of the present invention and are not intended to limit the scope of implementation of the present invention. Therefore, any changes made based on the shape and principle of the present invention should be included in the scope of protection of the present invention.

Claims

1. The X-ray security inspection dual-view detection method based on the improved FCOS network is characterized by: This method is based on an improved FCOS network to achieve accurate detection of contraband in dual-view X-ray security inspections. The improvement of the FCOS network lies in the addition of an image encoder, an HSV-guided encoder, and a multi-scale fusion encoder. The image encoder converts the color space of the input X-ray top view and X-ray side view from RGB to HSV before feature extraction and fusion, which is used to improve the recognition sensitivity of the FCOS network for the target to be detected; the HSV-guided encoder fuses the HSV-guided features extracted by the image encoder with the input original X-ray top view and side view respectively, reducing the aliasing interference of the background in the X-ray image and improving the FCOS network's recognition ability for contraband; the multi-scale fusion encoder is used to extract multi-scale features from the top view and side view, and fuse the multi-scale features of the two viewpoints to enhance the FCOS network's recognition ability for targets of different sizes to be detected; The specific implementation of the X-ray security inspection dual-viewing angle detection method includes: Collect X-ray top view and X-ray side view of the object to be inspected on the conveyor belt; The trained improved FCOS network is used to perform the following processing on the X-ray top view and X-ray side view: The X-ray top view and the X-ray side view are input into the image encoder and first subjected to color space transformation to obtain the top view HSV image T and the side view HSV image S. The top view HSV image T and the side view HSV image S are then subjected to feature extraction and fusion to obtain the HSV guided feature E. The HSV guidance encoder fuses and encodes the HSV guidance feature E with the X-ray top view to obtain the top-view guidance feature G; the HSV guidance feature E is fused and encoded with the X-ray side view to obtain the side-view guidance feature B. The multi-scale fusion encoder extracts and fuses the top-view guided feature G and the side-view guided feature B to obtain the global multi-scale feature M; The global multi-scale feature M is input into the feature extraction backbone network of the improved FCOS network to obtain the backbone feature, which is then input into the prediction module of the improved FCOS network for target positioning and classification, and finally the detection information is output and marked in the X-ray image.

2. The X-ray security inspection dual-viewing angle detection method based on the improved FCOS network according to claim 1 is characterized in that: Acquiring X-ray top and side views includes the following operations: The object to be inspected is placed on a tray. When the tray is transported to the inspection area by a conveyor belt, the multi-view X-ray instrument emits an X-ray beam from directly above the object and on the right side of the conveyor belt's travel direction to scan the object to be inspected. The X-ray beam passes through the object to be inspected and reaches the receiver, which is rendered by a computer program to obtain an X-ray top view and an X-ray side view of the object to be inspected.

3. The X-ray security inspection dual-viewing angle detection method based on the improved FCOS network according to claim 1 is characterized in that: The improved FCOS network includes: An image encoder is divided into a color space conversion module for calculating a top-view HSV image and a side-view HSV image based on the X-ray top view and the X-ray side view, and an image encoding module for calculating HSV guide features based on the top-view HSV image and the side-view HSV image; An HSV guidance encoder is divided into a top-view guidance module for fusing HSV guidance features with an X-ray top view and encoding and calculating top-view guidance features, and a side-view guidance module for fusing HSV guidance features with an X-ray side view and encoding and calculating side-view guidance features; a multi-scale fusion encoder comprising a top-view multi-scale module for extracting top-view multi-scale features based on top-view guided features, a side-view multi-scale module for extracting side-view multi-scale features based on side-view guided features, and a cross-attention module for calculating global multi-scale features based on the top-view multi-scale features and the side-view multi-scale features; Feature extraction backbone network, consisting of 5 computational layers and a residual structure; The prediction module consists of four linear convolutional layers, ReLU activation function and Group Normalization. The use of Group Normalization can effectively increase the robustness of the network.

4. The X-ray security inspection dual-viewing angle detection method based on the improved FCOS network according to claim 3 is characterized in that: The color space conversion module transforms the input X-ray top view into a top-view HSV image and the input X-ray side view into a side-view HSV image through color space transformation. The calculation process is as follows: T=trans(X) S=trans(Y) Where X is the input X-ray top view, Y is the input X-ray side view, trans() represents the color space transformation that transforms the RGB channels into HSV channels, T is the top view HSV image, and S is the side view HSV image; The image encoding module consists of multiple convolutional layers, gated convolutional layers, pooling layers, batch normalization layers, and nonlinear activation layers. It takes the top-view HSV image T and the side-view HSV image S output by the color space conversion module as input and finally obtains the HSV guided feature E through calculation. The calculation process is as follows: f ac =σ(W ac ·S+b ac ) E=σ(W e ·T+U e ·(f ac ⊙avgpool(T))+b e ) Where W ac and W e represents the convolution kernel parameters, b ac and b e represents the offset distance, σ represents the nonlinear activation, f ac represents the gated feature map, U e Represents the gated kernel of the gated convolution kernel, ⊙ represents element-wise multiplication, avgpool represents the average pooling operation, and the HSV guided feature E improves the recognition sensitivity of the FCOS network for contraband.

5. The X-ray security inspection dual-viewing angle detection method based on the improved FCOS network according to claim 3 is characterized in that: The top-view guidance module consists of multiple convolutional layers, gated convolutional layers, batch normalization layers, and nonlinear activation layers. After obtaining the HSV guidance features, the top-view guidance module inputs the HSV guidance features and the X-ray top view, and finally obtains the top-view guidance features G through calculation, which is calculated as follows: f gt1 =σ(W gt1 ·E+b gt1 ) f x =σ(W x ·X+b x ) G=σ(W g ·f x +U g ·(f gt1 ⊙f x )+b g ) Where W gt1 , W x and W g represents the convolution kernel parameters, b gt1 , b x and b g represents the offset distance, σ represents the nonlinear activation, f gt1 represents the gated activation map, X is the X-ray top view of the input, and f x Indicates the middle feature of the X-ray top view, U g represents the gated convolution kernel, ⊙ represents element-wise multiplication, and the top-down view guided feature G improves the recognition sensitivity of the FCOS network for contraband; The side view guidance module inputs the HSV guidance feature and the X-ray side view, and finally obtains the side view guidance feature B through calculation, which is calculated as follows: f gt2 =σ(W gt2 ·E+b gt2 ) f y =σ(W y ·Y+b y ) B=σ(W o ·f y +U o ·(f gt2 ⊙f y )+b o ) Where W gt2 、W y and W o represents the convolution kernel parameters, b gt2 、b y and b o represents the offset distance, σ represents the nonlinear activation, f gt2 represents the gated activation map, Y is the input X-ray side view, f y Indicates the middle feature of the X-ray side view, U o Represents the gated kernel of the gated convolution kernel, ⊙ represents element-by-element multiplication, and the side view guided feature B reduces the interference of the transmissive occlusion in the image on the FCOS network.

6. The X-ray security inspection dual-viewing angle detection method based on the improved FCOS network according to claim 3 is characterized in that: The top-down view multi-scale module takes the top-down view guided feature G as input and extracts the top-down view multi-scale feature I through multiple convolutional layers. The calculation is expressed as follows: I=σ(W i ·G+b i ) Where σ represents the Sigmoid function, W i represents the convolution kernel, b i Represents the offset distance. The multi-scale feature I of the top-down view improves the recognition sensitivity of the FCOS network for contraband of different sizes. The side view multi-scale module takes the side view guided feature B as input and extracts the side view multi-scale feature P through multiple convolutional layers. The calculation is expressed as follows: P=σ(W p ·B+b p ) Where σ represents the Sigmoid function, W p represents the convolution kernel, b p Represents the offset distance. The side view multi-scale feature P reduces the interference of non-contraband items of different sizes on the FCOS network. The cross attention module includes multiple convolutional layers, softmax layers, batch normalization layers, and nonlinear activation layers. First, the top-view multi-scale feature I is input into the cross attention module, and the query feature Q is calculated through the convolution layer. The calculation is expressed as follows: Q=W q ·I+b q Where W q represents the convolution kernel, b q Indicates the offset distance; Then the side view multi-scale feature P is input into the cross attention module, and the key feature K and value feature V are calculated through different convolutional layers. The calculation is expressed as follows: K=W k ·P+b k V=W v ·P+b v Where W k and W v represents the convolution kernel, b k and b v Indicates the offset distance; After the query feature Q is dot-multiplied with the key feature K, the attention weight A is obtained through the Softmax layer. The attention weight A is dot-multiplied with the value feature V and then added to the top-down multi-scale feature I to obtain the global multi-scale feature M. The calculation is expressed as follows: M=A·V+I Where, d k Represents the scaling factor of the attention weight, which is set to the feature dimension of the query feature Q.

Citation Information

Patent Citations

  • Image processing method and device thereof, storage medium and electronic device

    CN113506283A

  • Medical image segmentation using an integrated edge guidance module and object segmentation network

    US10482603B1