Security Inspection Image Detection Method, Device, Equipment and Product

Through multi-layer network processing of X-ray security images, the detection results are generated, and the problems of inefficient and high missed detection rates of traditional security inspection methods are solved, achieving fast, efficient and accurate security inspections.

CN119723473BActive Publication Date: 2025-06-20TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510246136.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The traditional X-ray security inspection method has problems such as low manual operation efficiency, insufficient professionalism, and high missed detection rate, which is difficult to meet the fast, efficient and accurate requirements of modern society for safety inspections.

Method used

A security image detection method is adopted to obtain X-ray images, and multi-layer network processing such as slice convolution network, multi-residual convolution network, attention pooling network, feature pyramid network, path aggregation network and detection head are used to generate detection results to characterize whether the image contains abnormal items.

Benefits of technology

It improves detection accuracy, improves the detection performance of small-sized targets and occluded targets, improves the generalization ability of the model, and meets the fast, efficient and accurate requirements of modern security inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723473B_ABST
    Figure CN119723473B_ABST
Patent Text Reader

Abstract

The present invention provides a security inspection image detection method, device, equipment and product, which can be applied to the field of target detection technology. The method includes: obtaining a security inspection image, where the security inspection image is obtained by scanning an item using a ray scanning device; processing the security inspection image using a slice convolutional network to obtain slice convolutional features; processing the slice convolutional features using a multi-residual convolutional network to obtain residual convolutional features; processing the residual convolutional features based on an attention pooling network to obtain pooled attention features; processing the pooled attention features and the residual convolutional features based on a feature pyramid network to obtain attention fusion features and convolutional attention features; processing the attention fusion features and the convolutional attention features based on a path aggregation network to obtain fusion features; and processing the attention fusion features and the fusion features based on a detection head to obtain a detection result of the security inspection image, where the detection result indicates whether the security inspection image contains abnormal items.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and specifically relates to a security inspection image detection method, device, equipment and product. Background Art

[0002] With the rapid development of the tourism and transportation industries, X-ray security inspection machines play an increasingly important role in modern cities. X-ray security inspection machines are a means of security inspection widely used in public places (such as airports, stations, etc.). Security inspectors need to carefully distinguish and observe X-ray images to ensure that there are no abnormal items in the passengers' luggage and effectively maintain the safety order of public places. However, the inventor has found that this traditional security inspection method has problems such as low manual operation efficiency, insufficient professionalism, and high missed detection rates, and it is difficult to meet the requirements of modern society for fast, efficient, and accurate security inspections. Summary of the Invention

[0003] In view of the above problems, the present invention provides a security inspection image detection method, device, equipment and product.

[0004] According to a first aspect of the present invention, there is provided a security inspection image detection method, including: obtaining a security inspection image, where the security inspection image is obtained by scanning an item using a ray scanning device; processing the security inspection image using a slice convolutional network to obtain slice convolutional features; processing the slice convolutional features using a multi-residual convolutional network to obtain residual convolutional features; processing the residual convolutional features based on an attention pooling network to obtain pooled attention features; processing the pooled attention features and the residual convolutional features based on a feature pyramid network to obtain attention fusion features and convolutional attention features; processing the attention fusion features and the convolutional attention features based on a path aggregation network to obtain fusion features; processing the attention fusion features and the fusion features based on a detection head to obtain a detection result of the security inspection image, where the detection result indicates whether the security inspection image contains abnormal items.

[0005] A second aspect of the present invention provides a security inspection image detection device, including:

[0006] An acquisition module, configured to obtain a security inspection image, where the security inspection image is obtained by scanning an item using a ray scanning device;

[0007] A slice convolutional feature obtaining module, configured to process the security inspection image using a slice convolutional network to obtain slice convolutional features;

[0008] A residual convolutional feature obtaining module, configured to process the slice convolutional features using a multi-residual convolutional network to obtain residual convolutional features;

[0009] A pooled attention feature obtaining module, configured to process the residual convolutional features based on an attention pooling network to obtain pooled attention features;

[0010] A convolution attention fusion module is used to process the pooled attention feature and the residual convolution feature based on the feature pyramid network to obtain an attention fusion feature and a convolution attention feature;

[0011] A fused feature module is used to process the attention fusion feature and the convolution attention feature based on the path aggregation network to obtain a fused feature;

[0012] A detection result module is used to process the attention fusion feature and the fused feature based on the detection head to obtain the detection result of the security inspection image, and the detection result represents whether the security inspection image contains abnormal items.

[0013] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the above one or more processors execute the above one or more computer programs to implement the steps of the above method.

[0014] The fourth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the above computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0015] According to the embodiments of the present invention, by using a slicing convolution network to process a security inspection image, a slicing convolution feature can be obtained. By using a multi-residual convolution network to process the slicing convolution feature, a residual convolution feature can be obtained. Based on an attention pooling network to process the residual convolution feature, the features at different scales corresponding to the residual convolution feature can be pooled, enriching the semantic information. And based on the attention mechanism to process the pooled feature, a pooled attention feature can be obtained, enhancing the receptive field and improving the detection of small-size targets. Based on the feature pyramid network to process the pooled attention feature and the residual convolution feature, an attention fusion feature and a convolution attention feature can be obtained. Based on the path aggregation network to process the attention fusion feature and the convolution attention feature, a fused feature can be obtained. Based on the detection head to process the attention fusion feature and the fused feature, the detection result of the security inspection image can be obtained, thereby improving the detection accuracy, improving the detection performance of small-size targets and occluded targets, and enhancing the generalization ability of the model. Description of the Drawings

[0016] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features and advantages of the present invention will become clearer. In the drawings:

[0017] Figure 1 The application scenario diagram of the security inspection image detection method and device according to the embodiments of the present invention is shown;

[0018] Figure 2Shows a flowchart of the security inspection image detection method according to an embodiment of the present invention;

[0019] Figure 3 Shows a structural diagram of a serial pooling layer according to an embodiment of the present invention;

[0020] Figure 4 Shows a structural diagram of a self-attention layer according to an embodiment of the present invention;

[0021] Figure 5 Shows a network structure diagram of a detection model according to an embodiment of the present invention;

[0022] Figure 6 Shows a flowchart of a detection model training and conversion method according to an embodiment of the present invention;

[0023] Figure 7 Shows an RKNN image inference flowchart of a real-time security inspection image detection system based on a target chip in the present invention;

[0024] Figure 8 Shows a flowchart of transmitting a converted RKNN model in a real-time security inspection image detection system based on a target chip to the target chip for inference in the present invention;

[0025] Figure 9 Shows a schematic diagram of a security inspection image detection result obtained by the security inspection image detection method in the present invention;

[0026] Figure 10 Shows a structural block diagram of a security inspection image detection device according to an embodiment of the present invention;

[0027] Figure 11 Shows a block diagram of an electronic device suitable for implementing the security inspection image detection method according to an embodiment of the present invention. Detailed implementation manners

[0028] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0029] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0031] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0032] With the rapid development of the tourism and transportation industries, X-ray security inspection machines are playing an increasingly important role in modern cities. X-ray security inspection machines are a means of security inspection widely used in public places (such as airports, railway stations, etc.). Security inspectors need to carefully distinguish and observe X-ray images to ensure that there are no abnormal items in the passengers' luggage, effectively maintaining the safety order of public places. However, the inventors have found that this traditional security inspection method has problems such as low efficiency of manual operation, lack of professionalism, and high missed detection rate, and it is difficult to meet the requirements of modern society for security inspection.

[0033] The X-ray image security inspection abnormal item detection technology based on computer vision and artificial intelligence technologies has received attention and research. This technology aims to automatically identify and mark abnormal items in X-ray images through technical means such as computer vision, image processing, and machine learning, using algorithms and models, and then conduct secondary verification by humans, thereby improving the accuracy and efficiency of security inspection work. At the same time, it can also reduce the influence of human factors on the results, lower the skill requirements for security inspectors, and training costs. However, the existing object detection networks are designed based on natural images, and the detection effect is not good when using the existing object detection networks in the X-ray image scenarios formed by security inspection machines, and problems such as missed detection and false detection are likely to occur.

[0034] In view of this, the present invention provides a security inspection image detection method, a security inspection image detection device, a device and a product. The method includes: obtaining a security inspection image, where the security inspection image is obtained by scanning an item using a ray scanning device; processing the security inspection image using a slice convolutional network to obtain slice convolutional features; processing the slice convolutional features using a multi-residual convolutional network to obtain residual convolutional features; processing the residual convolutional features based on an attention pooling network to obtain pooled attention features; processing the pooled attention features and the residual convolutional features based on a feature pyramid network to obtain attention fusion features and convolutional attention features; processing the attention fusion features and the convolutional attention features based on a path aggregation network to obtain fusion features; and processing the attention fusion features and the fusion features based on a detection head to obtain a detection result of the security inspection image, where the detection result indicates whether the security inspection image contains abnormal items.

[0035] It should be noted that the security inspection image detection method and the security inspection image detection device provided by the present invention can be used in the field of target detection technology, and can also be used in any field other than the field of target detection technology, such as the field of public security. Therefore, the application fields of the security inspection image detection method and the security inspection image detection device provided by the present invention are not limited.

[0036] In the technical solution of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. And the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure and application, etc., all comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0037] In the scenario of making automated decisions using personal information, the methods, devices and systems provided by the embodiments of the present invention all provide corresponding operation entrances for users to choose to agree or refuse the results of automated decisions; if the user chooses to refuse, the expert decision-making process will be entered. The expression "automated decision" here refers to the activity of automatically analyzing and evaluating an individual's behavior habits, hobbies or economic, health, credit status, etc. through a computer program and making a decision. The expression "expert decision" here refers to the activity of making a decision by a person who specializes in a certain field, has specialized experience, knowledge and skills and reaches a certain professional level.

[0038] Figure 1 Shows an application scenario diagram of the security inspection image detection method and device according to an embodiment of the present invention.

[0039] As Figure 1As shown, the application scenario according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0040] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0041] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0042] The server 105 may be a server providing various services, such as a background management server (only for example) that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0043] It should be noted that the security inspection image detection method provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the security inspection image detection device provided by the embodiments of the present invention can generally be set in the server 105. The security inspection image detection method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the security inspection image detection device provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0044] It should be understood, Figure 1The numbers of terminal devices, networks, and servers in [it] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0045] Figure 2 The flowchart of the security inspection image detection method according to an embodiment of the present invention is shown.

[0046] As Figure 2 shown, the security inspection image detection method of this embodiment includes operations S210 to S270, and this security inspection image detection method can be executed by an electronic device.

[0047] In operation S210, obtain a security inspection image.

[0048] In operation S220, process the security inspection image using a slice convolution network to obtain slice convolution features.

[0049] In operation S230, process the slice convolution features using a multi-residual convolution network to obtain residual convolution features.

[0050] In operation S240, process the residual convolution features based on an attention pooling network to obtain pooled attention features.

[0051] In operation S250, process the pooled attention features and the residual convolution features based on a feature pyramid network to obtain attention fusion features and convolutional attention features.

[0052] In operation S260, process the attention fusion features and the convolutional attention features based on a path aggregation network to obtain fusion features.

[0053] In operation S270, process the attention fusion features and the fusion features based on a detection head to obtain the detection result of the security inspection image, and the detection result indicates whether the security inspection image contains abnormal items.

[0054] According to an embodiment of the present invention, the security inspection image is obtained by scanning an item using a ray scanning device, and an open-source dataset of abnormal items in X-ray security inspection images can also be used. This dataset can include multiple X-ray security inspection images, and multiple types of abnormal items are included in the security inspection images. The abnormal items can include batons, pliers, hammers, power banks, scissors, wrenches, sprayers, lighters, etc., but are not limited thereto. The embodiments of the present disclosure do not make limitations in this regard.

[0055] According to an embodiment of the present invention, processing the security inspection image using a slice convolution network can obtain slice convolution features. Specifically, the security inspection image can be sliced using the slice convolution network and then a convolution operation is performed, thereby obtaining slice convolution features. The size of the slice convolution features can be , and the channel features .

[0056] According to an embodiment of the present invention, by using a multi-residual convolutional network to process the slice convolutional features, residual convolutional features can be obtained. Based on the attention pooling network to process the residual convolutional features, the features at different scales corresponding to the residual convolutional features can be pooled and the pooled features can be processed based on the attention mechanism to obtain the pooled attention features.

[0057] According to an embodiment of the present invention, the feature pyramid network is a top-down structure, and the high-level feature map is upsampled and fused with the low-level feature map. The pooled attention features and the residual convolutional features can be processed based on the feature pyramid network to obtain the attention fusion features and the convolutional attention features.

[0058] According to an embodiment of the present invention, the path aggregation network is a bottom-up structure and can transmit position information to the high level. The attention fusion features and the convolutional attention features can be processed based on the path aggregation network to obtain the fusion features.

[0059] According to an embodiment of the present invention, the detection head is responsible for converting the features into specific detection results. Based on the detection head to process the attention fusion features and the fusion features, the detection results of the security inspection image can be obtained. The detection results indicate whether the security inspection image contains abnormal items, and the detection results include the category of the target item, the bounding box information, and the width and height.

[0060] According to an embodiment of the present invention, by using a slice convolutional network to process the security inspection image, slice convolutional features can be obtained. By using a multi-residual convolutional network to process the slice convolutional features, residual convolutional features can be obtained. Based on the attention pooling network to process the residual convolutional features, the features at different scales corresponding to the residual convolutional features can be pooled and the pooled features can be processed based on the attention mechanism to obtain the pooled attention features, enriching the semantic information, enhancing the receptive field, and improving the detection of small-sized targets. Based on the feature pyramid network to process the pooled attention features and the residual convolutional features, the attention fusion features and the convolutional attention features can be obtained. Based on the path aggregation network to process the attention fusion features and the convolutional attention features, the fusion features can be obtained. Based on the detection head to process the attention fusion features and the fusion features, the detection results of the security inspection image can be obtained, thereby improving the detection accuracy, improving the detection performance of small-sized targets and occluded targets, and enhancing the generalization ability of the model.

[0061] According to an embodiment of the present invention, based on the attention pooling network to process the residual convolutional features to obtain the pooled attention features, it includes: using a serial pooling layer to process the residual convolutional features to obtain multi-scale fusion features; based on a self-attention layer to process the multi-scale fusion features to obtain the pooled attention features. The attention pooling network can include a serial pooling layer and a self-attention layer.

[0062] According to an embodiment of the present invention, by using a serial pooling layer to process residual convolution features, multi-scale fusion features can be obtained. By simplifying the pooling operation, the computing speed is significantly improved. Based on the self-attention layer to process the multi-scale fusion features, pooling attention features can be obtained, which can effectively capture global context information and can consider the relationships between all features in the image at the same time. This global information capture ability enables the model to better understand the overall structure in security inspection images and avoid misjudgments caused by the loss of local features.

[0063] According to an embodiment of the present invention, using a serial pooling layer to process residual convolution features to obtain multi-scale fusion features includes: using a convolution normalization block to process the residual convolution features to obtain a first residual normalization feature; performing a multi-scale pooling operation on the first residual normalization feature to obtain multi-scale pooling features; fusing the first residual normalization feature and the multi-scale pooling features to obtain multi-scale fusion features. The serial pooling layer includes a convolution normalization block.

[0064] Figure 3 The structural diagram of the serial pooling layer according to an embodiment of the present invention is shown.

[0065] As Figure 3 shown, the residual convolution features can be processed by using the convolution normalization block 310 to obtain a first residual normalization feature; based on the first pooling block 320 to process the first residual normalization feature, a first pooling feature can be obtained; based on the second pooling block 330 to process the first pooling feature, a second pooling feature can be obtained; based on the Nth pooling block 340 to process the (N - 1)th pooling feature, an Nth pooling feature can be obtained. By fusing the first residual normalization feature and the multi-scale pooling features, the multi-scale pooling features can include the first pooling feature to the Nth pooling feature, and finally multi-scale fusion features can be obtained, where N is an integer greater than 1, and the N pooling blocks all use the same pooling kernel. For example, the pooling kernel is max pooling, and the size of the pooling kernel is not limited in the embodiment of the present invention.

[0066] According to an embodiment of the present invention, by processing the residual convolution feature using a convolutional normalization block, a first residual normalization feature can be obtained. Performing a multi-scale pooling operation on the first residual normalization feature yields a multi-scale pooling feature. Among them, the same-sized pooling kernels are used in the multi-scale pooling operation, reducing the complexity of the pooling operation. Additionally, using the multi-scale pooling operation can more efficiently capture multi-scale information by gradually downsampling the first residual normalization feature, while reducing the computational amount. Fusing the first residual normalization feature and the multi-scale pooling feature can obtain a multi-scale fusion feature containing multi-scale information. This fusion method retains the context information of different scales while reducing the computational redundancy brought by traditional pooling operations. By reducing the number of pooling kernels and optimizing the pooling operation, the memory occupancy and computational complexity are significantly reduced, enhancing the robustness of the model.

[0067] According to an embodiment of the present invention, processing the multi-scale fusion feature based on a self-attention layer to obtain a pooling attention feature includes: processing the multi-scale fusion feature based on a convolutional block to obtain a query feature, a key feature, and a value feature; multiplying the query feature and the key feature to obtain a pixel correlation feature; processing the pixel correlation feature using a normalization algorithm to obtain an attention weight; obtaining a value attention feature according to the attention weight and the value feature; processing the value attention feature based on a convolutional block to obtain a value attention convolutional feature; performing a residual connection and summation operation on the value attention convolutional feature and the multi-scale fusion feature to obtain a pooling attention feature.

[0068] Figure 4 The structural diagram of the self-attention layer according to an embodiment of the present invention is shown.

[0069] As Figure 4 shown, processing the multi-scale fusion feature 410 based on the first convolutional block 420 to obtain a query feature Query, processing the multi-scale fusion feature 410 based on the second convolutional block 430 to obtain a key feature Key, and processing the multi-scale fusion feature 410 based on the third convolutional block 440 to obtain a value feature Value. Performing a reshaping operation on Query and Key and multiplying them to obtain a pixel correlation feature to calculate the relationship between each pixel point on Query and Key and other pixel points. The softmax function can be used to normalize the pixel correlation feature to obtain an attention weight. The attention weight can be weighted to Value to obtain a value attention feature; processing the value attention feature based on the fourth convolutional block 450 to obtain a value attention convolutional feature. Finally, performing a residual connection and summation operation on the value attention convolutional feature and the multi-scale fusion feature can obtain a pooling attention feature.

[0070] According to an embodiment of the present invention, by processing multi-scale fusion features based on a convolutional block, query features, key features, and value features can be obtained; multiplying the query features and the key features can obtain pixel correlation features, thereby evaluating the relationship between each element in the features and all other elements, enabling the model to capture the global information in the multi-scale fusion features, better understand long-range dependencies, processing the pixel correlation features using a normalization algorithm can obtain attention weights, and based on the attention weights and the value features, value attention features can be obtained, which can automatically adjust the focus of attention according to the input of different key information; processing the value attention features based on a convolutional block can obtain value attention convolutional features; performing a residual connection and summation operation on the value attention convolutional features and the multi-scale fusion features can obtain pooling attention features, in which the features of different scales can be dynamically fused, significantly improving the detection performance of the model for multi-scale targets, and significantly improving the detection accuracy and robustness when dealing with complex scenarios such as occlusion and distinguishing similar objects.

[0071] According to an embodiment of the present invention, the multi-residual convolutional network includes I feature extraction layers, where I is an integer greater than 1; processing the slice convolutional features using the multi-residual convolutional network to obtain residual convolutional features includes: inputting the slice convolutional features into the first feature extraction layer to output the first residual convolutional features; inputting the (i - 1)-th residual convolutional features into the i-th feature extraction layer to output the i-th residual convolutional features; in the case of i = I, determining I residual convolutional features based on the first residual convolutional features to the I-th residual convolutional features, where I ≥ i > 1 and I and i are integers.

[0072] According to an embodiment of the present invention, the multi-residual convolutional network may include i feature extraction layers, where i is an integer greater than 1. Using the CBS module and the C3 module of the feature extraction layer can perform feature extraction and reduce parameters. The CBS module may include a convolutional block, a batch normalization block, and a SiLU activation function. The SiLU activation function has the characteristics of being unbounded above and bounded below, smooth, and non-monotonic. This layer can output channel data . The C3 module may include three modules and multiple backbone residual blocks. The number of backbone residual blocks is specified by a parameter and is used for feature extraction. Finally, it can output channel data with a size of .

[0073] According to an embodiment of the present invention, inputting the slice convolutional features into the first feature extraction layer can output the first residual convolutional features; inputting the (i - 1)-th residual convolutional features into the i-th feature extraction layer can output the i-th residual convolutional features; in the case of i = I, I residual convolutional features can be determined based on the first residual convolutional features to the I-th residual convolutional features, where I ≥ i > 1.

[0074] According to an embodiment of the present invention, based on the detection head to process the attention fusion feature and the fusion feature to obtain the detection result of the security inspection image, including: based on the first detection head to process the attention fusion feature, the first detection result of the security inspection image can be obtained, and the first detection result characterizes whether there is an abnormal object with a first size in the security inspection image; based on the second detection head to process the fusion feature, the second detection result of the security inspection image can be obtained, and the second detection result characterizes whether there is an abnormal object with a second size in the security inspection image, and the size difference value between the first size and the second size is greater than a preset size threshold, and the detection head may include a first detection head and a second detection head.

[0075] According to an embodiment of the present invention, for example, the scale of the attention fusion feature is 20×20, and the scale of the fusion feature is 80×80. Each pixel in the feature map corresponding to the attention fusion feature covers a larger image area and is suitable for capturing large-size targets. Each pixel in the feature map corresponding to the fusion feature covers a smaller image area and is suitable for capturing small-size targets.

[0076] According to an embodiment of the present invention, by using different detection heads to detect articles of different scales, abnormal articles of different sizes can be detected specifically, significantly improving the efficiency and accuracy of security inspection image detection.

[0077] According to an embodiment of the present invention, before using the slice convolution network to process the security inspection image to obtain the slice convolution feature, the method further includes: using an image local adjustment algorithm to adjust the contrast and brightness of the initial security inspection image to obtain an intermediate image; performing image enhancement on the intermediate image based on the mosaic algorithm to obtain the security inspection image.

[0078] According to an embodiment of the present invention, the contrast and brightness of the initial security inspection image can be adjusted by using an image local adjustment algorithm to obtain an intermediate image, thereby increasing the detail and texture information of the image. The specific calculation is shown in formula (1):

[0079] (1);

[0080] where x and y represent the pixel coordinates of the image, represents the enhanced image brightness; represents the brightness value of the corresponding pixel point of the input image, represents the maximum value of the input image brightness. represents the logarithmic average value of the brightness values of the input image, m×n represents the size of the image, represents a non-zero and very small constant, to prevent encountering a brightness of 0 in the image and avoid numerical overflow when performing logarithmic calculation on pure black pixels.

[0081] According to an embodiment of the present invention, the intermediate image can be enhanced based on a mosaic algorithm to obtain a security inspection image. The image enhancement can include Hue Saturation Value (HSV) chromaticity, saturation, exposure transformation, random rotation and translation, etc. An image size processing method can also be used to scale the security inspection image with a size of security inspection image to the same height-width ratio as image . The edges that do not meet the conditions can be filled with white bars, and the obtained image is input into the slice convolutional network.

[0082] According to an embodiment of the present invention, by adjusting the attributes such as brightness, contrast, and color of the initial security inspection image, the details and texture information of the image are increased, so as to facilitate subsequent accurate target detection of the security inspection image.

[0083] According to an embodiment of the present invention, the original dataset picture format is Portable Network Graphics (png) format. The pictures are divided into four folders: training, simple, difficult, and occluded. The labels of the pictures are stored in four JavaScript Object Notation (json) files according to the folders respectively. The json files contain the category, bounding box information, width and height, file name, etc. of the abnormal items in each picture.

[0084] According to an embodiment of the present invention, the dataset can be divided into a training set and a test set in a ratio of 6:4. The training set has a total of 29,457 pictures, and the test set is further divided into three subsets according to the difficulty of abnormal item detection: simple, difficult, and occluded, with the number of pictures being 9,482, 3,733, and 5,005 respectively.

[0085] According to an embodiment of the present invention, small targets can refer to those less than 32×32 pixel points, medium targets can refer to those with pixels between 32×32 and 96×96, and large targets can refer to those with pixels greater than 96×96. The specific situation of small, medium, and large targets in the dataset is shown in Table 1.

[0086] Table 1 Situation table of small, medium, and large targets in the dataset

[0087] Dataset Small targets (pcs) Medium targets (pcs) Large targets (pcs) Simple 85 2278 7119 Difficult 65 1759 7068 Occlusion 130 1392 3486 Training 278 11480 27950

[0088] According to an embodiment of the present invention, the data set is collected from airports, subway stations, and railway stations, covering various scenarios of detecting abnormal items in reality, especially deliberately hidden items. The characteristic of this data set is that the hidden subset in its test set focuses on abnormal items deliberately hidden among cluttered objects. For example, by winding wires around the abnormal items to block them and interfere with the detection results; the occlusion of abnormal items is also reflected in that the items in personal luggage are usually randomly placed and overlap with each other. These characteristics make the detection of abnormal items more difficult.

[0089] Figure 5 The network structure diagram of the detection model according to an embodiment of the present invention is shown.

[0090] As Figure 5 shown, the security inspection image is processed by the slice convolution network 510 to obtain slice convolution features, and the slice convolution features are input into the (i - 1)-th feature extraction layer 521 to output the (i - 1)-th residual convolution features; the (i - 1)-th residual convolution features are input into the i-th feature extraction layer 522 to output the i-th residual convolution features; the multi-residual convolution network 520 may include I feature extraction layers.

[0091] According to an embodiment of the present invention, the residual convolution features are processed by the serial pooling layer 531 to obtain multi-scale fusion features; the multi-scale fusion features are processed based on the self-attention layer 532 to obtain pooled attention features, and the attention pooling network 530 includes the serial pooling layer 531 and the self-attention layer 532.

[0092] According to an embodiment of the present invention, the pooled attention features and the residual convolution features are processed based on the feature pyramid network 540 to obtain attention fusion features and convolutional attention features. The attention fusion features and the convolutional attention features are processed based on the path aggregation network 550 to obtain fusion features. The attention fusion features are processed based on the first detection head 561 to obtain the first detection result of the security inspection image, and the fusion features are processed based on the second detection head 562 to obtain the second detection result of the security inspection image. The detection head 560 includes the first detection head 561 and the second detection head 562.

[0093] Figure 6 The flowchart of the detection model training and conversion method according to an embodiment of the present invention is shown.

[0094] As Figure 6 shown, the detection model training and conversion method includes operations S610 to S619.

[0095] In operation S610, the sample security inspection image and label data are read.

[0096] In operation S611, the sample security inspection image is processed by the slice convolution network to obtain sample slice convolution features.

[0097] In operation S612, the sample slice convolution features are processed using a multi-residual convolution network to obtain sample residual convolution features.

[0098] In operation S613, the sample residual convolution features are processed based on an attention pooling network to obtain sample pooling attention features.

[0099] In operation S614, the sample pooling attention features and the sample residual convolution features are processed based on a feature pyramid network to obtain sample attention fusion features and sample convolution attention features.

[0100] In operation S615, the sample attention fusion features and the sample convolution attention features are processed based on a path aggregation network to obtain sample fusion features.

[0101] In operation S616, the sample attention fusion features and the sample fusion features are processed based on a detection head to obtain the detection results of the sample security inspection images.

[0102] In operation S617, the detection results of the sample security inspection images and the label data are processed according to a loss function to obtain a target loss value.

[0103] In operation S618, according to the target loss value, a detection model is trained to obtain a trained detection model.

[0104] In operation S619, the trained detection model is converted into a target format and embedded into a target chip for hardware-accelerated inference.

[0105] According to an embodiment of the present invention, the information of the json annotation file in the dataset can be read, and the annotation information is converted from the two-point coordinate form into the center point coordinate and height-width form The sample security inspection images and the label data can be read.

[0106] According to an embodiment of the present invention, the sample attention fusion features and the sample fusion features are decoded into corresponding position prediction boxes on the image, and then score sorting and non-maximum suppression (NMS) screening are performed to obtain the detection results of the sample security inspection images. Processing the detection results of the sample security inspection images and the label data according to a loss function can obtain a target loss value. According to the target loss value, training a detection model can obtain a trained detection model. Among them, the comprehensive intersection over union (CloU) and the intersection over union (IoU) can be used to evaluate the performance of the detection model, and the specific calculations of the comprehensive intersection over union and the intersection over union are shown in formula (2).

[0107] (2);

[0108] Among them, IoU represents the ratio of the intersection area to the union area of the predicted bounding box b and the ground truth bounding box , is the square of the distance between the center points of the predicted bounding box b and the ground truth bounding box , c is the diagonal length of the outer bounding box (the smallest rectangle enclosing the predicted bounding box and the ground truth bounding box), represents the weight function, which is used to measure the consistency of the aspect ratio, , are the width and height of the predicted bounding box b, , are the width and height of the ground truth bounding box.

[0109] The final loss function is defined as follows. The loss consists of three parts: class loss, object confidence loss, and bounding box localization loss. And an optimizer is used to minimize the loss function to train the network model. Both the class loss and the confidence loss adopt the binary cross-entropy loss function, and the calculation formula (3) is as follows:

[0110] (3);

[0111] Among them, n represents the total number of samples, represents the original value output by the model, represents the true label of the i-th sample, represents the sigmoid function, represents the exponential function.

[0112] According to the embodiments of the present invention, after obtaining the trained detection model, the target-related features of the abnormal item can be obtained. The trained detection model can be converted into a target format, such as the RKNN format, and embedded into the target chip for hardware-accelerated inference.

[0113] According to the embodiments of the present invention, converting the trained detection model into a target format and embedding it into the target chip for hardware-accelerated inference, and using the system with the target chip to process the security inspection image to obtain the detection result of the security inspection image may include the following steps: The system obtains an image containing abnormal security inspection items from the image acquisition unit and scales the image to a preset size proportionally. For example, , and sends it into the input queue ; Initialize the running environment of the Neural Network Processing Unit (NPU), continuously obtain images from the input queue , and the NPU with the highest computing performance is responsible for image inference, and sends the inference result into the inference queue ; Read the inference results from the output queue and select the detection box with the highest score , remove the output queue. If the remaining candidate boxes and 's value is greater than , then delete the candidate box; The system draws a box on the original image, displays the detection results in real time, and sends them to the output queue for saving; By using multiple processes to control specific hardware units or cores on the RKNN board respectively, they are assigned to different processes to handle multiple tasks to achieve parallel processing and improve the overall computing efficiency. Through the collaborative work of the Central Processing Unit (CPU) and the NPU, let the CPU handle complex logic and control processes, while the NPU focuses on executing deep learning algorithms, thereby maximizing the hardware performance.

[0114] Figure 7 shows the RKNN image inference flow chart of the security inspection image real-time detection system based on the target chip in the present invention.

[0115] As Figure 7 shown, the RKNN object S710 can be created first, the config interface can be called to set the preprocessing parameters of the model S720, the load_onnx interface can be called to import the ONNX model S730, and the build interface can be called to build the RKNN model S740. The running environment can be initialized by calling the init_runtime interface S750, specifying the target platform and the NPU core scheduling mode, and the inference interface can be used for inference to obtain the inference results S760. The input data needs to be preprocessed according to the model requirements. The RKNN model can also be exported by calling the export_RKNN interface S770. After the inference is completed, the release interface can be called to release the RKNN object S780 to release resources.

[0116] According to the embodiment of the present invention, the trained detection model with target weights can be converted into an RKNN model in a virtual machine environment through a model conversion tool, and an X-ray security inspection image detection system can be built on the target chip.

[0117] According to the embodiment of the present invention, the system uses a target chip to build an X-ray security inspection image detection system. The system obtains the image to be detected through input; The system performs multi-process real-time detection on the input X-ray security inspection image , and for each image It includes preprocessing, target inference, and postprocessing; the target chip transmits the detection results back, and the detection system interface displays and saves them. Specifically, it can include preprocessing, target inference, postprocessing display and saving, and utilize the multi-core CPU and NPU of the system to work together to speed up; the image acquisition unit acquires X-ray security inspection images and preprocesses them through the image preprocessing unit; the target inference unit performs inference on the preprocessed images based on the RKNN model to obtain reasonable detection frames; the postprocessing display unit calls NMS (Non-Maximum Suppression) to filter out duplicate frames, and finally marks the original image with frames and displays the inference results, realizing real-time detection of X-ray security inspection images on an embedded platform.

[0118] Figure 8 It shows the flowchart of transmitting the converted RKNN model in the real-time security inspection image detection system based on the target chip in the present invention to the target chip for inference execution.

[0119] As Figure 8 shown, it is possible to first create an RKNN object S810. Based on the RKNN object, use the load_rknn interface to import the RKNN model S820: initialize the running environment S830 by calling the init_runtime interface, specifying the target platform and the NPU core scheduling mode; use the inference interface to perform inference and obtain the inference result S840. The input data needs to be preprocessed according to the model requirements. After the inference is completed, call the release interface to release the RKNN object S850 to release resources.

[0120] Figure 9 It shows the schematic diagram of the security inspection image detection result obtained by the security inspection image detection method in the present invention.

[0121] As Figure 9 shown, the image postprocessing result diagram of the real-time security inspection image detection system based on the target chip contains abnormal items such as Hammer, Wrench, and Lighter.

[0122] Figure 10 It shows the structural block diagram of the security inspection image detection device according to an embodiment of the present invention.

[0123] As Figure 10 shown, the security inspection image detection device 1000 of this embodiment includes an acquisition module 1010, a sliced convolutional feature obtaining module 1020, a residual convolutional feature obtaining module 1030, a pooled attention feature obtaining module 1040, a convolutional attention fusion obtaining module 1050, a fused feature obtaining module 1060, and a detection result obtaining module 1070.

[0124] An acquisition module 1010, configured to acquire a security inspection image, where the security inspection image is obtained by scanning an item using a ray scanning device;

[0125] A sliced convolution feature obtaining module 1020, configured to process the security inspection image using a sliced convolution network to obtain sliced convolution features;

[0126] A residual convolution feature obtaining module 1030, configured to process the sliced convolution features using a multi-residual convolution network to obtain residual convolution features;

[0127] A pooled attention feature obtaining module 1040, configured to process the residual convolution features based on an attention pooling network to obtain pooled attention features;

[0128] A convolutional attention fusion obtaining module 1050, configured to process the pooled attention features and the residual convolution features based on a feature pyramid network to obtain attention fusion features and convolutional attention features;

[0129] A fused feature obtaining module 1060, configured to process the attention fusion features and the convolutional attention features based on a path aggregation network to obtain fused features;

[0130] A detection result obtaining module 1070, configured to process the attention fusion features and the fused features based on a detection head to obtain a detection result of the security inspection image, where the detection result indicates whether the security inspection image contains abnormal items.

[0131] According to an embodiment of the present invention, by processing a security inspection image using a sliced convolution network, sliced convolution features can be obtained. By processing the sliced convolution features using a multi-residual convolution network, residual convolution features can be obtained. By processing the residual convolution features based on an attention pooling network, features at different scales corresponding to the residual convolution features can be pooled and the pooled features can be processed based on an attention mechanism to obtain pooled attention features, enriching semantic information, enhancing the receptive field, and improving the detection of small-size targets. By processing the pooled attention features and the residual convolution features based on a feature pyramid network, attention fusion features and convolutional attention features can be obtained. By processing the attention fusion features and the convolutional attention features based on a path aggregation network, fused features can be obtained. By processing the attention fusion features and the fused features based on a detection head, a detection result of the security inspection image can be obtained, thereby improving the detection accuracy, improving the detection performance of small-size targets and occluded targets, and enhancing the generalization ability of the model.

[0132] According to an embodiment of the present invention, the pooled attention feature obtaining module 1040 includes: a multi-scale fusion feature obtaining sub-module and a pooled attention feature obtaining sub-module.

[0133] The multi-scale fusion feature obtaining sub-module is configured to process the residual convolution features using a serial pooling layer to obtain multi-scale fusion features.

[0134] The sub-module for obtaining pooled attention features is used to process multi-scale fusion features based on a self-attention layer to obtain pooled attention features. The attention pooling network includes a serial pooling layer and a self-attention layer.

[0135] According to an embodiment of the present invention, the sub-module for obtaining multi-scale fusion features includes: a unit for obtaining the first residual normalized feature, a unit for obtaining multi-scale pooling features, and a unit for obtaining multi-scale fusion features.

[0136] The unit for obtaining the first residual normalized feature is used to process the residual convolution feature by using a convolutional normalization block to obtain the first residual normalized feature.

[0137] The unit for obtaining multi-scale pooling features is used to perform a multi-scale pooling operation on the first residual normalized feature to obtain multi-scale pooling features.

[0138] The unit for obtaining multi-scale fusion features is used to fuse the first residual normalized feature and the multi-scale pooling features to obtain multi-scale fusion features. The serial pooling layer includes a convolutional normalization block.

[0139] According to an embodiment of the present invention, the sub-module for obtaining pooled attention features includes: a convolutional block processing unit, a pixel correlation feature obtaining unit, an attention weight obtaining unit, a value attention feature obtaining unit, a value attention convolutional feature obtaining unit, and a pooled attention feature obtaining unit.

[0140] The convolutional block processing unit is used to process the multi-scale fusion features based on a convolutional block to obtain query features, key features, and value features.

[0141] The pixel correlation feature obtaining unit is used to multiply the query features and the key features to obtain pixel correlation features.

[0142] The attention weight obtaining unit is used to process the pixel correlation features by using a normalization algorithm to obtain attention weights.

[0143] The value attention feature obtaining unit is used to obtain value attention features according to the attention weights and the value features.

[0144] The value attention convolutional feature obtaining unit is used to process the value attention features based on a convolutional block to obtain value attention convolutional features.

[0145] The pooled attention feature obtaining unit is used to perform a residual connection and summation operation on the value attention convolutional features and the multi-scale fusion features to obtain pooled attention features.

[0146] According to an embodiment of the present invention, the multi-residual convolution network includes i feature extraction layers, where i is an integer greater than 1.

[0147] According to an embodiment of the present invention, the residual convolution feature obtaining module 1030 includes: a first residual convolution feature obtaining sub-module, an i-th residual convolution feature obtaining sub-module, and a determination sub-module.

[0148] The first residual convolution feature obtaining sub-module is configured to input the slice convolution feature into the first feature extraction layer and output a first residual convolution feature.

[0149] The i-th residual convolution feature obtaining sub-module is configured to input the (i - 1)-th residual convolution feature into the i-th feature extraction layer and output an i-th residual convolution feature.

[0150] The determination sub-module is configured to, when i = I, determine I residual convolution features based on the first residual convolution feature to the I-th residual convolution feature, where I ≥ i > 1 and I and i are integers.

[0151] According to an embodiment of the present invention, the detection result obtaining module 1070 includes: a first detection result obtaining sub-module and a second detection result obtaining sub-module.

[0152] The first detection result obtaining sub-module is configured to process the attention fusion feature based on the first detection head to obtain a first detection result of the security inspection image, and the first detection result represents whether there is an abnormal object with a first size in the security inspection image.

[0153] The second detection result obtaining sub-module is configured to process the fusion feature based on the second detection head to obtain a second detection result of the security inspection image, and the second detection result represents whether there is an abnormal object with a second size in the security inspection image. The size difference value between the first size and the second size is greater than a preset size threshold. The detection head includes a first detection head and a second detection head.

[0154] According to an embodiment of the present invention, the security inspection image detection device 1000 further includes: an intermediate image obtaining module and a security inspection image obtaining module.

[0155] The intermediate image obtaining module is configured to adjust the contrast and brightness of the initial security inspection image by using an image local adjustment algorithm to obtain an intermediate image.

[0156] The security inspection image obtaining module is configured to perform image enhancement on the intermediate image based on a mosaic algorithm to obtain a security inspection image.

[0157] According to an embodiment of the present invention, any number of the acquisition module 1010, the sliced convolutional feature obtaining module 1020, the residual convolutional feature obtaining module 1030, the pooled attention feature obtaining module 1040, the convolutional attention fusion obtaining module 1050, the fused feature obtaining module 1060, and the detection result obtaining module 1070 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the acquisition module 1010, the sliced convolutional feature obtaining module 1020, the residual convolutional feature obtaining module 1030, the pooled attention feature obtaining module 1040, the convolutional attention fusion obtaining module 1050, the fused feature obtaining module 1060, and the detection result obtaining module 1070 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in a suitable combination of any several of them. Alternatively, at least one of the acquisition module 1010, the sliced convolutional feature obtaining module 1020, the residual convolutional feature obtaining module 1030, the pooled attention feature obtaining module 1040, the convolutional attention fusion obtaining module 1050, the fused feature obtaining module 1060, and the detection result obtaining module 1070 can be at least partially implemented as a computer program module, and when the computer program module runs, it can execute corresponding functions.

[0158] Figure 11 FIG. shows a block diagram of an electronic device suitable for implementing the security inspection image detection method according to an embodiment of the present invention.

[0159] As Figure 11 shown, the electronic device according to an embodiment of the present invention includes a processor 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage section 1108 into a random access memory (RAM) 1103. The processor 1101 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 1101 can also include on-board memory for caching purposes. The processor 1101 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0160] In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are stored. The processor 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. The processor 1101 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 1102 and / or the RAM 1103. It should be noted that the programs may also be stored in one or more memories other than the ROM 1102 and the RAM 1103. The processor 1101 may also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.

[0161] According to an embodiment of the present invention, the electronic device 1100 may further include an input / output (I / O) interface 1105, and the input / output (I / O) interface 1105 is also connected to the bus 1104. The electronic device 1100 may further include one or more of the following components connected to the input / output (I / O) interface 1105: an input part 1106 including a keyboard, a mouse, etc.; an output part 1107 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 1108 including a hard disk, etc.; and a communication part 1109 including a network interface card such as a LAN card, a modem, etc. The communication part 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the input / output (I / O) interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed so that a computer program read from it can be installed into the storage part 1108 as needed.

[0162] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0163] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 1102 and / or RAM 1103 described above and / or one or more memories other than the ROM 1102 and RAM 1103.

[0164] An embodiment of the present invention also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the security inspection image detection method provided by the embodiment of the present invention.

[0165] When the computer program is executed by the processor 1101, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0166] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 1109, and / or be installed from the removable medium 1111. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0167] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1109, and / or be installed from the removable medium 1111. When the computer program is executed by the processor 1101, it executes the above functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0168] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0170] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0171] The above describes the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A security inspection image detection method, characterized in that: include: Acquiring a security inspection image, wherein the security inspection image is obtained by scanning the object using a ray scanning device; Processing the security inspection image using a slice convolutional network to obtain slice convolutional features; Processing the slice convolution features using a multi-residual convolutional network to obtain residual convolution features; Processing the residual convolutional features based on an attention pooling network to obtain a pooled attention feature; Processing the pooled attention feature and the residual convolution feature based on a feature pyramid network to obtain an attention fusion feature and a convolution attention feature; Processing the attention fusion feature and the convolutional attention feature based on a path aggregation network to obtain a fusion feature; The detection head processes the attention fusion feature and the fusion feature to obtain a detection result of the security inspection image, wherein the detection result indicates whether the security inspection image contains abnormal items.

2. The method according to claim 1, characterized in that The processing of the residual convolutional features based on the attention pooling network to obtain the pooled attention features includes: Processing the residual convolution features using a serial pooling layer to obtain multi-scale fusion features; The multi-scale fusion features are processed based on the self-attention layer to obtain pooled attention features, and the attention pooling network includes the serial pooling layer and the self-attention layer.

3. The method according to claim 2, characterized in that The method of processing the residual convolution features by using a serial pooling layer to obtain a multi-scale fusion feature includes: Processing the residual convolution feature using a convolution normalization block to obtain a first residual normalized feature; Performing a multi-scale pooling operation on the first residual normalized feature to obtain a multi-scale pooling feature; The first residual normalized feature and the multi-scale pooling feature are fused to obtain the multi-scale fused feature, and the serial pooling layer includes the convolution normalization block.

4. The method according to claim 2, characterized in that: The step of processing the multi-scale fusion features based on the self-attention layer to obtain the pooled attention features includes: Processing the multi-scale fusion features based on a convolution block to obtain query features, key features, and value features; Multiplying the query feature and the key feature to obtain a pixel-related feature; Processing the pixel association features using a normalization algorithm to obtain an attention weight; Obtaining a value attention feature according to the attention weight and the value feature; Processing the value attention feature based on the convolution block to obtain a value attention convolution feature; Residual connection and addition operations are performed on the value attention convolution feature and the multi-scale fusion feature to obtain the pooled attention feature.

5. The method according to claim 1, characterized in that The multi-residual convolutional network includes I feature extraction layers, where I is an integer greater than 1; The using a multi-residual convolutional network to process the slice convolutional features to obtain residual convolutional features includes: Inputting the slice convolution feature into the first feature extraction layer, and outputting the first residual convolution feature; Input the i-1th residual convolution feature to the i-th feature extraction layer, and output the i-th residual convolution feature; and When i=I, I residual convolution features are determined based on the 1st residual convolution feature to the Ith residual convolution feature, and I≥i>1.

6. The method according to claim 1, characterized in that The detection head-based processing of the attention fusion feature and the fusion feature to obtain the detection result of the security inspection image includes: Processing the attention fusion feature based on the first detection head to obtain a first detection result of the security inspection image, wherein the first detection result indicates whether there is an abnormal object with a first size in the security inspection image; Based on the second detection head processing the fusion feature, a second detection result of the security inspection image is obtained, the second detection result characterizes whether there is an abnormal object of a second size in the security inspection image, and the size difference between the first size and the second size is greater than a preset size threshold. The detection head includes the first detection head and the second detection head.

7. The method according to claim 1, characterized in that Before using the slice convolution network to process the security inspection image to obtain the slice convolution feature, the method further includes: The contrast and brightness of the initial security inspection image are adjusted using the image local adjustment algorithm to obtain an intermediate image; The intermediate image is enhanced based on a mosaic algorithm to obtain the security inspection image.

8. A security inspection image detection device, characterized in that: include: An acquisition module, used for acquiring a security inspection image, wherein the security inspection image is obtained by scanning an object using a ray scanning device; A slice convolution feature acquisition module, used for processing the security inspection image using a slice convolution network to obtain a slice convolution feature; A residual convolution feature acquisition module, used for processing the slice convolution feature using a multi-residual convolution network to obtain a residual convolution feature; A pooled attention feature obtaining module, used for processing the residual convolutional features based on an attention pooling network to obtain a pooled attention feature; A convolutional attention fusion obtaining module, used for processing the pooled attention feature and the residual convolution feature based on a feature pyramid network to obtain an attention fusion feature and a convolutional attention feature; A fusion feature obtaining module, used for processing the attention fusion feature and the convolution attention feature based on a path aggregation network to obtain a fusion feature; The detection result obtaining module is used to process the attention fusion feature and the fusion feature based on the detection head to obtain the detection result of the security inspection image, and the detection result represents whether the security inspection image contains abnormal items.

9. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Target recognition network model training method, electronic equipment and program product

    CN115439693A

  • X-ray security check image forbidden article detection method based on improved YOLOv7

    CN118397303A