Intelligent security check method and system based on image recognition

Through the multi-modal security inspection data fusion and intelligent early warning system, the problem of existing security inspection equipment inaccurate identification of items with complex shapes and similar density is solved, and an efficient and accurate security inspection process is achieved, which improves security inspection efficiency and consistency.

CN120431469AInactive Publication Date: 2025-08-05SHENYANG ANFENG ELECTRONIC ENG CO LTD

Patent Information

Application Number
CN202510920059.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing security inspection equipment is difficult to accurately distinguish items with complex shapes and similar density, and it is prone to missed inspections, and the processing speed is insufficient in situations where personnel and items flow are large, resulting in congestion in security inspection channels.

Method used

Multimodal security data fusion method is adopted, and by obtaining X-ray transmission images, visible surface images, millimeter wave three-dimensional imaging data and infrared thermal sensing data, the wavelet transform-bilateral filtering combined noise reduction algorithm, ResNet-50 and Swin Transformer hybrid neural networks are used for feature extraction and weighting fusion, and combining Hungarian algorithms for correlation matching to identify and warn of suspicious areas.

Benefits of technology

It improves the accuracy of identification of items with complex shapes and similar density, reduces false inspections and missed inspections, improves security inspection efficiency and accuracy, reduces wait time for personnel, rationally allocate resources, and reduces the work intensity and subjectivity of security inspection personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431469A_ABST
    Figure CN120431469A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent security check method and system based on image recognition. The method comprises the following steps: acquiring initial multi-mode security check data in a security check system; performing noise reduction processing on the data, establishing a hybrid neural network based on local feature extraction and global context relation captured by a parallel module, and performing dynamic weighted fusion on multi-modal features through a cross attention mechanism; inputting the noise-reduced multi-modal security check data into a hybrid neural network for identification to obtain a security check image suspicious region; and carrying out association matching on the security check image suspicious area by using a Hungary algorithm, and carrying out early warning according to a matching result. Articles with complex shapes and similar densities are effectively distinguished, false detection and missing detection are reduced, and the security check accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to an intelligent security inspection method and system based on image recognition. Background Art

[0002] Existing security inspection methods primarily rely on single-mode detection equipment, and X-ray security inspection equipment and metal detectors have numerous limitations. While X-ray security inspection equipment can penetrate objects to reveal their internal structure, it struggles to accurately distinguish items with complex shapes and similar densities, leading to false detections and missed detections. In high-volume environments, the processing speed of traditional security inspection equipment is insufficient to meet actual requirements, easily leading to congestion in security inspection lanes and reduced traffic efficiency. Summary of the Invention

[0003] The purpose of the present invention is to solve the above problems and design an intelligent security inspection method and system based on image recognition.

[0004] To achieve the above-mentioned purpose, the technical solution of the present invention is a smart security inspection method based on image recognition. Furthermore, in the above-mentioned smart security inspection method based on image recognition, the smart security inspection method based on image recognition includes the following steps: Acquire X-ray transmission images, visible light surface images, millimeter wave three-dimensional imaging data, and infrared thermal data in the security inspection system to obtain multimodal security inspection data, and preprocess the multimodal security inspection data to obtain initial multimodal security inspection data; Dynamically adjust the filter parameters according to the image signal-to-noise ratio in the initial multimodal security inspection data using a wavelet transform-bilateral filtering joint noise reduction algorithm, and perform noise reduction processing on the data to obtain noise-reduced multimodal security inspection data; Based on ResNet-50 to extract local features, the parallel Swin Transformer module captures the global contextual relationship to establish a hybrid neural network. The multimodal features are dynamically weighted and fused through the cross-attention mechanism to obtain the ResNet-Transformer hybrid neural network. Inputting the denoised multimodal security inspection data into the ResNet-Transformer hybrid neural network for identification to obtain suspicious areas of the security inspection image; The Hungarian algorithm is used to perform correlation matching on the suspicious areas of the security inspection image to obtain a matching result, and an early warning is issued based on the matching result.

[0005] Furthermore, in the above-mentioned intelligent security inspection method based on image recognition, the preprocessing of the multimodal security inspection data to obtain initial multimodal security inspection data includes: The polynomial fitting algorithm is used to correct the grayscale of the X-ray transmission image, and the histogram equalization algorithm is used to expand the image grayscale dynamic range to obtain the corrected X-ray transmission image; The visible light surface image is denoised by a median filter algorithm to remove salt and pepper noise, and the color balance of the image is adjusted by calculating the RGB value of the white reference point in the image to obtain an adjusted visible light surface image; Converting the polar coordinate data of the millimeter wave three-dimensional imaging data into Cartesian coordinate data, and converting the polar coordinates of each three-dimensional data point into Cartesian coordinates using a coordinate transformation formula to obtain converted millimeter wave three-dimensional imaging data; The image registration algorithm based on feature points extracts feature points from infrared thermal images, and uses the corresponding feature point pairs to calculate the transformation matrix to obtain the registered infrared thermal data.

[0006] Furthermore, in the above-mentioned intelligent security inspection method based on image recognition, the wavelet transform-bilateral filtering joint noise reduction algorithm is used to dynamically adjust the filtering parameters according to the image signal-to-noise ratio in the initial multimodal security inspection data, and the data is subjected to noise reduction processing to obtain the noise-reduced multimodal security inspection data, including: The image signal-to-noise ratio (SNR) of the initial multimodal security inspection data is calculated using a local area statistics method, and the number of decomposition layers and threshold of the wavelet transform are dynamically adjusted according to the average SNR of the image. Adjusting the spatial domain standard deviation and the range standard deviation of the bilateral filter according to the image signal-to-noise ratio to balance the spatial filtering range and the sensitivity of the range difference; According to the adjusted denoising algorithm, each sub-band image is thresholded to remove high-frequency noise, and the processed wavelet coefficients are inversely transformed to obtain a preliminarily denoised image; The image after preliminary denoising is subjected to bilateral filtering to remove low-frequency noise and obtain denoised multimodal security inspection data.

[0007] Furthermore, in the above-mentioned intelligent security inspection method based on image recognition, the local features are extracted based on ResNet-50, and the global contextual relationship is captured in parallel with the Swin Transformer module to establish a hybrid neural network. The multimodal features are dynamically weighted and fused through the cross-attention mechanism to obtain the ResNet-Transformer hybrid neural network, including: ResNet-50 uses a five-layer convolutional block structure, and the Swin Transformer module has four layers, each of which contains multiple self-attention layers and feedforward neural network layers; the cross-attention mechanism dynamically weights and fuses features by calculating the similarity matrix between features of different modalities.

[0008] Furthermore, in the above-mentioned intelligent security inspection method based on image recognition, the inputting of the denoised multimodal security inspection data into the ResNet-Transformer hybrid neural network for identification to obtain suspicious areas of the security inspection image includes: The input data is passed through the convolutional layer of ResNet-50 for local feature extraction, and then passes through each convolution block in turn to obtain local feature maps at different levels; The local feature map is flattened into a sequence and input into the Swin Transformer module. The Swin Transformer module captures the global contextual relationship through multiple self-attention layers and feedforward neural network layers to obtain a global feature vector. Taking the local feature map and the global feature vector as input, performing dynamic weighted fusion of multimodal features through a cross-attention mechanism to obtain a fused feature vector; The fused feature vector passes through a convolution layer and a sigmoid activation function, outputting a probability map of the same size as the input image to obtain the suspicious area of the security inspection image.

[0009] Furthermore, in the above-mentioned intelligent security inspection method based on image recognition, the use of the Hungarian algorithm to perform correlation matching on the suspicious areas of the security inspection image to obtain a matching result, and issuing an early warning based on the matching result, includes: The Hungarian algorithm is used to calculate the Euclidean distance between the center position of the suspicious area of the current frame and the center position of the suspicious area of the previous frame; Extract the shape features of the suspicious area and calculate the Euclidean distance between the shape features of the suspicious area in the current frame and the previous frame. The smaller the distance, the higher the shape similarity. The feature vector of the suspicious area extracted by the hybrid neural network is used to calculate the cosine similarity between the feature vector of the suspicious area of the current frame and the previous frame. The higher the similarity, the higher the feature similarity.

[0010] Furthermore, in the above-mentioned intelligent security inspection method based on image recognition, the method of using the Hungarian algorithm to perform correlation matching on the suspicious areas of the security inspection image to obtain a matching result, and issuing an early warning based on the matching result, further includes: Perform weighted fusion of position similarity, shape similarity and feature similarity to obtain a comprehensive similarity score; Construct a similarity matrix, where the rows represent the suspicious areas of the current frame, the columns represent the suspicious areas of the previous frame, and the matrix elements are the comprehensive similarity scores of the corresponding suspicious areas. Solve the similarity matrix to obtain the matching results.

[0011] An intelligent security inspection system based on image recognition, comprising the following modules: A security inspection data acquisition module is used to acquire X-ray transmission images, visible light surface images, millimeter wave three-dimensional imaging data, and infrared thermal data in the security inspection system to obtain multimodal security inspection data, and preprocess the multimodal security inspection data to obtain initial multimodal security inspection data; a security inspection data processing module, configured to dynamically adjust filtering parameters according to the image signal-to-noise ratio in the initial multimodal security inspection data using a wavelet transform-bilateral filtering joint noise reduction algorithm, and perform noise reduction processing on the data to obtain noise-reduced multimodal security inspection data; A hybrid model building module is used to extract local features based on ResNet-50, capture global contextual relationships in parallel with the Swin Transformer module to establish a hybrid neural network, and dynamically weightedly fuse multimodal features through a cross-attention mechanism to obtain a ResNet-Transformer hybrid neural network; A security inspection image recognition module, configured to input the denoised multimodal security inspection data into the ResNet-Transformer hybrid neural network for recognition, thereby obtaining suspicious areas of the security inspection image; The result matching warning module is used to use the Hungarian algorithm to perform correlation matching on the suspicious area of the security inspection image to obtain a matching result, and issue an early warning based on the matching result.

[0012] Furthermore, in an intelligent security inspection system based on image recognition, the result matching warning module includes the following submodules: A calculation submodule, for calculating the Euclidean distance between the center positions of the suspicious area of the current frame and the suspicious area of the previous frame using the Hungarian algorithm; The extraction submodule is used to extract the shape features of the suspicious area and calculate the Euclidean distance between the shape features of the suspicious area in the current frame and the previous frame. The smaller the distance, the higher the shape similarity. The judgment submodule is used to calculate the cosine similarity between the suspicious area feature vectors of the current frame and the previous frame using the suspicious area feature vector extracted by the hybrid neural network. The higher the similarity, the higher the feature similarity.

[0013] Furthermore, in an intelligent security inspection system based on image recognition, the result matching warning module includes the following submodules: The fusion submodule is used to perform weighted fusion of position similarity, shape similarity and feature similarity to obtain a comprehensive similarity score; The construction submodule is used to construct a similarity matrix, where the rows represent the suspicious areas of the current frame, the columns represent the suspicious areas of the previous frame, and the matrix elements are the comprehensive similarity scores of the corresponding suspicious areas. Solving the similarity matrix obtains the matching results.

[0014] Its beneficial effects include: 1. It can more comprehensively and accurately describe object characteristics, effectively distinguishing items with complex shapes and similar densities, reducing false detections and missed detections, and improving security inspection accuracy. 2. It can rapidly process input multimodal security inspection data and output identification results for suspicious areas. It then quickly makes early warning decisions based on matching results and confidence levels, making the entire security inspection process efficient and smooth, effectively improving inspection efficiency and reducing waiting times for both personnel and items. 3. It enables security inspectors to conduct targeted inspections based on early warning levels, rationally allocating security inspection resources, and improving the scientific and effective nature of security inspections. Furthermore, this intelligent early warning system can reduce the workload and subjectivity of security inspectors and improve the consistency and accuracy of security inspection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Various other advantages and benefits will become apparent to those skilled in the art by reading the following detailed description of the preferred embodiment.The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present invention.

[0016] Figure 1 This is a schematic diagram of a first embodiment of an intelligent security inspection method based on image recognition in an embodiment of the present invention; Figure 2 Schematic diagram of a second embodiment of an intelligent security inspection method based on image recognition in an embodiment of the present invention; Figure 3 This is a schematic diagram of the first embodiment of an intelligent security inspection system based on image recognition in an embodiment of the present invention. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0018] Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a", "an", "" and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0019] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 As shown, a smart security inspection method based on image recognition includes the following steps: Step 101: Acquire X-ray transmission images, visible light surface images, millimeter wave three-dimensional imaging data, and infrared thermal data from a security inspection system to obtain multimodal security inspection data, and preprocess the multimodal security inspection data to obtain initial multimodal security inspection data. Specifically, in this embodiment, a polynomial fitting algorithm is used to perform grayscale correction on the X-ray transmission image, and a histogram equalization algorithm is used to expand the image grayscale dynamic range to obtain a corrected X-ray transmission image; The visible light surface image is denoised by a median filter algorithm to remove salt and pepper noise, and the color balance of the image is adjusted by calculating the RGB value of the white reference point in the image to obtain an adjusted visible light surface image; Converting the polar coordinate data of the millimeter wave three-dimensional imaging data into Cartesian coordinate data, and converting the polar coordinates of each three-dimensional data point into Cartesian coordinates using a coordinate transformation formula to obtain converted millimeter wave three-dimensional imaging data; The image registration algorithm based on feature points extracts feature points from infrared thermal images, and uses the corresponding feature point pairs to calculate the transformation matrix to obtain the registered infrared thermal data.

[0020] Specifically, 1. Multimodal Security Inspection Data Acquisition 1. X-ray transmission image acquisition Dual-energy X-ray security inspection equipment is used, which is equipped with two ray sources with different energies, namely a high-energy ray source and a low-energy ray source.

[0021] 2. Visible light surface image acquisition Multi-angle adjustable high-definition cameras are installed above and on both sides of the security check channel. The camera resolution is not less than 1920×1080 and the frame rate is 30fps.

[0022] 3. Millimeter-wave 3D imaging data acquisition It uses millimeter-wave three-dimensional imaging radar, which operates at a frequency of 77GHz and has high resolution and penetration capabilities.

[0023] 4. Infrared thermal data acquisition An infrared thermal imager with an operating wavelength of 8-14μm is used to detect the temperature distribution on the surface of an object.

[0024] 2. Preprocessing Operations 1. X-ray transmission image preprocessing Grayscale Correction: Non-uniformity in X-ray source intensity can lead to uneven grayscale distribution in the image. A polynomial fitting algorithm is used to correct image grayscale values. First, multiple regions of known uniform material are selected as reference areas within the image and the average grayscale value of these regions is calculated. Then, a polynomial fitting algorithm (typically a quadratic polynomial) is used to establish a mapping between X-ray source intensity and grayscale value. This correction is then applied to the grayscale values of the entire image, eliminating the effects of uneven X-ray source intensity.

[0025] Contrast enhancement: A histogram equalization algorithm is used to expand the image's grayscale dynamic range, improving the discernibility of object details. The specific steps are: calculating the image's grayscale histogram and counting the frequency of each grayscale level; calculating the cumulative distribution function (CDF) to map the original grayscale values to new grayscale values, ensuring that the new grayscale histogram is as evenly distributed as possible. To avoid excessive noise enhancement, the histogram equalization process employs contrast-limited adaptive histogram equalization (CLAHE). This divides the histogram into multiple small blocks, which are then equalized separately while limiting the amount of contrast enhancement applied to each block.

[0026] 1. Visible light surface image preprocessing Denoising: A median filter algorithm is used to denoise the image, removing salt and pepper noise while preserving edge details. The median filter window size is selected based on the intensity of image noise and the richness of image detail; a 3×3 or 5×5 window is typically used. For each pixel, the pixel values in its neighborhood are sorted and the median value is used as the new value for that pixel, effectively removing salt and pepper noise.

[0027] White balance adjustment: By calculating the RGB values of a white reference point in the image, the color balance of the entire image is adjusted to achieve realistic color reproduction. First, a known white object (a standard white calibration plate) is selected as a reference point in the image. The average values of the R, G, and B channels at this reference point are calculated. Based on these average values, the gain factor for each channel is calculated. Gain adjustment is then performed on each channel of the entire image to ensure that the RGB values at the white reference point are approximately equal (R=G=B), thus achieving white balance adjustment.

[0028] 1. Millimeter-wave 3D imaging data preprocessing Coordinate Conversion: Converts polar coordinate data acquired by the millimeter-wave 3D imaging radar into Cartesian coordinate data for subsequent processing and analysis. Based on the radar's installation location and scanning parameters, a conversion relationship between polar and Cartesian coordinates is established. The polar coordinates (radius, angle, height) of each 3D data point are converted to Cartesian coordinates (x, y, z) using a coordinate transformation formula.

[0029] Data normalization: Since the range of millimeter-wave 3D imaging data can be large, the data is normalized to facilitate neural network processing. Each coordinate component (x, y, z) is normalized to the range of [-1, 1].

[0030] 1. Infrared thermal data preprocessing Temperature calibration: Because infrared thermal imagers can have temperature errors, they require temperature calibration. Using a standard temperature source (blackbody radiation source), the thermal imager is calibrated to establish a calibration curve between measured and actual temperatures. When acquiring infrared thermal data, the temperature value of each pixel is corrected according to the calibration curve to improve temperature measurement accuracy.

[0031] Image Registration: To precisely align the infrared thermal data with the visible surface image in space, a feature point-based image registration algorithm is employed. First, feature points (SIFT feature points) are extracted from the visible surface image and the infrared thermal image. Then, a feature point matching algorithm (BF matching) is used to find corresponding feature point pairs in the two images. Finally, a transformation matrix (affine or perspective) is calculated using these corresponding feature point pairs, and the infrared thermal image is transformed to align it with the visible surface image.

[0032] Step 102: Using a wavelet transform-bilateral filtering joint noise reduction algorithm, the filter parameters are dynamically adjusted according to the image signal-to-noise ratio in the initial multimodal security inspection data, and the data is subjected to noise reduction processing to obtain noise-reduced multimodal security inspection data; Specifically, in this embodiment, the image signal-to-noise ratio in the initial multimodal security inspection data is calculated by a local area statistics method, and the number of decomposition layers and thresholds of the wavelet transform are dynamically adjusted according to the average signal-to-noise ratio of the image; Adjust the spatial domain standard deviation and range standard deviation of bilateral filtering according to the image signal-to-noise ratio to balance the spatial filtering range and range difference sensitivity; According to the adjusted denoising algorithm, each sub-band image is thresholded to remove high-frequency noise, and the processed wavelet coefficients are inversely transformed to obtain a preliminarily denoised image; The image after preliminary denoising is subjected to bilateral filtering to remove low-frequency noise and obtain denoised multimodal security inspection data.

[0033] Specifically, 1. Image Signal-to-Noise Ratio Calculation The image signal-to-noise ratio (SNR) is calculated using the local region statistics method. First, the image is divided into multiple non-overlapping local regions, each 16×16 pixels in size. For each local region, the signal mean and noise variance are calculated. The signal mean is the average of the pixel values within the region, and the noise variance is the average of the squares of the differences between the pixel values within the region and the signal mean. The average SNR for the entire image is then calculated by averaging the SNRs of all local regions.

[0034] 2. Wavelet transform-bilateral filtering joint noise reduction algorithm 1. Wavelet transform parameter adjustment Dynamically adjust the number of wavelet transform decomposition layers and threshold based on the image's average signal-to-noise ratio (SNR). When the SNR is high (SNR > 30dB), indicating low image noise, the number of decomposition layers is set to 3, and a soft threshold function is used with a threshold value of 3, where σ is the noise standard deviation (calculated from the local noise variance) to preserve more image detail. When the SNR is moderate (20dB ≤ SNR ≤ 30dB), the number of decomposition layers is set to 4, with a threshold value of 2.5, to balance noise reduction and detail preservation. When the SNR is low (SNR < 20dB), the number of decomposition layers is set to 5, with a threshold value of 2\sigma_n to enhance noise reduction.

[0035] 2. Bilateral filtering parameter adjustment The spatial domain standard deviation and range standard deviation of bilateral filtering are adjusted according to the signal-to-noise ratio. The spatial domain standard deviation determines the spatial range of the filter, and the range standard deviation determines the sensitivity of the filter to pixel value differences.

[0036] 3. Noise reduction process First, the image data in the initial multimodal security inspection data is subjected to a wavelet transform, decomposing it into different frequency subbands. Each subband image is then thresholded to remove high-frequency noise. The processed wavelet coefficients are then inverse-transformed to obtain a preliminarily denoised image. Finally, the preliminarily denoised image is subjected to bilateral filtering to further remove low-frequency noise, resulting in denoised multimodal security inspection data. For millimeter-wave 3D imaging data, since it is a 3D point cloud, a bilateral filtering method based on spatial distance is used to filter the neighborhood of each point cloud data point to remove outlier noise.

[0037] Step 103: Extract local features based on ResNet-50, connect the Swin Transformer module in parallel to capture the global contextual relationship to establish a hybrid neural network, and dynamically weight the multimodal features through the cross attention mechanism to obtain the ResNet-Transformer hybrid neural network; Specifically, in this embodiment, ResNet-50 uses a five-layer convolutional block structure, and the Swin Transformer module has 4 layers, each of which contains multiple self-attention layers and feedforward neural network layers; the cross-attention mechanism dynamically weights and fuses features by calculating the similarity matrix between different modal features.

[0038] Specifically, 1. ResNet-50 Network Structure ResNet-50 uses the classic five-layer convolutional block structure, as follows: 1. The first layer is a convolutional layer that uses a 7×7 kernel, a stride of 2, and a padding of 3. The number of input channels is determined by the type of multimodal data, and the number of output channels is 64. Batch normalization and ReLU activation functions are then performed.

[0039] 2. The second layer consists of three residual blocks, each of which contains two 3×3 convolutional layers with a stride of 1 and padding of 1. The first residual block has 64 input channels and 64 output channels. The second residual block increases the number of input channels to 128 using a 1×1 convolutional layer with a stride of 2 to reduce spatial resolution. The third residual block has 128 input channels and 128 output channels. Each residual block performs batch normalization and ReLU activation, and residual connections directly add the input to the output.

[0040] 3. The third layer consists of four residual blocks with a similar structure to the second layer. A 1×1 convolutional layer is used to increase the number of channels to 256 with a stride of 2, reducing the spatial resolution. Each residual block has 256 input channels and 256 output channels.

[0041] 4. The fourth layer consists of 6 residual blocks, with the number of channels increased to 512 and a stride of 2, reducing the spatial resolution. Each residual block has 512 input channels and 512 output channels.

[0042] 5. The fifth layer: the average pooling layer, which reduces the spatial size of the feature map to 1×1, and then connects to the fully connected layer. However, in this hybrid neural network, the output of ResNet-50 is the extracted local feature map, which is used for subsequent fusion with the Swin Transformer module.

[0043] 2. Swin Transformer module parameter settings The Swin Transformer module has 4 layers, each of which contains multiple self-attention layers and feedforward neural network layers. The specific parameters are as follows: 1. Number of layers: 4.

[0044] 2. Number of heads: 8 heads, each head has a dimension of 64, and the total dimension is 512.

[0045] 3. Window size: 7×7, using a sliding window mechanism to achieve cross-window information interaction.

[0046] 4. Relative position encoding: Introducing relative position encoding into self-attention calculation to improve the modeling ability of spatial position information.

[0047] 5. Feedforward neural network: Contains two fully connected layers, connected by a GELU activation function, and the hidden layer dimension is 2048.

[0048] 3. Implementation of Cross-Attention Mechanism The cross-attention mechanism calculates the similarity matrix between features of different modalities and dynamically weights the features. The specific steps are as follows: 1. Feature conversion: The local features extracted by ResNet-50 and the global context features captured by the Swin Transformer module are converted into query vector Q, key vector K, and value vector V through linear transformation. Feature conversion is performed separately for multimodal features.

[0049] 2. Similarity calculation: Calculate the dot product similarity between the query vector Q_i and other modal key vectors K_j (j=1,2,3,4) to obtain the similarity matrix. To stabilize the gradient, the similarity matrix is scaled.

[0050] 3. Weight calculation: Normalize the similarity matrix through the softmax function to obtain the attention weight matrix.

[0051] Feature fusion: The value vectors of each modality are weighted and summed according to the attention weight matrix to obtain the fused feature vector. In this way, dynamic weighted fusion of multimodal features is achieved, highlighting the features of other modalities related to the current modality and suppressing irrelevant features.

[0052] Step 104: Input the denoised multimodal security inspection data into a ResNet-Transformer hybrid neural network for identification to obtain suspicious areas in the security inspection image; Specifically, in this embodiment, the input data is subjected to local feature extraction by the convolutional layer of ResNet-50, and then passes through each convolution block in turn to obtain local feature maps at different levels; The local feature map is flattened into a sequence and input into the Swin Transformer module. The Swin Transformer module captures the global contextual relationship through multiple self-attention layers and feedforward neural network layers to obtain a global feature vector. Taking the local feature map and the global feature vector as input, the multimodal features are dynamically weighted fused through the cross-attention mechanism to obtain the fused feature vector. The fused feature vector passes through a convolution layer and a sigmoid activation function, outputting a probability map of the same size as the input image to obtain the suspicious area of the security inspection image.

[0053] Specifically, 1. Input Data Processing Convert denoised multimodal security inspection data into a format suitable for hybrid neural network input.

[0054] 2. Forward Propagation Process 1. ResNet-50 feature extraction: The input data first passes through the convolutional layer of ResNet-50 for local feature extraction. It then passes through each convolutional block in sequence to obtain local feature maps at different levels. The output feature map of the last convolutional block has a size of 7×7×512 (for an input size of 224×224).

[0055] 2. Global Feature Capture in the Swin Transformer Module: The local feature maps output by ResNet-50 are flattened into sequences and fed into the Swin Transformer module. The Swin Transformer module captures global contextual relationships through multiple self-attention layers and feedforward neural network layers, generating a global feature vector.

[0056] 3. Feature fusion using the cross-attention mechanism: The local features extracted by ResNet-50 and the global features captured by the Swin Transformer module are used as input, and multimodal features are dynamically weighted and fused through the cross-attention mechanism to obtain the fused feature vector.

[0057] Suspicious area prediction: The fused feature vector passes through a convolution layer and a sigmoid activation function, and outputs a probability map with the same size as the input image. The value of each pixel in the probability map represents the probability that the location is a suspicious area. By setting a threshold (0.5), the probability Figure 2 The binary mask of the suspicious area of the security inspection image is obtained by quantization. The area with a value of 1 in the mask represents the suspicious area, and the position coordinates and confidence score of the suspicious area are output at the same time.

[0058] Step 105: Use the Hungarian algorithm to perform correlation matching on the suspicious areas of the security inspection image to obtain a matching result, and issue an early warning based on the matching result.

[0059] Specifically, in this embodiment, the Hungarian algorithm is used to calculate the Euclidean distance between the center positions of the suspicious area of the current frame and the suspicious area of the previous frame; Extract the shape features of the suspicious area and calculate the Euclidean distance between the shape features of the suspicious area in the current frame and the previous frame. The smaller the distance, the higher the shape similarity. The feature vector of the suspicious area extracted by the hybrid neural network is used to calculate the cosine similarity between the feature vector of the suspicious area of the current frame and the previous frame. The higher the similarity, the higher the feature similarity.

[0060] Perform weighted fusion of position similarity, shape similarity and feature similarity to obtain a comprehensive similarity score; Construct a similarity matrix, where the rows represent the suspicious areas of the current frame, the columns represent the suspicious areas of the previous frame, and the matrix elements are the comprehensive similarity scores of the corresponding suspicious areas. Solve the similarity matrix to obtain the matching results.

[0061] Specifically, 1. Association Matching Basis Position similarity: Calculate the Euclidean distance between the center position of the suspicious area in the current frame and the center position of the suspicious area in the previous frame. The smaller the distance, the higher the position similarity.

[0062] Shape similarity: Extract the shape features of the suspicious area, such as area, perimeter, aspect ratio, etc., and calculate the Euclidean distance or cosine similarity between the shape features of the suspicious area in the current frame and the previous frame. The smaller the distance or the higher the similarity, the higher the shape similarity.

[0063] Feature similarity: The feature vector of the suspicious area extracted by the hybrid neural network is used to calculate the cosine similarity between the feature vector of the suspicious area of the current frame and the previous frame. The higher the similarity, the higher the feature similarity.

[0064] 2. Matching result calculation A weighted fusion of positional similarity, shape similarity, and feature similarity is performed to generate a comprehensive similarity score. Weighting coefficients are adjusted based on the actual security inspection scenario; typically, the weight for positional similarity is 0.4, the weight for shape similarity is 0.3, and the weight for feature similarity is 0.3. A similarity matrix is constructed, where rows represent suspicious areas in the current frame, columns represent suspicious areas in the previous frame, and the matrix elements represent the comprehensive similarity scores for the corresponding suspicious areas. The Hungarian algorithm is then used to solve this similarity matrix to obtain the optimal matching result and determine the correspondence between the suspicious areas in the current frame and the previous frame.

[0065] 3. Early Warning Strategy Issue an early warning based on the matching results and the confidence level of the suspicious area.

[0066] High-confidence warning (red warning): When the confidence level of a suspicious area is ≥0.9, and after correlation matching, the suspicious area appears in more than three consecutive frames and has dangerous characteristics (consistent with the characteristic patterns of objects such as explosives and weapons), a red warning is issued, prompting security personnel to conduct a focused inspection.

[0067] Medium confidence warning (yellow warning): When the confidence level of a suspicious area is between 0.7 and 0.9, or if the suspicious area appears in two consecutive frames after correlation matching, a yellow warning is issued, prompting security personnel to conduct further inspection.

[0068] Low confidence warning (blue warning): When the confidence level of a suspicious area is between 0.5 and 0.7, a blue warning is issued as a general warning to remind security personnel to pay attention to the area.

[0069] Its beneficial effects include: 1. It can more comprehensively and accurately describe object characteristics, effectively distinguishing items with complex shapes and similar densities, reducing false detections and missed detections, and improving security inspection accuracy. 2. It can rapidly process input multimodal security inspection data and output identification results for suspicious areas. It then quickly makes early warning decisions based on matching results and confidence levels, making the entire security inspection process efficient and smooth, effectively improving inspection efficiency and reducing waiting times for both personnel and items. 3. It enables security inspectors to conduct targeted inspections based on early warning levels, rationally allocating security inspection resources, and improving the scientific and effective nature of security inspections. Furthermore, this intelligent early warning system can reduce the workload and subjectivity of security inspectors and improve the consistency and accuracy of security inspection results.

[0070] See also Figure 2 In an intelligent security inspection method based on image recognition, a wavelet transform-bilateral filtering joint noise reduction algorithm is used to dynamically adjust the filter parameters according to the image signal-to-noise ratio in the initial multimodal security inspection data, and the data is subjected to noise reduction processing to obtain the noise-reduced multimodal security inspection data. The method includes the following steps: Step 201: Calculate the image signal-to-noise ratio (SNR) of the initial multimodal security inspection data using a local area statistics method, and dynamically adjust the number of decomposition layers and threshold of the wavelet transform based on the average SNR of the image. Step 202: Adjust the spatial domain standard deviation and range standard deviation of the bilateral filter according to the image signal-to-noise ratio to balance the spatial filtering range and the sensitivity of the range difference; Step 203: Perform threshold processing on each sub-band image according to the adjusted noise reduction algorithm to remove high-frequency noise, and perform inverse transformation on the processed wavelet coefficients to obtain a preliminarily noise-reduced image; Step 204: Perform bilateral filtering on the image after preliminary noise reduction to remove low-frequency noise, thereby obtaining noise-reduced multimodal security inspection data.

[0071] The above is an introduction to an embodiment of the intelligent security inspection method based on image recognition of the present invention. Figure 3 , an intelligent security inspection system based on image recognition, including the following modules: The security inspection data acquisition module is used to acquire X-ray transmission images, visible light surface images, millimeter wave three-dimensional imaging data, and infrared thermal data from the security inspection system to obtain multimodal security inspection data, and pre-process the multimodal security inspection data to obtain initial multimodal security inspection data; The security inspection data processing module is used to dynamically adjust the filtering parameters according to the image signal-to-noise ratio in the initial multimodal security inspection data using a wavelet transform-bilateral filtering joint noise reduction algorithm, and to perform noise reduction on the data to obtain noise-reduced multimodal security inspection data; A hybrid model building module is used to extract local features based on ResNet-50, capture global contextual relationships in parallel with the Swin Transformer module to establish a hybrid neural network, and dynamically weightedly fuse multimodal features through a cross-attention mechanism to obtain a ResNet-Transformer hybrid neural network; The security inspection image recognition module is used to input the denoised multimodal security inspection data into the ResNet-Transformer hybrid neural network for recognition, and obtain suspicious areas in the security inspection image; The result matching warning module is used to use the Hungarian algorithm to perform correlation matching on suspicious areas of the security inspection image, obtain matching results, and issue warnings based on the matching results.

[0072] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent security inspection method based on image recognition, characterized in that: The image recognition-based intelligent security inspection method includes the following steps: Acquire X-ray transmission images, visible light surface images, millimeter wave three-dimensional imaging data, and infrared thermal data in the security inspection system to obtain multimodal security inspection data, and preprocess the multimodal security inspection data to obtain initial multimodal security inspection data; Dynamically adjust the filter parameters according to the image signal-to-noise ratio in the initial multimodal security inspection data using a wavelet transform-bilateral filtering joint noise reduction algorithm, and perform noise reduction processing on the data to obtain noise-reduced multimodal security inspection data; Based on ResNet-50 to extract local features, the parallel Swin Transformer module captures the global contextual relationship to establish a hybrid neural network. The multimodal features are dynamically weighted and fused through the cross-attention mechanism to obtain the ResNet-Transformer hybrid neural network. Inputting the denoised multimodal security inspection data into the ResNet-Transformer hybrid neural network for identification to obtain suspicious areas of the security inspection image; The Hungarian algorithm is used to perform correlation matching on the suspicious areas of the security inspection image to obtain a matching result, and an early warning is issued based on the matching result.

2. The intelligent security inspection method based on image recognition according to claim 1, characterized in that: The preprocessing of the multimodal security inspection data to obtain initial multimodal security inspection data includes: The polynomial fitting algorithm is used to correct the grayscale of the X-ray transmission image, and the histogram equalization algorithm is used to expand the image grayscale dynamic range to obtain the corrected X-ray transmission image; The visible light surface image is denoised by a median filter algorithm to remove salt and pepper noise, and the color balance of the image is adjusted by calculating the RGB value of the white reference point in the image to obtain an adjusted visible light surface image; Converting the polar coordinate data of the millimeter wave three-dimensional imaging data into Cartesian coordinate data, and converting the polar coordinates of each three-dimensional data point into Cartesian coordinates using a coordinate transformation formula to obtain converted millimeter wave three-dimensional imaging data; The image registration algorithm based on feature points extracts feature points from infrared thermal images, and uses corresponding feature point pairs to calculate the transformation matrix to obtain the registered infrared thermal data.

3. The intelligent security inspection method based on image recognition according to claim 1, characterized in that: The wavelet transform-bilateral filtering combined noise reduction algorithm is used to dynamically adjust the filtering parameters according to the image signal-to-noise ratio in the initial multimodal security inspection data, and the data is subjected to noise reduction processing to obtain the noise-reduced multimodal security inspection data, including: The image signal-to-noise ratio (SNR) of the initial multimodal security inspection data is calculated using a local region statistics method, and the number of decomposition layers and threshold of the wavelet transform are dynamically adjusted according to the average SNR of the image. Adjusting the spatial domain standard deviation and the range standard deviation of the bilateral filter according to the image signal-to-noise ratio to balance the spatial filtering range and the sensitivity of the range difference; According to the adjusted denoising algorithm, each sub-band image is thresholded to remove high-frequency noise, and the processed wavelet coefficients are inversely transformed to obtain a preliminarily denoised image; The image after preliminary denoising is subjected to bilateral filtering to remove low-frequency noise and obtain denoised multimodal security inspection data.

4. The intelligent security inspection method based on image recognition according to claim 1, characterized in that: The method extracts local features based on ResNet-50, connects the Swin Transformer module in parallel to capture the global contextual relationship, establishes a hybrid neural network, and dynamically weights and fuses multimodal features through the cross-attention mechanism to obtain the ResNet-Transformer hybrid neural network, including: ResNet-50 uses a five-layer convolutional block structure, and the Swin Transformer module has four layers, each of which contains multiple self-attention layers and feedforward neural network layers; the cross-attention mechanism dynamically weights and fuses features by calculating the similarity matrix between features of different modalities.

5. The intelligent security inspection method based on image recognition according to claim 1, characterized in that: The step of inputting the denoised multimodal security inspection data into the ResNet-Transformer hybrid neural network for identification to obtain suspicious areas of the security inspection image includes: The input data is passed through the convolutional layer of ResNet-50 for local feature extraction, and then passes through each convolution block in turn to obtain local feature maps at different levels; The local feature map is flattened into a sequence and input into the Swin Transformer module. The Swin Transformer module captures the global contextual relationship through multiple self-attention layers and feedforward neural network layers to obtain a global feature vector. The local feature map and the global feature vector are used as input, and multimodal features are dynamically weighted fused through a cross-attention mechanism to obtain a fused feature vector; the fused feature vector passes through a convolutional layer and a sigmoid activation function, and a probability map with the same size as the input image is output to obtain the suspicious area of the security inspection image.

6. The intelligent security inspection method based on image recognition according to claim 1, characterized in that: The method of using the Hungarian algorithm to perform correlation matching on the suspicious areas of the security inspection image to obtain a matching result and issuing an early warning based on the matching result includes: The Hungarian algorithm is used to calculate the Euclidean distance between the center position of the suspicious area of the current frame and the center position of the suspicious area of the previous frame; Extract the shape features of the suspicious area and calculate the Euclidean distance between the shape features of the suspicious area in the current frame and the previous frame. The smaller the distance, the higher the shape similarity. The feature vector of the suspicious area extracted by the hybrid neural network is used to calculate the cosine similarity between the feature vector of the suspicious area of the current frame and the previous frame. The higher the similarity, the higher the feature similarity.

7. The intelligent security inspection method based on image recognition according to claim 1, characterized in that: The method further comprises: performing correlation matching on the suspicious areas of the security inspection image using the Hungarian algorithm to obtain a matching result, and issuing an early warning based on the matching result; Perform weighted fusion of position similarity, shape similarity and feature similarity to obtain a comprehensive similarity score; Construct a similarity matrix, where the rows represent the suspicious areas of the current frame, the columns represent the suspicious areas of the previous frame, and the matrix elements are the comprehensive similarity scores of the corresponding suspicious areas. Solve the similarity matrix to obtain the matching results.

8. An intelligent security inspection system based on image recognition, characterized in that: The image recognition-based intelligent security inspection system includes the following modules: A security inspection data acquisition module is used to acquire X-ray transmission images, visible light surface images, millimeter wave three-dimensional imaging data, and infrared thermal data in the security inspection system to obtain multimodal security inspection data, and preprocess the multimodal security inspection data to obtain initial multimodal security inspection data; a security inspection data processing module, configured to dynamically adjust filtering parameters according to the image signal-to-noise ratio in the initial multimodal security inspection data using a wavelet transform-bilateral filtering joint noise reduction algorithm, and perform noise reduction processing on the data to obtain noise-reduced multimodal security inspection data; A hybrid model building module is used to extract local features based on ResNet-50, capture global contextual relationships in parallel with the Swin Transformer module to establish a hybrid neural network, and dynamically weightedly fuse multimodal features through a cross-attention mechanism to obtain a ResNet-Transformer hybrid neural network; A security inspection image recognition module, configured to input the denoised multimodal security inspection data into the ResNet-Transformer hybrid neural network for recognition, thereby obtaining suspicious areas of the security inspection image; The result matching warning module is used to use the Hungarian algorithm to perform correlation matching on the suspicious area of the security inspection image to obtain a matching result, and issue an early warning based on the matching result.

9. The intelligent security inspection system based on image recognition as claimed in claim 8, characterized in that: The result matching warning module includes the following submodules: A calculation submodule, for calculating the Euclidean distance between the center positions of the suspicious area of the current frame and the suspicious area of the previous frame using the Hungarian algorithm; The extraction submodule is used to extract the shape features of the suspicious area and calculate the Euclidean distance between the shape features of the suspicious area in the current frame and the previous frame. The smaller the distance, the higher the shape similarity. The judgment submodule is used to calculate the cosine similarity between the suspicious area feature vectors of the current frame and the previous frame using the suspicious area feature vector extracted by the hybrid neural network. The higher the similarity, the higher the feature similarity.

10. The intelligent security inspection system based on image recognition according to claim 8, characterized in that: The result matching warning module includes the following submodules: The fusion submodule is used to perform weighted fusion of position similarity, shape similarity and feature similarity to obtain a comprehensive similarity score; The construction submodule is used to construct a similarity matrix, where the rows represent the suspicious areas of the current frame, the columns represent the suspicious areas of the previous frame, and the matrix elements are the comprehensive similarity scores of the corresponding suspicious areas. Solving the similarity matrix obtains the matching results.

Citation Information

Patent Citations

  • Circuit breaker fault operation and maintenance monitoring method and system based on image recognition analysis

    CN115131325A

  • Medical image segmentation method and system based on double-branch embedded attention mechanism

    CN116309650A

  • X-ray security inspection machine picture object detection method

    CN119600389A

  • Methods, systems, and apparatuses for inspecting goods

    EP3182335A2

Cited By

  • Security check method and system based on voice and smell fusion data recognition

    CN121598267A

  • Security check opening and checking strategy intelligent generation method and system based on multi-modal perception and deep reinforcement learning

    CN122510065A