Customs contraband detection method and system based on multi-modal adaptive fusion
Through the adaptive dynamic channel fusion module and AnyRes strategy to optimize feature expression, combined with bilateral soft matching and two-stage training, the problems of insufficient multi-level feature fusion, high computational complexity and poor correlation between modes in customs contraband detection are solved, and efficient and real-time contraband detection is achieved.
Patent Information
- Application Number
- CN202510492962.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing technology has insufficient multi-level feature fusion in customs contraband detection, which ignores low-level detailed information, making it difficult to identify dense or hidden contrabands; the calculation complexity is high, and the real-time requirements cannot be met; the correlation between modals is insufficient, resulting in low feature fusion efficiency and poor semantic consistency.
Adaptive dynamic channel fusion module is used to integrate shallow details and deep semantics, and optimize feature expression through adaptive weights; introduce AnyRes strategy to dynamically adjust resolution, and use bilateral soft matching strategy to reduce feature sequence length; two-stage training methods optimize feature alignment between modals.
It improves the feature resolution capability in complex scenarios, optimizes the computational complexity and real-timeness, improves detection accuracy and efficiency, and ensures the balance and accuracy of customs multimodal data.
Smart Images

Figure CN120411705A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of customs contraband detection, and particularly to a customs contraband detection method and system based on multi-modal adaptive fusion. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] With the increasing demands of customs supervision, the application of multi-modal data (such as text, images, videos, etc.) in cargo inspection, contraband identification, and customs clearance efficiency improvement has gradually received attention. In the prior art, the processing and feature extraction methods of multi-modal information have been applied to the customs supervision field to a certain extent, but there are still many deficiencies, which limit the performance and practical application effects of the system in complex scenarios.
[0004] In recent years, multi-modal large models have achieved cross-modal feature fusion through large-scale pre-training and fine-tuning based on pre-trained models, which can comprehensively understand cargo features and improve supervision capabilities. However, the prior art has limitations in multi-level feature fusion. Traditional visual encoders usually only extract deep semantic features and ignore low-level detail information, which is particularly disadvantageous in customs contraband detection. For densely arranged goods or concealed contraband, the lack of low-level features makes it difficult for the model to capture subtle differences, resulting in a decline in recognition ability; and when fusing multi-level features, the existing methods often use simple channel splicing, which is complex and inefficient, and cannot effectively balance performance and computational cost.
[0005] In terms of visual feature extraction, the prior art generally relies on high-resolution image input or a large number of visual markers to improve accuracy. However, this method significantly increases the computational complexity. Especially when processing multi-image or video data, the length of the feature sequence is too long, resulting in excessive consumption of computing resources and a significant decrease in the inference speed. For example, the feature extraction method based on the traditional Transformer model, the computational complexity of its attention mechanism is proportional to the square of the sequence length, making it difficult to meet the requirements of high-throughput customs inspection in real time. At the same time, in order to pursue performance improvement by increasing the resolution or the number of markers, complex preprocessing steps are often required, further increasing the operational complexity and system implementation cost.
[0006] When dealing with the sparsity and imbalance problems of customs multi-modal data, the existing technologies lack effective optimization strategies. In customs multi-modal data, there are significant differences in the sample size and quality between images and texts. For example, the amount of image data may far exceed the text description, and at the same time, the distribution of prohibited item categories is unbalanced (such as more data on common prohibited items and insufficient data on rare prohibited items), resulting in the amplification or weakening of the contributions of certain modalities or categories during the model training process, affecting the comprehensiveness and accuracy of the analysis results; if the model overly relies on visual information or data on common categories, it may ignore the key clues in the text and the characteristics of rare prohibited items, thereby reducing the overall detection performance. Summary of the Invention
[0007] To solve the above problems, the present invention proposes a customs prohibited item detection method and system based on multi-modal adaptive fusion, designs an adaptive dynamic channel fusion module, integrates shallow details and deep semantics from multiple hidden states, optimizes the feature representation through adaptive weights, realizes the effective fusion of low-level details and deep semantics, and improves the feature discrimination ability in complex scenarios.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] In the first aspect, the present invention provides a customs prohibited item detection method based on multi-modal adaptive fusion, including:
[0010] Obtain customs clearance image-text pairs, and extract the original features including multi-scale features of the customs clearance images.
[0011] Using the first N layers of features in the multi-scale features as shallow features and the remaining layer features as deep features, respectively perform weighted fusion on the shallow features and deep features to generate shallow fusion features and deep fusion features, perform multi-scale pooling operations on the deep fusion features to obtain pooled features, splice the shallow fusion features, deep fusion features and pooled features to generate enhanced features, and integrate the enhanced features and the original features to obtain visual features.
[0012] Perform modal alignment on the visual features and text features, based on the obtained aligned multi-modal features, perform two-stage training on the detection model, and for the customs clearance images to be measured, use the trained detection model to obtain the prohibited item detection results.
[0013] As an alternative implementation, after obtaining the customs clearance images, divide the customs clearance images into a×b cropped blocks, then the total number of visual Tokens is L = (a×b + 1)×T; when the total number of visual Tokens L exceeds the set Token number threshold τ, adjust the Token number and select the optimal spatial configuration.
[0014]
[0015] Among them, T new is the number of Tokens in each adjusted cropped block, and T is the original number of Tokens in each cropped block; a and b represent the number of columns and rows of the division; a×b + 1 is the total number of cropped blocks plus the original image, calculating the total number of cropping units.
[0016] As an alternative implementation, calculate the weights of the features of each layer for weighted fusion;
[0017]
[0018] After normalization, generate shallow-layer adaptive weights and deep-layer adaptive weights respectively :
[0019]
[0020] Among them, W s i is the weight of the i-th shallow-layer feature; is the weight of the j-th deep-layer feature; S i [:, l, d] is the value of the i-th shallow-layer feature at the l-th position and the d-th dimension; D j [:, l, d] is the value of the j-th deep-layer feature at the l-th position and the d-th dimension; L is the sequence length, and D is the hidden dimension; N s is the number of shallow layers, and N d is the number of deep layers.
[0021] As an alternative implementation, the process of performing multi-scale pooling operations on the deep fusion features includes:
[0022] Perform one-dimensional average pooling on the deep fusion feature F d , and permute the dimensions to change the shape to [B, D, L];
[0023] Apply the pooling operation and output with a shape of [B, D, L′], where
[0024] After adjusting the sequence length of the deep pooled feature to L through linear interpolation operation , adjust the dimension order to [B, L, D]; B is the batch size, L ′ and L are the sequence lengths, and D is the hidden dimension.
[0025] As an alternative implementation, for visual feature representation, a bilateral soft matching strategy is introduced. The cosine similarity is used to measure the similarity of Tokens, and the Token pair with the highest similarity is matched and fused to generate a new Token. Specifically, it includes: evenly dividing all Tokens into two sets A and B. For each Token in set A, using the cosine similarity, find the most similar Token in set B and connect them with an edge. After the matching is completed, the original Tokens are fused through weighted averaging to generate a new Token.
[0026] As an alternative implementation, pair the image-text pairs with the corresponding instruction templates to construct a visual instruction tuning dataset for model fine-tuning. And map the visual features to the semantic space consistent with the text features through a two-layer mapping layer to achieve modality alignment. Then the two-stage training process includes: The first stage is the cross-modal mapping optimization stage, freezing the parameters of the visual encoder and the language encoder, and training the two-layer mapping layer to optimize the mapping relationship between the visual features and the text features, achieving the initial alignment between modalities. The second stage is the visual instruction tuning stage, unfreezing the parameters of the visual encoder and the language encoder, and jointly training the overall model with the two-layer mapping layer.
[0027] As an alternative implementation, for a sequence of length L, the probability p(X a ) of the target answer X a |X v ,X q ) is defined as:
[0028]
[0029] where X v is the visual feature after modality alignment; X q is the text feature of the instruction template; X q,<i are all Tokens of the instruction template, X a,<i are the first i - 1 Tokens in the answer, respectively representing the instructions and answer context that the language encoder has processed when predicting the current answer Token x i ;
[0030] represents that the language encoder predicts each Token x i of the answer one by one in an autoregressive manner. Each time a prediction is made, referring to the visual feature X v , the instruction Token X q,<i and the generated answer Token X a,<i , calculate the probability p(x i |X v ,X q,<i ,X a,<i), finally multiply the probabilities of all Tokens to obtain the probability of the entire answer; the finally generated result includes the name of the contraband and the position of the bounding box.
[0031] In a second aspect, the present invention provides a customs contraband detection system based on multi-modal adaptive fusion, including:
[0032] A feature extraction module, configured to obtain a customs clearance image-text pair and extract the original features including multi-scale features of the customs clearance image;
[0033] An adaptive fusion module, configured to use the first N layers of features in the multi-scale features as shallow features, and the remaining layer features as deep features, respectively perform weighted fusion on the shallow features and the deep features to generate shallow fusion features and deep fusion features, perform multi-scale pooling operations on the deep fusion features to obtain pooled features, splice the shallow fusion features, the depth fusion features and the pooled features to generate enhanced features, and integrate the enhanced features and the original features to obtain visual features;
[0034] A training module, configured to perform modal alignment on the visual features and the text features, and based on the obtained aligned multi-modal features, perform two-stage training on the detection model, and for the to-be-tested customs clearance image, use the trained detection model to obtain the contraband detection result.
[0035] In a third aspect, the present invention provides an electronic device, including a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.
[0036] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by the processor, the method described in the first aspect is completed.
[0037] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by the processor, the method described in the first aspect is implemented.
[0038] Compared with the prior art, the beneficial effects of the present invention are:
[0039] The prior art fails to fully utilize the potential of pre-trained visual encoders in multi-level feature fusion, neglecting the extraction of low-level detail information, which makes it difficult to accurately identify the tiny features of densely distributed or concealed contraband in X-ray images, thus limiting the recognition ability in complex scenarios. Therefore, aiming at the problems of neglecting low-level details and being difficult to identify dense or concealed contraband in multi-level feature fusion of the prior art, the present invention designs an adaptive dynamic channel fusion module to integrate shallow details and deep semantics from multi-layer hidden states, optimize feature representation through adaptive weights, achieve effective fusion of low-level details and deep semantics, and improve the feature discrimination ability in complex scenarios. At the same time, it optimizes the balance and utilization rate of customs multi-modal data, solves the problems of sparsity and imbalance, improves the collaborative contribution of different modalities and different data types, and ensures the comprehensiveness and accuracy of the analysis results.
[0040] Compared with the prior art that relies on fixed high-resolution inputs or a large number of visual markers, resulting in high computational complexity and an inference speed insufficient to meet real-time requirements. The present invention introduces the AnyRes strategy to dynamically adjust the configuration according to the resolution of the input customs contraband visual data, enabling the model to adaptively receive the original resolution of the picture itself, supporting adaptive resolution. Thus, when inputting image and video data, it takes into account both the details of local visual features and the integrity of global visual information, comprehensively improves the model performance with a relatively small increase in the model scale, optimizes the processing ability for X-ray images of any resolution, ensures the efficiency of feature extraction, and solves the problem of insufficient adaptability of feature extraction caused by the prior art's reliance on fixed high-resolution inputs.
[0041] The present invention introduces a bilateral soft matching fusion strategy to effectively reduce the sequence length by matching and fusing similar Tokens, reduce resource consumption during the inference process, and at the same time retain key information. Thus, it optimizes the fusion structure, reduces redundant calculations, ensures that the system can achieve real-time response in high-throughput detection scenarios, and improves the balance between system efficiency and computational cost. It significantly reduces the computational complexity and improves real-time performance with a relatively low impact on model performance, and at the same time solves the problem of excessive computational burden caused by too long feature sequences, and optimizes the inference speed to meet the real-time requirements of customs high-throughput detection.
[0042] Aiming at the problems of low feature fusion efficiency and poor semantic consistency caused by insufficient inter-modal correlation in the prior art for customs contraband detection tasks. The inter-modal feature alignment is optimized through a two-stage training method. In the first stage, the visual encoder and the language encoder are frozen, and only the double-layer mapping layer is trained to accurately map visual features to the embedding space of the language encoder to improve the semantic consistency between visual features and text features. In the second stage, all parameters are unfrozen for overall fine-tuning, and through visual instruction tuning, the detection accuracy of the model in customs multi-modal tasks is improved.
[0043] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.
[0045] Figure 1 Flowchart of the customs contraband detection method based on multimodal adaptive fusion provided for Embodiment 1 of the present invention;
[0046] Figure 2 Schematic diagram of adaptive dynamic channel fusion provided for Embodiment 1 of the present invention;
[0047] Figure 3 Flowchart of the first-stage training provided for Embodiment 1 of the present invention;
[0048] Figure 4 Flowchart of the second-stage training provided for Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The present invention will be further described below in conjunction with the drawings and embodiments.
[0050] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0051] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0052] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0053] Embodiment 1
[0054] This embodiment provides a customs contraband detection method based on multimodal adaptive fusion, including:
[0055] Obtain customs clearance image-text pairs, and extract the original features including multi-scale features of the customs clearance images.
[0056] Using the first N layers of features in the multi-scale features as shallow features and the remaining layer features as deep features, respectively perform weighted fusion on the shallow features and deep features to generate shallow fusion features and deep fusion features, perform multi-scale pooling operations on the deep fusion features to obtain pooled features, splice the shallow fusion features, deep fusion features and pooled features to generate enhanced features, and integrate the enhanced features and the original features to obtain visual features.
[0057] Perform modal alignment on the visual features and text features. Based on the obtained aligned multimodal features, perform two-stage training on the detection model. For the customs clearance images to be measured, use the trained detection model to obtain the contraband detection results.
[0058] The following Figure 1 elaborates on the method of this embodiment in detail.
[0059] Step 1: Obtain customs clearance image-text pairs.
[0060] Specifically:
[0061] First, collect the original multimodal data related to contraband detection from the customs clearance supervision system, including customs clearance X-ray images and corresponding annotated texts.
[0062] Among them, the X-ray images cover various types of goods (such as non-metallic lighters, controlled knives, liquid containers, etc.), and the annotated texts include image IDs (such as "009004.jpg"), contraband labels (such as "straight knife") and contraband position coordinates [x1, y1, x2, y2] (such as "[711 523 800 598]"), where (x1, y1) is the upper left position of the contraband, and (x2, y2) is the lower right position of the contraband, providing a diverse input basis for the subsequent extraction of multimodal features.
[0063] Secondly, divide the customs clearance X-ray images into a global annotation group (annotating the object name and bounding box) and a local cropping group (cropping the contraband area), and adjust the category sample distribution through oversampling and undersampling to enhance the learning of spatial information and object features.
[0064] Specifically:
[0065] The local cropping group crops the contraband from the original image according to the bounding box, highlighting the object details, facilitating the model to better capture the information features of the object in visual feature extraction, and providing high-quality input for subsequent semantic space mapping.
[0066] The global annotation group generates a bounding box at the corresponding position in the X-ray image and annotates the corresponding object name according to the position coordinate information of the contraband in the annotation text, so that the model can better learn the global information and spatial information in the visual instruction tuning training.
[0067] By analyzing the sample distribution of contraband categories in the X-ray images, oversampling is performed on the categories with the number of samples less than the target threshold, and new samples are generated by image enhancement methods (horizontal flipping, vertical flipping, and rotation); undersampling is performed on the categories with the number of samples exceeding the target threshold, and the samples with higher resolution are retained until the target number is reached to balance the sample proportions of various categories and prevent the model from being biased towards the majority class.
[0068] Finally, pair the X-ray images of the global annotation group and the local cropping group with the corresponding descriptive texts extracted from the annotation text to form image-text pairs, ensuring that each image corresponds one-to-one with its related text content.
[0069] Step 2: According to the requirements of the customs contraband detection task, design an instruction template. The instruction template includes user instructions (such as "Identify and classify the contraband in the image", "Provide the bounding box position") and expected model responses (such as "Detected a knife, position [x, y, w, h]") to support task-oriented training;
[0070] Combine the paired image-text pairs with the corresponding instruction templates to construct a visual instruction tuning dataset for model fine-tuning; this dataset contains image IDs, image file paths, and conversation content. The conversation content includes multiple rounds of interactions between user instructions and model responses; for example, the user requests "Identify the contraband and provide the name and position", and the model returns "Found a straight knife, bounding box position [609, 506, 790, 549]", etc.
[0071] Step 3: Preprocess the visual instruction tuning dataset, including data balancing, data normalization, and data cleaning; among them, data normalization is performed by converting the image into a tensor and normalizing the pixel values of the image from the range [0, 255] to control the pixel values within [0, 1] to unify the numerical scale; data cleaning eliminates blurred images, incomplete pairings, or invalid texts to ensure the integrity and applicability of the data.
[0072] Step 4: Introduce the Adaptive Resolution Pooling (AnyRes) strategy, dynamically adjust the configuration according to the resolution, crop the customs clearance images, and process single-image, multi-image, and video data; then perform feature extraction through the Visual Encoder (SigLIP) to generate the original visual features.
[0073] Specifically, it includes:
[0074] Introduce the AnyRes strategy, dynamically adjust the configuration according to the resolution of the input visual data of customs contraband, and process single-image, multi-image, and video data; analyze the maximum number of tokens that the AnyRes strategy will generate in different input single-image, multi-image, and video scenarios, so that the model can dynamically adapt to the resolution and type of the input image.
[0075] The AnyRes strategy divides the customs clearance image into a×b cropping blocks according to the configuration (a, b), and the total number of visual tokens L=(a×b + 1)×T; when the total number of visual tokens L exceeds the threshold τ, bilinear interpolation is used to adjust the number of tokens and select the optimal configuration to adapt to different resolutions and aspect ratios:
[0076]
[0077] Among them, T new is the number of tokens in each cropping block after adjustment, T is the original number of tokens in each cropping block; a and b are the spatial configurations of the AnyRes strategy, representing the number of columns and rows divided respectively; a×b + 1 is the total number of cropping blocks plus the base image, calculating the total number of cropping units; τ is the preset token number threshold for controlling the total number of visual tokens.
[0078] The original number of tokens in the cropping block is the original number of tokens of the original image after processing, specifically calculated according to the resolution of the input image. Taking a single image as an example, if the input image has a resolution of 384×384, the image chunking strategy of the visual encoder is to divide the image into patches of size 14×14 each, then the original number of tokens=(384 / 14)×(384 / 14)=729; according to the AnyRes strategy, the number of tokens generated by this image is (1 + 9)×729 = 7290 <= 7290, so it does not exceed the token number threshold set for single images, so T new is originally 7290. If the resolution of the input image is greater than this threshold, it is adjusted to be less than the threshold 7290 to control the number of tokens.
[0079] For a single image, a large spatial resolution is set. Long sequence features are generated by cropping and adjusting the global thumbnail. Up to 729 visual tokens are allocated, and a single image can have up to 729×(1 + 9) = 7290 tokens. At the same time, local and global visual features are obtained through AnyRes pooling.
[0080] For multiple images, each image is scaled to 384x384 pixels, padded with 0 to maintain the aspect ratio, then input into the visual encoder and the padding is removed. The training data contains at most 12 images, with a total of at most 12×729 = 8748 tokens. High-resolution multi-block cropping is avoided to save computing resources.
[0081] For videos, each frame is adjusted to the base resolution and then feature maps are generated through the visual encoder. 2x2 bilinear interpolation is used to reduce the number of tokens to 196. At most 32 frames are sampled, with a total of at most 196×32 = 6272 tokens, in order to balance the number of frames and the computational cost.
[0082] In this embodiment, the subsequent bilateral soft matching strategy and probability calculation are based on tokens because tokens are the basic units of the multimodal large model. In this model, whether it is text or image (video), they will be uniformly encoded into tokens as the basic units of input and intermediate processing results.
[0083] 1. The control of the number of tokens by the AnyRes strategy is, on the one hand, to dynamically adapt to the resolution and type of the input image, so that the input image will not be uniformly adjusted to a fixed resolution like traditional models (for example, the input picture is adjusted to 224×224). On the other hand, it is to save computing resources because the transformers architecture has a limit on the input length.
[0084] 2. The subsequent bilateral soft matching strategy is to reduce the computational amount increased by adci while retaining most of the key information. Therefore, some merging operations need to be performed on the tokens.
[0085] [[ID=ID=18]]3. The probability calculation is because the model is based on the transformers architecture and is an autoregressive generation model. Therefore, the next token with the highest probability needs to be predicted based on the tokens.
[0086] Subsequently, the visual encoder (SigLIP) divides the input customs clearance image into fixed-size image patches (Patches), converts each image patch into a vector representation through a linear embedding layer, and adds position embeddings to form an image patch sequence. This image patch sequence is used as input and fed into a visual backbone network composed of multiple layers of Transformers. Through the multi-head self-attention mechanism, it is processed layer by layer, and the visual representation is continuously updated at each layer to extract increasingly rich image semantic features, namely the original image feature sequence F img .
[0087] Among them, the shallow features usually come from the first thirteen layers of the network, have a high spatial resolution and strong local structure expression ability, and are suitable for capturing low-level visual information such as edges and textures; the deep features come from the last thirteen layers of the network, have stronger semantic abstraction ability, and can model the context relationship and global dependency structure between objects.
[0088] Generally speaking, the visual encoder often uses deep features for the final output. To further explore semantic information at different levels, shallow features and deep features are extracted from all hidden layer outputs of the visual encoder and input into the ADCI (Adaptive Dynamic Channel Integration) module. The generated enhanced feature sequence is integrated with the original image feature sequence in the channel dimension to achieve multi-scale and multi-level feature fusion and enhancement, generating a richer and more structured high-dimensional visual feature sequence.
[0089] Step Five: As Figure 2 shown, an innovative ADCI (Adaptive Dynamic Channel Integration) module is proposed to extract shallow features and deep features from the hidden layers of the visual encoder.
[0090] Specifically, it includes:
[0091] To make more full use of the ability of the visual encoder in visual information extraction, the ADCI module is proposed to optimize the fusion effect of visual information by integrating multi-level visual features of the visual encoder.
[0092] Through a dense connection strategy and an adaptive feature fusion technique, from the multi-layer hidden layers H = {H0, H1, …, H N-1} (each layer has a shape of [B, L ′ , D]) of the visual encoder (SigLIP) and the original image feature sequence F img (with a shape of [B, L, D]) as input, where B is the batch size, L ′Let \(L\) be the sequence length, \(D\) be the hidden dimension, and \(N = 26\). In the ADCI module, shallow features and deep features are extracted from the hidden layer \(H\).
[0093] The steps for extracting shallow features are as follows: Shallow feature \(S\) i (\(i = 0, \ldots, N\) s - 1), where \(N\) s = 13 is the number of shallow layers. For each hidden layer \(H\) i keep all tokens (\(S\) i = \(H\) i [:,:]), without performing any slicing operations. Then stack \(S\) i along to get a shape of \([N\) s , B, L, D]\).
[0094] The steps for extracting deep features are as follows: Deep feature \(D\) j (\(j = N\) s , \ldots, N - 1), where \(N\) d = \(N - N\) s = 13 is the number of deep layers. For each hidden layer \(H\) j keep all tokens (\(D\) j = \(H\) j [:,:]), without performing any slicing operations. Then stack \(D\) j along to get a shape of \([N\) d , B, L, D]\).
[0095] Step 6: Calculate adaptive weights based on the L2 norm of the features in ADCI and perform weighted fusion to generate shallow fusion features and deep fusion features.
[0096] Specifically, it includes:
[0097] After extracting the shallow feature \(S\) and the deep feature \(D\) from the hidden layer \(H\), calculate the L2 norm for each layer of \(S\) and \(D\), which serves as the weight for each layer:
[0098]
[0099]
[0100] where, is the weight of the \(i\)-th shallow feature; is the weight of the \(j\)-th deep feature; \(S\) i [:, l, d] is the value of the \(i\)-th shallow feature at the \(l\)-th position and the \(d\)-th dimension; \(D\) j [:, l, d] is the value of the \(j\)-th deep feature at the \(l\)-th position and the \(d\)-th dimension; \(L\) is the sequence length, and \(D\) is the hidden dimension.
[0101] Then, shallow adaptive weights are generated through Softmax normalization and deep adaptive weights :
[0102]
[0103]
[0104] Among them, the weight shapes are [N s , B, 1, 1] and [N d , B, 1, 1], dynamically reflecting the importance of features in each layer.
[0105] Finally, these weights are used for weighted fusion to generate shallow fusion features and deep fusion features The output shapes are both [B, L ′ , D].
[0106] Step 7: In ADCI, a multi-scale pooling operation is additionally applied to the deep fusion features to generate deep pooled features.
[0107] Specifically, it includes:
[0108] First, one-dimensional average pooling is performed on F d , and the dimension is permuted to change the shape to [B, D, L].
[0109] Subsequently, a pooling operation is applied, and the output shape is [B, D, L ′ , where k is the size of the pooling window, indicating that in each pooling operation, the window covers k consecutive elements and takes the average of these elements; s is the stride of the pooling, indicating that the pooling window slides forward s elements each time.
[0110] Then, through a linear interpolation operation After adjusting the sequence length of the deep pooled features to L, finally, the dimension order is adjusted to [B, L, D]; size is specified to make the target sequence length the original L, that is, the length of the interpolated features should be L; mode is specified to be linear interpolation, and linear interpolation fills in the missing values by calculating the linear relationship between adjacent points, so that the feature sequence length is restored to L.
[0111] Step 8: In ADCI, the shallow fusion features, deep fusion features, and deep pooled features are concatenated in the channel dimension to generate an enhanced feature sequence with a shape of [B, L, 3D];
[0112] The formula for concatenating features is as follows:
[0113]
[0114] Among them, F s is the shallow fusion feature; F d is the deep fusion feature; is the deep pooling feature; dim=-1 means concatenation along the last dimension.
[0115] Step Nine: Integrate the enhanced feature sequence and the original image feature sequence along the channel dimension in ADCI, and output the final high-dimensional visual feature sequence with the shape of [B, L, 4D];
[0116] F output =Concat(F img , F ADCI , dim=-1)(7);
[0117] Among them, F img is the original image feature sequence; F ADCI is the enhanced feature sequence after passing through the ADCI module; dim=-1 means concatenation along the last dimension.
[0118] Step Ten: To ensure the model inference speed, for the high-dimensional visual feature sequence, introduce a bilateral soft matching strategy to reduce the length of the feature sequence. Use cosine similarity to measure the similarity of Tokens, match and fuse the Token pairs with the highest similarity, generate new Tokens, so as to reduce the length of the feature sequence while retaining important information.
[0119] Specifically, it includes:
[0120] Divide all Tokens input into the Token fusion module into two sets A and B; then for each Token in set A, find the most similar Token in set B and draw an edge to connect them. This matching process is based on cosine similarity measurement, and each matching pair has a similarity weight, reflecting the association strength between Tokens.
[0121] After the matching is completed, the features of multiple original Tokens are fused by weighted average to generate new Tokens. While reducing the number of Tokens, each Token also contains richer information. The information of the merged Tokens can still be propagated in the visual encoder.
[0122] When performing Token Merging between Attention and the multi-layer perceptron, the information of the merged Tokens can be propagated to other Tokens through the attention mechanism. The Token fusion algorithm can iteratively execute the above steps in each layer of the visual encoder, gradually reducing the number of Tokens. With each reduction in the number of layers, the computational complexity can be further reduced.
[0123] Step Eleven: Project the visual features into the text semantic space through a two-layer mapping layer (MLP) to achieve modality alignment, and the language encoder (LLM) generates customs task-related prediction results based on the aligned multi-modal features.
[0124] Specifically, it includes:
[0125] The two-layer mapping layer (MLP) maps the visual feature sequence to the semantic space consistent with the text features to achieve deep alignment between modalities; among them, the text features are encoded by the embedding layer of the language encoder (LLM) after being generated into text Tokens by the Tokenizer.
[0126] The language encoder (LLM) receives the visual features and text features aligned by the two-layer mapping layer (MLP), and generates structured results related to customs contraband detection based on the multi-modal joint input.
[0127] For a sequence of length L, the probability p(X a |X a , X v , X q ) is defined as follows:
[0128]
[0129] Among them, X v is the modality-aligned visual feature, generated by the visual encoder and the two-layer mapping layer; X q is the text feature of the instruction template, formed by encoding the Token sequence X q,<i of the instruction template through the embedding layer and the Transformers layer of the language encoder (LLM); X q,<i is the Token sequence of the instruction template, X a,<i is the first i - 1 Tokens in the answer, respectively representing the instructions and answer context that the language encoder (LLM) has processed when predicting the current answer Token x i .
[0130] Equation (8) means that the language encoder (LLM) predicts each Token x i of the answer one by one in an autoregressive manner. Each time a prediction is made, it refers to the visual feature X v, the Token sequence X of the instruction template q,<i and the generated answer Token X a,<i , calculate the probability p(x i |X v ,X q,<i ,X a,<i ) of the current Token, and finally multiply the probabilities of all Tokens to obtain the probability of the entire answer.
[0131] The final generated result includes the name of the prohibited item and the bounding box position (such as [x1, y1, x2, y2] coordinates), in the format of a text report.
[0132] Step Twelve: The detection model is trained using a two-stage training method.
[0133] Specifically, it includes:
[0134] As Figure 3 shown, the first stage is the cross-modal mapping optimization stage. Using an open-source pre-trained dataset, in this stage, the parameters of the visual encoder and the language encoder are frozen, and only the two-layer mapping layer is trained to optimize the mapping relationship between visual features and text features, achieving a preliminary alignment between modalities.
[0135] The cross-entropy loss function is adopted, the AdamW optimizer is used to update the parameters, the learning rate is set to η1 = 1e-3, the gradient accumulation step is 2, the cosine learning rate schedule is adopted (the warm-up ratio is 0.03), trained for 3 rounds, save a checkpoint every 1000 steps, and the maximum input sequence length of the model is 8192 to enhance the modality alignment ability.
[0136] As Figure 4 shown, the second stage is the visual instruction tuning stage. In this stage, the parameters of the visual encoder and the language encoder are unfrozen, and the overall model is trained together with the two-layer mapping layer to further improve the performance of the model on specific visual tasks.
[0137] Set the learning rate η2 of the language encoder to 1×10 -5 , and the learning rate η vision of the visual encoder to 2×10 -6 (Note that the learning rate of the language encoder is set to five times that of the visual encoder), the batch size is 1, the gradient accumulation step is 2, the cosine learning rate schedule is adopted (the warm-up ratio is 0.03), trained for 1 round, save a checkpoint every 1000 steps, and the maximum input sequence length of the model is 32768. Combine the gradient checkpoint technology to optimize memory usage and further improve the performance of the model on specific visual tasks such as customs prohibited item detection.
[0138] The customs contraband detection method based on multi-modal adaptive fusion provided in this embodiment designs an ADCI module to address the problems in the prior art of neglecting low-level details in multi-level feature fusion and difficulty in identifying dense or concealed contraband. It integrates shallow details and deep semantics from multiple hidden states, optimizes feature representation through adaptive weights, and improves the feature discrimination ability in complex scenarios.
[0139] Compared with the prior art that relies on fixed high-resolution inputs or a large number of visual markers, resulting in high computational complexity and an inference speed insufficient to meet real-time requirements. By introducing the AnyRes strategy, the configuration is dynamically adjusted according to the resolution of the input customs contraband visual data, enabling the model to adaptively receive the original resolution of the picture itself. Thus, when inputting image and video data, it takes into account both the details of local visual features and the integrity of global visual information, comprehensively improving the model performance with a relatively small increase in model scale; at the same time, the Token fusion strategy is introduced, and redundant Tokens are fused using bilateral soft matching clustering, significantly reducing the inference calculation amount of the model with a relatively low impact on model performance and enhancing the system processing efficiency.
[0140] Aiming at the problem in the prior art of low feature fusion efficiency and poor semantic consistency in the customs contraband detection task due to insufficient inter-modal correlation. The inter-modal feature alignment is optimized through a two-stage training method. In the first stage, the visual encoder and the language encoder are frozen, and only the double-layer mapping layer is trained to accurately map visual features to the embedding space of the language encoder to enhance the semantic consistency between visual features and text features. In the second stage, all parameters are unfrozen for overall fine-tuning, and through visual instruction tuning, the detection accuracy of the model in the customs multi-modal task is improved.
[0141] It should be noted that the acquisition of all data is based on compliance with laws, regulations, and user consent, and the data is legally applied.
[0142] Embodiment 2
[0143] This embodiment provides a customs contraband detection system based on multi-modal adaptive fusion, including:
[0144] A feature extraction module configured to obtain customs clearance image-text pairs and extract the original features including multi-scale features of the customs clearance images.
[0145] An adaptive fusion module configured to use the first N layers of features in the multi-scale features as shallow features and the remaining layer features as deep features, respectively perform weighted fusion on the shallow features and the deep features to generate shallow fusion features and deep fusion features, obtain pooled features after performing multi-scale pooling operations on the deep fusion features, splice the shallow fusion features, the deep fusion features, and the pooled features to generate enhanced features, and integrate the enhanced features and the original features to obtain visual features.
[0146] A training module, configured to perform modality alignment on visual features and text features, perform two-stage training on a detection model based on the obtained aligned multi-modal features, and use the trained detection model for a customs clearance image to be measured to obtain a contraband detection result.
[0147] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0148] In more embodiments, there is also provided:
[0149] An electronic device, including a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.
[0150] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0151] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random memory. For example, the memory may also store information about the device type.
[0152] A computer-readable storage medium, for storing computer instructions, which, when executed by a processor, complete the method described in Embodiment 1.
[0153] The method in Embodiment 1 can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0154] A computer program product, including a computer program, which, when executed by a processor, implements the method described in Embodiment 1.
[0155] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the processes / methods as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed. The machine-executable instructions for program modules can be executed within local or distributed devices. In a distributed device, program modules can be located in local and remote storage media.
[0156] The computer program code for implementing the method of the present invention can be written in one or more programming languages. The computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.
[0157] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier such that a device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, sound, or other forms of propagated signals, such as carrier waves, infrared signals, etc.
[0158] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0159] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A customs contraband detection method based on multi-modal adaptive fusion, characterized in that, Including: Obtain customs clearance image-text pairs, and extract the original features including multi-scale features of the customs clearance images; Use the first N layers of features in the multi-scale features as shallow features, and the remaining layer features as deep features. After weighted fusion of the shallow features and deep features respectively, generate shallow fusion features and deep fusion features. After performing multi-scale pooling operations on the deep fusion features, obtain pooling features. Concatenate the shallow fusion features, deep fusion features, and pooling features to generate enhanced features, and integrate the enhanced features and the original features to obtain visual features; Perform modality alignment on the visual features and the text features. Based on the obtained aligned multi-modal features, perform two-stage training on the detection model. For the customs clearance images to be measured, use the trained detection model to obtain the contraband detection results.
2. The method for detecting customs contraband based on multimodal adaptive fusion according to claim 1, wherein, After obtaining the customs clearance images, divide the customs clearance images into a×b cropped blocks, then the total number of visual tokens is L=(a×b + 1)×T; when the total number of visual tokens L exceeds the set token number threshold τ, adjust the number of tokens and select the optimal spatial configuration; Among them, T new is the number of Tokens in each cropped block after adjustment, and T is the original number of Tokens in each cropped block; a and b represent the number of columns and rows divided; a×b + 1 is the total number of cropped blocks plus the original image, calculating the total number of cropping units.
3. A customs contraband detection method based on multi-modal adaptive fusion as described in claim 1, characterized in that, Calculate the weights of each layer of features for weighted fusion; After normalization, shallow adaptive weights and deep adaptive weights are generated respectively. Among them, is the weight of the i-th shallow feature; is the weight of the j-th deep feature; S i [:, l, d] is the value of the i-th shallow feature at the l-th position and the d-th dimension; D j [:, l, d] is the value of the j-th deep feature at the l-th position and the d-th dimension; L is the sequence length, and D is the hidden dimension; N s is the number of shallow layers, N d is the number of deep layers.
4. The customs contraband detection method based on multimodal adaptive fusion according to claim 1, characterized in that, The process of performing multi-scale pooling operations on the deep fusion features includes: Perform one-dimensional average pooling on the deep fusion feature F d and permute the dimensions to change the shape to [B, D, L]; Apply pooling operation The output shape is [B, D, L′], where Through linear interpolation operation After adjusting the sequence length of the deep pooling features to L, adjust the dimension order to [B, L, D]; B is the batch size, L' and L are the sequence lengths, and D is the hidden dimension.
5. The customs contraband detection method based on multi-modal adaptive fusion according to claim 1, characterized in that, For the visual feature representation, introduce a bilateral soft matching strategy, use cosine similarity to measure the token similarity, match and fuse the token pairs with the highest similarity, and generate new tokens; Specifically include: evenly divide all tokens into two sets A and B. For each token in set A, use cosine similarity to find the most similar token in set B and connect them with an edge; after completion of the matching, fuse the original tokens through weighted averaging to generate new tokens; Combine the paired image-text pairs with the corresponding instruction templates to construct a visual instruction tuning dataset for model fine-tuning; and map the visual features to the semantic space consistent with the text features through a double-layer mapping layer to achieve modality alignment; then the two-stage training process includes: the first stage is the cross-modal mapping optimization stage, freeze the parameters of the visual encoder and the language encoder, and train the double-layer mapping layer to optimize the mapping relationship between the visual features and the text features to achieve preliminary alignment between modalities; the second stage is the visual instruction tuning stage, unfreeze the parameters of the visual encoder and the language encoder, and jointly perform overall model training with the double-layer mapping layer.
6. The customs contraband detection method based on multimodal adaptive fusion according to claim 1, wherein, For a sequence of length L, the target answer X a The probability p(X a |X v ,X q ) is defined as: Among them, X v is the visual feature after modal alignment; X q is the text feature of the instruction template; X q,<i are all Tokens of the instruction template, X a,<i are the first i - 1 Tokens in the answer, respectively representing the instruction and answer context that the language encoder has processed when predicting the current answer Token x i ; It is expressed that the language encoder predicts each token of the answer one by one in an autoregressive manner. i , and when making each prediction, it refers to the visual features X v , the instruction token X q,<i and the generated answer tokens X a,<i , and calculates the probability p(x i | X v , X q,<i , X a,<i ) of the current token. Finally, the probabilities of all tokens are multiplied to obtain the probability of the entire answer; the final generated result includes the name of the contraband and the bounding box position.
7. A customs contraband detection system based on multi-modal adaptive fusion, characterized in that, Including: A feature extraction module, configured to obtain customs clearance image-text pairs, and extract the original features including multi-scale features of the customs clearance images; An adaptive fusion module, configured to use the first N layers of features in the multi-scale features as shallow features, and the remaining layer features as deep features. After weighted fusion of the shallow features and deep features respectively, generate shallow fusion features and deep fusion features. After performing multi-scale pooling operations on the deep fusion features, obtain pooling features. Concatenate the shallow fusion features, deep fusion features, and pooling features to generate enhanced features, and integrate the enhanced features and the original features to obtain visual features; A training module, configured to perform modality alignment on visual features and text features, and based on the obtained aligned multi-modal features, perform two-stage training on a detection model, and use the trained detection model for a customs clearance image to be measured to obtain a contraband detection result.
8. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in any one of claims 1-6 is completed.
9. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the method described in any one of claims 1-6 is completed.
10. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by the processor, the method described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Contraband detection method, device and system for X-ray security check image
CN114943907A
Weak supervision saliency target detection method and system for unmanned aerial vehicle video data
CN117173394A
Rolling bearing unknown fault detection method based on multi-modal feature fusion enhancement
CN117516937A
Method and device for detecting forbidden articles in complex environment based on multi-scale feature fusion
CN117765378A
Abnormal behavior detection method and device based on improved YOLOv8 and storage medium
CN119152344A