A customs risk early warning method and system based on cross-modal retrieval enhancement
By combining the multimodal large model SAVA with SC-DAG and AFS modules, efficient fusion and risk scoring of cross-modal evidence were achieved, solving the problem of low risk identification efficiency in customs clearance supervision and improving identification and handling capabilities.
Patent Information
- Application Number
- CN202511500449.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-10-21
AI Technical Summary
When faced with complex and ever-changing cross-border trade scenarios, the existing customs clearance and supervision system suffers from several problems. Single-modal models are unable to achieve unified modeling of cross-modal evidence, and multimodal large models lack the ability to express features in structured tabular data. This results in low efficiency in risk identification and early warning, and serious misjudgments and omissions.
The multimodal large model SAVA is adopted, combined with the spatial-channel decoupled attention gating module (SC-DAG) and the adaptive hierarchical feature selection module (AFS). Cross-modal understanding is achieved through hierarchical weighted table retrieval enhancement method, realizing efficient fusion of multi-level feature representations, and performing risk scoring and classification.
It significantly improves the identification and handling of high-risk goods and abnormal declarations in customs clearance supervision, enhances cross-modal understanding and reasoning capabilities, and ensures robustness and effectiveness in identification under complex scenarios.
Smart Images

Figure CN120975681B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of risk warning technology, specifically relating to a customs risk warning method and system based on cross-modal retrieval enhancement. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Cross-border logistics and customs clearance operations are increasingly facing the challenges of diverse goods, large quantities, and varied risks due to the rapid expansion of global trade. Customs oversight requires not only reviewing structured import and export declarations but also comprehensively assessing relevant evidence. Traditional oversight models rely heavily on manual experience and rule-based matching, which is effective for fixed-pattern risks but often proves inefficient and prone to misjudgments and omissions in complex and ever-changing cross-border trade scenarios.
[0004] To enhance the automation of customs supervision, machine learning and deep learning methods can be introduced, such as using statistical methods and classifiers to predict risks in declaration data. However, these methods are typically limited to single-modal data, lack cross-modal information fusion capabilities, and struggle to simultaneously process heterogeneous data from multiple sources, such as tables and text, making it difficult to comprehensively capture potential risk signals. Furthermore, existing early warning models mostly rely on static thresholds and fixed algorithms, making it difficult to cope with new risk patterns and dynamically changing customs clearance environments.
[0005] The emergence of Retrieval Enhanced Generation (RAG) methods has provided new possibilities for intelligent reasoning in customs supervision under complex data scenarios. RAG establishes connections between multi-source information by dynamically retrieving external knowledge bases and combining them with large model reasoning, thereby improving interpretability and accuracy. However, existing RAG research is mostly focused on document retrieval and question answering tasks, and the retrieval enhancement of import and export declaration data with tables as the core is still imperfect. It lacks cross-table evidence fusion and hierarchical weighting mechanisms, making it difficult to directly adapt to the needs of customs risk identification.
[0006] Meanwhile, the development of Multimodal Large Language Models (MLLMs) has significantly improved cross-modal understanding and reasoning capabilities, enabling unified processing of multiple modalities such as text and tables, and demonstrating excellent performance in general tasks. However, MLLMs treat the feature encoder as a frozen module, taking only the features from the last or penultimate layer as feature input, ignoring the fine-grained information contained in shallow features, resulting in the underutilization of multi-level feature representation. To address this challenge, multi-level feature fusion strategies can be adopted, but most methods rely on simple concatenation or linear weighting, which cannot effectively distinguish the differences between different levels of features in spatial and channel dimensions, easily introducing redundancy and computational overhead.
[0007] In summary, the existing customs clearance supervision has the following shortcomings:
[0008] (1) Single-modal models are difficult to achieve unified modeling of cross-modal evidence and cannot fully adapt to the hierarchical retrieval and fusion requirements of structured tabular data;
[0009] (2) The multimodal large model has insufficient feature representation ability in structured tabular data, which limits its effectiveness in fine-grained field partitioning and association modeling. Summary of the Invention
[0010] To address the aforementioned issues, this invention proposes a customs risk warning method and system based on cross-modal retrieval enhancement. It employs a multimodal large model (Spatial-Adaptive Vision Assistant, SAVA) and utilizes a spatial-channel decoupled attention gate (SC-DAG) and an adaptive feature selector (AFS) module. This achieves efficient fusion of multi-layer feature representations while maintaining information fidelity, providing more discriminative feature representations for multimodal reasoning and further enhancing the automated identification and warning capabilities for prohibited and high-risk goods during customs clearance.
[0011] According to some embodiments, the first aspect of the present invention provides a customs risk warning method based on cross-modal retrieval enhancement, employing the following technical solution:
[0012] A customs risk early warning method based on cross-modal retrieval enhancement includes:
[0013] Obtain the customs task instruction dataset;
[0014] The hierarchical weighted table retrieval enhancement method was used to perform cross-table retrieval on the acquired customs task instruction dataset to obtain a structured evidence package of the dataset.
[0015] Based on the multimodal large model, cross-modal understanding is performed on the obtained structured evidence package to obtain an early warning middleware for the dataset;
[0016] The obtained early warning middleware is scored to complete the customs risk warning based on cross-modal retrieval enhancement.
[0017] As a further technical limitation, the multimodal large model includes at least a spatial-channel decoupled attention gating module and an adaptive hierarchical feature selection module; the spatial-channel decoupled attention gating module refines and decouples the obtained structured evidence package, models it in spatial location and channel dimensions respectively, obtains spatial attention weights and channel attention weights, and performs additive fusion of the obtained spatial attention weights and channel attention weights to obtain the gating features of the dataset.
[0018] Furthermore, the gating features of the obtained dataset are multi-dimensional features. The obtained multi-dimensional features are hierarchically grouped, and the grouped gating features are evaluated by global pooling and lightweight networks to obtain the importance of features at different levels. Based on the importance of the obtained features at different levels, the grouped gating features are weighted and summed to obtain multiple sets of adaptive fusion features. Based on the obtained adaptive weighted multiple sets of fusion features, the features are concatenated to obtain the enhanced features, that is, the early warning middleware of the dataset.
[0019] As a further technical limitation, in the early warning scoring process, the structural consistency score, statistical residual score, and evidence coverage score are calculated sequentially based on the obtained early warning middleware. The obtained structural consistency score, statistical residual score, and evidence coverage score are then weighted and fused to obtain a comprehensive risk score. Based on the comprehensive risk score and rolling statistical threshold, hierarchical processing and closed-loop management are performed to complete the customs risk early warning based on cross-modal retrieval enhancement.
[0020] As a further technical limitation, during the cross-table retrieval process, the customs task instruction dataset is transformed into a query through a hierarchical weighted table retrieval enhancement method. Cross-table retrieval is performed by combining vector similarity and hierarchical weights, and statistical results and representative samples are aggregated to form a structured evidence package of the dataset that includes at least statistical intervals, representative sample summaries, and traceable identifiers.
[0021] As a further technical limitation, the resulting early warning middleware shall include at least the image conclusions, key attributes and evidence entries of the customs declaration data.
[0022] According to some embodiments, a second aspect of the present invention provides a customs risk warning system based on cross-modal retrieval enhancement, employing the following technical solution:
[0023] A customs risk early warning system based on cross-modal retrieval enhancement includes:
[0024] The acquisition module is configured to acquire a dataset of customs task instructions.
[0025] The retrieval module is configured to perform cross-table retrieval on the acquired customs task instruction dataset based on the hierarchical weighted table retrieval enhancement method to obtain a structured evidence package of the dataset;
[0026] The processing module is configured to perform cross-modal understanding of the obtained structured evidence package based on the multimodal large model, and obtain an early warning middleware for the dataset;
[0027] The early warning module is configured to score the obtained early warning middleware and complete customs risk warning based on cross-modal retrieval enhancement.
[0028] According to some embodiments, a third aspect of the present invention provides a computer-readable storage medium, employing the following technical solution:
[0029] A computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the customs risk warning method based on cross-modal retrieval enhancement as described in the first aspect of the present invention.
[0030] According to some embodiments, the fourth aspect of the present invention provides an electronic device, which adopts the following technical solution:
[0031] An electronic device includes a memory, a processor, and a program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps in the customs risk warning method based on cross-modal retrieval enhancement as described in the first aspect of the present invention.
[0032] According to some embodiments, the fifth aspect of the present invention provides a computer program product, which adopts the following technical solution:
[0033] A computer program product includes software code, wherein the program in the software code performs the steps of the customs risk warning method based on cross-modal retrieval enhancement as described in the first aspect of the present invention.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] (1) This invention significantly improves the level of identification and handling of high-risk goods and abnormal declarations in customs clearance supervision by using an integrated risk warning method that combines cross-modal understanding, table-enhanced retrieval, refined feature fusion, risk identification and intelligent risk warning capabilities.
[0036] (2) This invention adopts a hierarchical weighted table retrieval enhancement (HCRAG) linked with “text retrieval + selective table execution”, and introduces commodity coding hierarchy and routing weight and sparse fallback strategy to significantly improve cross-table evidence hit rate and coverage; it integrates an innovative SAAC connector (composed of SC-DAG module and AFS module) in the multimodal model SAVA to implement spatial-channel decoupling and hierarchical adaptive fusion of multi-dimensional features, and simultaneously preserves shallow details and deep semantics without significantly increasing computational overhead, thereby significantly improving the robustness and robustness of recognition in complex scenarios such as occlusion and disguise. Attached Figure Description
[0037] The accompanying drawings, which form part of this embodiment, are used to provide a further understanding of this embodiment. The illustrative embodiments and their descriptions are used to explain this embodiment and do not constitute an improper limitation of this embodiment.
[0038] Figure 1 This is a flowchart of the customs risk warning method based on cross-modal retrieval enhancement in Embodiment 1 of the present invention;
[0039] Figure 2 The diagram shows the detailed steps of the customs risk warning method based on cross-modal retrieval enhancement in Embodiment 1 of the present invention.
[0040] Figure 3 This is an architecture diagram of the SAVA model in Embodiment 1 of the present invention;
[0041] Figure 4 This is a schematic diagram of the SC-DAG structure in Embodiment 1 of the present invention;
[0042] Figure 5 This is a schematic diagram of the AFS structure in Embodiment 1 of the present invention;
[0043] Figure 6 This is a structural block diagram of the customs risk warning system based on cross-modal retrieval enhancement in Embodiment 2 of the present invention. Detailed Implementation
[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0045] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0046] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0047] In this invention, terms such as "upper," "lower," "left," "right," "front," "back," "vertical," "horizontal," "side," and "bottom" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used only to facilitate the description of the structural relationships of the various components or elements of this invention and do not specifically refer to any component or element in this invention. They should not be construed as limiting the invention.
[0048] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0049] Example 1
[0050] Embodiment 1 of this invention introduces a customs risk warning method based on cross-modal retrieval enhancement.
[0051] To address the challenge of establishing a unified modeling framework between structured import / export declaration forms and unstructured declaration texts, and to overcome the limitations of traditional methods that rely solely on a single modality, this paper proposes a table retrieval enhancement mechanism that supports hierarchical indexing and weighted aggregation to overcome the shortcomings of existing retrieval enhancement generation methods in processing table data. This mechanism enables dynamic evidence discovery and fusion across tables and multiple dimensions. Furthermore, existing multimodal models typically rely on single-layer feature representations, resulting in insufficient detailed information and difficulty in supporting accurate risk assessment in complex and ever-changing declaration data scenarios.
[0052] This embodiment innovatively proposes the SAAC (Spatial-Channel Aware Adaptive Connector) module, which fully leverages the complementarity between shallow details and deep semantics through spatial-channel decoupling attention and adaptive feature selection mechanisms, thereby obtaining a more complete and discriminative multimodal representation. Existing early warning mechanisms rely on static thresholds and lack flexible hierarchical capabilities. By introducing a lightweight intelligent early warning model, anomaly detection and risk classification prompts are performed on multimodal inference results, ensuring efficient and stable early warning capabilities in complex and ever-changing customs clearance environments. This embodiment constructs an integrated method and system that combines cross-modal understanding, table-enhanced retrieval, refined feature fusion, risk identification, and intelligent risk early warning capabilities, significantly improving the identification and handling of high-risk goods and abnormal declarations in customs clearance supervision.
[0053] like Figure 1 and Figure 2 The customs risk warning method based on cross-modal retrieval enhancement shown includes:
[0054] Obtain the customs task instruction dataset;
[0055] The hierarchical weighted table retrieval enhancement method was used to perform cross-table retrieval on the acquired customs task instruction dataset to obtain a structured evidence package of the dataset.
[0056] Based on the multimodal large model, cross-modal understanding is performed on the obtained structured evidence package to obtain an early warning middleware for the dataset;
[0057] The obtained early warning middleware is scored to complete the customs risk warning based on cross-modal retrieval enhancement.
[0058] As one or more implementation methods, the customs import declaration dataset is standardized, including field verification, null value imputation, and standardization of currency and weight units to ensure structural consistency. To avoid interference from extreme values, quantile rules are used for truncation to obtain a standardized numerical range. Simultaneously, derived indicators (such as unit value) are calculated, and time windows and "origin → destination" combination features are generated to support subsequent retrieval and risk analysis. A unique number and hash value are generated for each record to ensure traceability.
[0059] In the process of cleaning and standardizing the customs import declaration dataset, this embodiment performs field verification and null value filling, and fixes a unified pattern for 22 key attributes, including commodity code, country of origin, destination, declared amount, net weight, quantity, trade terms, declaring company identifier, and declaration date. The declared amount is uniformly converted to US dollars, and the weight is uniformly converted to kilograms. Quantile rules are used to truncate extreme outliers to ensure a reasonable numerical range. Derived indicators (unit value) are calculated, and a time window field and a "country of origin → destination" combined feature are generated. A unique number and hash value are generated for each record to ensure data traceability.
[0060] As one or more implementation methods, based on structured declaration data as the core, auxiliary unstructured multi-source heterogeneous data (such as manual inspection notes, risk warning items, review comments, related image materials, etc.) can be collected. Task templates can be designed according to regulatory needs, and these auxiliary data can be combined with tabular text to uniformly construct a multi-task instruction-based dataset. By attaching task identifiers, different task types such as field verification, risk assessment, and anomaly detection can be explicitly distinguished, improving the model's adaptability and generalization ability in multi-task scenarios.
[0061] This embodiment constructs a multimodal instruction fine-tuning dataset for customs tasks. Specifically, it includes: based on structured declaration data as the core, optional collection of auxiliary unstructured multi-source heterogeneous data (such as manual inspection notes, risk warning items, review comments, related image data, and other multi-source heterogeneous information); designing task-oriented instruction templates according to regulatory needs; combining the above auxiliary data with table field text to uniformly construct a multi-task instruction dataset; and using additional task identifiers (such as [verify] for field verification, [detect] for anomaly identification, and [assess] for risk classification) to distinguish different task categories, forming a unified multi-task instruction set to improve the model's adaptability and generalization ability in complex regulatory scenarios.
[0062] As one or more implementation methods, the new declaration records are transformed into queries based on the Hierarchical Weighted Table Retrieval Enhancement (HCRAG) method. Cross-table retrieval is performed by combining vector similarity and hierarchical weights, and statistical results and representative samples are aggregated to form a structured evidence package E (including statistical intervals, representative sample summaries, and traceable identifiers) as evidence input for subsequent reasoning.
[0063] The hierarchical weighted table retrieval enhancement method in this embodiment, referring to the publicly available process of the table retrieval enhancement framework, transforms new declaration records into queries and generates a structured evidence package E through an iterative process combining text retrieval and table execution. Specifically, it includes:
[0064] (1) Offline construction of dual databases, standardization of mapping and query
[0065] Import the customs import declaration dataset into a relational database, fixing the table structure and field types. Simultaneously generate two types of searchable corpora: one is linearized text of table chunks (including column names, examples, and meanings); the other is a table schema description library (table name, column name, type, and examples). Establish a mapping between "chunks → table schemas" and "table rows → row-level identifiers," construct dense vector indexes and keyword inverted indexes, and record timestamps and source information. On the online side, unify the units and currencies of the records to be evaluated, serialize them into query text according to the offline template, explicitly annotate missing fields, and attach contextual anchors such as commodity code level, origin to destination, time segment, and declaring company identifier. Generate query vectors and retain the correspondence with the schema library.
[0066] (2) Two-stage text retrieval and selective table execution
[0067] In the text library (table segments and descriptive text), a dense vector coarse recall is performed to obtain candidate segments. Then, a stronger relevance model is used for fine ranking and noise reduction, retaining the segments most relevant to the target table and target fields while maintaining traceable links to the table schema. When the fine ranking results involve table source segments, relevant table schemas are summarized, and a natural language to SQL conversion tool is used to generate executable statements and run them in the database to obtain a structured result set (statistical aggregation and row-level details). If the current iteration only requires textual evidence, the execution phase is skipped and the merging phase begins.
[0068] (3) Subproblem decomposition, iterative solution and consistency verification
[0069] The overall query is broken down into sub-problems that can be solved sequentially (similar historical intervals, comparisons with the same subject, trends in adjacent time periods, etc.). The "retrieve-execute-merge" process is repeated for each sub-problem, and intermediate answers from the previous round are used as hints and constraints for the next round to avoid redundant calculations. Consistency checks are performed on the text search results and SQL execution results. Conflicting fields are marked with source tags and confidence levels. Weighted selection or parallel presentation is performed based on data recentity, product code level matching degree, and source reliability. The reasons for ignored evidence are recorded for traceability.
[0070] (4) Layered weighting, rollback and quality control
[0071] When merging evidence, commodity coding hierarchy and routing weights are introduced, prioritizing data with six-digit codes and consistent routing. When samples are sparse, a gradual regression to four-digit and two-digit codes or an expanded time window is implemented in a predetermined order, with the regression process recorded. The weight of lower-level evidence is reduced to control bias. Simultaneously, duplicate rows from the same invoice or batch are removed, samples with missing key fields or outliers are eliminated, and samples with excessive concentration in a single enterprise are corrected to ensure diverse evidence sources and controllable quality.
[0072] (5) Generation and integration of evidence package E
[0073] The merged results are organized into a structured evidence package E, which includes at least a field statistical summary and typical range, representative sample key points and row-level backtracking identifiers, source mapping and table schema identifiers or SQL fragment summaries, coverage and recency assessment, rollback records, generation time and configuration version. The output is in JSON or an equivalent structured format as a retrieval enhancement context. Evidence package E and the query question are input into SAVA for inference. When generating the early warning middleware, SAVA writes the evidence item number, field summary and citation relationship into the evidence field of EWI to ensure that the subsequent warning module directly consumes and completes interpretable scoring and classification.
[0074] As one or more implementation methods, the Spatial-Channel Decoupled Attention Gating Module (SC-DAG) models the "importance of spatial location" and "importance of channel content" in multidimensional features respectively. By using a decoupling gating mechanism, it highlights key regions and key semantics to avoid confusion between information from different dimensions, allowing the model to more accurately distinguish "where is important" and "what is important", resulting in cleaner feature input.
[0075] In this embodiment, as Figure 4 As shown, the SC-DAG module refines the raw features output by the feature encoder. By decoupling and modeling the spatial importance of features and their semantic weights in the channel dimension, it provides a more optimized and information-density-rich feature representation for subsequent adaptive hierarchical feature selection. Specifically, this includes:
[0076] Spatial Attention Gate: This aims to identify "where" is more important in a feature representation. It operates on each feature token one by one, assigning each token a scalar attention score to measure its relative importance in the overall semantics. Specifically, input features... First, a multilayer perceptron (MLP) consisting of two linear transformations and a single ReLU activation function performs a nonlinear mapping to capture complex spatial dependencies. Finally, the output is compressed to the (0,1) interval using a sigmoid function to generate the final spatial attention weights. This process can be formalized as follows:
[0077]
[0078] in, These are the features of the input. and These are the learnable weight matrices for the first and second layers of a spatial MLP, respectively. It is the reduction ratio of spatial dimensions. It is a modified linear unit activation function; It is the Sigmoid activation function. It is the final calculated spatial attention weight.
[0079] Channel Attention Gate: Parallel to spatial attention gate, channel attention gate aims to identify "which features are more important." To obtain global channel context information, it first processes the input features... Global average pooling is performed on the spatial dimension (N). from Compressed into a global feature descriptor .
[0080]
[0081] Then, The input consists of a bottlenecked multilayer perceptron (MLP) with two layers of linear transformation and ReLU activation to capture nonlinear dependencies between channels. Finally, channel attention weights are generated through a sigmoid activation function. .
[0082]
[0083] in, It is the global feature descriptor calculated by formula (2). and These are the learnable weight matrices for the first and second layers of a channel MLP, respectively. It is the reduction rate of the channel dimension. The final calculated channel attention weights are expanded along the spatial dimension during computation. , with original features Multiply element-wise along the channel dimension.
[0084] Additive fusion, which involves obtaining spatial attention weights separately... and channel attention weights Then, compare them with the original features respectively. Performing the Hadamard product operation yields two independently optimized feature representations. Unlike many attention mechanisms that use cascaded multiplication for fusion, a key design feature of SC-DAG is to fuse these two feature representations through addition to obtain the final gated feature. :
[0085]
[0086] in, Broadcasting at the channel dimension Broadcasting in the spatial dimension, in order to... Align elements one by one.
[0087] Additive fusion has a significant robustness advantage over multiplicative fusion. In multiplicative fusion (i.e., Low weighting in any dimension can stifle feature transmission, potentially leading to unexpected information loss. Additive fusion, however, allows spatial and channel attention to act as two independent enhancement signals on the original feature. Even if a feature is not spatially prominent (…), Smaller), but if its channel content is very critical ( Even with a relatively large feature size, its information can still be effectively preserved and transmitted. This design makes the feature extraction process more stable and the information retention more complete. After SC-DAG processing, the output features... It has been optimized in both spatial and channel dimensions, providing high-quality "raw materials" for the subsequent Adaptive Hierarchical Fusion (AFS) module.
[0088] As one or more implementation methods, such as Figure 5 As shown, the Adaptive Hierarchical Feature Selection (AFS) module groups multidimensional features hierarchically and evaluates the importance of each layer's features through global pooling and a lightweight network, dynamically allocating weights so that the model can automatically balance shallow details and deep semantics in different scenarios, avoiding information dilution.
[0089] Let the L-layer feature representation after processing by the SC-DAG module be denoted as set. Then, according to SAAC's hierarchical grouping strategy, these L layers of features are divided into G groups, each group including... Each group consists of adjacent features. The features within each group are summed, represented as... This ultimately produces G fused feature representations:
[0090]
[0091] The AFS module considers the features of each layer within any group g. The system learns to assign semantic importance weights to the feature representations of each layer, thereby guiding inter-layer fusion. This process is implemented through a two-stage lightweight network. The first stage assigns semantic importance weights to the feature representations of each layer. Global Average Pooling (GAP) is applied to the spatial dimension to extract the global context semantic descriptor. The descriptor captures the macroscopic semantics of the features at this layer:
[0092]
[0093] In the second stage, these global descriptors are fed into a lightweight MLP network shared between layers, which is responsible for calculating a scalar layer importance score for the features of each layer. :
[0094]
[0095] in, denoted as the layer importance score for the features of the i-th layer. and These represent the weight matrices of the two fully connected layers in an MLP, and and These are the corresponding bias vectors, and these parameters are learned during model training.
[0096] Adaptive Weighted Fusion: After obtaining the importance scores of all M layers within a group, the Softmax function is used to normalize these scores along the layer dimension, thereby generating a set of attention weights with a sum of 1. This hierarchical weighting mechanism allows the model to dynamically model the importance of features at different levels in the current task in the form of probability distributions.
[0097]
[0098] Ultimately, the adaptive fusion features within the group By analyzing the input features of each layer We obtain the weighted summation based on the corresponding attention weights:
[0099]
[0100] As one or more implementation methods, the adaptively weighted fused features are concatenated with the final layer features of the feature encoder and then fed into the projection module to generate the final feature embedding, which represents higher information density, clearer hierarchy, and provides high-quality semantic support for subsequent language model inference. After completing the adaptive fusion of all groups, this embodiment obtains the adaptive fused features of G semantic groups. With the final layer features of the feature encoder The data is concatenated along the channel dimension and input into the projection module to generate the final enhanced feature embedding. :
[0101]
[0102] As one or more implementation methods, such as Figure 3 The multimodal large model SAVA (Spatial-AdaptiveVision Assistant) shown employs a two-stage training method:
[0103] The first stage is cross-modal mapping optimization. In this stage, the main parameters of the feature encoder and language model are frozen, and only the projection module and necessary output reduction parameters are trained to optimize the alignment of features from different modalities (such as table fields, text descriptions, and auxiliary data), achieving a preliminary unified representation between modalities. The optimizer used is AdamW, with a learning rate of [value missing]. The system uses 4 GPUs, has a global batch size of 256, trains for 1 epoch, has a maximum input sequence length of 8192, and saves checkpoints every 1000 steps to obtain stable category recognition, question answering, and basic reference capabilities.
[0104] The second stage is instruction tuning. In this stage, the feature encoder and language encoder are unfrozen and trained together with the projection module to further improve the model's inference and adaptation performance in specific tasks. Specifically, the parameters of the feature encoder, projection module, and large language model are trained simultaneously using a multi-task instruction dataset built based on customs scenarios. Multi-task joint optimization further enhances the model's inference ability and adaptation performance in complex regulatory tasks. During training, the input and output formats are standardized, and the interface and error handling strategies are solidified to ensure stable integration with the evidence retrieval and early warning modules. The optimizer used is AdamW, and the language model learning rate is [not specified]. Feature encoder learning rate The system uses 4 GPUs, has a global batch size of 128, and trains for one epoch. At the end of the phase, the input / output interfaces and error handling strategies are fixed to ensure that the model's output corresponds one-to-one with the EWI field specifications, guaranteeing stable integration with the evidence retrieval and early warning modules.
[0105] As one or more implementation methods, the early warning module reads the EWI (Early Warning Middleware, which includes image conclusions, key attributes, and evidence entries) generated by the multimodal large model SAVA, calculates the structural consistency score, statistical residual score, and evidence coverage score in sequence, and merges them into a comprehensive risk score according to preset weights; the evidence number and citation are retained simultaneously during output to ensure that the results are interpretable and traceable; specifically:
[0106] Read and verify the structure and version of the EWI, check the required fields in the key attribute fields and evidence fields, form a list of missing fields and mark the verification results, complete the consistency processing of currency weight and timestamp and record the conversion source and time, while maintaining the full association between evidence entries and row-level backtracking identifiers to calculate the structural consistency score, evaluate the semantic consistency between the declaration fields and commodity code categories using cross-modal embedding similarity, verify the matching relationship between the origin-to-destination trade terms and the declaration scenario using the rule base and generate traceable trigger records, and check whether key values exceed the limits using empirical intervals and deduct points based on the distance of exceeding the limits. The structural consistency score is calculated using the following formula:
[0107] (11)
[0108] in, The structural consistency score ranges from 0 to 100. The semantic similarity has been normalized to 0 to 1. The penalty for rule conflict is a non-negative value. The penalty for out-of-bounds values is a non-negative value. Non-negative weights can be determined by offline calibration or learning and can be versioned. This indicates that the input is truncated to a lower bound of 0 and an upper bound of 100 to calculate the statistical residual score. Using recent samples from the customs import declaration dataset as a baseline, a rolling distribution is established according to the commodity code level and routing. When the sample is sparse, the time window is expanded sequentially or back to the four-digit and two-digit levels, and the backtracking path is recorded. For fields such as declared amount, unit value, and quantity, robust deviation is measured using the median and MAD or IQR, and quantile mapping is performed. Simultaneously, anomaly scoring from an open-source anomaly detection library is performed and normalized to the same dimension. The weighted average or maximum of the two is taken as the output of this channel, and reverse normalized to 0 to 100 according to the principle that the larger the residual, the higher the risk. A composition description is output to calculate the evidence coverage score. Based on evidence package E, the recentity and source diversity of key fields are measured. Penalties are set for insufficient coverage, excessive time, or single source cases, and a list of gap fields and evidence timeliness is listed. The score fusion stage uses weighted linear fusion to ensure that the sum of the weights is one. The comprehensive risk score is calculated using the following formula:
[0109]
[0110] in, The overall risk score ranges from 0 to 100. To statistically analyze the residuals, values from 0 to 100 are used. The data is divided into values from 0 to 100. The non-negative fusion weights sum to 1 can be calibrated by the development set or updated online adaptively, and the version output is recorded, including details of the score composition, trigger basis evidence number, and uncertainty prompts. At the same time, the algorithm and threshold version input summary generation time and row-level backtracking identifier are saved in the database to support auditing and traceability.
[0111] As one or more implementation methods, the system performs tiered processing based on a comprehensive risk score and rolling statistical thresholds. These thresholds are dynamically calculated from the quantile statistics (e.g., P95, P99) of the import / export risk table dataset over the past 30 days. High-risk samples trigger secondary inspection and reassessment; medium-risk samples are recommended for manual sampling; and low-risk samples can be released directly. The system also saves the tiered results, evidence citations, and handling measures to support subsequent traceability and regulatory closure.
[0112] This embodiment dynamically generates thresholds based on the quantile statistics of the customs import declaration dataset over the past thirty days, prioritizing the distribution of the six-digit commodity code consistent with the routing as the baseline; when the sample is sparse, the time window is successively expanded or the data is rolled back to the four-digit and two-digit levels, and the rollback path and effective interval are recorded. The threshold calculation uses the conditional quantile method:
[0113]
[0114]
[0115] in, The comprehensive risk score (0–100) is defined in Formula 12; HS6 represents the six-digit commodity code condition; route represents the origin-destination routing condition; win=30d represents the rolling sample of the past thirty days; It is a quantile function; For high and median quantiles (e.g.) , To suppress short-term fluctuations, the above thresholds are exponentially smoothed:
[0116]
[0117] in, The threshold for taking effect on day t. This is the original quantile threshold for the day. For example, a smoothing coefficient. The data is stored in the database based on the threshold version and the effective time.
[0118] The tiered handling shall be carried out according to the following rules: When Determined to be high risk; when Determined as medium risk; when It was determined to be low risk.
[0119] in, , These are the smoothed high and median thresholds, conforming to the same (HS6, route) conditional caliber; when the sample size over the past thirty days under this caliber is lower than the minimum sample threshold. If the time window is expanded to sixty or ninety days and still insufficient, the system will revert to the HS4 or HS2 level while maintaining consistent routing. The system will record the revert path and effective interval.
[0120] High-risk scenarios automatically trigger secondary security checks and re-evaluation or commodity code verification, generating a supplementary document list and processing time limit based on the evidence gaps outlined in Article 10. Medium-risk scenarios suggest manual spot checks or supplementary documentation and enter a time-limited review process. Low-risk scenarios allow passage or include in random spot checks. The classification results, handling measures, and evidence citations are archived together, retaining the model version, threshold version, and effective date. Manual review conclusions are written back to update the rolling baseline and weight parameters, monitor category and route distribution drift, and automatically re-evaluate thresholds periodically and generate change records, thus forming a traceable and auditable closed-loop early warning mechanism.
[0121] To verify the effectiveness of the method, this embodiment uses the customs import declaration dataset shown in Table 1. To evaluate the model's multimodal capabilities, this embodiment tests it on six representative benchmarks, as shown in Tables 2 and 3. GQA mainly examines the model's multi-step reasoning and relational understanding; POPE is used to evaluate the model's performance in preventing illusions and maintaining semantic consistency; SQAI combines subject knowledge and multi-source information to test the model's cross-modal knowledge fusion and reasoning capabilities; TextVQA emphasizes the recognition and semantic understanding of textual information; VizWiz is geared towards complex and low-quality inputs to test the model's robustness and answerability judgment in complex environments; AI2D tests the model's capabilities in structured graphics and element parsing through illustrative diagrams and process relational question answering. These benchmarks collectively cover core dimensions such as reasoning, semantic consistency, text recognition, robustness, and structured graph understanding, providing a comprehensive perspective for comparing the performance of multimodal models.
[0122] Table 1 Customs Import Declaration Data Set
[0123]
[0124] Table 2 Test Results 1
[0125]
[0126] Table 3 Test Results 2
[0127]
[0128] This embodiment employs a hierarchical weighted table retrieval enhancement (HCRAG) combined with "text retrieval + selective table execution," and introduces product coding hierarchy, routing weights, and a sparse backoff strategy to significantly improve cross-table evidence hit rate and coverage. Image conclusions, location coordinates, and table statistics are uniformly encapsulated into an early warning middleware (EWI), achieving decoupling of inference and warning and supporting end-to-end traceability. The multimodal model SAVA integrates an innovative SAAC connector (composed of an SC-DAG module and an AFS module) to implement spatial-channel decoupling and hierarchical adaptive fusion of multi-dimensional features, simultaneously preserving shallow details and deep semantics without significantly increasing computational overhead, thereby significantly improving recognition robustness and robustness in complex scenarios such as occlusion and camouflage. A unified task-oriented input and output reduction is adopted to ensure the structured and standardized generation results, thereby improving the stability of engineering integration. A two-stage training method is used for SAVA, with cross-modal pre-training alignment and instruction tuning completed using a publicly available general dataset. The warning module uses a lightweight and interpretable three-part scoring head (…). The Header component integrates structural consistency, statistical residuals, and evidence coverage through a weighted fusion process, outputting evidence numbers and triggering criteria for business review and auditing. It employs offline quantile thresholds based on HS six-digit numbers and routing layers, providing corresponding high / medium / low handling suggestions and backtracking information. Demonstration and calibration are completed using publicly available customs import declaration data, avoiding sensitive data exposure and facilitating migration to private business data. The results retain metadata such as rollback paths, SQL summaries, evidence sources, and model / threshold versions, supporting auditing, maintenance, and closed-loop governance.
[0129] Example 2
[0130] Embodiment 2 of the present invention introduces a customs risk early warning system based on cross-modal retrieval enhancement.
[0131] like Figure 6 The customs risk early warning system shown includes:
[0132] The acquisition module is configured to acquire a dataset of customs task instructions.
[0133] The retrieval module is configured to perform cross-table retrieval on the acquired customs task instruction dataset based on the hierarchical weighted table retrieval enhancement method to obtain a structured evidence package of the dataset;
[0134] The processing module is configured to perform cross-modal understanding of the obtained structured evidence package based on the multimodal large model, and obtain an early warning middleware for the dataset;
[0135] The early warning module is configured to score the obtained early warning middleware and complete customs risk warning based on cross-modal retrieval enhancement.
[0136] The detailed steps are the same as those of the customs risk warning method based on cross-modal retrieval enhancement provided in Example 1, and will not be repeated here.
[0137] Example 3
[0138] Embodiment 3 of the present invention provides a computer-readable storage medium.
[0139] A computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the customs risk warning method based on cross-modal retrieval enhancement as described in Embodiment 1 of the present invention.
[0140] The detailed steps are the same as those of the customs risk warning method based on cross-modal retrieval enhancement provided in Example 1, and will not be repeated here.
[0141] Example 4
[0142] Embodiment 4 of the present invention provides an electronic device.
[0143] An electronic device includes a memory, a processor, and a program stored in the memory and running on the processor. When the processor executes the program, it implements the steps in the customs risk warning method based on cross-modal retrieval enhancement as described in Embodiment 1 of the present invention.
[0144] The detailed steps are the same as those of the customs risk warning method based on cross-modal retrieval enhancement provided in Example 1, and will not be repeated here.
[0145] Example 5
[0146] Embodiment 5 of the present invention provides a computer program product.
[0147] A computer program product includes software code, wherein the program in the software code performs the steps of the customs risk warning method based on cross-modal retrieval enhancement as described in Embodiment 1 of the present invention.
[0148] The detailed steps are the same as those of the customs risk warning method based on cross-modal retrieval enhancement provided in Example 1, and will not be repeated here.
[0149] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0150] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0151] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0152] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0153] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0154] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0155] The above description is merely a preferred embodiment of this practice and is not intended to limit the scope of this practice. Various modifications and variations can be made to this practice by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this practice should be included within the protection scope of this practice.
Claims
1. A customs risk early warning method based on cross-modal retrieval enhancement, characterized in that, include: Obtain the customs task instruction dataset; The hierarchical weighted table retrieval enhancement method was used to perform cross-table retrieval on the acquired customs task instruction dataset to obtain a structured evidence package of the dataset. Based on the multimodal large model, cross-modal understanding is performed on the obtained structured evidence package to obtain an early warning middleware for the dataset; The obtained early warning middleware is scored to complete the customs risk warning based on cross-modal retrieval enhancement; Based on the structured customs import declaration dataset, auxiliary unstructured multi-source heterogeneous data is collected. Task templates are designed according to regulatory requirements. The collected auxiliary unstructured multi-source heterogeneous data is combined with table text to uniformly construct a multi-task instruction dataset. The multimodal large model includes at least a spatial-channel decoupled attention gating module and an adaptive hierarchical feature selection module; The spatial-channel decoupled attention gating module refines and decouples the obtained structured evidence package, models it in spatial location and channel dimensions respectively, and obtains spatial attention weights and channel attention weights. The obtained spatial attention weights and channel attention weights are additively fused to obtain the gating features of the dataset. The gating features of the obtained dataset are multidimensional features. The obtained multidimensional features are hierarchically grouped, and the grouped gating features are evaluated by global pooling and lightweight network to obtain the importance of features at different levels. The grouped gating features are weighted and summed according to the importance of features at different levels to obtain multiple sets of adaptive fusion features. The multiple sets of adaptively weighted fusion features are concatenated to obtain enhanced features, which are the early warning middleware of the dataset. During the cross-table retrieval process, the new customs task declaration records are transformed into queries through a hierarchical weighted table retrieval enhancement method. Cross-table retrieval is performed by combining vector similarity and hierarchical weights, and statistical results and representative samples are aggregated to form a structured evidence package of the dataset that includes at least statistical intervals, representative sample summaries, and traceable identifiers. The resulting early warning middleware includes at least image conclusions, key attributes, and evidence entries from customs declaration data.
2. The customs risk early warning method based on cross-modal retrieval enhancement as described in claim 1, characterized in that, In the early warning scoring process, the structural consistency score, statistical residual score, and evidence coverage score are calculated sequentially based on the obtained early warning middleware. The obtained structural consistency score, statistical residual score, and evidence coverage score are then weighted and fused to obtain a comprehensive risk score. Based on the comprehensive risk score and rolling statistical threshold, a graded processing and closed-loop management system is implemented to complete customs risk early warning based on cross-modal retrieval enhancement.
3. A customs risk early warning system based on cross-modal retrieval enhancement, employing the customs risk early warning method based on cross-modal retrieval enhancement as described in claim 1, characterized in that, include: The acquisition module is configured to acquire a dataset of customs task instructions. The retrieval module is configured to perform cross-table retrieval on the acquired customs task instruction dataset based on the hierarchical weighted table retrieval enhancement method to obtain a structured evidence package of the dataset; The processing module is configured to perform cross-modal understanding of the obtained structured evidence package based on the multimodal large model, and obtain an early warning middleware for the dataset; The early warning module is configured to score the obtained early warning middleware and complete customs risk warning based on cross-modal retrieval enhancement.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the steps of the customs risk warning method based on cross-modal retrieval enhancement as described in any one of claims 1-2.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the steps of the customs risk warning method based on cross-modal retrieval enhancement as described in any one of claims 1-2.
6. A computer program product, comprising software code, characterized in that, The program in the software code performs the steps of the customs risk warning method based on cross-modal retrieval enhancement as described in any one of claims 1-2.
Citation Information
Patent Citations
Custom clearance risk identification method and device, equipment, medium and program product
CN116911591A
Field medical sharp instrument recovery management method and system based on module anti-permeation structure
CN119784283A