Customs clearance risk detection method based on customs clearance document identification

By conducting deep learning and dynamic benchmarking model construction of customs declaration data, the problems of inefficiency and high misjudgment rate of existing systems in identifying complex and hidden abnormal behaviors are solved, efficient abnormal detection and early warning are achieved, and customs supervision capabilities are improved.

CN120030446AInactive Publication Date: 2025-05-23CHINA XIAMEN OCEAN SHIPPING AGENCY CO LTD

Patent Information

Application Number
CN202510522693.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the existing customs supervision system processes massive customs declaration documents, it is difficult to accurately identify complex and hidden abnormal behaviors, resulting in high inefficiency and misjudgment rates, which cannot meet the needs of dynamic supervision.

Method used

By obtaining historical data from customs declaration documents, performing data cleaning and time series segmentation, using long and short-term memory network to extract customs declaration behavior characteristics, constructing a dynamic benchmark model, calculating abnormal deviation degrees, and combining a random forest algorithm for abnormal classification and early warning.

Benefits of technology

It has achieved efficient identification and early warning of abnormal behaviors during customs declaration, and improved customs supervision efficiency and risk prevention and control capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030446A_ABST
    Figure CN120030446A_ABST
Patent Text Reader

Abstract

The invention provides a customs clearance risk detection method based on customs clearance document identification, which comprises the following steps: acquiring historical data from customs clearance documents, removing noise data through data cleaning and preprocessing, dividing continuous historical data into a plurality of sub-data sets according to time windows by adopting a time sequence segmentation method, and obtaining a structured customs clearance behavior sequence; according to the preliminary behavior mode representation, grouping normal customs declaration behaviors in historical data through a clustering analysis method, calculating a behavior characteristic mean value and variance in each time window in combination with statistical distribution characteristics, and determining parameters of a dynamic reference model; and multi-level early warning output is obtained, so that high-risk abnormal behaviors can be found in time and corresponding early warning can be triggered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a customs clearance risk detection method based on customs declaration document identification. Background Art

[0002] Against the backdrop of the rapid development of global trade, the intelligence of customs supervision has become an indispensable key to ensuring trade security and efficiency. With the rapid expansion of cross-border e-commerce and international logistics, customs need to process a huge amount of customs declaration document data every day. How to accurately identify anomalies and issue timely warnings is directly related to the core needs of combating violations and maintaining economic order. The importance of this field is self-evident, and its research and application are of strategic significance to improving the level of modern customs management. However, the commonly used methods in customs supervision currently rely on manual review or simple rule matching. Although they can detect explicit problems to a certain extent, they often seem powerless in the face of complex and changeable trade scenarios and highly concealed abnormal behaviors. Especially with the surge in data volume, the inefficiency and misjudgment rate of traditional methods have gradually been exposed, making it difficult to meet the needs of dynamic supervision.

[0003] The limitations of existing solutions are mainly reflected in the lack of deep mining and adaptive capabilities for historical data. Most systems are based only on fixed rules or shallow statistical analysis, making it difficult to capture the potential patterns of customs declaration behavior, and unable to cope with the rapid evolution of new types of violations. This results in limited accuracy and real-time performance of anomaly detection, especially when it comes to cross-validation of multi-dimensional data, the system is prone to missed reports or false alarms. The core challenges focus on how to establish a dynamic normal behavior benchmark through technical means, and how to efficiently identify anomalies that deviate from this benchmark in massive data. Specifically, deep learning modeling of historical data, differentiated identification of multiple types of anomalies, and real-time triggering of early warning mechanisms are technical challenges that need to be overcome. These factors have not been resolved, making it difficult for the system to accurately distinguish between normal fluctuations and real risks, and it is impossible to provide customs staff with graded and actionable decision support.

[0004] Therefore, how to build a system based on deep learning of historical customs declaration data that can dynamically update the baseline model and implement multi-level early warning for different types of anomalies has become a key issue that needs to be urgently addressed in this study. Summary of the invention

[0005] The purpose of the present invention is to solve the above-mentioned problems and provide a customs clearance risk detection method based on customs declaration document identification.

[0006] The technical solution of the present invention is achieved in this way: The present invention provides a customs clearance risk detection method based on customs declaration document identification, comprising: Historical data is obtained from customs declaration documents, and noise data is removed through data cleaning and preprocessing. The continuous historical data is divided into multiple sub-datasets according to the time window using the time series segmentation method to obtain a structured customs declaration behavior sequence; for the structured customs declaration behavior sequence, the long short-term memory network in deep learning is used to extract the features of the customs declaration behavior within the time window, and the dynamic change trend of the customs declaration behavior is generated through multi-layer neural network training to obtain a preliminary behavior pattern representation; based on the preliminary behavior pattern representation, the normal customs declaration behaviors in the historical data are grouped through the clustering analysis method, and the mean and variance of the behavior characteristics in each time window are calculated in combination with the statistical distribution characteristics to determine the parameters of the dynamic benchmark model; in the dynamic Based on the parameters of the dynamic benchmark model, the Euclidean distance between the current customs declaration behavior and the dynamic benchmark is calculated to quantify the abnormal deviation. If the abnormal deviation exceeds the preset threshold, it is marked as a potential abnormal behavior to obtain the initial screening result of anomaly detection; for the initial screening result of anomaly detection, the random forest algorithm is used to classify the multi-dimensional features of potential abnormal behaviors, and the classifier is trained in combination with known violation cases in historical data to determine the type of abnormal behavior and obtain a graded abnormal classification result; high-risk abnormal types are extracted from the graded abnormal classification results, and the current customs declaration documents are continuously monitored through the real-time data stream processing framework. If a high-risk abnormal type is detected, the corresponding warning signal is triggered to obtain a real-time warning trigger status.

[0007] As a further improvement, the method further comprises: According to the real-time warning trigger status, the message queue technology is used to distribute the warning signal to different processing channels according to priority. The signal is processed in layers through the pre-established multi-level warning rules to generate multi-level warning outputs for different abnormal types.

[0008] As a further improvement, the method further comprises: After obtaining the multi-level warning output, the warning results are compared with the behavior patterns in the historical data through the feedback mechanism, and the parameters of the deep learning model are dynamically adjusted using the online learning method to obtain an updated dynamic benchmark model. As a further improvement, the method further includes: For the updated dynamic benchmark model, anomaly detection and early warning processing are performed on the new round of customs declaration document data through a cyclic iterative method, and sliding window technology is used to continuously track the behavioral change trends in trade scenarios to obtain continuously optimized anomaly early warning capabilities.

[0009] The technical solution provided by the embodiment of the present invention may have the following beneficial effects: The method of the present invention cleans and segments the historical customs declaration data in time series, extracts the characteristics of customs declaration behavior using a long short-term memory network, and constructs a dynamic benchmark model. On this basis, the present invention uses Euclidean distance to calculate the abnormal deviation degree, and combines the random forest algorithm to classify potential abnormal behaviors. Through real-time data stream processing and multi-level early warning rules, the present invention can timely detect high-risk abnormal behaviors and trigger corresponding early warnings. In addition, the present invention can also introduce feedback mechanisms and online learning methods to continuously optimize the dynamic benchmark model and improve the accuracy of anomaly detection. This method can effectively identify and warn abnormal behaviors in the customs declaration process, and improve customs supervision efficiency and risk prevention and control capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 The present invention is a flow chart of a method for detecting customs clearance risks based on customs declaration document identification. DETAILED DESCRIPTION

[0011] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.

[0012] like Figure 1 As shown, a customs clearance risk detection method based on customs declaration document identification in an embodiment of the present invention may specifically include: Step S101, historical data is obtained from the customs declaration documents, noise data is removed through data cleaning and preprocessing, and the continuous historical data is divided into multiple sub-data sets according to the time window using the time series segmentation method to obtain a structured customs declaration behavior sequence.

[0013] Historical data is obtained from customs declaration documents, and the complete data set is extracted through scanning and parsing technology to obtain the initial data set. The initial data set is preprocessed through data cleaning technology, and the outlier detection method is used to remove noise data to obtain the cleaned data set. For the cleaned data set, the time series analysis method is used to identify the characteristics of continuous data and determine the time series pattern. The time series pattern is segmented through the time window partitioning technology to obtain multiple sub-data sets to obtain the segmented data set. According to the segmented data set, the customs declaration behavior sequence characteristics are extracted, and the behavior patterns in the customs declaration behavior sequence characteristics are classified by clustering algorithm to obtain the classified behavior sequence. For the classified behavior sequence, a structured data format is generated, and the data is integrated through field mapping technology to obtain a structured behavior data set. If there are missing values ​​in the structured behavior data set, they are supplemented by interpolation method to obtain the final behavior sequence data set.

[0014] Exemplarily, when obtaining historical data from customs declaration documents, firstly, the customs declaration records of the past year are extracted through database query, for example, the customs declaration data from January 1, 2022 to December 31, 2022 are extracted, a total of 365 days of records, each record contains fields such as customs declaration time, commodity code, declared amount, import and export type, etc. Then, data cleaning and preprocessing are carried out to remove noise data, such as deleting abnormal records with declared amounts of 0 or negative values, eliminating invalid data with missing commodity codes, and standardizing the time field to ensure that all time formats are unified as "YYYY-MM-DD HH:MM:SS". In the cleaning process, a statistical outlier detection method is adopted, such as using the Z-score algorithm, and data points with Z-score absolute values ​​greater than 3 are regarded as outliers and eliminated, and finally about 95% of valid data are retained. Subsequently, the time series segmentation method is used to divide the continuous historical data into multiple sub-datasets according to the time window. For example, with 7 days as a time window, the 365-day data is divided into 52 sub-datasets, each of which contains a 7-day customs declaration behavior sequence. During the segmentation process, the sliding window technology is used to ensure that there is a 1-day overlap between each window to capture the changes in customs declaration behavior in continuous time. Finally, the customs declaration behavior sequence in each sub-dataset is structured, for example, the customs declaration behavior in each window is arranged in chronological order, and key features are extracted, such as the number of daily customs declarations, the average declared amount, the main commodity categories, etc., to form a structured customs declaration behavior sequence for subsequent analysis and modeling.

[0015] Step S102, for the structured customs declaration behavior sequence, the long short-term memory network in deep learning is used to extract features of the customs declaration behavior within the time window, and the dynamic change trend of the customs declaration behavior is generated through multi-layer neural network training to obtain a preliminary behavior pattern representation.

[0016] The long short-term memory network is used to extract features within the time window for the structured customs declaration behavior sequence to obtain a preliminary feature set. The feature set is trained through a multi-layer neural network to generate a dynamic change trend of the customs declaration behavior and obtain trend data. For the trend data, if the change amplitude exceeds the preset threshold, it is marked as abnormal behavior to obtain an abnormal tag set. According to the abnormal tag set, the corresponding customs declaration behavior sequence within the time window is obtained to obtain the abnormal behavior sequence. The abnormal behavior sequence is extracted through sequence analysis technology to obtain the abnormal behavior pattern representation. The clustering algorithm is used to group the abnormal behavior pattern representation to obtain the classified behavior pattern set. Through the classified behavior pattern set, the time distribution law of the customs declaration behavior is determined to obtain the dynamic behavior distribution result.

[0017] In one possible implementation, when a long short-term memory network is used to extract features within a time window for a structured sequence of customs declarations, it can be considered as a method of capturing time dependencies. Long short-term memory networks are good at processing sequence data, retaining key information and ignoring irrelevant noise through forget gates and input gate mechanisms. For example, from a 7-day sequence of customs declarations, the network can identify the fluctuation pattern of the number of customs declarations and the declared amount per day, extract features such as "frequent customs declarations at the beginning of the week" or "concentrated declarations of specific commodity categories", and form a preliminary feature set. The advantage of this method is that it can extract the core pattern of dynamic changes from a continuous time window. Specifically, when training a feature set through a multi-layer neural network, the features can be input into a network structure composed of fully connected layers to gradually learn the deep laws of customs declaration behavior.

[0018] In one embodiment, assuming that the input features include the number of customs declarations per day and the average declared amount, the network may output a trend vector through multiple layers of transformation, indicating the change of customs declaration behavior over time. For example, after training, it is found that the number of customs declarations in a certain time window gradually decreases from 50 times per day to 20 times per day. This dynamic trend data provides a basic basis for subsequent abnormal judgment.

[0019] For example, for abnormal judgment of trend data, if the preset threshold is a change of more than 30%, each time window can be checked one by one. For example, if the average declared amount in a window drops from 100,000 yuan to 50,000 yuan, a decrease of 50%, exceeding the threshold, it will be marked as abnormal behavior. From another perspective, if the number of customs declarations at a window surges from 100 to 200 times, an increase of 100%, it will also be marked. This multi-angle threshold design can capture potential anomalies more comprehensively.

[0020] In one embodiment, when the corresponding customs declaration behavior sequence is obtained according to the abnormal mark set, the specific record can be traced back. For example, a certain abnormal window shows that the declared amount has dropped abnormally. After extraction, it is found that the record set involving a certain type of commodity is missing, which may be an underreporting or data error. This tracing method helps to identify the root cause of the problem. Preferably, when extracting patterns from abnormal behavior sequences through sequence analysis technology, frequent pattern mining can be used. For example, it is found that "frequent changes in a certain commodity code" appear multiple times in the abnormal sequence, forming an abnormal behavior pattern representation, which is convenient for subsequent classification.

[0021] It should be noted that when clustering algorithms are used to group abnormal behavior patterns, methods such as K-means can be used to group similar patterns into one category. For example, one category may be "abnormal fluctuations in declared amounts" and another category may be "abnormal concentration of customs declaration times." In a specific grouping, suppose there are 10 abnormal patterns, which are clustered into 3 groups, each reflecting different types of abnormal behavior characteristics. This grouping helps to more clearly identify behavioral patterns.

[0022] It is understandable that when determining the time distribution law through the classified behavior pattern set, the time concentration of abnormal behavior can be analyzed. For example, it is found that a certain type of anomaly occurs more at the end of the month, which may be related to the settlement cycle. Another type of anomaly is concentrated before the holidays, which may be related to the logistics peak. This distribution result provides a basis for optimizing the supervision strategy. For example, for the end-of-month anomaly, the review can be strengthened in advance. This kind of dynamic behavior distribution mining not only improves the accuracy of anomaly detection, but also provides a practical reference for business decision-making.

[0023] Step S103, based on the preliminary behavior pattern representation, the normal customs declaration behaviors in the historical data are grouped by cluster analysis method, and the mean and variance of the behavior characteristics in each time window are calculated in combination with the statistical distribution characteristics to determine the parameters of the dynamic benchmark model.

[0024] Process the customs declaration behaviors in the historical data through cluster analysis to obtain the grouping results. Extract the behavioral features in each time window from the grouping results to obtain the feature set. Calculate the mean and variance for the feature set to determine the statistical distribution characteristics. Build a dynamic benchmark based on the statistical distribution characteristics to obtain the benchmark description. If the benchmark description exceeds the preset threshold, adjust the grouping results to obtain an updated grouping. Optimize the benchmark model through the updated grouping and determine the model parameters. Use the model parameters to verify the behavioral features in the time window and determine the abnormal state.

[0025] For example, in the historical customs declaration data, key behavioral features are first extracted, such as the declared amount, frequency of commodity types, and declaration time interval, to construct a feature matrix. Assuming that a company has 5,000 customs declaration records in the past 6 months, each record contains the declared amount (unit: 10,000 yuan) and the number of declarations per day, the data is divided into 3 groups by the K-means clustering algorithm. The initial cluster centers are set to [10,2], [50,5], and [200,10]. After 10 iterations, the final cluster centers converge to [15.3±2.1,1.8±0.6], [48.7±5.3,4.9±1.2], and [195.6±18.7,9.7±2.4], corresponding to normal, medium, and high-frequency behavior patterns. Then, with 1 hour as the time window, the mean and variance of the declared amount in each window are calculated. For example, the mean declared amount in window W1 is 325,000 yuan, the variance is 4.7, and the mean declared amount in window W2 is 281,000 yuan, the variance is 3.9. The sliding window method is used to update the benchmark model parameters, and the attenuation factor α is set to 0.9, and the new parameter = α×old parameter+(1-α)×current window statistic. When it is detected that the mean declared amount in a certain window exceeds the historical mean ±3 times the variance range (such as the mean of window W3 suddenly increases to 1.2 million yuan, while the historical benchmark mean is 350,000±60,000 yuan), an abnormal warning is triggered. At the same time, the Mahalanobis distance is combined to calculate the degree of deviation between the current behavior characteristics and each cluster center, and the threshold is set to 5.99 (95% quantile of the chi-square distribution). When the distance exceeds the threshold, it is determined to be abnormal behavior.

[0026] Step S104, based on the parameters of the dynamic benchmark model, the Euclidean distance between the current customs declaration behavior and the dynamic benchmark is calculated to quantify the abnormal deviation. If the abnormal deviation exceeds a preset threshold, it is marked as a potential abnormal behavior to obtain a preliminary screening result of abnormal detection.

[0027] By obtaining the parameter basis of the current customs declaration behavior data and the dynamic benchmark model, the Euclidean distance between the two is calculated to obtain the distance deviation value. According to the distance deviation value, the abnormal deviation is calculated to determine the degree of deviation. If the abnormal deviation exceeds the preset threshold, the potential abnormal behavior is determined through marking judgment. For the behavior marked as potentially abnormal, its quantitative characteristics are obtained to obtain quantitative description data. By comparing the quantitative description data with the dynamic benchmark, the persistence of the abnormal deviation is judged to obtain the deviation trend. Based on the deviation trend, the random forest algorithm is used to determine the classification result of the potential abnormal behavior. The classification result is obtained, combined with the initial screening result, and the final abnormal detection output is obtained through logical verification.

[0028] Exemplarily, based on the parameters of the dynamic benchmark model, the dynamic benchmark model is first constructed through historical customs declaration data. It is assumed that the benchmark model contains parameters such as the customs declaration amount, commodity type, and customs declaration time, where the average value of the customs declaration amount is 10,000 yuan, the standard deviation is 2,000 yuan, the commodity types are 5 categories, and the customs declaration time is concentrated between 9 am and 11 am. Next, the Euclidean distance between the current customs declaration behavior and the dynamic benchmark is calculated. For example, if the customs declaration amount is 15,000 yuan, the commodity types are 7 categories, and the customs declaration time is 2 pm, the distance value is calculated by the Euclidean distance formula √[(15000-10000)² / 2000² + (7-5)² +(14-10)²]≈3.16. Then, the threshold of abnormal deviation is set to 2.5. If the calculated Euclidean distance exceeds the threshold, it is marked as a potential abnormal behavior.

[0029] For example, the calculated value of 3.16 exceeds 2.5, so the declaration is marked as a potential anomaly. Finally, the initial screening results are output through the automated system for further analysis or manual review to ensure the accuracy and timeliness of anomaly detection.

[0030] Step S105, based on the initial screening results of anomaly detection, a random forest algorithm is used to classify the multi-dimensional features of potential abnormal behaviors, and a classifier is trained in combination with known violation cases in historical data to determine the type of abnormal behavior and obtain a graded anomaly classification result.

[0031] Through the initial screening results, a data set of potential abnormal behaviors is obtained, and the multi-dimensional features are separated by feature extraction methods to obtain a feature set. For the feature set, the random forest algorithm is applied for classification training, and the trained classifier is obtained by combining the violation cases in the historical data. The trained classifier is used to judge the type of abnormal behavior and obtain a preliminary classification label. According to the preliminary classification label, the violation cases related to the label in the historical data are obtained to determine whether there is a matching pattern and obtain the matching result. If the matching result shows consistency, the matching result is fused with the multi-dimensional features through data combination technology to obtain an enhanced feature set. For the enhanced feature set, the logistic regression algorithm is used to classify the abnormal behavior and obtain the final classification result. Through the final classification result, the priority sequence of abnormal behavior is determined, and the sorted abnormal classification list is output.

[0032] Exemplarily, in the initial screening stage of anomaly detection, the original data is first cleaned and standardized by the data preprocessing module. For example, the timestamp is converted into a unified format, missing values are removed, and numerical features are normalized to between 0 and 1. Then, the random forest algorithm is used to classify multi-dimensional features. Assuming the input data contains 100 features, the random forest model sets 100 decision trees, with the maximum depth of each tree being 10. Using the Gini coefficient as the splitting criterion, the features are sorted by importance, and the top 20 key features are selected. Combining with the known violation cases in historical data, such as 5000 violation behavior data recorded in the past year, these data are divided into a training set and a test set in a ratio of 8:2. The training set is used to build a classifier, and the test set is used to evaluate the model performance. During the training process, the cross-validation method is adopted, and 5-fold cross-validation is set to ensure the generalization ability of the model. After training, the model predicts the test set to obtain the classification results of abnormal behaviors. For example, the abnormal behaviors are divided into three levels: high risk, medium risk, and low risk, corresponding to prediction results with confidence levels greater than 0.9, between 0.7 and 0.9, and less than 0.7 respectively. Finally, the model performance is evaluated through the confusion matrix and the ROC curve. Assuming the accuracy of the model on the test set reaches 95% and the AUC value is 0.98, it indicates that the model has high classification ability. The whole process is implemented through an automated script to ensure efficiency and repeatability.

[0033] Step S106, extract high-risk anomaly types from the classified anomaly classification results, continuously monitor the current customs declaration documents through the real-time data stream processing framework. If a high-risk anomaly type is detected, trigger the corresponding warning signal to obtain the real-time warning trigger status.

[0034] Extract high-risk anomaly types through the classification results to obtain an anomaly classification set. Use real-time data stream technology to process the customs declaration documents to obtain continuous monitoring data. If a high-risk anomaly type appears in the continuous monitoring data, judge the detection anomaly status through the processing framework. Generate a warning signal according to the detection anomaly status to determine the trigger condition. Update the real-time monitoring process through the trigger condition to obtain the trigger status data. Adjust the data processing logic for the trigger status data to obtain the optimized monitoring results. Extract new high-risk anomaly types from the optimized monitoring results to update the anomaly classification set.

[0035] Exemplarily, in the classified anomaly classification results, high-risk anomaly types usually include amount anomalies, inconsistent commodity categories, abnormal declaration times, etc. in the customs declaration documents. Through the real-time data stream processing framework, such as Apache Flink, the current customs declaration documents can be continuously monitored.

[0036] For example, the threshold for abnormal amount is set to a single customs declaration amount exceeding RMB 1 million, the criterion for mismatch of commodity category is that the declared commodity is inconsistent with the actual commodity HS code, and the criterion for abnormal declaration time is that the deviation between the declaration form submission time and the historical average submission time exceeds 30 minutes. When the system detects that the amount of a customs declaration is RMB 1.2 million, the commodity HS code is inconsistent with the declaration, and the submission time is 40 minutes later than the historical average time, the system will immediately trigger a high-risk abnormal warning signal. The warning signal is sent to the warning processing module through the message queue. The warning processing module realizes the real-time warning trigger status according to the preset rules, such as sending SMS notifications to relevant personnel or automatically generating abnormal processing work orders. The entire process ensures the timeliness and accuracy of abnormal detection and warning through the efficient computing and low latency characteristics of the real-time data stream processing framework.

[0037] In other embodiments, as a further improvement, after step S106, the following may be further included: Step S107, according to the real-time warning trigger status, the warning signal is distributed to different processing channels according to priority using message queue technology, and the signal is processed in layers using pre-established multi-level warning rules to generate multi-level warning outputs for different abnormal types.

[0038] The trigger status of real-time warning is obtained through the monitoring system, and the trigger status is sorted by message queue technology to determine the priority distribution sequence. The signal features are extracted from the sorted trigger status, and the preset queue technology is used to distribute it to the corresponding processing channel to obtain the channel signal set. For the channel signal set, the preset multi-level rules are used for matching. If the signal features meet the rules of a certain level, it is marked as the corresponding abnormal type. According to the marked abnormal type, the decision tree algorithm is used to process the signal in layers to determine the warning level of each layer. The warning level after layered processing is obtained, and the abnormal type and level are integrated through the signal processing module to generate a structured warning signal. The key fields are extracted from the structured warning signal, and the mapping table technology is used to generate multi-level warning outputs for different abnormal types. The multi-level warning output is transmitted to the downstream system through the output interface to complete the signal processing process.

[0039] For example, RabbitMQ can be used as a message queue tool to sort the trigger status by timestamp and abnormal severity. For example, a customs declaration triggers an amount abnormality and a time abnormality. The system determines that the amount abnormality has a higher priority than the time abnormality according to the preset rules, and the amount abnormality is ranked first after sorting. This sorting method ensures that high-risk signals are processed first. Signal features are extracted from the sorted trigger state. In a possible implementation method, the signal features include fields such as abnormality type, triggering time, and amount involved. The preset queue technology is used to distribute the signal to the corresponding processing channel. For example, the amount abnormality signal enters the financial review channel, and the time abnormality enters the timeliness monitoring channel, and the channel signal set is generated. This distribution mechanism ensures the pertinence of signal processing. For the channel signal set, the preset multi-level rules are matched. Specifically, the amount abnormality rule can be set to mark a single transaction exceeding 800,000 yuan as medium risk and a single transaction exceeding 1.5 million yuan as high risk. If a signal feature shows that the customs declaration amount is 1.6 million yuan, it is marked as a high-risk amount abnormality. The multi-level rule design makes the abnormality classification more refined. According to the marked abnormality type, the decision tree algorithm is used for hierarchical processing. In one embodiment, the first layer of the decision tree determines whether the anomaly involves an amount, the second layer determines whether the amount exceeds the standard, and the third layer combines historical data to assess risk trends. For example, if the amount of a customs declaration is 1.6 million yuan and has frequently exceeded the standard recently, the warning level is set to level three. This hierarchical logic improves the comprehensiveness of the judgment. After obtaining the warning level after hierarchical processing, the signal processing module integrates the anomaly type and level to generate a structured warning signal. It can be understood that the structured signal contains fields such as "anomaly type: amount exceeds the standard", "level: level three", and "trigger time: 2025-04-08 14:30". This format is easy for downstream systems to identify. Key fields are extracted from the structured warning signal, and the mapping table technology generates multi-level warning outputs. Preferably, the mapping table maps the amount exceeding the standard to "red alarm" and the time anomaly to "yellow reminder". For example, a customs declaration of 1.6 million yuan generates a "red alarm" output. This intuitive mapping facilitates rapid response. The multi-level warning output is transmitted to the downstream system through the output interface, for example, the "red alarm" is pushed to the mobile phone of the manager, and an exception handling work order is generated at the same time. This method ensures that the warning signal reaches the relevant parties in a timely manner. It should be noted that the combination of message queue sorting and decision tree layering can effectively improve the priority and accuracy of exception handling. In one embodiment, if a customs declaration triggers multiple exceptions at the same time, the system gives priority to processing the signal of excessive amount to avoid waste of resources. For example, for a customs declaration with an amount of 1.6 million yuan and a time delay of 40 minutes, the amount issue is given priority. This design optimizes processing efficiency. In one possible implementation, parallel processing of channel signal sets can also shorten response time. For example, the financial channel and the timeliness channel operate at the same time, processing amount and time exceptions respectively, without interfering with each other. This parallel mechanism ensures real-time performance. Specifically, the diversity of multi-level warning outputs supports the needs of different scenarios.For example, high-risk signals trigger SMS notifications, while medium-risk signals only record logs. This flexibility makes the system more adaptable.

[0040] In other embodiments, as a further improvement, after step S107, the following may be further included: Step S108, after obtaining the multi-level warning output, the warning result is compared with the behavior pattern in the historical data through the feedback mechanism, and the parameters of the deep learning model are dynamically adjusted using the online learning method to obtain an updated dynamic benchmark model.

[0041] After obtaining the multi-level warning output, the behavior pattern is extracted from the historical data through the feedback mechanism to obtain the preliminary comparison result. Based on the difference between the comparison result and the behavior pattern, the adjustment direction of the deep learning model is calculated by the online learning method to obtain the parameter update value. If the parameter update value exceeds the preset threshold, the model parameters of the deep learning model are dynamically adjusted to obtain the adjusted temporary model. The multi-level warning output is reprocessed according to the adjusted temporary model to obtain the updated warning result. The updated warning result is compared with the behavior pattern in the historical data for the second time through the feedback mechanism to obtain the comparison deviation value. The temporary model is optimized according to the comparison deviation value by the online learning method to obtain the updated dynamic baseline model. The subsequent input data is processed by the dynamic baseline model to obtain the real-time warning output.

[0042] Exemplarily, after obtaining the multi-level early warning output, the system first extracts the abnormal records of the user login behavior in the past 30 days through the time series database, such as features like the number of single-day logins exceeding the threshold of 50 times or the sudden change of the login region, and calculates the Z-score value of the behavior frequency at intervals of 5 minutes using the sliding window algorithm. When a real-time early warning is triggered, the feedback mechanism will perform a similarity match between the current early warning event and 2000 typical patterns in the historical behavior library, calculate the morphological distance of the time series using the improved DTW algorithm (Dynamic Time Warping), and determine that it is a known attack pattern if the similarity threshold is above 0.85. For new patterns that do not match, the online learning module will initiate incremental parameter updates, adopt the stochastic gradient descent algorithm with momentum term (β = 0.9), adjust the weights of the 128-dimensional hidden layer of the LSTM model at a learning rate of 0.001, and at the same time analyze the difference in the output distributions of the new and old models through KL divergence. When the difference value is greater than 0.3, it triggers the adaptive adjustment of the model structure. The updated dynamic benchmark model will be deployed to the inference service cluster in real time. When processing 5000 real-time data streams per second, the confidence of the model is evaluated through Monte Carlo sampling. When the standard deviation of 100 consecutive predictions exceeds the preset threshold of 0.05, it will automatically roll back to the previous stable version. Throughout the process, the feature engineering link uses the automatic binning technique to discretize the numerical features after IP address conversion into 20 intervals, and filters out the key features with a P-value less than 0.01 through the chi-square test to participate in model training.

[0043] In other embodiments, as a further improvement, after step S108, it may further include: Step S109, for the updated dynamic benchmark model, perform anomaly detection and early warning processing on the new round of customs declaration document data through a loop iteration method, adopt the sliding window technique to continuously track the behavior change trend in the trade scenario, and obtain continuously optimized anomaly early warning capabilities.

[0044] Through loop iteration processing of the customs declaration document data, obtain the initial anomaly detection result. Use the sliding window technique to analyze the detection result and determine the behavior change trend in the trade scenario. Update the dynamic benchmark model according to the behavior change trend to obtain the adjusted benchmark parameters. For the adjusted benchmark parameters, determine whether the customs declaration document data exceeds the preset threshold to obtain an anomaly mark. Compare the anomaly mark with the trend tracking result to determine whether the early warning capability needs to be optimized. If the optimization requirement is established, adjust the sliding window parameters through loop iteration to obtain optimized anomaly detection capabilities. Process the new round of customs declaration document data with the optimized detection capabilities to obtain the updated early warning output.

[0045] For example, in the process of updating the dynamic benchmark model, the sliding window technology is first used to set a 30-day window period, and the customs declaration document data is processed in a daily incremental update manner. The window sliding step is 1 day to ensure real-time performance. By calculating the Z-score value of the declared amount in the window, the threshold is set to ±2.5 standard deviations, and the primary warning is triggered when the Z-score of the declared price of a batch of goods reaches 3.2. Subsequently, the isolation forest algorithm is used to jointly analyze multidimensional features (including cargo weight, category, origin, etc.), and the number of sub-sampling is set to 256. When the sample anomaly score exceeds 0.65, a secondary warning is generated. For the continuously triggered warnings, the time series ARIMA model is introduced for trend prediction, and the parameters are set to (p=2, d=1, q=1). When the prediction deviation exceeds 15%, the manual review process is started. After each warning processing, the system automatically feeds back the confirmed abnormal case features (such as the unit price deviation >40% common in the false price mode) to the benchmark model, and updates the weight matrix with a learning rate of 0.01 through the stochastic gradient descent algorithm. For seasonal trade fluctuations, the STL decomposition algorithm is used to split the data into trend, season and residual terms. When the residual term exceeds the historical mean of the same period by 2 times, the seasonal anomaly mark is triggered. In the iterative optimization phase, the current model version is locked when the accuracy of the confusion matrix calculation reaches 92.3%, and the model snapshots of the previous 5 versions are retained for rollback. Throughout the process, the Holt-Winters exponential smoothing algorithm continuously corrects the warning sensitivity with parameter settings of α=0.2, β=0.1, and γ=0.3 to ensure that the response delay to emerging smuggling methods does not exceed 8 hours.

[0046] The above is only a preferred embodiment of the present invention. It should be pointed out that ordinary technicians in this technical field can make several improvements and supplements without departing from the principle of the present invention. These improvements and supplements should also be regarded as the scope of protection of the present invention.

Claims

1. A customs clearance risk detection method based on customs declaration document identification, characterized in that: The method comprises: Historical data is obtained from customs declaration documents, and noise data is removed through data cleaning and preprocessing. The continuous historical data is divided into multiple sub-datasets according to the time window using the time series segmentation method to obtain a structured customs declaration behavior sequence; for the structured customs declaration behavior sequence, the long short-term memory network in deep learning is used to extract the features of the customs declaration behavior within the time window, and the dynamic change trend of the customs declaration behavior is generated through multi-layer neural network training to obtain a preliminary behavior pattern representation; based on the preliminary behavior pattern representation, the normal customs declaration behaviors in the historical data are grouped through the clustering analysis method, and the mean and variance of the behavior characteristics in each time window are calculated in combination with the statistical distribution characteristics to determine the parameters of the dynamic benchmark model; in the dynamic Based on the parameters of the dynamic benchmark model, the Euclidean distance between the current customs declaration behavior and the dynamic benchmark is calculated to quantify the abnormal deviation. If the abnormal deviation exceeds the preset threshold, it is marked as a potential abnormal behavior to obtain the initial screening result of anomaly detection; for the initial screening result of anomaly detection, the random forest algorithm is used to classify the multi-dimensional features of potential abnormal behaviors, and the classifier is trained in combination with known violation cases in historical data to determine the type of abnormal behavior and obtain a graded abnormal classification result; high-risk abnormal types are extracted from the graded abnormal classification results, and the current customs declaration documents are continuously monitored through the real-time data stream processing framework. If a high-risk abnormal type is detected, the corresponding warning signal is triggered to obtain a real-time warning trigger status.

2. The method according to claim 1, characterized in that The historical data is obtained from the customs declaration documents, and the noise data is removed through data cleaning and preprocessing. The continuous historical data is divided into multiple sub-data sets according to the time window by using the time series segmentation method to obtain a structured customs declaration behavior sequence, including: Obtain historical data from customs declaration documents, extract the complete data set through scanning and parsing technology, and obtain the initial data set; The initial data set is preprocessed by data cleaning technology, and the outlier detection method is used to remove noise data to obtain the cleaned data set; For the cleaned data set, time series analysis methods are used to identify the characteristics of continuous data and determine the time series pattern; The time series pattern is segmented by time window partitioning technology to obtain multiple sub-data sets and obtain the segmented data set; According to the segmented data set, the characteristics of the customs declaration behavior sequence are extracted, and the behavior pattern is classified using a clustering algorithm to obtain the classified behavior sequence; Generate structured data format for classified behavior sequences, integrate data through field mapping technology, and obtain structured behavior data sets; If there are missing values ​​in the structured behavior dataset, they are supplemented through interpolation methods to obtain the final behavior sequence dataset.

3. The method according to claim 1, characterized in that For the structured customs declaration behavior sequence, the long short-term memory network in deep learning is used to extract the features of the customs declaration behavior within the time window, and the dynamic change trend of the customs declaration behavior is generated through multi-layer neural network training to obtain a preliminary behavior pattern representation, including: The long short-term memory network is used to extract features within the time window of the structured customs declaration behavior sequence to obtain a preliminary feature set. The feature set is trained through a multi-layer neural network to generate the dynamic change trend of customs declaration behavior and obtain trend data; For trend data, if the change range exceeds the preset threshold, it is marked as abnormal behavior to obtain an abnormal mark set; According to the abnormal mark set, the corresponding customs declaration behavior sequence in the time window is obtained to obtain the abnormal behavior sequence; The abnormal behavior sequence is pattern extracted through sequence analysis technology to obtain the abnormal behavior pattern representation; Clustering algorithms are used to group abnormal behavior pattern representations to obtain a set of classified behavior patterns; Through the classified behavior pattern set, the time distribution law of customs declaration behavior is determined and the dynamic behavior distribution result is obtained.

4. The method according to claim 1, characterized in that: According to the preliminary behavior pattern representation, the normal customs declaration behaviors in the historical data are grouped by cluster analysis method, and the mean and variance of the behavior characteristics in each time window are calculated in combination with the statistical distribution characteristics to determine the parameters of the dynamic benchmark model, including: Process the customs declaration behaviors in historical data through cluster analysis to obtain grouping results; Extract the behavior features in each time window from the grouping results to obtain a feature set; Calculate the mean and variance for the feature set and determine the statistical distribution characteristics; Build a dynamic benchmark based on statistical distribution characteristics and obtain a benchmark description; If the benchmark description exceeds the preset threshold, the grouping result is adjusted to obtain an updated grouping; Optimize the benchmark model through the updated grouping and determine the model parameters; Model parameters are used to verify the behavioral characteristics within the time window and determine the abnormal state.

5. The method according to claim 1, characterized in that Based on the parameters of the dynamic benchmark model, the Euclidean distance between the current customs declaration behavior and the dynamic benchmark is calculated to quantify the abnormal deviation. If the abnormal deviation exceeds the preset threshold, it is marked as a potential abnormal behavior, and the initial screening results of the abnormal detection are obtained, including: By obtaining the current customs declaration behavior data and the parameter basis of the dynamic benchmark model, the Euclidean distance between the two is calculated to obtain the distance deviation value; According to the distance deviation value, calculate the abnormal deviation and determine the degree of deviation; If the abnormal deviation exceeds the preset threshold, the potential abnormal behavior is determined through marking judgment; For behaviors marked as potentially abnormal, obtain the quantitative characteristics of the behaviors and obtain quantitative description data; By comparing quantitative description data with dynamic benchmarks, we can determine the persistence of abnormal deviations and obtain deviation trends; By deviating from the trend, the random forest algorithm is used to determine the classification results of potential abnormal behaviors; Get the classification results, combine them with the initial screening results, and perform logical verification to get the final anomaly detection output.

6. The method according to claim 1, characterized in that The initial screening results for anomaly detection use the random forest algorithm to classify the multi-dimensional features of potential abnormal behaviors, and train the classifier in combination with known violation cases in historical data to determine the type of abnormal behavior and obtain graded anomaly classification results, including: A data set of potential abnormal behaviors is obtained through the initial screening results, and multi-dimensional features are separated using feature extraction methods to obtain a feature set; For the feature set, the random forest algorithm is applied for classification training, and the trained classifier is obtained by combining the violation cases in the historical data; Use the trained classifier to determine the type of abnormal behavior and obtain a preliminary classification label; Based on the preliminary classification labels, obtain the violation cases related to the labels in the historical data, determine whether there is a matching pattern, and obtain the matching results; If the matching results show consistency, the matching results are fused with the multi-dimensional features through data combination technology to obtain an enhanced feature set; Based on the enhanced feature set, a logistic regression algorithm is used to classify abnormal behaviors and obtain the final classification results; Through the final classification results, the priority sequence of abnormal behaviors is determined, and a sorted abnormal classification list is output.

7. The method according to claim 1, characterized in that The high-risk anomaly type is extracted from the graded anomaly classification results, and the current customs declaration documents are continuously monitored through the real-time data stream processing framework. If a high-risk anomaly type is detected, the corresponding warning signal is triggered to obtain a real-time warning trigger state, including: Extract high-risk anomaly types through classification results and obtain anomaly classification set; Use real-time data stream technology to process customs declaration documents and obtain continuous monitoring data; If a high-risk anomaly type appears in the continuous monitoring data, the abnormal state of the detection is determined through the processing framework; Generate early warning signals based on detected abnormal conditions and determine trigger conditions; Update the real-time monitoring process through trigger conditions and obtain trigger status data; Adjust the data processing logic according to the trigger status data to obtain optimized monitoring results; Extract new high-risk anomaly types from the optimized monitoring results and update the anomaly classification set.

8. The method according to claim 1, characterized in that The method further comprises: According to the real-time warning trigger status, the message queue technology is used to distribute the warning signal to different processing channels according to priority. The signal is processed in layers through the pre-established multi-level warning rules to generate multi-level warning outputs for different abnormal types.

9. The method according to claim 8, characterized in that The method further comprises: After obtaining the multi-level warning output, the warning results are compared with the behavior patterns in the historical data through the feedback mechanism, and the parameters of the deep learning model are dynamically adjusted using the online learning method to obtain the updated dynamic benchmark model.

10. The method according to claim 9, characterized in that The method further comprises: For the updated dynamic benchmark model, anomaly detection and early warning processing are performed on the new round of customs declaration document data through a cyclic iterative method, and sliding window technology is used to continuously track the behavioral change trends in trade scenarios to obtain continuously optimized anomaly early warning capabilities.

Citation Information

Patent Citations

  • Customs import and export commodity risk identification method based on declaration quality assessment

    CN115617979A

  • Enterprise big data analysis system based on artificial intelligence

    CN117993737A

  • Artificial intelligence (AI) driven estimated shipment delivery time advisor

    US20220237432A1

Cited By

  • Zone area energy storage operation dynamic early warning control method and system

    CN120613765A

  • A transformer area energy storage operation dynamic early warning control method and system

    CN120613765B

  • Experimental data processing method and device, AI analysis module and computer equipment

    CN120705604A

  • Cross-border electronic customs declaration inspection early warning method and system based on dynamic risk learning

    CN120746425A