Intelligent port-based international trade information automatic verification and risk early warning method

CN122243518APending Publication Date: 2026-06-19CHANGSHA JIAOTONG LOGISTICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGSHA JIAOTONG LOGISTICS CO LTD
Filing Date
2026-03-04
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing international trade information verification and risk warning technologies suffer from insufficient data integration capabilities, limited risk identification accuracy, and poor dynamic adaptability, making it difficult to meet the needs of smart port construction for efficient, accurate, and intelligent supervision.

Method used

By establishing a unified data dictionary to achieve standardized integration of multi-source heterogeneous document data, a rule engine and weighted scoring mechanism are used for intelligent verification across document fields. A spatiotemporal graph of trade behavior is constructed and an improved random walk algorithm is used to mine associated risk characteristics. By combining historical case similarity matching and regional risk benchmark adjustment, risk thresholds are dynamically adjusted to trigger adaptive early warnings.

Benefits of technology

It has significantly improved the intelligence level of port supervision, the accuracy of risk identification, and the efficiency of regulatory resource allocation, effectively preventing illegal and irregular activities such as smuggling, false declarations, and tax evasion, and safeguarding national economic security and trade order.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122243518A_ABST
    Figure CN122243518A_ABST
Patent Text Reader

Abstract

This invention discloses an automatic verification and risk warning method for international trade information based on smart ports. It collects multi-source document data, including customs declarations, bills of lading, transportation routes, payment vouchers, and certificates of origin, and converts them into a standardized format through field mapping rules. A cross-document verification rule base is established, and rule matching is performed on key fields such as invoice amount and declared price, and bill of lading quantity and packing list, using a weighted scoring mechanism to calculate consistency scores. A spatiotemporal graph of trade behavior is constructed, and risk characteristics are statistically analyzed using an improved random walk algorithm, outputting a dynamic risk feature matrix. This matrix is ​​matched with a historical risk case database, and the probability of risks such as smuggling, false declarations, and tax evasion is calculated in conjunction with regional risk benchmarks. Risk thresholds are dynamically adjusted based on port clearance flow, triggering adaptive warnings and generating warning information including warning level, risk type, and inspection measures, thereby improving the intelligence level of port supervision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of trade supervision, and in particular to a method for automatic verification and risk warning of international trade information based on smart ports. Background Technology

[0002] With the deepening of global trade liberalization and regional economic integration, the scale of international trade continues to expand, and cross-border cargo flows are becoming increasingly frequent. Against this backdrop, customs and other port regulatory departments face unprecedented regulatory pressure. They must ensure trade facilitation and improve customs clearance efficiency while effectively preventing smuggling, misdeclaration, tax evasion, and other illegal and irregular activities to safeguard national economic security and public interests. Traditional manual document review and random inspection models are no longer sufficient to meet the regulatory needs of massive amounts of trade data, necessitating the use of information technology and intelligent methods to improve the accuracy and effectiveness of port supervision.

[0003] Current international trade activities involve a wide variety of documents, including customs declarations, invoices, bills of lading, packing lists, certificates of origin, licenses, and dozens of other types of documents. These documents are scattered across different information systems such as customs declaration systems, electronic port platforms, logistics tracking systems, and bank settlement systems, with varying data formats, field definitions, and coding standards, creating serious information silos. Meanwhile, to evade supervision, trade entities often resort to false declarations, under-declaration, over-declaration, and document forgery to commit illegal acts, and their methods are constantly evolving, exhibiting characteristics of high concealment, advanced technology, and organized crime. Traditional regulatory models mainly rely on the experience of customs officers for manual review and on-site inspections, resulting in significant problems such as low efficiency, inconsistent standards, and high under-declaration rates, making it difficult to effectively address complex and ever-changing trade risks.

[0004] In existing technologies, some scholars and institutions have attempted to use information technology to assist trade supervision. For example, rule-based document review systems use preset logical rules to verify the consistency of document data. However, this method can only identify explicit data contradictions and cannot detect hidden counterfeiting patterns and associated risks. Furthermore, the rule base has high maintenance costs and is difficult to adapt to the dynamic evolution of risk characteristics. Some studies use machine learning methods to build risk classification models and identify high-risk goods by training on historical case data. However, these methods often analyze each shipment as an isolated sample, ignoring the complex relationships between trading entities and failing to reveal group risk characteristics such as organized smuggling and fraudulent trade networks. In addition, most existing methods use static risk thresholds and fail to dynamically adjust warning standards based on real-time factors such as port clearance volume and inspection resource availability. This leads to higher underreporting rates during peak periods or excessively high false alarm rates during off-peak periods, making it difficult to optimize the allocation of regulatory resources.

[0005] In cross-document verification, existing systems primarily employ precise field-level matching or simple numerical comparisons, lacking a deep understanding of business rules such as commodity attributes, trade practices, and exchange rate fluctuations, which easily leads to numerous false alarms. For example, in verifying the invoice amount against the declared price on the customs declaration, existing methods typically set fixed deviation thresholds but fail to consider the price fluctuation characteristics of different commodity categories, resulting in insufficient accuracy for verifying price-sensitive commodities such as bulk goods. Regarding risk feature extraction, existing methods mainly rely on the static attribute features of individual shipments, such as commodity codes, declared prices, and trading party names, failing to fully explore the spatiotemporal correlation and network topology features of trade activities, and thus cannot effectively identify complex illegal activities that evade supervision through multi-layered subcontracting, fictitious transaction chains, etc.

[0006] In summary, existing international trade information verification and risk warning technologies suffer from insufficient data integration capabilities, limited risk identification accuracy, and poor dynamic adaptability, making it difficult to meet the demands of smart port construction for efficient, precise, and intelligent supervision. Therefore, there is an urgent need to develop a comprehensive solution that integrates multi-source heterogeneous data integration, cross-document intelligent verification, spatiotemporal graph correlation analysis, and dynamic risk assessment technologies to improve the intelligence level and risk prevention and control capabilities of port supervision. Summary of the Invention

[0007] In view of this, the present invention provides an automatic verification and risk warning method for international trade information based on smart ports. The purpose is to achieve standardized integration of multi-source heterogeneous document data by establishing a unified data dictionary, to achieve intelligent verification across document fields by using a rule engine and weighted scoring mechanism, to construct a spatiotemporal graph of trade behavior and to use an improved random walk algorithm to mine associated risk features, to achieve accurate risk type identification by combining historical case similarity matching and regional risk benchmark adjustment, and to dynamically adjust risk thresholds according to the real-time operation status of the port to trigger adaptive warnings. This comprehensively improves the intelligence level of port supervision, the accuracy of risk identification, and the efficiency of regulatory resource allocation, effectively prevents illegal and irregular activities such as smuggling, false declarations, and tax evasion, and safeguards national economic security and trade order.

[0008] To achieve the above objectives, this invention provides a method for automatic verification and risk warning of international trade information based on smart ports, comprising the following steps: A1: Collect multi-source document data from international trade activities. The multi-source document data includes customs declaration data from the customs declaration system, bill of lading data from the electronic port platform, transportation route data from the logistics tracking system, payment voucher data from the bank settlement system, and certificate of origin data from third-party certification bodies. Establish a unified data dictionary, and convert data fields from different sources into standardized formats through preset field mapping rules to construct a structured trade information set containing commodity codes, declared prices, declared quantities, declared weights, trader information, transportation route information, and timestamp information. A2: Input the structured trade information set into the rule engine-based cross-document verification module. In the rule engine-based cross-document verification module, establish a cross-document verification rule base. The cross-document verification rule base defines matching rules, value range constraint rules, and time logic constraint rules for key fields. Perform rule matching operation on each pair of corresponding key fields in the structured trade information set. Use a weighted scoring mechanism to calculate the overall consistency score of the documents. Output a verification result vector containing inconsistency field identifiers, deviation values, and consistency scores. A3: Extract trade feature information from the verification result vector. The trade feature information includes information on the inconsistency field identifier set and the document consistency score. A spatiotemporal graph of trade behavior is constructed. In the spatiotemporal graph of trade behavior, trade entities are used as nodes and transaction relationships are used as edges. A time series feature vector is established for each trade entity node. A dual decay mechanism is used to weight the historical transaction features. An improved random walk algorithm is used to perform biased multi-hop traversal on the spatiotemporal graph of trade behavior. The risk feature indicators on the traversal path are counted, and a dynamic risk feature matrix containing node risk degree, path risk degree, network topology features and spatiotemporal decay coefficient is output. A4: Perform similarity matching between the dynamic risk feature matrix and the feature vectors in the historical risk case database, and select the top [cases] with the highest similarity. A historical risk case, regarding the aforementioned Historical risk cases are normalized and weighted according to similarity, the weighted votes of each risk type label are counted, the port area risk benchmark value is introduced as an adjustment factor, the probability value of each risk type is calculated, and a risk assessment result set including smuggling risk probability value, false declaration risk probability value, tax evasion risk probability value and prohibited and restricted items risk probability value is output. A5: Based on the current customs clearance volume, inspection resource availability, and historical early warning accuracy, the high-risk threshold, medium-risk threshold, and low-risk threshold are dynamically calculated. The probability values ​​of each risk type in the risk assessment result set are compared with their corresponding risk thresholds, and the highest risk level is taken as the early warning level. When the early warning level reaches medium or high risk, the corresponding risk warning is triggered, and risk warning information containing the warning level, main risk types, abnormal field list, and suggested inspection measures is generated.

[0009] As a further improvement of the present invention: Optionally, step A1 further includes: A101: Collect customs declaration data from the customs declaration system, bill of lading data and invoice data from the electronic port platform, transportation route data from the logistics tracking system, payment voucher data from the bank settlement system, and certificate of origin data from third-party certification bodies; A102: Establish a unified data dictionary, which defines standard field names, data types, and field mapping rules; A103: Based on the field mapping rules, convert data fields from different sources into a standardized format, fill missing fields with preset default values, perform unified unit conversion for numeric fields, and convert time fields into a unified timestamp format; A104: Constructing a Structured Trade Information Set The structured trade information set The form is ,in Indicates the total number of document records. Indicates the first One document record, It contains information in a standard format.

[0010] Optionally, step A2 further includes: A201: Establish a cross-document verification rule base The cross-document verification rule base Include Verification rules, denoted as Each verification rule Define matching rules between a pair of key fields for a structured trade information set. The first in Document Records Execute the first Verification rules The matching operation determines whether key field pairs meet the matching conditions; if they do, a consistency flag is set. If the condition is not met, a consistency flag is set. And record the deviation value. ; A202: For the verification of the invoice amount and the declared price on the customs declaration, calculate the relative deviation between the two, and the deviation value. The calculation formula is: ; in, This indicates the deviation between the invoice amount and the declared price on the customs declaration form. This indicates the invoice amount in the invoice data. This represents the declared price in the customs declaration data. This indicates the absolute value operation; A203: For verifying the quantity on the bill of lading and the quantity on the packing list, calculate the absolute deviation between the two, and the deviation value. The calculation formula is: ; in, This indicates the deviation between the quantity on the bill of lading and the quantity on the packing list. This indicates the number of bills of lading in the bill of lading data. This indicates the number of boxes recorded in the packing list data; A204: For each verification rule Set severity weights The severity weight The range of values ​​is And the sum of the severity weights of all verification rules satisfies Calculate the first Document Records Overall consistency score The calculation formula is: ; in, Indicates the first Document Records Overall consistency score, Indicates the first The severity weight of each verification rule. Indicates the first Document record in the first Consistency indicators under the verification rules; A205: Constructing the Verification Result Vector The verification result vector It includes a set of inconsistency field identifiers, the deviation values ​​of each inconsistency field, and the overall consistency score. .

[0011] This step utilizes a rule engine mechanism to systematize and refine cross-document consistency verification. Traditional document verification methods often rely on manual sampling, which suffers from low efficiency, narrow coverage, and strong subjectivity. In contrast, the rule engine in this step can automatically perform multi-dimensional cross-verification of all documents, ensuring that the key fields of every trade are rigorously checked.

[0012] This step establishes a verification rule base that includes matching rules, value range constraints, and time logic constraints, achieving comprehensive coverage of the logical relationships between documents. By assigning differentiated severity weights to different verification rules, the verification system can perform weighted scoring based on the risk level of different inconsistency types, thereby generating a more discriminative consistency score index. This weighted scoring mechanism can not only identify deviations in a single field but also comprehensively assess the cumulative effect of inconsistencies across multiple fields, providing a reliable data foundation for subsequent risk feature extraction.

[0013] Optionally, step A3 further includes: A301: From the verification result vector Extract information from the inconsistency field identifier set and document consistency score. Constructing a spatiotemporal map of trade behavior ,in Represents the set of trade entity nodes. This represents the set of edges representing transaction relationships, and the set of nodes representing trade entities. Each node in Establish time series feature vectors The time series feature vector Record the historical transaction characteristics of the trading entity, including the frequency of historical transactions. Fluctuation range of amount Diversity of commodity category distribution and the number of historical violations The fluctuation range of the amount By calculating the past of this node The standard deviation of transaction amount within a time period is obtained, and the diversity of commodity category distribution is also considered. This is obtained by calculating the ratio of the number of different product codes involved in the node to the total number of transactions; A302: Define the time decay factor and spatial decay factor The time decay factor is based on the time of the transaction. With current time The time intervals between nodes decrease exponentially, and the spatial decay factor is based on the graph distance between nodes. It decreases in a power-law manner, with a time decay factor. The calculation formula is: ; in, Indicates the time when it occurs The time decay factor of the transaction, This represents the time decay rate parameter. Indicates the current moment. Indicates the time when a historical transaction occurred. The base of the natural logarithm; spatial decay factor The calculation formula is: ; in, The graph distance is represented as Spatial decay factor over time Spatiotemporal graph representing trade behavior The shortest path length between two nodes is calculated using a breadth-first search algorithm. This represents the spatial decay index parameter; A303: Employing an improved random walk algorithm in the spatiotemporal graph of trade behavior Perform a traversal operation, starting from the trade entity node corresponding to the current transaction record to be verified. Starting from the current node, perform a multi-hop traversal with bias, during which the traversal begins from the current node. Transfer to adjacent nodes transition probability The calculation formula is: ; in, Indicates starting from the current node Transfer to adjacent nodes The transition probability, This represents the time decay factor of the edges between the current node and its neighboring nodes. This indicates the time when the transaction corresponding to that edge occurred. Indicates the starting node With neighboring nodes Spatial attenuation factor between Indicates the starting node With neighboring nodes Graph distance between them This represents the weight value of the edge. Calculated based on the transaction amount and frequency. Indicates the current node The set of all adjacent nodes, Indicates the current node's... Adjacent nodes Indicates the relationship between the current node and the first node. The edges between adjacent nodes correspond to the times when transactions occur. Indicates the starting node With the Graph distance between adjacent nodes Indicates the relationship between the current node and the first node. The weight values ​​of the edges between adjacent nodes; A304: During the traversal, record the... traversal path ,in This represents the maximum number of hops in the traversal path, and it calculates the risk characteristic indicators on that path, including the risk node density. Repeatability of abnormal transaction patterns and the complexity of the associated network Among them, the density of risk nodes By counting the number of historical violations on the path Greater than the preset threshold The percentage of nodes in the path relative to the total number of nodes is used to determine the repetition rate of abnormal transaction patterns. By extracting the feature vectors of transactions corresponding to each edge on the path, a clustering algorithm is used to identify clusters of abnormal transaction patterns. The proportion of transaction edges belonging to abnormal pattern clusters to the total number of edges on the path is used to obtain the correlation network complexity. The value is obtained by calculating the product of the average degree centrality of the nodes on the path and the average clustering coefficient of the path. The calculation of the... traversal path Path aggregation risk score The calculation formula is: ; in, Indicates the first The path aggregation risk score of the traversal path. The weighting coefficients representing the density of risk nodes. Indicates the first The density of risk nodes on the traversal path The weighting coefficients representing the repetition rate of abnormal transaction patterns. Indicates the first The repetition rate of abnormal transaction patterns on each traversal path Weight coefficients representing the complexity of the network. Indicates the first The complexity of the network of traversal paths, and the weight coefficients satisfy... ; Optionally, in step A304, the repeatability of abnormal transaction patterns... Specific methods for obtaining them include: A3041: Regarding the first traversal path Each transaction edge on Extract the feature vector of the transaction The feature vector This includes the transaction amount, the commodity code, the transaction time interval, the document consistency score, and the number of historical violations by both parties in the transaction; A3042: Path The set of feature vectors of all transaction edges As a sample set, This represents the total number of transaction edges in the path, and the K-means clustering algorithm is used to cluster the feature vector set. Perform clustering, with the number of clusters set to [number]. ; A3043: For each cluster Calculate the mean cosine similarity between the samples within this cluster and the transaction feature vectors in the historical normal transaction pattern database. If the mean cosine similarity Less than the preset normal pattern similarity threshold If so, then the cluster is marked as an anomalous pattern cluster; A3044: Statistical Path The number of transaction edges belonging to the abnormal pattern cluster Calculate the repeatability of abnormal transaction patterns The calculation formula is: ; in, Indicates the first The repetition rate of abnormal transaction patterns on each traversal path Representing a path The number of transaction edges belonging to the abnormal pattern cluster; A305: Yes The path aggregation risk scores of the traversal paths are weighted and combined, and the first path is the first path. Path fusion weight Based on path length and the centrality median of the nodes visited on the path The node centrality is determined jointly by calculating the degree centrality of a node, where degree centrality is defined as the number of a node's neighbors, and the median of the centrality is also considered. The median of the degree centrality values ​​of all nodes on the path, sorted by size, is calculated using the following formula: ; in, Indicates the first The fusion weight of the traversal paths Indicates the first The path length of the traversal path. Indicates the first The median of the centrality values ​​of all visited nodes along the traversal path. This indicates the total number of traversed paths. Indicates the first The path length of the traversal path. Indicates the first The median of the centrality values ​​of all visited nodes along the traversal path; A306: Computational Fusion Risk level of nodes after traversing the path Path risk and network topology characteristics The node risk level Based on the starting node Number of historical violations Historical transaction frequency The calculated network topology features This includes the degree centrality, proximity centrality, and betweenness centrality of the starting node; and the generation of a dynamic risk feature matrix. The dynamic risk feature matrix Includes node risk level Path risk Network topology characteristics and the spatiotemporal decay coefficient vector Among them, path risk The calculation formula is: ; in, This indicates the path risk level after weighted fusion. Indicates the first The fusion weight of the traversal paths Indicates the first The path aggregation risk score of the traversal path. This indicates the total number of traversal paths.

[0014] This step utilizes a spatiotemporal decay graph mechanism to deeply mine and accurately capture the dynamic correlation characteristics of trade activities. Traditional risk assessment methods often focus only on the static characteristics of individual transactions, neglecting the relationships between trading entities and historical behavioral patterns, making it difficult to identify systemic violations committed through interconnected networks. This step constructs a spatiotemporal graph of trade activities, mapping trading entities and their transaction relationships into a graph structure, providing a systematic data organization form for correlation risk analysis.

[0015] This step introduces a dual decay mechanism to fully consider the evolution of risk characteristics in both time and space dimensions. The time decay factor ensures that recent transactions have a greater weight in the current risk assessment, consistent with the actual law of risk timeliness; the spatial decay factor reflects the inhibitory effect of correlation distance on risk propagation, with nodes farther apart having weaker risk correlation. By simultaneously applying time and spatial decay factors in the random walk algorithm, the traversal process can adaptively focus on risk propagation paths with high correlation and high timeliness, thereby extracting more representative dynamic risk characteristics. Combining a weighted fusion strategy of path length and node centrality allows the model to comprehensively consider the breadth and depth of risk propagation, significantly improving its ability to identify risks in complex correlation networks.

[0016] Optionally, step A4 further includes: A401: Establish a historical risk case database The historical risk case database Include Each historical risk case record contains a dynamic risk characteristic matrix for that case. and the corresponding risk type labels The risk type label From the preset set of risk types Select from the dynamic risk characteristic matrix of the current transaction to be evaluated. Historical risk case library Dynamic risk feature matrix for each case Similarity is calculated using Euclidean distance as a metric. Euclidean distance between historical cases and current transactions The calculation formula is: ; in, Indicates the current transaction and the first Euclidean distance between historical cases This represents the number of feature dimensions in the dynamic risk feature matrix. Represents the dynamic risk characteristic matrix of the current transaction In the The values ​​in each feature dimension Indicates the first The dynamic risk characteristic matrix of historical cases in the first Numerical values ​​in each feature dimension; A402: Based on Euclidean distance Sort by size from smallest to largest, and select the smallest distance. These historical cases are recorded as a set of similar cases. For similar case sets In Calculate the normalized similarity weight for each historical case, and then... Normalized similarity weights for similar cases The calculation formula is: ; in, Indicates the first Normalized similarity weights for similar cases, Indicates the current transaction and the first Euclidean distance between similar cases This represents the total number of similar cases selected. Indicates the current transaction and the first Euclidean distance between similar cases; A403: Analyze the risk type labels in a set of similar cases. The weighted votes in the ranking, for risk type Its weighted votes The calculation formula is: ; in, Indicates risk type The weighted votes, Indicates the first Normalized similarity weights for similar cases, Indicates the first Risk type labels for similar cases Indicates the preset risk type. This indicates an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. A404: Introducing a risk benchmark value for port areas As a regulating factor, the port area risk benchmark value Based on historical risk case incidence statistics for the port area, specifically targeting risk types... Its port area risk benchmark value By statistical analysis of the port's past Risk types within a time period The proportion of cases occurring to the total number of customs clearances is used to calculate the risk type. probability value The calculation formula is: ; in, Indicates risk type The probability value, Indicates risk type The weighted votes, This represents the adjustment weighting coefficient for the port area risk benchmark value. Indicates risk type The port area risk benchmark value, This indicates the total number of preset risk types. Indicates risk type The weighted votes, Indicates risk type The port area risk benchmark value; A405: Constructing a Risk Assessment Results Set The risk assessment result set Includes smuggling risk probability value False reporting risk probability value Tax evasion risk probability value and the probability value of prohibited and restricted items .

[0017] This step achieves accuracy and interpretability in risk type identification through a weighted voting mechanism. Traditional risk assessment methods often employ fixed threshold rules or simple statistical models, which are difficult to adapt to complex and ever-changing trade risk patterns. In contrast, this step, based on the similarity matching of historical cases, can fully utilize the empirical knowledge of known risk cases and effectively identify unknown risk patterns by finding historical cases most similar to the current transaction.

[0018] This step ensures the objectivity and accuracy of similar case selection by calculating Euclidean distance to measure similarity in the feature space. The design of normalized similarity weights allows closer cases to have a greater impact on risk assessment, reflecting the "nearest neighbor priority" decision-making principle. Introducing a port area risk benchmark as a moderating factor integrates regional risk characteristics into the probability calculation process, enabling risk assessment results to simultaneously reflect individual transaction characteristics and regional environmental characteristics, significantly improving the comprehensiveness and adaptability of risk identification. This method, combining local similarity matching with global regional characteristics, provides a more reliable technical means for risk type identification in complex trade scenarios.

[0019] Optionally, step A5 further includes: A501: Collects the current port throughput data. Check resource availability and historical early warning accuracy The customs clearance flow The availability of inspection resources is obtained by counting the number of customs clearance documents processed within the current time period. The historical early warning accuracy rate is obtained by calculating the ratio of the currently available inspection personnel to the total number of inspection personnel. By statistical analysis of the past The high-risk threshold is calculated by determining the percentage of cases that were triggered and confirmed as genuine risk cases within a given time period, out of the total number of warnings. Medium risk threshold and low risk threshold High risk threshold The calculation formula is: ; in, This indicates the high-risk threshold for the current time period. This represents the benchmark value for the high-risk threshold. The adjustment coefficient representing the customs clearance flow. This indicates the current customs clearance volume at the port. This indicates the reference customs clearance volume. This represents an adjustment coefficient indicating the availability of resources. This indicates the availability of verification resources during the current time period. An adjustment coefficient representing the historical accuracy of early warnings. Indicates the historical accuracy rate of early warnings; A502: Medium Risk Threshold The calculation formula is: ; in, This indicates the medium-risk threshold for the current period. This represents the baseline value for the medium-risk threshold. Low risk threshold The calculation formula is: ; in, This indicates the low-risk threshold for the current period. This represents the baseline value for the low-risk threshold. A503: Set up the risk assessment results The probability values ​​of each risk type are compared with their corresponding risk thresholds. For each risk type... If its probability value If the risk level of this type of risk is high, then the risk level is determined to be high risk; if If so, the risk level is determined to be medium risk; if If the risk level is low, then the risk level is determined to be low; if If so, the risk level is determined to be normal; A504: Traverse all risk types and take the highest risk level as the warning level. When the warning level reaches medium or high risk, trigger the corresponding level of risk warning and generate risk warning information. The risk warning information includes the warning level, the set of main risk types, the list of abnormal fields, and the suggested verification measures. The set of main risk types includes all risk types that reach the medium or high risk level. The list of abnormal fields is extracted from the verification result vector of step A2. The suggested verification measures are obtained by matching the main risk types and abnormal fields from the preset verification measures knowledge base.

[0020] Compared with the prior art, the present invention has at least the following beneficial effects: This invention achieves standardized processing of multi-source trade data by establishing a unified data dictionary and field mapping rules, significantly improving the automation level and data quality of document verification. Traditional methods typically rely on manual verification of each document, which is not only inefficient but also prone to missing key information. The standardized mapping mechanism of this invention can convert heterogeneous data from customs declaration systems, electronic port platforms, logistics tracking systems, bank settlement systems, and third-party certification bodies into a unified format, providing a high-quality data foundation for subsequent cross-document verification and risk analysis. Through preset field mapping rules and missing value imputation strategies, data integrity and consistency are ensured, avoiding verification errors caused by inconsistent data formats.

[0021] This invention employs a rule-engine-based cross-document verification mechanism to achieve comprehensive automated inspection of trade documents, effectively overcoming the limitations of traditional methods in document consistency verification. By establishing a verification rule base that includes matching rules, value range constraint rules, and time logic constraint rules, this invention can systematically check the consistency between key field pairs such as invoice amount and customs declaration price, bill of lading quantity and packing list quantity, and the correspondence between certificate of origin and commodity code. A weighted scoring mechanism is introduced, assigning differentiated weights based on the severity of different inconsistency types to generate a comprehensive consistency score index, making the verification results more discriminative and interpretable. This rule-driven automated verification method not only improves verification efficiency but also ensures the uniformity and objectivity of verification standards.

[0022] This invention achieves in-depth mining of risks associated with trade activities through a dynamic risk feature extraction method based on spatiotemporal decay graphs, significantly improving the ability to identify systemic violations. By mapping trade entities and their transaction relationships to a spatiotemporal graph structure and introducing a dual decay mechanism of time and space decay factors, this invention can fully consider the evolution of risk features in both time and space dimensions. An improved random walk algorithm is used to perform biased multi-hop traversal on the graph. By statistically analyzing the density of risk nodes, the repetition of abnormal transaction patterns, and the complexity of the associated network along the traversal path, more representative dynamic risk features are extracted. A weighted fusion strategy combining path length and node centrality enables risk assessment to comprehensively consider the breadth and depth of risk propagation, effectively identifying complex violation patterns implemented through the associated network. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating an embodiment of the present invention of an automatic verification and risk warning method for international trade information based on a smart port. Detailed Implementation

[0024] The present invention will be further described below with reference to the accompanying drawings, but this is not intended to limit the present invention in any way. Any modifications or substitutions made based on the teachings of the present invention shall fall within the protection scope of the present invention.

[0025] Example 1: A method for automatic verification and risk warning of international trade information based on smart ports, such as... Figure 1 As shown, it includes the following steps: A1: Collect multi-source documentary data from international trade activities. This multi-source documentary data includes customs declaration data from the customs declaration system, bills of lading data from the electronic port platform, transportation route data from the logistics tracking system, payment voucher data from the bank settlement system, and certificate of origin data from third-party certification bodies. Establish a unified data dictionary, and convert data fields from different sources into standardized formats through preset field mapping rules. Construct a structured trade information set containing commodity codes, declared prices, declared quantities, declared weights, trading party information, transportation route information, and timestamp information, including: A101: Data is collected from the customs declaration system (customs declaration system), bill of lading and invoice data from the electronic port platform, transportation route data from the logistics tracking system, payment voucher data from the bank settlement system, and certificate of origin data from third-party certification bodies. In this embodiment, customs declaration data is collected from the customs H2018 declaration system by calling the data interface provided by the customs system. The collected customs declaration data fields include customs declaration number, declaring company name, declaring company customs registration code, commodity HS code, declared price, declared quantity, declared weight, country of origin code, mode of transport code, bill of lading number, and declaration date. Bill of lading and invoice data are collected from the single window system of the electronic port platform by using the data query interface provided by the platform. Bill of lading data fields include bill of lading number, carrier name, port of loading code, port of discharge code, cargo description, number of pieces, gross weight, and bill of lading date. Invoice data... The data includes: invoice number, invoice date, invoice amount, currency code, buyer's name, seller's name, and goods details; transportation route data is collected from the logistics tracking system, obtained through GPS positioning devices and the logistics company's cargo tracking system, and includes: waybill number, origin coordinates, destination coordinates, sequence of route node coordinates, arrival timestamp sequence, and means of transport identifier; payment voucher data is collected from the bank settlement system, obtained through the cross-border payment data interface provided by the bank, and includes: payment voucher number, payer's account information, payee's account information, payment amount, payment currency, payment time, and description of payment purpose; certificate of origin data is collected from third-party certification bodies such as the China Council for the Promotion of International Trade or the Customs Certification Center, obtained through the institution's data sharing interface, and includes: certificate number, issuance date, exporter's name, consignee's name, commodity description, HS code, origin determination criteria, and issuing institution code. A102: Establish a unified data dictionary, which defines standard field names, data types, and field mapping rules. In this embodiment, the data dictionary is stored using a relational database table structure. The table structure includes a field identifier column, a standard field name column, a data type column, a source system identifier column, a source field name column, and a mapping rule column. Standard field names include, but are not limited to: trade_id (unique identifier for trade records), hs_code (HS code for commodities), declared_price (declared price), declared_quantity (declared quantity), declared_weight (declared weight), exporter_name (exporter name), importer_name (importer name), origin_country (country of origin code), transport_route (transport route), timestamp_declare (declaration timestamp), timestamp_invoice (invoice issuance timestamp), timestamp_transport (transport timestamp), invoice_amount (invoice amount), bill_quantity (bill of lading quantity), packing_quantity (packing list quantity), and certificate_number (certificate of origin number). A103: Based on the field mapping rules, convert data fields from different sources into a standardized format, fill missing fields with preset default values, perform unified unit conversion for numeric fields, and convert time fields into a unified timestamp format; A104: Constructing a Structured Trade Information Set The structured trade information set The form is ,in Indicates the total number of document records. Indicates the first One document record, It contains standardized formatted information; in this embodiment, it is a structured trade information set. The database uses relational tables for storage, with a table structure including a primary key column and standardized field columns; total number of document records. In the test scenario of this embodiment, the value is 10000.

[0026] A2: Input the structured trade information set into the rule engine-based cross-document verification module. In this module, a cross-document verification rule base is established. This rule base defines matching rules for key fields, value range constraints, and time logic constraints. Rule matching is performed on each pair of corresponding key fields in the structured trade information set. A weighted scoring mechanism is used to calculate the overall document consistency score. The output is a verification result vector containing inconsistency field identifiers, deviation values, and consistency scores, including: A201: Establish a cross-document verification rule base The cross-document verification rule base Include Verification rules, denoted as Each verification rule Define matching rules between a pair of key fields for a structured trade information set. The first in Document Records Execute the first Verification rules The matching operation determines whether key field pairs meet the matching conditions; if they do, a consistency flag is set. If the condition is not met, a consistency flag is set. And record the deviation value. In this embodiment, cross-document verification rule base Each verification rule Includes attributes such as rule identifier, rule name, key field pairs, matching conditions, deviation threshold, and severity weight; total number of verification rules. The system is set to 15 rules, specifically including: verification of the relative deviation between the invoice amount and the declared price on the customs declaration; verification of the absolute deviation between the quantity on the bill of lading and the quantity on the packing list; verification of the consistency between the HS code of the goods on the certificate of origin and the HS code on the customs declaration; verification of the logical constraints between the invoice issuance date and the declaration date on the customs declaration; verification of the logical constraints between the bill of lading date and the declaration date on the customs declaration; verification of the consistency between the buyer's name on the invoice and the importer's name on the customs declaration; verification of the consistency between the seller's name on the invoice and the exporter's name on the customs declaration; verification of the consistency between the goods description on the bill of lading and the commodity name on the customs declaration; verification of the relative deviation between the declared weight on the customs declaration and the gross weight on the bill of lading; verification of the consistency between the country of origin on the certificate of origin and the country of origin code on the customs declaration; verification of the relative deviation between the amount on the payment voucher and the invoice amount; verification of the logical constraints between the payment voucher time and the invoice issuance time; verification of the consistency between the origin of the transportation route and the port of export on the customs declaration; verification of the consistency between the destination of the transportation route and the port of import on the customs declaration; and verification of the consistency between the mode of transport on the customs declaration and the type of transport vehicle in the logistics tracking system. For the first... Document Records The system executes the above steps sequentially. Verification rules are set, and judgments are made based on the matching conditions of the rules. And calculate the deviation value ; A202: For the verification of the invoice amount and the declared price on the customs declaration, calculate the relative deviation between the two, and the deviation value. The calculation formula is: ; in, This indicates the deviation between the invoice amount and the declared price on the customs declaration form. This indicates the invoice amount in the invoice data. This represents the declared price in the customs declaration data. This indicates the absolute value operation; in this embodiment, if the calculated result is... If the deviation is less than or equal to the preset deviation threshold of 0.05, then Set to 1; if If it is greater than 0.05, then Set to 0 and record ; A203: For verifying the quantity on the bill of lading and the quantity on the packing list, calculate the absolute deviation between the two, and the deviation value. The calculation formula is: ; in, This indicates the deviation between the quantity on the bill of lading and the quantity on the packing list. This indicates the number of bills of lading in the bill of lading data. This indicates the number of boxes recorded in the packing list data; in this embodiment, and All have been standardized to the same unit of measurement; if If it equals 0, then Set to 1; if If greater than 0, then Set to 0 and record ; A204: For each verification rule Set severity weights The severity weight The range of values ​​is And the sum of the severity weights of all verification rules satisfies Calculate the first Document Records Overall consistency score The calculation formula is: ; in, Indicates the first Document Records Overall consistency score, Indicates the first The severity weight of each verification rule. Indicates the first Document record in the first Consistency indicators under the 15 verification rules; in this embodiment, the severity weights of the 15 verification rules. The impact of each rule on trade compliance is set as follows: 0.15, 0.12, 0.1, 0.08, 0.08, 0.07, 0.07, 0.06, 0.09, 0.06, 0.05, 0.03, 0.02, 0.01, 0.01. As an adaptive optimization scheme, the system periodically (quarterly) uses a logistic regression model to refit the relationship coefficients between each verification rule and the actual probability of risk occurrence based on historical confirmed risk case data, and automatically updates the severity weights. This ensures that the weighting configuration always reflects the latest risk patterns and characteristics; A205: Constructing the Verification Result Vector The verification result vector It includes a set of inconsistency field identifiers, the deviation values ​​of the inconsistency fields, and the overall consistency score. .

[0027] A3: Extract trade feature information from the verification result vector. The trade feature information includes information on the inconsistency field identifier set and the document consistency score. A spatiotemporal graph of trade behavior is constructed. In this graph, trade entities are used as nodes, and transaction relationships are used as edges. A time-series feature vector is established for each trade entity node. A dual decay mechanism is used to weight historical transaction features. An improved random walk algorithm is used to perform a biased multi-hop traversal on the spatiotemporal graph of trade behavior. Risk feature indicators on the traversal path are statistically analyzed, and a dynamic risk feature matrix is ​​output, including node risk, path risk, network topology features, and spatiotemporal decay coefficients. A301: From the verification result vector Extract information from the inconsistency field identifier set and document consistency score. Constructing a spatiotemporal map of trade behavior ,in Represents the set of trade entity nodes. This represents the set of edges representing transaction relationships, and the set of nodes representing trade entities. Each node in Establish time series feature vectors The time series feature vector Record the historical transaction characteristics of the trading entity, including the frequency of historical transactions. Fluctuation range of amount Diversity of commodity category distribution and the number of historical violations The fluctuation range of the amount By calculating the past of this node The standard deviation of transaction amount within a time period is obtained, and the diversity of commodity category distribution is also considered. This is obtained by calculating the ratio of the number of different product codes involved in the node to the total number of transactions; A302: Define the time decay factor and spatial decay factor The time decay factor is based on the time of the transaction. With current time The time intervals between nodes decrease exponentially, and the spatial decay factor is based on the graph distance between nodes. It decreases in a power-law manner, with a time decay factor. The calculation formula is: ; in, Indicates the time when it occurs The time decay factor of the transaction, This represents the time decay rate parameter. Indicates the current moment. Indicates the time when a historical transaction occurred. The base of the natural logarithm; spatial decay factor The calculation formula is: ; in, The graph distance is represented as Spatial decay factor over time Spatiotemporal graph representing trade behavior The shortest path length between two nodes is calculated using a breadth-first search algorithm. This represents the spatial decay exponent parameter; in this embodiment, it represents the time decay rate parameter. Set to 0.0005; Spatial decay index parameter Set to 1.5; map distance The breadth-first search algorithm is used for calculation; A303: Employing an improved random walk algorithm in the spatiotemporal graph of trade behavior Perform a traversal operation, starting from the trade entity node corresponding to the current transaction record to be verified. Starting from the current node, perform a multi-hop traversal with bias, during which the traversal begins from the current node. Transfer to adjacent nodes transition probability The calculation formula is: ; in, Indicates starting from the current node Transfer to adjacent nodes The transition probability, This represents the time decay factor of the edges between the current node and its neighboring nodes. This indicates the time when the transaction corresponding to that edge occurred. Indicates the starting node With neighboring nodes Spatial attenuation factor between Indicates the starting node With neighboring nodes Graph distance between them This represents the weight value of the edge. Calculated based on the transaction amount and frequency. Indicates the current node The set of all adjacent nodes, Indicates the current node's... Adjacent nodes Indicates the relationship between the current node and the first node. The edges between adjacent nodes correspond to the times when transactions occur. Indicates the starting node With the Graph distance between adjacent nodes Indicates the relationship between the current node and the first node. The weight values ​​of the edges between adjacent nodes; in this embodiment, the edge weights The calculation formula is: ,in For the amount of the transaction, For the maximum transaction amount, The last transaction time; A304: During the traversal, record the... traversal path ,in This represents the maximum number of hops in the traversal path, and it calculates the risk characteristic indicators on that path, including the risk node density. Repeatability of abnormal transaction patterns and the complexity of the associated network Among them, the density of risk nodes By counting the number of historical violations on the path Greater than the preset threshold The percentage of nodes in the path relative to the total number of nodes is used to determine the repetition rate of abnormal transaction patterns. By extracting the feature vectors of transactions corresponding to each edge on the path, a clustering algorithm is used to identify clusters of abnormal transaction patterns. The proportion of transaction edges belonging to abnormal pattern clusters to the total number of edges on the path is used to obtain the correlation network complexity. The value is obtained by calculating the product of the average degree centrality of the nodes on the path and the average clustering coefficient of the path. The calculation of the... traversal path Path aggregation risk score The calculation formula is: ; in, Indicates the first The path aggregation risk score of the traversal path. The weighting coefficients representing the density of risk nodes. Indicates the first The density of risk nodes on the traversal path The weighting coefficients representing the repetition rate of abnormal transaction patterns. Indicates the first The repetition rate of abnormal transaction patterns on each traversal path Weight coefficients representing the complexity of the network. Indicates the first The complexity of the network of traversal paths, and the weight coefficients satisfy... In this embodiment, the weighting coefficient is set to... , =0.35、 ; Optionally, in step A304, the repeatability of abnormal transaction patterns... Specific methods for obtaining them include: A3041: Regarding the first traversal path Each transaction edge on Extract the feature vector of the transaction The feature vector This includes the transaction amount, the commodity code, the transaction time interval, the document consistency score, and the number of historical violations by both parties in the transaction; A3042: Path The set of feature vectors of all transaction edges As a sample set, This represents the total number of transaction edges in the path, and the K-means clustering algorithm is used to cluster the feature vector set. Perform clustering, with the number of clusters set to [number]. In this embodiment, it is 3; A3043: For each cluster Calculate the mean cosine similarity between the samples within this cluster and the transaction feature vectors in the historical normal transaction pattern database. If the mean cosine similarity Less than the preset normal pattern similarity threshold If the value is 0.75, then the cluster is marked as an abnormal pattern cluster. A3044: Statistical Path The number of transaction edges belonging to the abnormal pattern cluster Calculate the repeatability of abnormal transaction patterns The calculation formula is: ; in, Indicates the first The repetition rate of abnormal transaction patterns on each traversal path Representing a path The number of transaction edges belonging to the abnormal pattern cluster; A305: Yes The path aggregation risk scores of the traversal paths are weighted and combined, and the first path is the first path. Path fusion weight Based on path length and the centrality median of the nodes visited on the path The node centrality is determined jointly by calculating the degree centrality of a node, where degree centrality is defined as the number of a node's neighbors, and the median of the centrality is also considered. The median of the degree centrality values ​​of all nodes on the path, sorted by size, is calculated using the following formula: ; in, Indicates the first The fusion weight of the traversal paths Indicates the first The path length of the traversal path. Indicates the first The median of the centrality values ​​of all visited nodes along the traversal path. This indicates the total number of traversed paths. Indicates the first The path length of the traversal path. Indicates the first The median of the centrality values ​​of all visited nodes along the traversal path; A306: Computational Fusion Risk level of nodes after traversing the path Path risk and network topology characteristics The node risk level Based on the starting node Number of historical violations Historical transaction frequency The calculated network topology features This includes the degree centrality, proximity centrality, and betweenness centrality of the starting node; and the generation of a dynamic risk feature matrix. The dynamic risk feature matrix Includes node risk level Path risk Network topology characteristics and the spatiotemporal decay coefficient vector Among them, path risk The calculation formula is: ; in, This indicates the path risk level after weighted fusion. Indicates the first The fusion weight of the traversal paths Indicates the first The path aggregation risk score of the traversal path. This indicates the total number of traversal paths.

[0028] A4: Perform similarity matching between the dynamic risk feature matrix and the feature vectors in the historical risk case database, and select the top [cases] with the highest similarity. A historical risk case, regarding the aforementioned Historical risk cases are normalized and weighted according to similarity. The weighted votes for each risk type label are counted. The port area risk benchmark value is introduced as an adjustment factor to calculate the probability value of each risk type. The output is a risk assessment result set including the probability values ​​of smuggling risk, false declaration risk, tax evasion risk, and prohibited and restricted goods risk, including: A401: Establish a historical risk case database The historical risk case database Include Each historical risk case record contains a dynamic risk characteristic matrix for that case. and the corresponding risk type labels The risk type label From the preset set of risk types Select from the dynamic risk characteristic matrix of the current transaction to be evaluated. Historical risk case library Dynamic risk feature matrix for each case Similarity is calculated using Euclidean distance as a metric. Euclidean distance between historical cases and current transactions The calculation formula is: ; in, Indicates the current transaction and the first Euclidean distance between historical cases This represents the number of feature dimensions in the dynamic risk feature matrix. Represents the dynamic risk characteristic matrix of the current transaction In the The values ​​in each feature dimension Indicates the first The dynamic risk characteristic matrix of historical cases in the first Numerical values ​​in each feature dimension; A402: Based on Euclidean distance Sort by size from smallest to largest, and select the smallest distance. A historical case, which is the case in this embodiment. , denoted as the set of similar cases For similar case sets In Calculate the normalized similarity weight for each historical case, and then... Normalized similarity weights for similar cases The calculation formula is: ; in, Indicates the first Normalized similarity weights for similar cases, Indicates the current transaction and the first Euclidean distance between similar cases This represents the total number of similar cases selected. Indicates the current transaction and the first Euclidean distance between similar cases; A403: Analyze the risk type labels in a set of similar cases. The weighted votes in the ranking, for risk type Its weighted votes The calculation formula is: ; in, Indicates risk type The weighted votes, Indicates the first Normalized similarity weights for similar cases, Indicates the first Risk type labels for similar cases Indicates the preset risk type. This indicates an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. A404: Introducing a risk benchmark value for port areas As a regulating factor, the port area risk benchmark value Based on historical risk case incidence statistics for the port area, specifically targeting risk types... Its port area risk benchmark value By statistical analysis of the port's past Risk types within a time period The proportion of cases occurring to the total number of customs clearances is used to calculate the risk type. probability value The calculation formula is: ; in, Indicates risk type The probability value, Indicates risk type The weighted votes, This represents the adjustment weighting coefficient for the port area risk benchmark value, which is 0.1 in this embodiment. Indicates risk type The port area risk benchmark value, This indicates the total number of preset risk types. Indicates risk type The weighted votes, Indicates risk type The port area risk benchmark value; A405: Constructing a Risk Assessment Results Set The risk assessment result set Includes smuggling risk probability value False reporting risk probability value Tax evasion risk probability value and the probability value of prohibited and restricted items .

[0029] A5: Based on the current port clearance volume, inspection resource availability, and historical early warning accuracy, dynamically calculate high-risk, medium-risk, and low-risk thresholds. Compare the probability values ​​of each risk type in the risk assessment result set with their corresponding risk thresholds, and take the highest risk level as the early warning level. When the early warning level reaches medium or high risk, trigger the corresponding risk warning and generate risk warning information that includes the warning level, main risk types, a list of abnormal fields, and suggested inspection measures, including: A501: Collects the current port throughput data. Check resource availability and historical early warning accuracy The customs clearance flow The availability of inspection resources is obtained by counting the number of customs clearance documents processed within the current time period. The historical early warning accuracy rate is obtained by calculating the ratio of the currently available inspection personnel to the total number of inspection personnel. By statistical analysis of the past The high-risk threshold is calculated by determining the percentage of cases that were triggered and confirmed as genuine risk cases within a given time period, out of the total number of warnings. Medium risk threshold and low risk threshold High risk threshold The calculation formula is: ; in, This indicates the high-risk threshold for the current time period. This represents the benchmark value for the high-risk threshold. The adjustment coefficient representing the customs clearance flow. This indicates the current customs clearance volume at the port. This indicates the reference customs clearance volume. This represents an adjustment coefficient indicating the availability of resources. This indicates the availability of verification resources during the current time period. An adjustment coefficient representing the historical accuracy of early warnings. This represents the historical early warning accuracy rate; the high-risk threshold benchmark value in this embodiment. Set to 0.75; the adjustment coefficient for customs clearance flow. Set to 0.2 to check the adjustment coefficient for resource availability. The adjustment coefficient for historical early warning accuracy is set to 0.3. Set to 0.15; A502: Medium Risk Threshold The calculation formula is: ; in, This indicates the medium-risk threshold for the current period. This represents the baseline value for the medium-risk threshold, which is set to 0.5 in this embodiment. Low risk threshold The calculation formula is: ; in, This indicates the low-risk threshold for the current period. This represents the low-risk threshold benchmark value, which is set to 0.3 in this embodiment; A503: Set up the risk assessment results The probability values ​​of each risk type are compared with their corresponding risk thresholds. For each risk type... If its probability value If the risk level of this type of risk is high, then the risk level is determined to be high risk; if If so, the risk level is determined to be medium risk; if If the risk level is low, then the risk level is determined to be low; if If so, the risk level is determined to be normal; A504: Traverse all risk types and take the highest risk level as the warning level. When the warning level reaches medium or high risk, trigger the corresponding level of risk warning and generate risk warning information. The risk warning information includes the warning level, the set of main risk types, the list of abnormal fields, and the suggested verification measures. The set of main risk types includes all risk types that reach the medium or high risk level. The list of abnormal fields is extracted from the verification result vector of step A2. The suggested verification measures are obtained by matching the main risk types and abnormal fields from the preset verification measures knowledge base.

[0030] It should be noted that the sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0031] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0032] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A method for automatic verification and risk warning of international trade information based on smart ports, characterized in that: Includes the following steps: A1: Collect multi-source document data from international trade activities. The multi-source document data includes customs declaration data from the customs declaration system, bill of lading data from the electronic port platform, transportation route data from the logistics tracking system, payment voucher data from the bank settlement system, and certificate of origin data from third-party certification bodies. Establish a unified data dictionary, and convert data fields from different sources into standardized formats through preset field mapping rules to construct a structured trade information set containing commodity codes, declared prices, declared quantities, declared weights, trader information, transportation route information, and timestamp information. A2: Input the structured trade information set into the rule engine-based cross-document verification module. In the rule engine-based cross-document verification module, establish a cross-document verification rule base. The cross-document verification rule base defines matching rules, value range constraint rules, and time logic constraint rules for key fields. Perform rule matching operation on each pair of corresponding key fields in the structured trade information set. Use a weighted scoring mechanism to calculate the overall consistency score of the documents. Output a verification result vector containing inconsistency field identifiers, deviation values, and consistency scores. A3: Extract trade feature information from the verification result vector. The trade feature information includes information on the inconsistency field identifier set and the document consistency score. A spatiotemporal graph of trade behavior is constructed. In the spatiotemporal graph of trade behavior, trade entities are used as nodes and transaction relationships are used as edges. A time series feature vector is established for each trade entity node. A dual decay mechanism is used to weight the historical transaction features. An improved random walk algorithm is used to perform biased multi-hop traversal on the spatiotemporal graph of trade behavior. The risk feature indicators on the traversal path are counted, and a dynamic risk feature matrix containing node risk degree, path risk degree, network topology features and spatiotemporal decay coefficient is output. A4: Perform similarity matching between the dynamic risk feature matrix and the feature vectors in the historical risk case database, and select the top [cases] with the highest similarity. A historical risk case, regarding the aforementioned Historical risk cases are normalized and weighted according to similarity, the weighted votes of each risk type label are counted, the port area risk benchmark value is introduced as an adjustment factor, the probability value of each risk type is calculated, and a risk assessment result set including smuggling risk probability value, false declaration risk probability value, tax evasion risk probability value and prohibited and restricted items risk probability value is output. A5: Based on the current customs clearance volume, inspection resource availability, and historical early warning accuracy, the high-risk threshold, medium-risk threshold, and low-risk threshold are dynamically calculated. The probability values ​​of each risk type in the risk assessment result set are compared with their corresponding risk thresholds, and the highest risk level is taken as the early warning level. When the early warning level reaches medium or high risk, the corresponding risk warning is triggered, and risk warning information containing the warning level, main risk types, abnormal field list, and suggested inspection measures is generated.

2. The method for automatic verification and risk warning of international trade information based on smart ports according to claim 1, characterized in that, Step A1 includes: A101: Collect customs declaration data from the customs declaration system, bill of lading data and invoice data from the electronic port platform, transportation route data from the logistics tracking system, payment voucher data from the bank settlement system, and certificate of origin data from third-party certification bodies; A102: Establish a unified data dictionary, which defines standard field names, data types, and field mapping rules; A103: Based on the field mapping rules, convert data fields from different sources into a standardized format, fill missing fields with preset default values, perform unified unit conversion for numeric fields, and convert time fields into a unified timestamp format; A104: Constructing a Structured Trade Information Set The structured trade information set The form is ,in Indicates the total number of document records. Indicates the first One document record, It contains information in a standard format.

3. The method for automatic verification and risk warning of international trade information based on smart ports according to claim 2, characterized in that, Step A2 includes: A201: Establish a cross-document verification rule base The cross-document verification rule base Include Verification rules, denoted as Each verification rule Define matching rules between a pair of key fields for a structured trade information set. The first in Document Records Execute the first Verification rules The matching operation determines whether key field pairs meet the matching conditions; if they do, a consistency flag is set. If the condition is not met, a consistency flag is set. And record the deviation value. ; A202: For the verification of the invoice amount and the declared price on the customs declaration, calculate the relative deviation between the two, and the deviation value. The calculation formula is: ; in, This indicates the deviation between the invoice amount and the declared price on the customs declaration form. This indicates the invoice amount in the invoice data. This represents the declared price in the customs declaration data. This indicates the absolute value operation; A203: For verifying the quantity on the bill of lading and the quantity on the packing list, calculate the absolute deviation between the two, and the deviation value. The calculation formula is: ; in, This indicates the deviation between the quantity on the bill of lading and the quantity on the packing list. This indicates the number of bills of lading in the bill of lading data. This indicates the number of boxes recorded in the packing list data; A204: For each verification rule Set severity weights The severity weight The range of values ​​is And the sum of the severity weights of all verification rules satisfies Calculate the first Document Records Overall consistency score The calculation formula is: ; in, Indicates the first Document Records Overall consistency score, Indicates the first The severity weight of each verification rule. Indicates the first Document record in the first Consistency indicators under the verification rules; A205: Constructing the Verification Result Vector The verification result vector It includes a set of inconsistency field identifiers, the deviation values ​​of the inconsistency fields, and the overall consistency score. .

4. The method for automatic verification and risk warning of international trade information based on smart ports according to claim 3, characterized in that, Step A3 includes: A301: From the verification result vector Extract information from the inconsistency field identifier set and document consistency score. Constructing a spatiotemporal map of trade behavior ,in Represents the set of trade entity nodes. This represents the set of edges representing transaction relationships, and the set of nodes representing trade entities. Each node in Establish time series feature vectors The time series feature vector Record the historical transaction characteristics of the trading entity, including the frequency of historical transactions. Fluctuation range of amount Diversity of commodity category distribution and the number of historical violations The fluctuation range of the amount By calculating the past of this node The standard deviation of transaction amount within a time period is obtained, and the diversity of commodity category distribution is also considered. This is obtained by calculating the ratio of the number of different product codes involved in the node to the total number of transactions; A302: Define the time decay factor and spatial decay factor The time decay factor is based on the time of the transaction. With current time The time intervals between nodes decrease exponentially, and the spatial decay factor is based on the graph distance between nodes. It decreases in a power-law manner, with a time decay factor. The calculation formula is: ; in, Indicates the time when it occurs The time decay factor of the transaction, This represents the time decay rate parameter. Indicates the current moment. Indicates the time when a historical transaction occurred. The base of the natural logarithm; spatial decay factor The calculation formula is: ; in, The graph distance is represented as Spatial decay factor over time Spatiotemporal graph representing trade behavior The shortest path length between two nodes is calculated using a breadth-first search algorithm. This represents the spatial decay index parameter; A303: Employing an improved random walk algorithm in the spatiotemporal graph of trade behavior Perform a traversal operation, starting from the trade entity node corresponding to the current transaction record to be verified. Starting from the current node, perform a multi-hop traversal with bias, during which the traversal begins from the current node. Transfer to adjacent nodes transition probability The calculation formula is: ; in, Indicates starting from the current node Transfer to adjacent nodes The transition probability, This represents the time decay factor of the edges between the current node and its neighboring nodes. This indicates the time when the transaction corresponding to that edge occurred. Indicates the starting node With neighboring nodes Spatial attenuation factor between Indicates the starting node With neighboring nodes Graph distance between them This represents the weight value of the edge. Calculated based on the transaction amount and frequency. Indicates the current node The set of all adjacent nodes, Represents the current node's... Adjacent nodes Indicates the relationship between the current node and the first node. The edges between adjacent nodes correspond to the times when transactions occur. Indicates the starting node With the Graph distance between adjacent nodes Indicates the relationship between the current node and the first node. The weight values ​​of the edges between adjacent nodes; A304: During the traversal, record the... traversal path ,in This represents the maximum number of hops in the traversal path, and it calculates the risk characteristic indicators on that path, including the risk node density. Repeatability of abnormal transaction patterns and the complexity of the associated network Among them, the density of risk nodes By counting the number of historical violations on the path Greater than the preset threshold The percentage of nodes in the path relative to the total number of nodes is used to determine the repetition rate of abnormal transaction patterns. By extracting the feature vectors of transactions corresponding to each edge on the path, a clustering algorithm is used to identify clusters of abnormal transaction patterns. The proportion of transaction edges belonging to abnormal pattern clusters to the total number of edges on the path is used to obtain the correlation network complexity. The value is obtained by calculating the product of the average degree centrality of the nodes on the path and the average clustering coefficient of the path. The calculation of the... traversal path Path aggregation risk score The calculation formula is: ; in, Indicates the first The path aggregation risk score of the traversal path. The weighting coefficients representing the density of risk nodes. Indicates the first The density of risk nodes on the traversal path The weighting coefficients representing the repetition rate of abnormal transaction patterns. Indicates the first The repetition rate of abnormal transaction patterns on each traversal path Weight coefficients representing the complexity of the network. Indicates the first The complexity of the network of traversal paths, and the weight coefficients satisfy... ; A305: Yes The path aggregation risk scores of the traversal paths are weighted and combined, and the first path is the first path. Fusion weights of each path Based on path length and the centrality median of the nodes visited on the path The node centrality is determined jointly by calculating the degree centrality of a node, where degree centrality is defined as the number of a node's neighbors, and the median of the centrality is also considered. The median of the degree centrality values ​​of all nodes on the path, sorted by size, is calculated using the following formula: ; in, Indicates the first The fusion weight of the traversal paths Indicates the first The path length of the traversal path. Indicates the first The median of the centrality values ​​of all visited nodes along the traversal path. This indicates the total number of traversed paths. Indicates the first The path length of the traversal path. Indicates the first The median of the centrality values ​​of all visited nodes along the traversal path; A306: Computational Fusion Risk level of nodes after traversing the path Path risk and network topology characteristics The node risk level Based on the starting node Number of historical violations Historical transaction frequency The calculated network topology features This includes the degree centrality, proximity centrality, and betweenness centrality of the starting node; and the generation of a dynamic risk feature matrix. The dynamic risk feature matrix Includes node risk level Path risk Network topology characteristics and the spatiotemporal decay coefficient vector Among them, path risk The calculation formula is: ; in, This indicates the path risk level after weighted fusion. Indicates the first The fusion weight of the traversal paths Indicates the first The path aggregation risk score of the traversal path. This indicates the total number of traversal paths.

5. The method for automatic verification and risk warning of international trade information based on smart ports according to claim 4, characterized in that, The repeatability of abnormal transaction patterns in step A304 The calculation method specifically includes: A3041: Regarding the first traversal path Each transaction edge on Extract the feature vector of the transaction The feature vector This includes the transaction amount, the commodity code, the transaction time interval, the document consistency score, and the number of historical violations by both parties in the transaction; A3042: Path The set of feature vectors of all transaction edges As a sample set, This represents the total number of transaction edges in the path, and the K-means clustering algorithm is used to cluster the feature vector set. Perform clustering, with the number of clusters set to [number]. ; A3043: For each cluster Calculate the mean cosine similarity between the samples within this cluster and the transaction feature vectors in the historical normal transaction pattern database. If the mean cosine similarity Less than the preset normal pattern similarity threshold If so, then the cluster is marked as an anomalous pattern cluster; A3044: Statistical Path The number of transaction edges belonging to the abnormal pattern cluster Calculate the repeatability of abnormal transaction patterns The calculation formula is: ; in, Indicates the first The repetition rate of abnormal transaction patterns on each traversal path Representing a path The number of transaction edges belonging to the abnormal pattern cluster.

6. The method for automatic verification and risk warning of international trade information based on smart ports according to claim 4, characterized in that, Step A4 includes: A401: Establish a historical risk case database The historical risk case database Include Each historical risk case record contains a dynamic risk characteristic matrix for that case. and the corresponding risk type labels The risk type label From the preset set of risk types Select from the dynamic risk characteristic matrix of the current transaction to be evaluated. Historical Risk Case Library Dynamic risk feature matrix for each case Similarity is calculated using Euclidean distance as a metric. Euclidean distance between historical cases and current transactions The calculation formula is: ; in, Indicates the current transaction and the first Euclidean distance between historical cases This represents the number of feature dimensions in the dynamic risk feature matrix. Represents the dynamic risk characteristic matrix of the current transaction. In the The values ​​in each feature dimension Indicates the first The dynamic risk characteristic matrix of historical cases in the first Numerical values ​​in each feature dimension; A402: Based on Euclidean distance Sort by size from smallest to largest, and select the smallest distance. These historical cases are recorded as a set of similar cases. For similar case sets In Calculate the normalized similarity weight for each historical case, and then... Normalized similarity weights for similar cases The calculation formula is: ; in, Indicates the first Normalized similarity weights for similar cases, Indicates the current transaction and the first Euclidean distance between similar cases This represents the total number of similar cases selected. Indicates the current transaction and the first Euclidean distance between similar cases; A403: Analyze the risk type labels in a set of similar cases. The weighted votes in the ranking, for risk type Its weighted votes The calculation formula is: ; in, Indicates risk type The weighted votes, Indicates the first Normalized similarity weights for similar cases, Indicates the first Risk type labels for similar cases Indicates the preset risk type. This indicates an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. A404: Introducing Port Area Risk Benchmark Values As a regulating factor, the port area risk benchmark value Based on historical risk case incidence statistics for the port area, specifically targeting risk types... Its port area risk benchmark value By statistical analysis of the port's past Risk types within a time period The proportion of cases occurring to the total number of customs clearances is used to calculate the risk type. probability value The calculation formula is: ; in, Indicates risk type The probability value, Indicates risk type The weighted votes, This represents the adjustment weighting coefficient for the port area risk benchmark value. Indicates risk type The port area risk benchmark value, This indicates the total number of preset risk types. Indicates risk type The weighted votes, Indicates risk type The port area risk benchmark value; A405: Constructing a Risk Assessment Results Set The risk assessment result set Includes smuggling risk probability value False reporting risk probability value Tax evasion risk probability value and the probability value of prohibited and restricted items .

7. The method for automatic verification and risk warning of international trade information based on smart ports according to claim 4, characterized in that, Step A5 includes: A501: Collects the current port throughput data. Check resource availability and historical early warning accuracy The customs clearance flow The availability of inspection resources is obtained by counting the number of customs clearance documents processed within the current time period. The historical early warning accuracy rate is obtained by calculating the ratio of the currently available inspection personnel to the total number of inspection personnel. By statistical analysis of the past The high-risk threshold is calculated by determining the percentage of cases that were triggered and confirmed as genuine risk cases within a given time period, out of the total number of warnings. Medium risk threshold and low risk threshold High risk threshold The calculation formula is: ; in, This indicates the high-risk threshold for the current time period. This represents the benchmark value for the high-risk threshold. The adjustment coefficient representing the customs clearance flow. This indicates the current port throughput. This indicates the reference customs clearance volume. This represents an adjustment coefficient indicating the availability of resources. This indicates the availability of verification resources during the current time period. An adjustment coefficient representing the historical accuracy of early warnings. Indicates the historical accuracy rate of early warnings; A502: Medium Risk Threshold The calculation formula is: ; in, This indicates the medium-risk threshold for the current period. This represents the baseline value for the medium-risk threshold. Low risk threshold The calculation formula is: ; in, This indicates the low-risk threshold for the current period. This represents the baseline value for the low-risk threshold. A503: Set up the risk assessment results The probability values ​​of each risk type are compared with their corresponding risk thresholds. For each risk type... If its probability value If the risk level of this type of risk is high, then the risk level is determined to be high risk; if If so, the risk level is determined to be medium risk; if If the risk level is low, then the risk level is determined to be low; if If so, the risk level is determined to be normal; A504: Traverse all risk types and take the highest risk level as the warning level. When the warning level reaches medium or high risk, trigger the corresponding level of risk warning and generate risk warning information. The risk warning information includes the warning level, the set of main risk types, the list of abnormal fields, and the suggested verification measures. The set of main risk types includes all risk types that reach the medium or high risk level. The list of abnormal fields is extracted from the verification result vector of step A2. The suggested verification measures are obtained by matching the main risk types and abnormal fields from the preset verification measures knowledge base.