A method and system for identifying foreign trade risks based on big data analysis

CN122779902APending Publication Date: 2026-09-18GUIZHOU BUSINESS SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611265119.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-20
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0003]在现有技术中,对外贸易风险识别通常以单份贸易单证或单一业务环节为对象,对商品名称、货物数量、交易金额、收发货主体及结算信息进行独立校验,然而,不同贸易单证分别由交易企业、物流企业、金融机构及海关业务系统生成,其数据格式、记录口径及业务语义存在差异,现有方法难以沿同一笔跨境交易连续核验结算账户归属、货物交付、报关代理及物流凭证之间的关联关系

Benefits of technology

[0054]The process involves: acquiring various cross-border trade documents; performing entity recognition and semantic analysis on the business content of all cross-border trade documents to generate a cross-document business semantic graph; conducting cross-domain risk verification on the settlement account location and goods delivery location of the trading parties, and identifying cross-border transactions with settlement and delivery anomalies based on the risk verification results and the cross-document business semantic graph; using the customs declaration information of cross-border transactions with settlement and delivery anomalies as constraints, constructing a routine business profile based on the constraints and historical agency behavior data of customs brokerage agencies, and identifying cross-border transactions with customs brokerage anomalies through the routine business profile; performing topological closed-loop determination on payment and receipt vouchers and logistics delivery vouchers of cross-border transactions with settlement and delivery anomalies based on cross-domain mutual verification to generate an abnormal evidence chain in the foreign trade process, and identifying concealed documents with abnormal business content through the abnormal evidence chain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122779902A_ABST
    Figure CN122779902A_ABST
Patent Text Reader

Abstract

This application provides a method and system for identifying foreign trade risks based on big data analysis, relating to the field of trade risk identification technology. It generates a cross-documentary business semantic graph; performs cross-domain risk verification on the settlement account location and goods delivery location of the trading parties, thereby identifying cross-border transactions with settlement and delivery anomalies; constructs a routine business profile using customs declaration information of cross-border transactions with settlement and delivery anomalies as constraints, and identifies cross-border transactions with abnormal customs brokerage through this profile; performs topological closed-loop determination based on cross-domain mutual verification on payment and receipt vouchers and logistics delivery vouchers of cross-border transactions with settlement and delivery anomalies, generating an abnormal evidence chain in the foreign trade process, and identifies concealed documents with abnormal business content through this abnormal evidence chain. This application can realize consistency verification of business content among multiple types of trade documents, thereby improving the ability to identify concealed document anomalies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of trade risk identification technology, and more specifically, to a method and system for identifying foreign trade risks based on big data analysis. Background Technology

[0002] Trade risk identification technology is a crucial technology for identifying risks in cross-border transactions involving anomalies in entities, funds, logistics, and documentation. This technology aggregates business data such as contracts, invoices, packing lists, bills of lading, customs declarations, certificates of origin, and payment vouchers to verify the correlation between trading entities, goods flow, fund settlement, and customs brokerage activities. It provides technical support for customs supervision, financial risk control, trade compliance review, and corporate internal control, thereby improving the efficiency of foreign trade risk assessment and the ability to supervise cross-border transactions.

[0003] In existing technologies, foreign trade risk identification typically focuses on a single trade document or a single business process, independently verifying the commodity name, quantity, transaction amount, consignor / consignee, and settlement information. However, different trade documents are generated by trading companies, logistics companies, financial institutions, and customs business systems, resulting in differences in data format, recording standards, and business semantics. Existing methods struggle to continuously verify the relationships between settlement account ownership, goods delivery, customs brokerage, and logistics documents along the same cross-border transaction. When risk entities evade review by using name variations, quantity splitting, amount allocation, settlement entity replacement, or logistics route changes, while individual contents within each document may be within reasonable limits, inconsistencies arise across document business content. This makes it difficult to establish cross-verification relationships for anomalous evidence scattered across different business processes, hindering the accurate identification of hidden document anomalies. Therefore, achieving consistency verification of business content across multiple types of trade documents to improve the ability to identify hidden document anomalies has become a challenge for the industry. Summary of the Invention

[0004] This application provides a method and system for identifying foreign trade risks based on big data analysis, which can realize the consistency verification of business content among multiple types of trade documents, thereby improving the ability to identify hidden document anomalies.

[0005] Firstly, this application provides a method for identifying foreign trade risks based on big data analysis, comprising the following steps:

[0006] Obtain different cross-border trade documents;

[0007] Entity identification and semantic analysis are performed on the business content of all cross-border trade documents to generate a cross-document business semantic graph.

[0008] Cross-domain risk verification is performed on the settlement account location and goods delivery location of the trading parties, and cross-border transactions with settlement and delivery anomalies are identified based on the risk verification results and the cross-document business semantic graph.

[0009] Using the customs declaration behavior characteristics of abnormal cross-border transactions in settlement and delivery as constraints, a routine business profile is constructed based on the constraints and the historical agency behavior data of the customs brokerage agency. The abnormal cross-border transactions of customs brokerage agency are identified through the routine business profile.

[0010] A topological closed-loop determination is performed on payment and receipt vouchers for abnormal cross-border transactions and logistics delivery vouchers for abnormal cross-border customs brokerage transactions to generate an abnormal evidence chain in the foreign trade process, and the hidden documents with abnormal business content are identified through the abnormal evidence chain.

[0011] In this embodiment, entity recognition and semantic analysis are performed on the business content of all cross-border trade documents to generate a cross-document business semantic graph, specifically including:

[0012] Optical character recognition and section classification are performed on all cross-border trade documents. Business entities are extracted using a trade term attention bias named entity recognition model, and local semantic edges based on section affiliation are established within the cross-border trade documents.

[0013] By using a multilingual thesaurus of trade entities and cross-document co-occurrence constraint rules, business entities that refer to the same transaction object are linked and merged to generate a unified entity node across documents.

[0014] Based on the unified entity node across documents, the trade terms, timestamps and amount fields in the associated cross-border trade documents are extracted to construct the execution time relationship edge and the fund and goods association edge across documents.

[0015] Using the cross-document unified entity node as the node, and the local semantic edge, the performance time sequence relation edge, and the fund and goods association edge as the edge, a cross-document business semantic graph is generated.

[0016] In this embodiment, cross-regional risk verification of the settlement account location and goods delivery location of the trading parties specifically includes:

[0017] By querying the offshore financial regulatory list and the trade embargo zone list, we can obtain the financial regulatory attribute label of the settlement account location of the trade transaction parties and the geopolitical risk attribute label of the goods delivery location.

[0018] Cross-matching is performed between the financial regulatory attribute tags and the geopolitical risk attribute tags to generate rule hit identifiers;

[0019] The risk verification result is obtained by performing risk verification on the rule hit identifier through preset cross-domain risk rules.

[0020] In this embodiment, identifying cross-border transactions with settlement and delivery anomalies based on risk verification results and the cross-document business semantic graph specifically includes:

[0021] Extract the trading parties marked as having cross-domain risks from the risk verification results, and locate the settlement account entity node associated with the trading parties in the cross-document business semantic graph;

[0022] Starting from the settlement account entity node, traverse the cross-document business semantic graph along the fund and goods association edge to construct a settlement and delivery association subgraph;

[0023] The settlement and delivery association subgraph is input into a pre-trained graph isomorphism discriminant network, and the topological deviation between the settlement and delivery association subgraph and the normal settlement and delivery mode subgraph is calculated.

[0024] When the deviation of the topology exceeds the normal deviation threshold, it is determined that there is an anomaly in the settlement and delivery of the cross-border transaction in which the trading parties participate.

[0025] In this embodiment, the customs declaration behavior characteristics of abnormal cross-border transactions in settlement and delivery are used as constraints. The construction of a routine business profile based on these constraints and the historical agency behavior data of the customs brokerage agency specifically includes:

[0026] Extract customs declaration behavior characteristics from customs declaration information of abnormal cross-border transactions with settlement and delivery;

[0027] Using the aforementioned customs declaration behavior characteristics as constraints in the data retrieval process, the customs declaration records of the customs brokerage agency within the historical period are retrieved to obtain historical agency behavior data.

[0028] Multi-peak distribution fitting is performed on the behavioral data in the historical agent behavior data to obtain different dense behavioral intervals and outlier boundaries;

[0029] Generate a routine business profile based on all densely populated behavioral areas and outlier boundaries.

[0030] In this embodiment, identifying abnormal cross-border transactions involving customs brokerage through the routine business profile specifically includes:

[0031] Based on the usual business profile, positive and negative examples of customs brokerage behavior are generated, and a behavior determination rule tree is constructed using the generated positive and negative examples.

[0032] The behavior determination rule tree is used to identify abnormal cross-border transactions by customs brokerage agents.

[0033] In this embodiment, a topological closed-loop determination is performed on payment and receipt vouchers for abnormal cross-border transactions involving settlement and delivery irregularities, and on logistics delivery vouchers for abnormal cross-border transactions involving customs brokerage irregularities, to generate an abnormal evidence chain in the foreign trade process. Specifically, this includes:

[0034] The fund transfer records and goods transfer records in the payment and receipt vouchers and the logistics delivery vouchers are projected onto a unified topological space according to the transaction subject and timestamp to form a discrete vector field of fund flow and goods flow;

[0035] The discrete vector field is subjected to Hodge decomposition to obtain the divergence-free curl field component, and the vortex region in the divergence-free curl field component with curl magnitude exceeding zero is extracted as the topological breakpoint.

[0036] For each topological breakpoint, calculate the loop integral of the capital flow vector and the cargo flow vector on the boundary loop of the topological breakpoint, and mark the boundary loop where the difference of the loop integral exceeds the preset closure tolerance as a closed loop break loop.

[0037] Based on all payment and receipt voucher identifiers and logistics delivery voucher identifiers involved in the area enclosed by each closed-loop break, the falsification relationship between vouchers is arranged according to the loop direction of the closed-loop break, thereby generating an abnormal evidence chain in the foreign trade process.

[0038] In this embodiment, projecting the fund transfer records and goods transfer records in the payment vouchers and logistics delivery vouchers onto a unified topological space according to the transaction subject and timestamp to form a discrete vector field of fund flow and goods flow specifically includes:

[0039] Based on the fund transfer records and goods transfer records in the payment and receipt vouchers and the logistics delivery vouchers, a cross-domain transfer graph is constructed with the transaction entity as the node and the fund flow and goods flow as the transfer edges;

[0040] The cross-domain flow graph is input into a pre-trained graph attention network to generate cross-domain fusion embedding vectors for each transaction entity;

[0041] The vector space spanned by the cross-domain fusion embedding vectors of all trading entities is used as the unified topological space;

[0042] Using the cross-domain fusion embedding vectors of each trading entity in a unified topological space as node coordinates, the direction of capital flow and the direction of goods flow as edge directions, and the timestamp interval as edge weights, all directed edges are mapped to directed vectors in a unified topological space, thus forming a discrete vector field of capital flow and goods flow.

[0043] In this embodiment, the concealed documents used to identify abnormal business content through the abnormal evidence chain specifically include:

[0044] Extract the voucher identifier associated with the closed-loop break loop from the abnormal evidence chain, and perform distributed index matching in the cross-border trade document data lake to obtain the business content text that matches the cross-border trade document.

[0045] The business content text is input into a pre-trained trade domain language model to generate a semantic vector, and an anomaly pattern query vector is constructed based on the falsification correlation of the anomaly evidence chain. The semantic representation of anomaly perception is generated through cross-attention fusion.

[0046] The semantic representation of anomaly perception is input into a preset anomaly detection network to calculate the anomaly score of cross-border trade documents. Cross-border trade documents with anomaly scores exceeding a preset judgment threshold are identified as hidden documents with abnormal business content.

[0047] Secondly, this application provides a foreign trade risk identification system based on big data analysis, used to execute a foreign trade risk identification method based on big data analysis, the foreign trade risk identification system comprising:

[0048] The cross-border trade document collection module is used to acquire different cross-border trade documents;

[0049] The cross-document semantic graph construction module is used to perform entity recognition and semantic analysis on the business content of all cross-border trade documents, and generate a cross-document business semantic graph.

[0050] The settlement and delivery risk screening module is used to perform cross-domain risk verification on the settlement account location and the goods delivery location of the trading parties, and to identify cross-border transactions with settlement and delivery abnormalities based on the risk verification results and the cross-document business semantic graph.

[0051] The customs brokerage agent profile recognition module is used to construct a routine business profile based on the customs declaration information of abnormal cross-border transactions with settlement and delivery as constraints, and the historical agency behavior data of the customs brokerage agent. The abnormal cross-border transactions of the customs brokerage agent are identified through the routine business profile.

[0052] The cross-domain document closed-loop judgment module is used to perform topological closed-loop judgment based on cross-domain mutual verification on payment and receipt documents for abnormal cross-border transactions and logistics delivery documents for abnormal cross-border customs brokerage transactions. It generates an abnormal evidence chain in the foreign trade process and identifies hidden documents with abnormal business content through the abnormal evidence chain.

[0053] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:

[0054] The process involves: acquiring various cross-border trade documents; performing entity recognition and semantic analysis on the business content of all cross-border trade documents to generate a cross-document business semantic graph; conducting cross-domain risk verification on the settlement account location and goods delivery location of the trading parties, and identifying cross-border transactions with settlement and delivery anomalies based on the risk verification results and the cross-document business semantic graph; using the customs declaration information of cross-border transactions with settlement and delivery anomalies as constraints, constructing a routine business profile based on the constraints and historical agency behavior data of customs brokerage agencies, and identifying cross-border transactions with customs brokerage anomalies through the routine business profile; performing topological closed-loop determination on payment and receipt vouchers and logistics delivery vouchers of cross-border transactions with settlement and delivery anomalies based on cross-domain mutual verification to generate an abnormal evidence chain in the foreign trade process, and identifying concealed documents with abnormal business content through the abnormal evidence chain.

[0055] Therefore, this application demonstrates that the abnormal evidence chain can identify concealed documents with abnormal business content. Firstly, by acquiring cross-border trade documents from different sources and performing entity identification and semantic association on the trading entities, commodity information, transaction amounts, settlement information, customs declaration information, and logistics information in each document, interference caused by differences in name spelling, field structure, and record standards can be eliminated, allowing document content scattered across different business systems to establish semantic connections based on the same cross-border transaction. Secondly, by performing cross-domain risk verification on the settlement account location and goods delivery location of the trading parties, and combining this with cross-document business semantic graph verification of the relationships between settlement entities, consignor / consignee entities, and goods flow, abnormal transactions where the fund settlement location and the actual goods delivery location do not conform to normal trade logic can be identified, preventing discrepancies between fund flow and goods flow caused by partial information in a single document being within a reasonable range. Furthermore, by limiting the verification object to customs declaration information of abnormal cross-border transactions with settlement and delivery, and constructing a customary [system / mechanism] based on the historical agency behavior data of customs brokerage agencies, [the application can be further refined]. Business profiling allows for the correlation and comparison of current customs brokerage activities with existing agents, declared goods, agency regions, and business frequencies. This provides further verification of anomalies discovered during the settlement and delivery process at the customs brokerage level, while reducing misjudgments of abnormal transactions caused by occasional regional differences. Finally, by cross-domain mutual verification and topological closed-loop determination of payment and receipt vouchers for abnormal cross-border transactions and logistics delivery vouchers for abnormal cross-border customs brokerage transactions, it is possible to verify whether various vouchers have a complete business continuity relationship along the stages of fund settlement, customs brokerage, cargo transportation, and actual delivery. Mutually corroborating or contradictory abnormal information is linked into an abnormal evidence chain. Based on this, the document identifiers and business content associated with each abnormal relationship can be traced through the abnormal evidence chain. This allows for the identification of specific documents with mismatched entities, inconsistent amounts, inconsistent cargo quantities, or broken logistics delivery relationships from relevant cross-border trade documents. This enables the identification of hidden documents that conceal business contradictions through methods such as name variations, quantity splitting, amount allocation, replacement of settlement entities, and changes in logistics routes.

[0056] In summary, the technical solution adopted in this application can realize the consistency verification of business content among multiple types of trade documents, thereby improving the ability to identify hidden document anomalies. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this embodiment of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This is an exemplary flowchart of a foreign trade risk identification method based on big data analysis provided in this application;

[0059] Figure 2 This is a comparison chart of the semantic association effects across document business provided in this application;

[0060] Figure 3 This is a comparison chart of the closed-loop identification effect of the abnormal evidence chain provided in this application;

[0061] Figure 4 This is a module structure diagram of a foreign trade risk identification system based on big data analysis provided in this application;

[0062] Figure 5 This is an application scenario diagram of the foreign trade risk identification method based on big data analysis provided in this application. Detailed Implementation

[0063] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0064] This application provides a method and system for identifying foreign trade risks based on big data analysis. The core of this method is to acquire different cross-border trade documents; perform entity recognition and semantic analysis on the business content of all cross-border trade documents to generate a cross-document business semantic graph; conduct cross-domain risk verification on the settlement account location and goods delivery location of the trading parties, and identify cross-border transactions with settlement and delivery anomalies based on the risk verification results and the cross-document business semantic graph; use the customs declaration information of cross-border transactions with settlement and delivery anomalies as constraints, construct a habitual business profile based on the constraints and historical agency behavior data of customs brokerage agencies, and identify cross-border transactions with abnormal customs brokerage through the habitual business profile; perform topological closed-loop determination based on cross-domain mutual verification on the payment and receipt vouchers of cross-border transactions with settlement and delivery anomalies and the logistics delivery vouchers of cross-border transactions with abnormal customs brokerage, generate an abnormal evidence chain in the foreign trade process, and identify concealed documents with abnormal business content through the abnormal evidence chain.

[0065] Example 1

[0066] To better understand the above technical solutions, a detailed explanation will be provided below with reference to the accompanying drawings and specific implementation methods. Figure 1As shown in the figure, this is an exemplary flowchart of a foreign trade risk identification method based on big data analysis according to this embodiment of the application, which includes the following steps:

[0067] In step S1, different cross-border trade documents are obtained.

[0068] In practice, data exchange interfaces are used to connect to the customs declaration system, trade business system, settlement business system, and logistics and transportation system. Electronic customs declarations, trade contracts, commercial invoices, packing lists, bills of lading, payment receipts, and logistics delivery receipts are retrieved according to cross-border transaction numbers, customs declaration numbers, contract numbers, or bills of lading numbers. For cross-border trade documents stored as scanned copies, the corresponding image files are read using a batch file upload method. When receiving each cross-border trade document, the document type, source system, original document identifier, transaction entity, timestamp, and transaction association identifier of the cross-border trade document are read simultaneously. Duplicate cross-border trade documents are removed based on the original document identifier, and the retained electronic files, image files, and their transaction association identifiers are stored as different cross-border trade documents in the cross-border trade document data lake.

[0069] In step S2, entity identification and semantic analysis are performed on the business content of all cross-border trade documents to generate a cross-document business semantic graph.

[0070] In this embodiment, entity identification and semantic analysis are performed on the business content of all cross-border trade documents to generate a cross-document business semantic graph, which can be achieved through the following steps:

[0071] Optical character recognition and section classification are performed on all cross-border trade documents. Business entities are extracted using a trade term attention bias named entity recognition model, and local semantic edges based on section affiliation are established within the cross-border trade documents.

[0072] By using a multilingual thesaurus of trade entities and cross-document co-occurrence constraint rules, business entities that refer to the same transaction object are linked and merged to generate a unified entity node across documents.

[0073] Based on the unified entity node across documents, the trade terms, timestamps and amount fields in the associated cross-border trade documents are extracted to construct the execution time relationship edge and the fund and goods association edge across documents.

[0074] Using the cross-document unified entity node as the node, and the local semantic edge, the performance time sequence relation edge, and the fund and goods association edge as the edge, a cross-document business semantic graph is generated.

[0075] It should be noted that, in this application, the business entity refers to an independent semantic unit in a cross-border trade document that carries information such as the transaction subject, settlement account, goods, amount, time, or document identification; the trade term attention-biased named entity recognition model refers to a named entity recognition model that identifies cross-border trade business entities by increasing the attention weight of corresponding words in the trade term; the local semantic edge refers to the semantic association relationship established between business entities in the same section within a single cross-border trade document based on field membership; and the multilingual trade entity thesaurus refers to an entity dictionary that stores the mapping relationship between different language names, standardized abbreviations, and official abbreviations and unified entity codes. The cross-document co-occurrence constraint rule represents the association determination rule used to limit whether business entities in different cross-border trade documents can point to the same transaction object; the cross-document unified entity node represents the graph node obtained by merging business entities referring to the same transaction object in multiple cross-border trade documents; the performance time sequence relationship edge represents the directed association relationship established by different performance stages in the same cross-border transaction according to the time of business occurrence; the funds and goods association edge represents the directed association relationship established by the settlement account entity and the goods entity in the same cross-border transaction based on the flow of funds and the flow of goods; the cross-document business semantic graph represents the semantic network reflecting the business association relationship of multiple cross-border trade documents.

[0076] In specific implementation, firstly, grayscale conversion, adaptive binarization, and tilt correction methods are used to preprocess the character regions of the page images of all cross-border trade documents. Then, connected component labeling is used to extract candidate character regions, and the text content, text box coordinates, and font size in the candidate character regions are read by an optical character recognition engine. A section classification method combining horizontal projection, vertical projection, and section title rule matching is used to divide sections based on the row and column spacing between text boxes, font size variations, and title keywords corresponding to the document header, body, signature, and seal areas. A named entity recognition model with trade term attention bias, including a Transformer encoding layer and a conditional random field decoding layer, is used. After self-attention weight normalization, the attention weights corresponding to the words that hit the preset trade term vocabulary are multiplied by 1.2 and re-normalized. The 1.2 is selected based on the bias coefficient corresponding to the highest entity recognition F1 value in the historical cross-border trade document verification set. The text content and text box coordinates of each section are input. A trade term attention-biased named entity recognition model reads the start and end positions and entity categories of the entities output by the model. It treats continuous text fragments corresponding to the transaction entity, settlement account, goods name, amount, time, and document identifier as business entities. Then, assuming consistent section numbers, it connects business entities according to the nearest neighbor relationship between field names and field values, and the same row or column relationship between table cells, treating the established connections as local semantic edges. Next, it uses enterprise registration names, customs commodity code names, international trade term standard terms, and verified historical entity aliases as term sources, mapping different language names, standardized abbreviations, and official abbreviations to unique entity codes and writing them into a multilingual trade entity thesaurus. It calls the multilingual trade entity thesaurus to convert business entities in different cross-border trade documents into unified names. Then, it uses an edit distance similarity algorithm to compare business entity names that cannot be directly mapped to the same entity code, using a name similarity of not less than 0.85 as the name matching condition.85. Based on the historical business entity matching verification, the threshold corresponding to the lowest mismatch rate and missed match rate is selected; according to the cross-document co-occurrence constraint rules, the contract number, customs declaration number, invoice number, bill of lading number, settlement account number, and customs commodity code are further checked. When two business entities are mapped to the same entity code, or the name matching condition is met and they have at least one identical transaction association identifier and one identical entity attribute, the two business entities are linked to the same graph node. The graph node that has completed attribute deduplication and source document identifier retention is used as the cross-document unified entity node; then, the associated cross-border trade documents are read according to the source document identifier retained by the cross-document unified entity node, the trade terms are matched word by word using the trade term thesaurus, the year, month, day, and hour and minute fields are matched using the date format template and uniformly converted to timestamps, and the amount field is read using the joint matching method of amount field name and currency symbol; according to the document type and trade terms, the contract signing, customs declaration, goods shipment, payment, and goods delivery are mapped to the corresponding performance stages. The process involves connecting cross-document unified entity nodes related to the performance stage according to their timestamps, using connections established by pointing earlier nodes to later nodes as performance sequence edges. Within the scope of cross-border transactions limited by the same contract number, customs declaration number, invoice number, or bill of lading number, directed connections are established between the cross-document unified entity nodes corresponding to the settlement accounts and the cross-document unified entity nodes corresponding to the goods, based on the payer and payee recorded in the amount field and the goods delivery party and recipient defined in the trade terms. These directed connections serve as funds-goods association edges. Finally, the cross-document unified entity nodes are written into the node set of the attribute graph, recording the entity category, unified entity code, and source document identifier. Local semantic edges, performance sequence edges, and funds-goods association edges are written into the edge set of the attribute graph, recording the edge type, edge direction, timestamp, and amount fields. Graph indexes are then created according to entity category, unified entity code, source document identifier, and edge type. The attribute graph, with completed node writing, edge writing, and graph index creation, serves as the cross-document business semantic graph.

[0077] For example, to verify the effectiveness of the cross-document business semantic graph in supporting the association of multiple types of document business content, verification data conforming to the actual cross-border trade field range can be used for processing. The verification data includes 1200 sets of cross-border transaction records. Each set of records contains customs declarations, trade contracts, commercial invoices, packing lists, bills of lading, payment and receipt vouchers, and logistics delivery vouchers. In some records, multilingual expressions of company names, partial missing document numbers, differences in time formats, and differences in currency expressions for amounts are included. Business entity links and relationship construction are completed using both the intra-document field matching method and the cross-document business semantic graph method of this application. (Reference) Figure 2As shown in the figure, this is a comparison of the semantic association effects across documents provided in this application. The business entity alignment rate, performance time sequence relationship coverage rate, fund-goods association coverage rate, and complete transaction chain restoration rate of the method in this application are 95.4%, 93.1%, 91.7%, and 88.9%, respectively, which are 12.8, 18.3, 22.2, and 27.7 percentage points higher than the intra-document field matching method. It can be seen that by connecting the same transaction object in documents from different sources through a unified entity node across documents, and further constructing performance time sequence relationship edges and fund-goods association edges, the interruption of business association caused by name variations, missing fields, and format differences can be reduced.

[0078] In step S3, cross-domain risk verification is performed on the settlement account location and goods delivery location of the trading parties, and cross-border transactions with settlement and delivery anomalies are identified based on the risk verification results and the cross-document business semantic graph.

[0079] In this embodiment, cross-regional risk verification of the settlement account location and goods delivery location of the trading parties can be achieved through the following steps:

[0080] By querying the offshore financial regulatory list and the trade embargo zone list, we can obtain the financial regulatory attribute label of the settlement account location of the trade transaction parties and the geopolitical risk attribute label of the goods delivery location.

[0081] Cross-matching is performed between the financial regulatory attribute tags and the geopolitical risk attribute tags to generate rule hit identifiers;

[0082] The risk verification result is obtained by performing risk verification on the rule hit identifier through preset cross-domain risk rules.

[0083] It should be noted that, in this application, the offshore financial regulatory list refers to a list of rules recording the regulatory status set by different countries or regions for the opening of offshore accounts, cross-border fund settlement, and review of fund flows; the trade embargo zone list refers to a list of rules recording the prohibition, restriction, or licensed delivery requirements for goods applicable to different countries or regions; the financial regulatory attribute label refers to an attribute mark reflecting the applicable financial regulatory category, financial regulatory status, and applicable period of the settlement account of the trading parties; the geopolitical risk attribute label refers to an attribute mark reflecting the applicable trade embargo category, trade restriction status, and applicable period of the goods delivery location; the rule hit identifier refers to the rule code corresponding to when the financial regulatory attribute label and the geopolitical risk attribute label meet the same cross-border risk rule triggering condition; the cross-border risk rule refers to a judgment rule pre-established based on the financial regulatory status in the offshore financial regulatory list and the trade restriction status in the trade embargo zone list, used to determine whether there is a cross-border risk association between the settlement account's location and the goods delivery location; and the risk verification result refers to the verification information recording whether the trading parties are marked as having cross-border risk and whether they have hit the cross-border risk rule.

[0084] In practice, firstly, the settlement account information and goods delivery address of the trading parties are retrieved from the cross-border trade documents. The country code and region code corresponding to the settlement account opening institution are then retrieved using a bank identification code parsing method, and these codes are used as the settlement account's place of origin. Next, a country name dictionary, administrative region code table, and port code table are used to match the country name, region name, and port name in the goods delivery address level by level, and the matched country code and region code are used as the goods delivery location. Finally, based on the cross-border transaction date, the offshore financial regulatory list and the trade embargo list are queried using the settlement account's place of origin, the goods delivery location, and the list's validity period as search fields. The financial regulatory information is then retrieved from the offshore financial regulatory list. The system first extracts the categories, financial regulatory status, and applicable periods from the list of prohibited trade areas, encoding these attributes as financial regulatory attribute labels. Then, it extracts the categories, status, and applicable periods from the offshore financial regulatory list and the list of prohibited trade areas, encoding these attributes as geopolitical risk attribute labels. Next, before performing risk verification, it extracts the financial regulatory categories, status, applicable regions, and applicable periods from the offshore financial regulatory list and the list of prohibited trade areas. The system uses combinations of attributes where both financial regulatory status and trade restriction status can simultaneously apply to the same cross-border transaction as rule conditions. Risk triggering states are set according to the restrictions recorded in the regulatory lists, where the settlement account's location is considered. When restrictions on settlement, enhanced review, or prohibition of settlement are applied, and the place of goods delivery is subject to prohibition, restriction of delivery, or licensing review requirements, the corresponding attribute combination is written into the rule record with cross-regional risk. When the settlement account's location and the place of goods delivery do not simultaneously meet the aforementioned restrictions, the corresponding attribute combination is written into the rule record without cross-regional risk. A unique rule code is assigned to each rule record, and the rule record is written with the financial regulatory category, financial regulatory status, trade embargo category, trade restriction status, applicable region, applicable period, and risk trigger status. All rule records are combined into a cross-regional risk rule. A field-by-field matching method is used to match the financial regulatory category, financial regulatory status, and applicable period in the financial regulatory attribute tag with the geopolitical risk. The cross-border risk rules are entered into the trade embargo category, trade restriction status, and applicable period in the risk attribute label. Rule records with consistent applicable regions, applicable periods covering the cross-border transaction dates, and consistent attribute values ​​are filtered out. The unique rule code obtained from the filtering is read and used as the rule hit identifier. Finally, the cross-border risk rule corresponding to the rule hit identifier is indexed, and the risk trigger status in the cross-border risk rule is read. When the risk trigger status is that cross-border risk exists, the trade parties are written into the record indicating that cross-border risk exists; when the risk trigger status is that cross-border risk does not exist, the trade parties are written into the record indicating that cross-border risk does not exist. The trade party identifier, rule hit identifier, and corresponding record are used together as the risk verification result.

[0085] In this embodiment, identifying cross-border transactions with settlement and delivery anomalies based on risk verification results and the cross-document business semantic graph can be achieved through the following steps:

[0086] Extract the trading parties marked as having cross-domain risks from the risk verification results, and locate the settlement account entity node associated with the trading parties in the cross-document business semantic graph;

[0087] Starting from the settlement account entity node, traverse the cross-document business semantic graph along the fund and goods association edge to construct a settlement and delivery association subgraph;

[0088] The settlement and delivery association subgraph is input into a pre-trained graph isomorphism discriminant network, and the topological deviation between the settlement and delivery association subgraph and the normal settlement and delivery mode subgraph is calculated.

[0089] When the deviation of the topology exceeds the normal deviation threshold, it is determined that there is an anomaly in the settlement and delivery of the cross-border transaction in which the trading parties participate.

[0090] It should be noted that, in this application, the settlement account entity node refers to the settlement account of the corresponding trade transaction party in the cross-document business semantic graph, and is used to connect the fund settlement relationship and the goods delivery relationship; the settlement delivery association subgraph represents a local graph structure reflecting the fund settlement relationship and the goods delivery relationship of the same cross-border transaction; the normal settlement delivery mode subgraph represents a reference graph structure in which the fund settlement relationship and the goods delivery relationship correspond to each other; the pre-trained graph isomorphism discriminant network represents a graph neural network that pre-trains network parameters using topologically consistent graph pairs and topologically inconsistent graph pairs, and outputs a graph-level structure representation based on node connection relationships, edge directions, and edge types; the topological deviation degree represents the degree of difference between the settlement delivery association subgraph and the normal settlement delivery mode subgraph in terms of node connection methods, edge directions, and edge types; the normal deviation threshold represents the upper limit of allowable deviation for the normal settlement delivery mode set according to the topological deviation degree distribution of historical normal cross-border transactions; the cross-border transaction with settlement delivery anomalies represents a cross-border transaction in which the topological deviation degree between the fund settlement relationship and the goods delivery relationship exceeds the normal deviation threshold.

[0091] It should also be noted that the pre-trained graph isomorphism discriminant network in this application can be obtained in the following way: From historical cross-border transactions that have been verified by business to have no settlement and delivery anomalies, normal settlement and delivery pattern subgraphs are obtained according to the same traversal range and fund and goods related edge extraction rules as the settlement and delivery related subgraphs. These normal settlement and delivery pattern subgraphs are then divided into training, validation, and test sets in an 8:1:1 ratio. The 8:1:1 ratio is set based on the principle that the training set needs to cover the main node connection methods and edge types, while retaining independent validation and test samples. The node arrangement order of the same normal settlement and delivery pattern subgraph is replaced, and the two subgraphs before and after the replacement are compared... To create topologically consistent graph pairs, while maintaining the number and types of nodes, the connection endpoints of the fund-goods related edges in the normal settlement and delivery mode subgraph are swapped, edge directions are reversed, or edge types are replaced. The modified subgraph and the original normal settlement and delivery mode subgraph are then considered as topologically inconsistent graph pairs, and the number of topologically consistent and inconsistent graph pairs is kept consistent. A graph isomorphism discrimination network is constructed using two graph isomorphism coding branches with shared network parameters. The entity type of cross-document unified entity nodes is converted into node input vectors, and the edge type and edge direction of the fund-goods related edges are converted into edge input vectors. Each graph isomorphism message passing layer in the graph isomorphism coding branch... The node input vectors of adjacent cross-document unified entity nodes are concatenated with their corresponding edge input vectors, and the concatenated vectors are summed and aggregated. The summed and aggregated results, along with the current node input vector of the cross-document unified entity node, are then input into the multilayer perceptron to update the node input vector of the current cross-document unified entity node. The number of graph isomorphic message passing layers is set by rounding up the 95th percentile of the shortest path length between the settlement account entity node in the training set and the cross-document unified entity node corresponding to the farthest goods, ensuring that the graph isomorphic coding branches can cover the main settlement and delivery associations in historical normal cross-border transactions. After completing all graph isomorphic message passing layer operations, a summed pooling method is used to aggregate the input vectors of each node. The output vectors of unified entity nodes across documents are aggregated into graph-level vectors. Contrastive loss is used to reduce the graph-level vector distance between graph pairs with consistent topologies, while increasing the graph-level vector distance between graph pairs with inconsistent topologies. The Adam optimization algorithm and backpropagation algorithm are used to update the network parameters. 0.0001, 0.0005, and 0.001 are selected as candidate initial learning rates, respectively. The candidate initial learning rate corresponding to the lowest validation set loss is used as the training parameter. When the validation set loss does not decrease for 10 consecutive training epochs, the network parameter update is stopped. The network parameters corresponding to the training epoch with the lowest validation set loss are written into the graph isomorphism discriminant network to complete the pre-training of the graph isomorphism discriminant network.The normal deviation threshold can be set as follows: Input each normal settlement and delivery pattern subgraph in the validation set into the pre-trained graph isomorphism discriminant network, read the graph-level vector corresponding to each normal settlement and delivery pattern subgraph, and calculate the Euclidean distance between each normal settlement and delivery pattern subgraph and each normal settlement and delivery pattern subgraph in the training set. Take the minimum Euclidean distance corresponding to each normal settlement and delivery pattern subgraph as the topological deviation degree of historical normal cross-border transactions. Arrange the topological deviation degrees of all historical normal cross-border transactions in the validation set in ascending order of value, and take the 95th percentile value as the normal deviation threshold to cover the main topological fluctuations in historical normal cross-border transactions and reduce the impact of a small number of extreme samples on the normal deviation threshold setting result.

[0092] In specific implementation, firstly, the system uses a verification mark field matching method to read the trade transaction party identifiers and settlement account numbers marked as having cross-domain risks from the risk verification results. It then calls the graph database node index to match the cross-document business semantic graph corresponding to the trade transaction party identifiers with cross-document unified entity nodes. Next, it reads the associated nodes connected by local semantic edges between the cross-document unified entity nodes, filters the associated nodes whose entity category is settlement account and whose account numbers are consistent, and uses the filtered associated nodes as settlement account entity nodes. Secondly, it uses a breadth-first traversal method to write the settlement account entity nodes into the access queue. It reads the cross-document unified entity nodes sequentially from the head of the access queue, only accessing adjacent cross-document unified entity nodes along the funds-goods association edge, and using the contract number, customs declaration number, invoice number, and bill of lading number as transaction association identifiers. When an adjacent cross-document unified entity node has at least one identical transaction association identifier with the current traversal range, it is written into the access queue. When all transaction association identifiers are inconsistent, the corresponding traversal branch is terminated until the access queue is empty. The settlement account entity nodes, cross-document unified entity nodes corresponding to the trading parties, cross-document unified entity nodes corresponding to the goods, and the fund-goods association edges between the cross-document unified entity nodes are all written into the same local graph structure, which is used as the settlement delivery association subgraph. Next, the settlement delivery association subgraph and each normal settlement delivery mode subgraph in the training set are input into the pre-trained graph isomorphism discrimination network. The graph isomorphism message passing layer aggregates the adjacency information layer by layer according to the entity category of the cross-document unified entity node, the edge type and edge direction of the fund-goods association edge, and outputs the graph-level vectors corresponding to the settlement delivery association subgraph and each normal settlement delivery mode subgraph through summation pooling. The distance between the graph-level vector of the settlement delivery association subgraph and the graph-level vector of each normal settlement delivery mode subgraph is calculated one by one using Euclidean distance. The Euclidean distance with the smallest value is taken as the topology deviation degree. Finally, the topology deviation degree is compared with the normal deviation threshold. When the topology deviation degree is greater than the normal deviation threshold, the cross-border transaction in which the trading parties participate is regarded as a cross-border transaction with settlement delivery abnormality.

[0093] In step S4, the customs declaration behavior characteristics of abnormal cross-border transactions in settlement and delivery are used as constraints. A routine business profile is constructed based on the constraints and the historical agency behavior data of the customs brokerage agency. Abnormal cross-border transactions of customs brokerage agency are identified through the routine business profile.

[0094] In this embodiment, using the customs declaration behavior characteristics of abnormal cross-border transactions as constraints, the following steps can be taken to construct a routine business profile based on the constraints and the historical agency behavior data of the customs brokerage agency:

[0095] Extract customs declaration behavior characteristics from customs declaration information of abnormal cross-border transactions with settlement and delivery;

[0096] Using the aforementioned customs declaration behavior characteristics as constraints in the data retrieval process, the customs declaration records of the customs brokerage agency within the historical period are retrieved to obtain historical agency behavior data.

[0097] Multi-peak distribution fitting is performed on the behavioral data in the historical agent behavior data to obtain different dense behavioral intervals and outlier boundaries;

[0098] Generate a routine business profile based on all densely populated behavioral areas and outlier boundaries.

[0099] It should be noted that the customs declaration behavior characteristics mentioned in this application represent a set of features reflecting the customs brokerage agency's customs declaration habits in terms of declaration scope, declaration time, and declaration scale; the constraints are derived from the customs brokerage agency identifier and declaration scope in the customs declaration behavior characteristics, and are used to limit the scope of retrieval of customs declaration records; the historical period is selected backward from the declaration date of abnormal cross-border transactions, and is used to cover the time interval of the customs brokerage agency's usual customs declaration activities; the historical agency behavior data represents the customs brokerage agency's customs declaration records and behavior dataset that meet the constraints within the historical period. The behavioral data refers to quantifiable data reflecting the timing and scale of customs declarations in historical agency behavior data; the multi-peak distribution fitting refers to the processing method of continuously fitting multiple high-frequency value segments in the behavioral data with probability density; the behavioral dense interval refers to the continuous value range of behavioral data that appears in the same distribution peak and covers the main historical customs declaration records; the outlier boundary refers to the numerical boundary at both ends of the behavioral dense interval used to distinguish between habitual customs declaration behavior and low-frequency deviation behavior; the habitual business profile refers to the business behavior description reflecting the concentrated distribution interval of the historical customs declaration behavior of the customs brokerage agency and its deviation boundary.

[0100] In practice, the process begins by using a data dictionary matching method to read customs declaration information for cross-border transactions with settlement and delivery anomalies. This involves reading the customs brokerage agency identifier from the customs brokerage agency field, the port of declaration from the customs declaration location field, the supervision method from the supervision method field, the commodity code from the commodity number field, the transportation method from the transportation method field, the declaration date from the declaration date field, and the declared amount from the total price field. The number of valid commodity records in the customs declaration commodity details is then counted as the commodity item count. Finally, the data is combined with the customs brokerage agency identifier, port of declaration, supervision method, commodity code, transportation method, declaration date, declared amount, and commodity item count. The set of features is used as the characteristics of customs declaration behavior. Secondly, the customs brokerage agency identifier, port of declaration, supervision method, commodity code, and mode of transport are extracted from these characteristics. Consistency of the customs brokerage agency identifier is set as the main search condition, while consistency of the port of declaration, supervision method, first four digits of the commodity code, and mode of transport are set as the business scope search conditions. These main search conditions and business scope search conditions are used together as constraints in the data retrieval process. A historical period of 12 months is selected, using the date corresponding to the declaration time in the customs declaration behavior characteristics as the time base. This 12-month period covers the possible occurrences of customs brokerage agencies within a full year. To track seasonal changes in customs brokerage, when fewer than 30 customs brokerage records meet the constraints, the historical period is extended forward by one month until 30 records are reached or the historical period is extended to 24 months. The 30 records provide a basic sample size for fitting a multi-peaked distribution, while the 24-month period limits the impact of earlier records on recent customs brokerage practices. A database composite index is used to retrieve customs brokerage records that simultaneously meet the constraints and the historical period. Each record is then read, including the customs brokerage agency identifier, port of entry, regulatory method, commodity code, mode of transport, declaration time, declared amount, and number of commodity items. The retrieved customs declaration records and their behavioral data are used as historical agency behavior data. Next, business data groups are formed based on the port of declaration, regulatory method, the first four digits of the commodity code, and the mode of transport within the historical agency behavior data. Within each business data group, the declaration time is converted into the cumulative number of minutes since 00:00 of the current day. The declaration amount is arranged by currency, and the integer value of the number of commodity items is retained. The Gaussian kernel density estimation method is used to fit the continuous probability density of the declaration time, declaration amount, and number of commodity items respectively. First, the Scott rule is used to obtain the initial bandwidth, and then 0.5 times, 0.75 times, 1 times, 1.25 times, and 1 times the initial bandwidth are applied.A candidate bandwidth of 5x was used. Leave-one-out cross-validation was employed to calculate the sample likelihood value corresponding to each candidate bandwidth. The candidate bandwidth with the highest sample likelihood value was used as the bandwidth for Gaussian kernel density estimation. This was to reduce the possibility of false peaks in individual samples due to excessively small bandwidth or the merging of adjacent distribution peaks due to excessively large bandwidth. Local peaks were marked according to the position where the fitted density changed from rising to falling. Different distribution peaks were divided according to the position with the lowest fitted density between adjacent local peaks. Within each distribution peak, the fitted density was accumulated from small to large based on the behavioral data values. The continuous value range covered by the cumulative probability from 2.5% to 97.5% was defined as the behavioral dense region. The lower and upper limits of the range are used as outlier boundaries. A 95% cumulative probability coverage area is used to retain the main historical customs brokerage records within the distribution peak while excluding a small number of low-frequency values ​​at both ends. Finally, a profile index is established using the customs brokerage agency identifiers in the historical agency behavior data. The declaration business scope corresponding to the customs brokerage behavior characteristics is recorded by the port of declaration, regulatory method, the first four digits of the commodity code, and the mode of transport. Under each declaration business scope, all densely populated intervals and outlier boundaries corresponding to the declaration time, declaration amount, and number of commodity items are written. The structured record containing the profile index, declaration business scope, densely populated intervals, and outlier boundaries is used as the routine business profile.

[0101] In this embodiment, identifying abnormal cross-border transactions involving customs brokerage through the routine business profile can be achieved using the following steps:

[0102] Based on the usual business profile, positive and negative examples of customs brokerage behavior are generated, and a behavior determination rule tree is constructed using the generated positive and negative examples.

[0103] The behavior determination rule tree is used to identify abnormal cross-border transactions by customs brokerage agents.

[0104] It should be noted that, in this application, positive examples of samples indicate that the customs declaration behavior characteristics are within the dense behavior range defined by the usual business profile, reflecting the behavior samples of the customs brokerage agency's usual customs declaration activities; negative examples of samples indicate that the customs declaration behavior characteristics exceed the outlier boundary defined by the usual business profile, reflecting the behavior samples of the customs brokerage agency that deviate from the usual customs declaration activities; the behavior judgment rule tree represents a tree structure that classifies customs declaration agency behavior according to the behavior judgment rules; and abnormal cross-border transactions of customs brokerage agency indicate cross-border transactions where the customs declaration behavior characteristics deviate from the concentrated distribution range of the customs brokerage agency's historical customs declaration activities.

[0105] In specific implementation, firstly, the system retrieves profile records from the routine business profile that match the customs brokerage agency identifier, declaration port, regulatory method, the first four digits of the commodity code, and transportation method. It then reads the behavioral density intervals and outlier boundaries corresponding to the declaration time, declaration amount, and number of commodity items, respectively. Next, it reads the historical agency behavior data used when constructing the routine business profile line by line. Historical agency behavior data where the declaration time, declaration amount, and number of commodity items all fall within the corresponding behavioral density interval are labeled as normal, and these are used as positive examples. Data where at least one of the declaration time, declaration amount, or number of commodity items exceeds the corresponding outlier boundary is excluded. Historical proxy behavior data that did not fall into other dense behavior intervals were labeled as anomalies, and these labeled historical proxy behavior data were used as negative examples. A stratified random sampling method was used to divide positive and negative examples into training and validation samples in an 8:2 ratio. This 8:2 ratio was used to retain the majority of samples for rule learning while providing independent samples to verify the classification results. A classification regression tree algorithm was used to construct a behavior determination rule tree, using the declaration time, declaration amount, and number of items as splitting fields, and the lower limit, upper limit, and outlier boundary of each dense behavior interval as candidate splitting values. The splitting field was calculated layer by layer under different candidate splitting values. The decrease in Gini impurity under a given value is used to determine the splitting field and candidate splitting value with the largest decrease in Gini impurity, which are then written into the splitting node. The maximum tree depth is set to 2 to 8, and the minimum number of samples per leaf node is set to 3%, 5%, and 10% of the total training samples, respectively. Different parameter combinations are validated group by group using validation samples. The parameter combination with the highest classification accuracy and the fewest nodes in the validation samples is used as the structural parameters of the behavior decision rule tree. When the number of positive samples in a leaf node is not less than the number of negative samples, a normal category label is written; when the number of negative samples in a leaf node is greater than the number of positive samples, an abnormal category label is written. The splitting node and splitting path are then defined. The tree-like rule structure containing path and category labels serves as the behavior determination rule tree. Secondly, the declaration time, declaration amount, and number of goods are read from the customs declaration behavior characteristics of abnormal cross-border transactions involving settlement and delivery. The conversion method for the declaration time, the currency classification method for the declaration amount, and the statistical method for the number of goods are kept consistent with those used when constructing the routine business profile. The declaration time, declaration amount, and number of goods are sequentially input into the behavior determination rule tree. Starting from the root node, the corresponding splitting path is entered according to the comparison results between the splitting field and the candidate splitting value, until the leaf node is reached. When an abnormal category label is written to the leaf node, the abnormal cross-border transaction involving settlement and delivery is classified as a cross-border transaction involving abnormal customs brokerage.

[0106] In step S5, a topological closed-loop determination is performed on the payment and receipt vouchers for abnormal cross-border transactions and the logistics delivery vouchers for abnormal cross-border customs brokerage transactions to generate an abnormal evidence chain in the foreign trade process, and the hidden documents with abnormal business content are identified through the abnormal evidence chain.

[0107] In this embodiment, the following steps can be used to perform topological closed-loop determination on payment and receipt vouchers for abnormal cross-border transactions involving settlement and delivery irregularities, and logistics delivery vouchers for abnormal cross-border transactions involving customs brokerage irregularities, to generate an abnormal evidence chain in the foreign trade process:

[0108] The fund transfer records and goods transfer records in the payment and receipt vouchers and the logistics delivery vouchers are projected onto a unified topological space according to the transaction subject and timestamp to form a discrete vector field of fund flow and goods flow;

[0109] The discrete vector field is subjected to Hodge decomposition to obtain the divergence-free curl field component, and the vortex region in the divergence-free curl field component with curl magnitude exceeding zero is extracted as the topological breakpoint.

[0110] For each topological breakpoint, calculate the loop integral of the capital flow vector and the cargo flow vector on the boundary loop of the topological breakpoint, and mark the boundary loop where the difference of the loop integral exceeds the preset closure tolerance as a closed loop break loop.

[0111] Based on all payment and receipt voucher identifiers and logistics delivery voucher identifiers involved in the area enclosed by each closed-loop break, the falsification relationship between vouchers is arranged according to the loop direction of the closed-loop break, thereby generating an abnormal evidence chain in the foreign trade process.

[0112] It should be noted that, in this application, the divergence-free curl field component refers to a vector field component in a discrete vector field that does not have net flow inflow or outflow and reflects circumferential flow transformation; the curl magnitude refers to a value reflecting the degree of circumferential deflection of the divergence-free curl field component within a local loop; the vortex region refers to a circumferential flow region in the divergence-free curl field component that has a continuous non-zero curl distribution; the topological breakpoint refers to a region where the capital flow relationship and the goods flow relationship exhibit local circumferential deviation during topological closure; the boundary loop refers to a closed path formed by connecting the head and tail of the directed edges around the topological breakpoint; and the loop product... The term "component" represents the cumulative flow of a vector along the boundary loop direction; the preset closure tolerance represents the upper limit of the allowable difference between the loop integrals of the fund flow vector and the goods flow vector in normal cross-border transactions; the closed-loop breakage loop represents the boundary loop where the difference between the loop integrals of the fund flow vector and the goods flow vector exceeds the preset closure tolerance; the falsification correlation represents the contradictory correlation between the fund flow content and the goods flow content recorded on different vouchers; the abnormal evidence chain represents the chain evidence structure connecting the payment voucher identifier, the logistics delivery voucher identifier, and the falsification correlation according to the loop direction of the closed-loop breakage loop.

[0113] It should also be noted that the preset closure tolerance in this application can be set in the following way: Select historical normal cross-border transactions that have been verified by business and have no conflict between the content of fund flow and the content of goods flow, form discrete vector fields of fund flow and goods flow according to the same topological projection method as the cross-border transactions to be identified, and process the historical normal cross-border transactions using the same Hodge decomposition, topological breakpoint extraction and boundary loop identification methods; calculate the loop integral difference between the fund flow vector and the goods flow vector on each boundary loop, arrange all loop integral differences in ascending order of value, and use the 95th percentile value as the preset closure tolerance. The 95th percentile value is used to cover the main closure fluctuations in historical normal cross-border transactions and reduce the impact of a small number of extreme records on the preset closure tolerance.

[0114] In specific implementation, firstly, the fund transfer records and goods transfer records in the payment and receipt vouchers and the logistics delivery vouchers are projected onto a unified topological space according to the transaction subject and timestamp, forming a discrete vector field of fund flow and goods flow; secondly, a node-edge association matrix is ​​established according to the connection relationship between each transaction subject and directed vector in the unified topological space, and the minimum cycle basis is extracted from the directed graph corresponding to the discrete vector field using the Horton minimum cycle basis algorithm, and an edge-cycle association matrix is ​​established according to the relationship that the direction of each directed vector is consistent with or opposite to the direction of the minimum cycle basis cycle; the node potential function corresponding to the node-edge association matrix is ​​solved using the sparse least squares method, and the node potential function is mapped to the directed vector and then from the discrete vector... The gradient field component is subtracted from the field, and the remaining vector is projected onto the ring space corresponding to the edge-ring correlation matrix. The projection result of the ring space is taken as the divergence-free curl field component. The direction values ​​of the divergence-free curl field component on each directed vector are accumulated along the ring direction of each minimum ring basis. The absolute value of the accumulated result is taken as the curl magnitude. Curl magnitudes lower than floating-point arithmetic precision are set to zero, and the minimum ring basis whose curl magnitude is still greater than 0 after being set to zero is marked as a vortex ring. Adjacent vortex rings sharing a transaction subject or a directed vector are merged using the connected component marking method. The merged continuous vortex region is taken as a topological breakpoint. Then, the outer directed vector occupied by only one vortex ring in each topological breakpoint is extracted, and the outer directed vectors are classified according to the starting transaction subject and the ending transaction subject. The transaction entities are connected sequentially to form a closed boundary loop. The capital flow vector and goods flow vector are read sequentially along the boundary loop direction. When the direction of either the capital flow vector or the goods flow vector is consistent with the direction of the boundary loop, the corresponding direction value is accumulated as a positive value; when the direction is opposite to the direction of the boundary loop, the corresponding direction value is accumulated as a negative value. The accumulated result of the capital flow vector is used as the loop integral of the capital flow vector, and the accumulated result of the goods flow vector is used as the loop integral of the goods flow vector. The absolute difference between the loop integrals of the capital flow vector and the goods flow vector is taken as the loop integral difference. The loop integral difference is numerically compared with a preset closure tolerance. When the loop integral difference is large... When the preset closure tolerance is met, the corresponding boundary loop is marked as a closed-loop break loop. Finally, according to the loop direction of the closed-loop break loop, the payment and receipt voucher identifier or logistics delivery voucher identifier associated with each directed vector is read sequentially, and adjacent voucher identifiers are connected in the order of the previous voucher identifier pointing to the next voucher identifier. The voucher connection relationship where the fund transfer record lacks a corresponding goods transfer record, the goods transfer record lacks a corresponding fund transfer record, or the fund transfer direction is opposite to the goods transfer direction is used as the falsification association relationship. The closed-loop break loop identifier, all payment and receipt voucher identifiers, all logistics delivery voucher identifiers, and falsification association relationships are written into the chain structure in the loop direction, and the chain structure is used as the abnormal evidence chain in the foreign trade process.

[0115] In this embodiment, projecting the fund transfer records and goods transfer records in the payment vouchers and logistics delivery vouchers onto a unified topological space according to the transaction subject and timestamp to form a discrete vector field of fund flow and goods flow can be achieved by the following steps:

[0116] Based on the fund transfer records and goods transfer records in the payment and receipt vouchers and the logistics delivery vouchers, a cross-domain transfer graph is constructed with the transaction entity as the node and the fund flow and goods flow as the transfer edges;

[0117] The cross-domain flow graph is input into a pre-trained graph attention network to generate cross-domain fusion embedding vectors for each transaction entity;

[0118] The vector space spanned by the cross-domain fusion embedding vectors of all trading entities is used as the unified topological space;

[0119] Using the cross-domain fusion embedding vectors of each trading entity in a unified topological space as node coordinates, the direction of capital flow and the direction of goods flow as edge directions, and the timestamp interval as edge weights, all directed edges are mapped to directed vectors in a unified topological space, thus forming a discrete vector field of capital flow and goods flow.

[0120] It should be noted that, in this application, the cross-domain flow graph refers to a graph structure used to associate transaction entities and their flow relationships in fund flow records and goods flow records; the pre-trained graph attention network refers to a graph neural network that has completed network parameter training in advance through historical cross-domain flow graphs and assigns adjacency information weights according to the flow relationships between transaction entities; the cross-domain fusion embedding vector refers to a vector representation of the fund flow, goods flow, and timestamp interval associated with the transaction entities; the unified topological space refers to a vector space spanned by the cross-domain fusion embedding vectors of all transaction entities, used to uniformly carry the topological relationships of fund flow and goods flow; the node coordinates refer to the position vectors of the transaction entities in the unified topological space; the edge direction refers to the flow direction of fund flow or goods flow from the outflowing transaction entity to the inflowing transaction entity; the timestamp interval refers to the time difference between the outflow time and the inflow time in the same fund flow record or goods flow record; the edge weight refers to the edge attribute used to reflect the time required for the fund flow or goods flow to complete the corresponding flow process; and the discrete vector field refers to a set of discrete vectors reflecting the distribution of the flow direction and time interval of fund flow and goods flow in the unified topological space.

[0121] It should also be noted that the pre-trained graph attention network in this application can be obtained in the following way: Select historical cross-border transactions for which fund transfer records and goods transfer records have been verified, construct a historical cross-domain transfer graph according to the same node merging method and transfer edge connection method as the cross-border transactions to be identified, and divide the historical cross-domain transfer graph into training set, validation set and test set in a ratio of 8:1:1; read the number of fund inflows, fund outflows, goods inflows and goods outflows associated with the transaction entity from each historical cross-domain transfer graph, normalize the reading results and use them as the node input vector of the transaction entity, and encode the fund flow or goods flow category, flow direction and timestamp interval corresponding to the transfer edge as the edge input vector; use a graph attention network containing a graph attention layer and an edge reconstruction layer, randomly occlude some transfer edges in the training set, and have the graph attention layer calculate attention weights according to the node input vector of the current transaction entity, the node input vectors of adjacent transaction entities and the corresponding edge input vectors, and then cluster according to the attention weights. The node input vectors of adjacent trading entities are aggregated, and the edge reconstruction layer restores the existence state, fund flow or goods flow category, and flow direction of the obscured flow edges based on the aggregated node vectors. The obscuration ratio of the flow edges is set to 10%, 15%, and 20%, the number of graph attention layers is set to 2, 3, and 4, and the number of attention heads is set to 2, 4, and 8. The reconstruction error rate of the obscured flow edges is compared group by group using the validation set. The flow edge obscuration ratio, the number of graph attention layers, and the number of attention heads corresponding to the lowest reconstruction error rate are used as network parameters. The Adam optimization algorithm and the backpropagation algorithm are used to update the network parameters. The candidate initial learning rates are 0.0001, 0.0005, and 0.001, respectively. The candidate initial learning rate corresponding to the lowest reconstruction error rate of the validation set is used as the training parameter. When the reconstruction error rate of the validation set does not decrease for 10 consecutive training rounds, the network parameter update is stopped. The network parameters corresponding to the training round with the lowest reconstruction error rate of the validation set are written into the graph attention network to complete the pre-training of the graph attention network.

[0122] In specific implementation, firstly, the payment transaction entity, receiving transaction entity, payment time, arrival time, and payment / receipt voucher identifier are read from the payment / receipt voucher using a voucher field mapping method. Secondly, the goods delivery transaction entity, goods receiving transaction entity, dispatch time, receipt time, and logistics delivery voucher identifier are read from the logistics delivery voucher. Thirdly, duplicate transaction entities are merged according to their transaction entity identifiers. The merged transaction entities are then used as nodes. Connections from the payment transaction entity to the receiving transaction entity are used as the flow edges corresponding to the fund flow, and connections from the goods delivery transaction entity to the goods receiving transaction entity are used as the flow edges corresponding to the goods flow. Finally, the payment time, arrival time, and payment / receipt voucher identifier are written into the flow edges corresponding to the fund flow. First, the dispatch time, receipt time, and logistics delivery certificate identifier are written into the flow edges corresponding to the goods flow. The graph structure with completed node connections and flow edge attribute writing is used as a cross-domain flow graph. Second, the number of fund inflows, fund outflows, goods inflows, and goods outflows associated with each transaction entity are counted from the cross-domain flow graph. The statistical results are scaled to a value range of 0 to 1 according to the maximum corresponding number in the cross-domain flow graph, and the scaled results are used as the node input vector of the transaction entity. The fund flow or goods flow category, flow direction, and timestamp interval are converted into edge input vectors. The timestamp interval corresponding to the fund flow is the time difference between the receipt time and the payment time, and the timestamp interval corresponding to the goods flow is the time difference between the receipt time and the payment time. The time difference between receiving and sending times is calculated. The node input vector, edge input vector, and node connections in the cross-domain flow graph are input into a pre-trained graph attention network. Each graph attention layer aggregates the node input vectors of adjacent trading entities layer by layer according to the attention weights corresponding to the flow edges. The node vector output by the last graph attention layer is read, and this node vector is used as the cross-domain fusion embedding vector for the corresponding trading entity. Next, the cross-domain fusion embedding vectors of all trading entities are arranged according to the same vector dimension. Each dimension of the cross-domain fusion embedding vector is used as the corresponding coordinate component, and the vector range that can be linearly represented by all cross-domain fusion embedding vectors is used as a unified topological space. Finally, the cross-domain fusion of each trading entity is... The embedded vector serves as the node coordinates of the transaction entities in a unified topological space. For the flow edges corresponding to the fund flow, the node coordinates of the paying entity are used as the starting coordinates, and the node coordinates of the receiving entity are used as the ending coordinates. The coordinate difference vector from the starting coordinates to the ending coordinates is used as the directed vector corresponding to the fund flow, and the time difference between the arrival time and the payment time is used as the edge weight. For the flow edges corresponding to the goods flow, the node coordinates of the goods delivery entity are used as the starting coordinates, and the node coordinates of the goods receiving entity are used as the ending coordinates. The coordinate difference vector from the starting coordinates to the ending coordinates is used as the directed vector corresponding to the goods flow, and the time difference between the receipt time and the dispatch time is used as the edge weight.Retain the payment / receipt voucher or logistics delivery voucher corresponding to each directed vector, and combine all directed vectors corresponding to fund flows and goods flows into a discrete vector field for fund flows and goods flows.

[0123] In this embodiment, the identification of concealed documents indicating abnormal business content through the abnormal evidence chain can be achieved through the following steps:

[0124] Extract the voucher identifier associated with the closed-loop break loop from the abnormal evidence chain, and perform distributed index matching in the cross-border trade document data lake to obtain the business content text that matches the cross-border trade document.

[0125] The business content text is input into a pre-trained trade domain language model to generate a semantic vector, and an anomaly pattern query vector is constructed based on the falsification correlation of the anomaly evidence chain. The semantic representation of anomaly perception is generated through cross-attention fusion.

[0126] The semantic representation of anomaly perception is input into a preset anomaly detection network to calculate the anomaly score of cross-border trade documents. Cross-border trade documents with anomaly scores exceeding a preset judgment threshold are identified as hidden documents with abnormal business content.

[0127] It should be noted that, in this application, the cross-border trade document data lake refers to a data storage set used to uniformly store cross-border trade documents, business content text, and voucher identifier index relationships; the business content text refers to the text content in the cross-border trade documents that carries information on the transaction subject, goods, settlement, customs declaration, and delivery; the pre-trained trade domain language model refers to a language model that has been pre-trained using cross-border trade document text to extract the semantics of the business content; the semantic vector refers to a vector representation reflecting the transaction object, business matter, and semantic relationship in the business content text; the abnormal pattern query vector refers to a query vector reflecting the voucher connection direction and the type of contradiction in the business content in the falsification relationship; and the cross-attention fusion refers to the fusion of cross-attention according to the correlation between the abnormal pattern query vector and the semantic vector. The processing method involves weighted fusion of business content semantics; the semantic representation of anomaly perception is a vector representation reflecting the abnormal semantic features of business content text under the influence of falsification association; the preset anomaly detection network is a detection network that is pre-trained using parameters from normal cross-border trade documents and cross-border trade documents with abnormal business content, and is used to output the degree of business content anomaly; the anomaly score is a numerical value reflecting the degree of deviation of the business content of cross-border trade documents from the normal business semantic distribution; the preset judgment threshold is an anomaly score boundary used to distinguish between normal cross-border trade documents and concealed documents with abnormal business content; the concealed documents with abnormal business content refer to cross-border trade documents whose business content is not easily abnormal when verified alone, but presents business semantic contradictions under the association of abnormal evidence chains.

[0128] It should also be noted that the pre-trained trade domain language model in this application can be obtained in the following way: Historical cross-border trade documents with complete business content text and verified business data are selected from the cross-border trade document data lake. The text is segmented according to document sections and field boundaries, and duplicate text and unrecognizable characters are removed. The processed business content text is then divided into training, validation, and test sets in an 8:1:1 ratio. A BERT-based structure, including 12 Transformer encoding layers, 768-dimensional hidden vectors, and 12 attention heads, is used as the trade domain language model. 15% of the tokens are randomly selected from the training set as mask objects, of which 80%... The mask objects are replaced with mask markers, 10% of the mask objects are replaced with random words, and the remaining 10% of the mask objects retain their original words. The mask objects are recovered using context words, and the network parameters are updated using the backpropagation algorithm. 0.00001, 0.00003, and 0.00005 are selected as candidate initial learning rates, respectively. The candidate initial learning rate corresponding to the lowest mask word prediction loss on the validation set is used as the training parameter. When the mask word prediction loss on the validation set does not decrease for 10 consecutive training epochs, the network parameter update is stopped. The network parameters corresponding to the training epoch with the lowest mask word prediction loss on the validation set are written into the trade domain language model to complete the pre-training of the trade domain language model.

[0129] It should also be noted that the anomaly detection network preset in this application can be obtained in the following way: Historical cross-border trade documents that have been verified to have no abnormal business content are selected as normal samples, and historical cross-border trade documents that have been verified to have inconsistencies in business content between documents are selected as abnormal samples. Normal samples and abnormal samples are divided into training set, validation set, and test set in a ratio of 8:1:1. The semantic representations of anomaly perception corresponding to normal samples and abnormal samples are obtained respectively, using the same trade domain language model invocation method and cross-attention fusion method as the cross-border trade documents to be identified. A feedforward neural network including an input layer, two fully connected layers, and an output layer is used as the anomaly detection network. The output dimension of the first fully connected layer is set to half the dimension of the semantic representation of anomaly perception, and the output dimension of the second fully connected layer is set to half the dimension of the semantic representation of anomaly perception. The output dimension of the connecting layer is set to 1 / 4 of the semantic representation dimension of the anomaly detection layer. Both fully connected layers use the ReLU activation function, and the output layer uses the Sigmoid activation function to output anomaly scores in the range of 0 to 1. The anomaly detection network is trained using class-weighted binary cross-entropy loss. The class weights corresponding to normal and abnormal samples are set according to the reciprocal of their respective sample numbers to reduce the impact of sample number differences on network parameter updates. The parameters of the anomaly detection network, the linear mapping layer corresponding to the anomaly pattern query vector, and the cross-attention fusion layer are updated synchronously using the Adam optimization algorithm and the backpropagation algorithm. When the validation set loss does not decrease for 10 consecutive training epochs, parameter updates are stopped, and the parameters corresponding to the lowest validation set loss in the training epochs are written into the anomaly detection network to obtain the preset anomaly detection network. The preset judgment threshold can be set in the following way: normal samples and abnormal samples in the validation set are sequentially input into the preset anomaly detection network, and the corresponding anomaly scores are read respectively. The range of 0.01 to 0.99 is set as candidate thresholds in intervals of 0.01. The accuracy, precision, and recall corresponding to each candidate threshold are calculated one by one. The candidate threshold with the highest F1 value is taken as the preset judgment threshold. When there are multiple candidate thresholds with the same F1 value, the candidate threshold with the highest precision is taken as the preset judgment threshold, so as to reduce the proportion of normal cross-border trade documents being misidentified as hidden documents with abnormal business content.

[0130] In practical implementation, firstly, all document identifiers sequentially connected in the abnormal evidence chain are read according to the loop direction of the closed-loop broken loop. The document identifier is used as a distributed index key to calculate the data partition number. The document identifier is then routed to the corresponding data partition in the cross-border trade document data lake. The storage address of the cross-border trade document associated with the document identifier is read using the hash index within the data partition. Then, the main text and field text are read based on the cross-border trade document storage address. Finally, the main text and field text associated with each document identifier are arranged according to the loop direction of the closed-loop broken loop. The arranged text content is used as the business content for matching cross-border trade documents. First, the text content is divided into segments according to document sections and field boundaries. Then, the segments are converted into word sequences using the same word encoding method as the pre-training process of the trade domain language model. These word sequences are input into the pre-trained trade domain language model. All word vectors output from the last Transformer encoding layer are read, and mean pooling is performed on the effective word vectors. The pooling result is used as the semantic vector corresponding to each text segment. Second, the falsification relationships associated with each document identifier are read from the abnormal evidence chain. Third, the contradiction types of business content and documents within the falsification relationships are identified. The connection direction, forward document type, and backward document type are each converted into one-hot encoding. All one-hot encodings are concatenated and input into a linear mapping layer. An output vector with the same dimension as the semantic vector is read and used as the anomaly pattern query vector. This anomaly pattern query vector is then input into the query projection end of the cross-attention fusion layer. All semantic vectors corresponding to the same matching cross-border trade document are input into the key projection end and value projection end, respectively. Attention weights are calculated using the dot product of the anomaly pattern query vector and each semantic vector. All semantic vectors are then weighted and summed according to these attention weights. The weighted summation result is then compared with the anomaly pattern query vector. The concatenated pattern query vectors are input into a linear mapping layer. The vector output by the linear mapping layer is used as the semantic representation of the anomaly detection corresponding to the cross-border trade document. Finally, the semantic representation of the anomaly detection corresponding to each cross-border trade document is sequentially input into a preset anomaly detection network. After passing through two fully connected layers and an output layer, the output values ​​in the range of 0 to 1 are read and used as the anomaly score of the corresponding cross-border trade document. Each anomaly score is compared with a preset judgment threshold. When the anomaly score is greater than the preset judgment threshold, the cross-border trade document corresponding to the anomaly score is regarded as a hidden document with anomalies in business content.

[0131] For example, to verify the role of abnormal evidence chains in identifying documents with concealed business content, 240 sets of samples containing business inconsistencies between documents can be set from the aforementioned verification data. These samples are divided into four categories: subject-related conflicts, amount-related conflicts, performance sequence conflicts, and fund-goods flow conflicts. Each type of anomaly is jointly manifested by at least two cross-border trade documents of different types. Identification is performed using both a direct document comparison method without topological loop determination and the abnormal evidence chain method of this application. (Reference) Figure 3 As shown in the figure, this is a comparison of the closed-loop identification effect of the abnormal evidence chain provided in this application. After constructing the abnormal evidence chain, the closed-loop identification rates for the four types of anomalies are 91.5%, 94.2%, 89.8%, and 93.4%, respectively, while those without constructing the abnormal evidence chain are 68.7%, 72.4%, 65.9%, and 59.6%. Among them, the identification rate of conflicts in the flow of funds and goods increased by 33.8 percentage points, indicating that by projecting the fund transfer records and goods transfer records onto a unified topological space, and arranging the payment voucher identifiers, logistics delivery voucher identifiers, and falsification associations based on the closed-loop broken loops, it is possible to connect local contradictions scattered in different documents into traceable chain evidence, thereby improving the ability to identify hidden anomalies that are not easily discovered based on the business content of a single document.

[0132] Therefore, this application demonstrates that the abnormal evidence chain can identify concealed documents with abnormal business content. Firstly, by acquiring cross-border trade documents from different sources and performing entity identification and semantic association on the trading entities, commodity information, transaction amounts, settlement information, customs declaration information, and logistics information in each document, interference caused by differences in name spelling, field structure, and record standards can be eliminated, allowing document content scattered across different business systems to establish semantic connections based on the same cross-border transaction. Secondly, by performing cross-domain risk verification on the settlement account location and goods delivery location of the trading parties, and combining this with cross-document business semantic graph verification of the relationships between settlement entities, consignor / consignee entities, and goods flow, abnormal transactions where the fund settlement location and the actual goods delivery location do not conform to normal trade logic can be identified, preventing discrepancies between fund flow and goods flow caused by partial information in a single document being within a reasonable range. Furthermore, by limiting the verification object to customs declaration information of abnormal cross-border transactions with settlement and delivery, and constructing a customary [system / mechanism] based on the historical agency behavior data of customs brokerage agencies, [the application can be further refined]. Business profiling allows for the correlation and comparison of current customs brokerage activities with existing agents, declared goods, agency regions, and business frequencies. This provides further verification of anomalies discovered during the settlement and delivery process at the customs brokerage level, while reducing misjudgments of abnormal transactions caused by occasional regional differences. Finally, by cross-domain mutual verification and topological closed-loop determination of payment and receipt vouchers for abnormal cross-border transactions and logistics delivery vouchers for abnormal cross-border customs brokerage transactions, it is possible to verify whether various vouchers have a complete business continuity relationship along the stages of fund settlement, customs brokerage, cargo transportation, and actual delivery. Mutually corroborating or contradictory abnormal information is linked into an abnormal evidence chain. Based on this, the document identifiers and business content associated with each abnormal relationship can be traced through the abnormal evidence chain. This allows for the identification of specific documents with mismatched entities, inconsistent amounts, inconsistent cargo quantities, or broken logistics delivery relationships from relevant cross-border trade documents. This enables the identification of hidden documents that conceal business contradictions through methods such as name variations, quantity splitting, amount allocation, replacement of settlement entities, and changes in logistics routes.

[0133] In summary, the technical solution adopted in this application can realize the consistency verification of business content among multiple types of trade documents, thereby improving the ability to identify hidden document anomalies.

[0134] Example 2

[0135] This application provides a foreign trade risk identification system based on big data analysis, with reference to... Figure 4 As shown in the figure, this is a module structure diagram of a foreign trade risk identification system based on big data analysis according to this embodiment of the present application. The foreign trade risk identification system includes:

[0136] Cross-border trade document acquisition module 100 is used to acquire different cross-border trade documents;

[0137] The cross-document semantic graph construction module 200 is used to perform entity recognition and semantic analysis on the business content of all cross-border trade documents and generate a cross-document business semantic graph.

[0138] The settlement and delivery risk screening module 300 is used to perform cross-domain risk verification on the settlement account location and the goods delivery location of the trading parties, and to identify cross-border transactions with settlement and delivery abnormalities based on the risk verification results and the cross-document business semantic graph.

[0139] The customs brokerage agent profile recognition module 400 is used to construct a routine business profile based on the customs declaration information of abnormal cross-border transactions with settlement and delivery as constraints, and the historical agency behavior data of the customs brokerage agent. The abnormal cross-border transactions of the customs brokerage agent are identified through the routine business profile.

[0140] The cross-domain document closed-loop judgment module 500 is used to perform topological closed-loop judgment based on cross-domain mutual verification on payment and receipt vouchers for abnormal cross-border transactions and logistics delivery vouchers for abnormal cross-border customs brokerage transactions, generate abnormal evidence chains in the foreign trade process, and identify hidden documents with abnormal business content through the abnormal evidence chains.

[0141] It should also be noted that the reference Figure 5 As shown in the figure, this is an application scenario diagram of the foreign trade risk identification method based on big data analysis provided in this application. Cross-border trade data sources include commercial invoices, packing lists, bills of lading, certificates of origin, customs declarations, payment and receipt vouchers, and logistics delivery vouchers. These cross-border trade documents are integrated into a cross-border trade document data lake after data access, and data processing is supported by distributed storage, computing engines, graph computing engines, and model training services. The foreign trade risk identification system first performs optical character recognition, business entity extraction, and semantic association processing on each cross-border trade document to construct a cross-document business semantic graph. Subsequently, it performs cross-domain risk verification by combining the settlement account's location, the place of goods delivery, the offshore financial regulatory list, and the trade embargo zone list, and identifies abnormal settlement and delivery transactions through graph isomorphism discrimination. Furthermore, the system constructs a profile of routine business based on the characteristics of abnormal customs declaration behavior and historical agency behavior data to identify abnormal customs declaration agency transactions; then it performs topological closed-loop determination on payment and receipt vouchers and logistics delivery vouchers to generate an abnormal evidence chain, and retrieves related documents in the cross-border trade document data lake based on the abnormal evidence chain, outputting settlement and delivery abnormal transactions, abnormal customs declaration agency transactions, and hidden documents with abnormal business content, for customs supervisors, risk analysts, and management decision-makers to conduct risk assessment.

[0142] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine such that, when executed on the processor of the computer or other programmable data processing apparatus, they produce means for implementing the functions specified by one or more blocks in the flowchart illustrations and / or block diagrams.

[0143] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compactdisc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0144] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

Claims

1. A method for identifying foreign trade risks based on big data analysis, characterized in that, Includes the following steps: Obtain different cross-border trade documents; Perform entity identification and semantic analysis on the business content of all cross-border trade documents to generate a cross-document business semantic graph. Cross-domain risk verification is performed on the settlement account location and goods delivery location of the trading parties, and cross-border transactions with settlement and delivery anomalies are identified based on the risk verification results and the cross-document business semantic graph. Using the customs declaration behavior characteristics of abnormal cross-border transactions in settlement and delivery as constraints, a routine business profile is constructed based on the constraints and the historical agency behavior data of the customs brokerage agency. The abnormal cross-border transactions of customs brokerage agency are identified through the routine business profile. A topological closed-loop determination is performed on payment and receipt vouchers for abnormal cross-border transactions and logistics delivery vouchers for abnormal cross-border customs brokerage transactions to generate an abnormal evidence chain in the foreign trade process, and the hidden documents with abnormal business content are identified through the abnormal evidence chain.

2. The method for identifying foreign trade risks based on big data analysis as described in claim 1, characterized in that, Entity identification and semantic analysis are performed on the business content of all cross-border trade documents to generate a cross-document business semantic graph, specifically including: Optical character recognition and section classification are performed on all cross-border trade documents. Business entities are extracted using a trade term attention bias named entity recognition model, and local semantic edges based on section affiliation are established within the cross-border trade documents. By using a multilingual thesaurus of trade entities and cross-document co-occurrence constraint rules, business entities that refer to the same transaction object are linked and merged to generate a unified entity node across documents. Based on the unified entity node across documents, the trade terms, timestamps and amount fields in the associated cross-border trade documents are extracted to construct the execution time relationship edge and the fund and goods association edge across documents. Using the cross-document unified entity node as the node, and the local semantic edge, the performance time sequence relation edge, and the fund and goods association edge as the edge, a cross-document business semantic graph is generated.

3. The method for identifying foreign trade risks based on big data analysis as described in claim 1, characterized in that, Cross-regional risk verification of the settlement account location and goods delivery location of trading parties specifically includes: By querying the offshore financial regulatory list and the trade embargo zone list, we can obtain the financial regulatory attribute label of the settlement account location of the trade transaction parties and the geopolitical risk attribute label of the goods delivery location. Cross-matching is performed between the financial regulatory attribute tags and the geopolitical risk attribute tags to generate rule hit identifiers; The risk verification result is obtained by performing risk verification on the rule hit identifier through preset cross-domain risk rules.

4. The method for identifying foreign trade risks based on big data analysis as described in claim 1, characterized in that, Based on the risk verification results and the cross-document business semantic graph, the specific cross-border transactions with settlement and delivery anomalies identified include: Extract the trading parties marked as having cross-domain risks from the risk verification results, and locate the settlement account entity node associated with the trading parties in the cross-document business semantic graph; Starting from the settlement account entity node, traverse the cross-document business semantic graph along the fund and goods association edge to construct a settlement and delivery association subgraph; The settlement and delivery association subgraph is input into a pre-trained graph isomorphism discriminant network, and the topological deviation between the settlement and delivery association subgraph and the normal settlement and delivery mode subgraph is calculated. When the deviation of the topology exceeds the normal deviation threshold, it is determined that there is an anomaly in the settlement and delivery of the cross-border transaction in which the trading parties participate.

5. The method for identifying foreign trade risks based on big data analysis as described in claim 1, characterized in that, Using the customs declaration behavior characteristics of abnormal cross-border transactions in settlement and delivery as constraints, and constructing a routine business profile based on the constraints and historical agency behavior data of customs brokerage agencies, specifically includes: Extract customs declaration behavior characteristics from customs declaration information of abnormal cross-border transactions with settlement and delivery; Using the aforementioned customs declaration behavior characteristics as constraints in the data retrieval process, the customs declaration records of the customs brokerage agency within the historical period are retrieved to obtain historical agency behavior data. Multi-peak distribution fitting is performed on the behavioral data in the historical agent behavior data to obtain different dense behavioral intervals and outlier boundaries; Generate a routine business profile based on all densely populated behavioral areas and outlier boundaries.

6. The method for identifying foreign trade risks based on big data analysis as described in claim 1, characterized in that, The cross-border transactions identified as abnormal by customs brokerage through the aforementioned routine business profile include: Based on the usual business profile, positive and negative examples of customs brokerage behavior are generated, and a behavior determination rule tree is constructed using the generated positive and negative examples. The behavior determination rule tree is used to identify abnormal cross-border transactions by customs brokerage agents.

7. The method for identifying foreign trade risks based on big data analysis as described in claim 1, characterized in that, A topological closed-loop determination is performed on payment and receipt vouchers for abnormal settlement and delivery in cross-border transactions and logistics delivery vouchers for abnormal customs brokerage transactions to generate an abnormal evidence chain in the foreign trade process, specifically including: The fund transfer records and goods transfer records in the payment and receipt vouchers and the logistics delivery vouchers are projected onto a unified topological space according to the transaction subject and timestamp to form a discrete vector field of fund flow and goods flow; The discrete vector field is subjected to Hodge decomposition to obtain the divergence-free curl field component, and the vortex region in the divergence-free curl field component with curl magnitude exceeding zero is extracted as the topological breakpoint. For each topological breakpoint, calculate the loop integral of the capital flow vector and the cargo flow vector on the boundary loop of the topological breakpoint, and mark the boundary loop where the difference of the loop integral exceeds the preset closure tolerance as a closed loop break loop. Based on all payment and receipt voucher identifiers and logistics delivery voucher identifiers involved in the area enclosed by each closed-loop break, the falsification relationship between vouchers is arranged according to the loop direction of the closed-loop break, thereby generating an abnormal evidence chain in the foreign trade process.

8. The method for identifying foreign trade risks based on big data analysis as described in claim 7, characterized in that, Projecting the fund transfer records and goods transfer records in the payment vouchers and logistics delivery vouchers onto a unified topological space according to the transaction subject and timestamp to form a discrete vector field of fund flow and goods flow specifically includes: Based on the fund transfer records and goods transfer records in the payment and receipt vouchers and the logistics delivery vouchers, a cross-domain transfer graph is constructed with the transaction entity as the node and the fund flow and goods flow as the transfer edges; The cross-domain flow graph is input into a pre-trained graph attention network to generate cross-domain fusion embedding vectors for each transaction entity; The vector space spanned by the cross-domain fusion embedding vectors of all trading entities is used as the unified topological space; Using the cross-domain fusion embedding vectors of each trading entity in a unified topological space as node coordinates, the direction of capital flow and the direction of goods flow as edge directions, and the timestamp interval as edge weights, all directed edges are mapped to directed vectors in a unified topological space, thus forming a discrete vector field of capital flow and goods flow.

9. The method for identifying foreign trade risks based on big data analysis as described in claim 1, characterized in that, The specific documents used to identify concealed evidence of abnormal business content through the aforementioned abnormal evidence chain include: Extract the voucher identifier associated with the closed-loop break loop from the abnormal evidence chain, and perform distributed index matching in the cross-border trade document data lake to obtain the business content text that matches the cross-border trade document. The business content text is input into a pre-trained trade domain language model to generate a semantic vector, and an anomaly pattern query vector is constructed based on the falsification correlation of the anomaly evidence chain. The semantic representation of anomaly perception is generated through cross-attention fusion. The semantic representation of anomaly perception is input into a preset anomaly detection network to calculate the anomaly score of cross-border trade documents. Cross-border trade documents with anomaly scores exceeding a preset judgment threshold are identified as hidden documents with abnormal business content.

10. A foreign trade risk identification system based on big data analysis, used to execute a foreign trade risk identification method based on big data analysis as described in any one of claims 1 to 9, characterized in that, The foreign trade risk identification system includes: The cross-border trade document collection module is used to acquire different cross-border trade documents; The cross-document semantic graph construction module is used to perform entity recognition and semantic analysis on the business content of all cross-border trade documents, and generate a cross-document business semantic graph. The settlement and delivery risk screening module is used to perform cross-domain risk verification on the settlement account location and the goods delivery location of the trading parties, and to identify cross-border transactions with settlement and delivery abnormalities based on the risk verification results and the cross-document business semantic graph. The customs brokerage agent profile recognition module is used to construct a routine business profile based on the customs declaration information of abnormal cross-border transactions with settlement and delivery as constraints, and the historical agency behavior data of the customs brokerage agent. The abnormal cross-border transactions of the customs brokerage agent are identified through the routine business profile. The cross-domain document closed-loop judgment module is used to perform topological closed-loop judgment based on cross-domain mutual verification on payment and receipt documents for abnormal cross-border transactions and logistics delivery documents for abnormal cross-border customs brokerage transactions. It generates an abnormal evidence chain in the foreign trade process and identifies hidden documents with abnormal business content through the abnormal evidence chain.