Data verification method and device for cross-border documents, electronic device and storage medium
By using cross-border knowledge graph mapping and real-time tariff monitoring, the problems of low efficiency and high error rate in manual review of cross-border documents have been solved, realizing intelligent and automated processing of cross-border documents, reducing compliance risks, and improving the efficiency and security of global trade.
Patent Information
- Application Number
- CN202511141240.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-15
AI Technical Summary
The processing of cross-border warehousing documents relies on manual review, resulting in low efficiency and a high error rate. Furthermore, traditional technologies lack the ability to dynamically integrate and intelligently verify tariff data from various regions, making it impossible to adapt to tariff changes or regulatory updates in real time and increasing compliance risks.
By extracting key paragraphs from cross-border documents and mapping them to a pre-defined cross-border knowledge graph, a three-level mapping relationship (commodity description - harmonized system code - local extension code) is obtained. Real-time monitoring of tariff updates is performed, mandatory authentication requests are obtained, and change logs are saved through blockchain to build an intelligent cross-border document verification system.
It has improved the accuracy and efficiency of cross-border document processing, reduced missed detections and misjudgments, realized the intelligent, automated and standardized processing of cross-border documents, reduced compliance risks, and improved the efficiency and security of global trade.
Smart Images

Figure CN120653648B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data verification technology for cross-border documents, and in particular to a data verification method, apparatus, electronic device, and storage medium for cross-border documents. Background Technology
[0002] Currently, the processing of cross-border warehousing documents mainly relies on manual review. However, the tariffs, HS codes, and certification requirements of different regions are complex and constantly changing. Operators are prone to errors during manual matching and verification, leading to customs clearance delays and increased fines. Manual review has obvious drawbacks such as low efficiency and high error rate. Traditional technologies lack the ability to dynamically integrate and intelligently verify tariff data from various regions, and cannot adapt to tariff changes or regulatory updates in real time, resulting in verification errors and increased compliance risks. Summary of the Invention
[0003] Therefore, it is necessary to address the existing data verification problem of cross-border documents by proposing a data verification method, device, electronic equipment, and storage medium for cross-border documents.
[0004] A data verification method for cross-border documents, the method comprising:
[0005] Extract key paragraphs from specified cross-border documents;
[0006] The key paragraphs are mapped into a pre-defined cross-border knowledge graph to obtain the corresponding three-level mapping relationship; wherein, the three-level mapping relationship is product description - harmonized system code - local extension code;
[0007] The cross-border documents are verified based on the aforementioned three-level mapping relationship;
[0008] Before the step of obtaining the corresponding three-level mapping relationship in the preset cross-border knowledge graph through the key paragraph mapping, the method further includes:
[0009] Obtain tariff data from multiple preset regions;
[0010] Extract the preset three-level mapping relationship from the tariff data;
[0011] The preset cross-border knowledge graph is constructed based on the preset three-level mapping relationship.
[0012] Furthermore, after the step of constructing the preset cross-border knowledge graph based on the preset three-level mapping relationship, the method further includes:
[0013] Real-time monitoring of the tariff update status of each of the preset regions;
[0014] If the tariff update status of the target preset region is detected to be updated, then the updated tariff of the target preset region is obtained;
[0015] Based on the updated tariff, the target three-level mapping relationship is extracted and the preset three-level mapping relationship corresponding to the preset cross-border knowledge graph is updated.
[0016] Furthermore, after the step of constructing the preset cross-border knowledge graph based on the preset three-level mapping relationship, the method further includes:
[0017] Obtain the mandatory authentication requests for each of the preset regions;
[0018] Based on the mandatory authentication request, obtain the three-level mapping relationship to be authenticated corresponding to the preset cross-border knowledge graph;
[0019] Determine whether the three-level mapping relationship to be authenticated satisfies the mandatory authentication request;
[0020] If the three-level mapping relationship to be authenticated does not meet the mandatory authentication request, a change prompt is generated based on the three-level mapping relationship to be authenticated and the mandatory authentication request;
[0021] Obtain the changed content based on the change prompt;
[0022] Based on the changes, the three-level mapping relationship to be certified is modified to form a compliant three-level mapping relationship;
[0023] The compliance three-level mapping relationship is input into the preset cross-border knowledge graph to update the corresponding three-level mapping relationship to be certified.
[0024] Furthermore, after the step of modifying the three-level mapping relationship to be authenticated based on the changed content to form a compliant three-level mapping relationship, the method further includes:
[0025] Obtain the identity of the operator who modified the three-level mapping relationship to be authenticated based on the modified content, as well as the time of the modification.
[0026] A change log is generated based on the change time, operator identity, the three-level mapping relationship to be authenticated, the three-level mapping relationship for compliance, and the change prompt.
[0027] The change log is uploaded to a pre-defined blockchain for storage.
[0028] Furthermore, prior to the step of extracting key paragraphs from specified cross-border documents, the following steps are also included:
[0029] Collect multiple abnormal scenarios in advance to build an abnormal pattern library;
[0030] Construct structured rule templates based on the aforementioned anomaly pattern library;
[0031] The specified cross-border documents are inspected based on the structured rule template to obtain non-compliant data and corresponding anomaly levels;
[0032] Based on the correspondence table between the anomaly level and the preset level and the data acquisition method, the target data acquisition method is obtained;
[0033] Obtain the compliant data corresponding to the non-compliant data based on the target data acquisition method, and replace the non-compliant data in the cross-border document with the compliant data.
[0034] Furthermore, the step of extracting key paragraphs from a specified cross-border document includes:
[0035] The specified cross-border documents are divided into multilingual regions using a preset YOLOv5 model;
[0036] Obtain the corresponding OCR processing model based on the language type of each of the aforementioned language regions;
[0037] The corresponding language regions are processed by each of the OCR processing models to obtain sub-text data corresponding to each language region;
[0038] The individual sub-text data are concatenated to obtain the cross-border document text data;
[0039] Key paragraphs are extracted from the cross-border document text data using a pre-defined large language model.
[0040] Furthermore, before the step of concatenating the various sub-text data to obtain the cross-border document text data, the method further includes:
[0041] Each of the sub-text data is input into a preset context-aware model to correct each part of the sub-text data, resulting in corrected sub-text data.
[0042] A data verification device for cross-border documents, the device comprising:
[0043] The first extraction module is used to extract key paragraphs from specified cross-border documents;
[0044] The mapping module is used to map the key paragraphs into a preset cross-border knowledge graph to obtain the corresponding three-level mapping relationship; wherein the three-level mapping relationship is product description - harmonized system code - local extension code;
[0045] The verification module is used to verify the cross-border documents based on the three-level mapping relationship;
[0046] The acquisition module is used to acquire tariff data from multiple preset regions;
[0047] The second extraction module is used to extract the preset three-level mapping relationship from the tariff data;
[0048] A construction module is used to construct the preset cross-border knowledge graph based on the preset three-level mapping relationship.
[0049] An electronic device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps:
[0050] Extract key paragraphs from specified cross-border documents;
[0051] The key paragraphs are mapped into a pre-defined cross-border knowledge graph to obtain the corresponding three-level mapping relationship; wherein, the three-level mapping relationship is product description - harmonized system code - local extension code;
[0052] The cross-border documents are verified based on the aforementioned three-level mapping relationship;
[0053] Before the step of obtaining the corresponding three-level mapping relationship in the preset cross-border knowledge graph through the key paragraph mapping, the method further includes:
[0054] Obtain tariff data from multiple preset regions;
[0055] Extract the preset three-level mapping relationship from the tariff data;
[0056] The preset cross-border knowledge graph is constructed based on the preset three-level mapping relationship.
[0057] A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the following steps:
[0058] Extract key paragraphs from specified cross-border documents;
[0059] The key paragraphs are mapped into a pre-defined cross-border knowledge graph to obtain the corresponding three-level mapping relationship; wherein, the three-level mapping relationship is product description - harmonized system code - local extension code;
[0060] The cross-border documents are verified based on the aforementioned three-level mapping relationship;
[0061] Before the step of obtaining the corresponding three-level mapping relationship in the preset cross-border knowledge graph through the key paragraph mapping, the method further includes:
[0062] Obtain tariff data from multiple preset regions;
[0063] Extract the preset three-level mapping relationship from the tariff data;
[0064] The preset cross-border knowledge graph is constructed based on the preset three-level mapping relationship.
[0065] The beneficial effects of this invention are as follows: By extracting key paragraphs from specified cross-border documents and mapping them to a preset cross-border knowledge graph, the verification process can quickly obtain the three-level mapping relationship between product description, harmonized system code, and local extension code, thereby improving the accuracy of document processing and making compliance verification more efficient. It also reduces omissions and misjudgments caused by manual review, realizing the intelligent, automated, and standardized processing of cross-border documents, and giving global trade higher efficiency and security. Attached Figure Description
[0066] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0067] in:
[0068] Figure 1 This is an application environment diagram of a data verification method for cross-border documents in one embodiment;
[0069] Figure 2 This is a flowchart of a data verification method for cross-border documents in one embodiment;
[0070] Figure 3 This is a structural block diagram of a data verification device for cross-border documents in one embodiment;
[0071] Figure 4 This is a structural block diagram of an electronic device in one embodiment. Detailed Implementation
[0072] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0073] Figure 1 This is a diagram illustrating a data validation application environment for cross-border documents in one embodiment. (Refer to...) Figure 1This cross-border document data verification method is applied to a cross-border document data verification system. The system includes a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal; a mobile terminal can be at least one of a mobile phone, tablet, or laptop. The server 120 can be a standalone server or a server cluster consisting of multiple servers. The terminal 110 is used to retrieve specified cross-border documents, and the server 120 is used to verify these documents.
[0074] like Figure 2 As shown, in one embodiment, a data verification method for cross-border documents is provided. This method can be applied to both terminals and servers; this embodiment uses terminal application as an example. The data verification method for cross-border documents specifically includes the following steps:
[0075] S1: Extract key paragraphs from specified cross-border documents;
[0076] S2: Obtain the corresponding three-level mapping relationship by mapping the key paragraphs into the preset cross-border knowledge graph; wherein, the three-level mapping relationship is product description - harmonized system code - local extension code;
[0077] S3: Verify the cross-border documents based on the aforementioned three-level mapping relationship;
[0078] Before step S2, which involves mapping the key paragraphs to a preset cross-border knowledge graph to obtain the corresponding three-level mapping relationship, the method further includes:
[0079] S101: Obtain tariff data for multiple preset regions;
[0080] S102: Extract the preset three-level mapping relationship from the tariff data;
[0081] S103: Construct the preset cross-border knowledge graph based on the preset three-level mapping relationship.
[0082] As described in step S1, key paragraphs of the specified cross-border documents are extracted. The input cross-border documents are parsed to identify important information. Key paragraphs typically include product descriptions, quantities, invoice numbers, shipper and consignee information, etc. Specifically, optical character recognition (OCR) technology can be used to convert paper or scanned documents into editable text format. In addition, combined with natural language processing (NLP) technology, the system can identify and extract key fields related to finance and trade through specific semantic models, such as named entity recognition (NER) models. To ensure the accuracy of extraction, the recognition results can be post-processed, such as data cleaning and format correction, thereby improving data consistency.
[0083] As described in step S2 above, the key paragraphs are mapped into a preset cross-border knowledge graph to obtain the corresponding three-level mapping relationship. After extracting the key paragraphs, this information is mapped into the preset cross-border knowledge graph. The preset cross-border knowledge graph typically contains data structures for relevant data and trade compliance in various regions, mainly including a three-level mapping relationship of product description, Harmonized System Code (HS Code), and local extension code. First, the corresponding entry is searched in the knowledge graph based on the extracted product description. This process involves querying the knowledge graph and fine-grained comparison to ensure accurate matching of the corresponding product's HS Code and its local extension code. In this mapping process, machine learning techniques can be used to optimize the matching algorithm to improve accuracy. In addition, to address the differences in language and regulations between different regions, the system needs to dynamically adapt to the compliance requirements of each region to ensure the effectiveness and timeliness of the mapping relationship.
[0084] As described in step S3 above, the cross-border documents are verified based on the three-level mapping relationship. After obtaining the three-level mapping relationship, the system will verify the cross-border documents based on these relationships. The system compares the pre-set compliance logic with the extracted data to check whether the key information in the documents complies with the tariff requirements of different regions. If the product description and corresponding HS code and local extension code in the document correctly match the relevant regulations, the system will determine that the document is valid and proceed to the next processing step. If a discrepancy is found, the system will automatically generate a warning message to prompt relevant personnel to supplement or correct the information.
[0085] As described in step S101 above, tariff data from multiple preset regions is obtained. In practical applications, to ensure accurate verification of cross-border documents, the system first needs to obtain the latest tariff data from multiple preset regions. When the tariff data is updated, it triggers an automatic update of the knowledge graph, involving the collection of relevant information from official data agencies or trade databases in each region. This process can be carried out through API interfaces or by periodically downloading relevant information from officially released databases. This tariff data should include the latest HS codes, commodity classifications, tariff rates, etc. In addition, the system can also format and standardize the acquired data so that subsequent steps can effectively utilize this information. In this way, a comprehensive and timely updated tariff data foundation is established to ensure sufficient data support for subsequent knowledge graph construction.
[0086] As described in step S102 above, the preset three-level mapping relationship is extracted from the tariff data. After obtaining the tariff data, the system needs to extract the required preset three-level mapping relationship from it. Specifically, this can be done by parsing the tariff data obtained from various regions and identifying structured information such as commodity descriptions, corresponding Harmonized System codes, HS codes, and their local extension codes. Text mining and data column extraction techniques, combined with regular expressions or pattern matching algorithms, are used to quickly extract and classify the required information.
[0087] As described in step S103 above, a preset cross-border knowledge graph is constructed based on the preset three-level mapping relationship. This preset cross-border knowledge graph will serve as the data foundation for the entire verification system, providing a structured model of the relationship between various commodities and their applicable regional tariffs. During construction, graph database technology may be used, storing data in the form of nodes and edges. Nodes represent commodity descriptions, HS codes, and local extension codes, while edges represent the connections between elements. This process needs to consider flexibility and scalability to accommodate the addition of new regional tariffs or commodity categories in the future. Simultaneously, to improve query efficiency, the graph design also needs to conform to data indexing and retrieval optimization strategies. Establishing an effective index structure improves the response speed of subsequent queries. Ultimately, the resulting cross-border knowledge graph will provide real-time and accurate data support for the intelligent verification of subsequent documents.
[0088] In one embodiment, after step S103 of constructing the preset cross-border knowledge graph based on the preset three-level mapping relationship, the method further includes:
[0089] S1041: Real-time monitoring of the tariff update status of each of the preset regions;
[0090] S1042: If the tariff update status of the target preset area is detected to be updated, then the updated tariff of the target preset area is obtained;
[0091] S1043: Extract the target three-level mapping relationship based on the updated tariff and update the preset three-level mapping relationship corresponding to the preset cross-border knowledge graph.
[0092] As described in steps S1041-S1043 above, a real-time monitoring mechanism is established to continuously track and acquire tariff update information from multiple preset regions. This process can be achieved in various ways. For example, the system can periodically access the official customs or relevant data bureau websites of each region to view the latest announcements and dynamic information. To improve monitoring efficiency, API interfaces can be used to automatically acquire tariff-related data from each region. In addition to directly retrieving web page information, the system can also utilize email subscriptions or RSS subscription services to receive notifications of relevant update information. Furthermore, databases published by each region can also serve as data sources. The frequency of updates monitored typically depends on the characteristics of tariff changes in each region; therefore, the inspection frequency for different regions can be customized to capture change information more efficiently. Through this real-time monitoring mechanism, the system can ensure the acquisition of the latest tariff dynamics, preparing for subsequent update processing and thereby reducing compliance risks and the possibility of misjudgments caused by tariff changes.
[0093] The system identifies specific target regions related to monitored tariff updates. When a region's tariff status is confirmed as "updated," the system initiates an automatic process to acquire the latest tariff data. This process typically involves calling pre-defined APIs or data scraping modules to access the latest tariff information provided by that country. It's crucial to ensure the acquired data format is compatible with the system's internal data structure for effective parsing and storage. Furthermore, the system must ensure data integrity and accuracy when acquiring updated tariffs. This includes comparing changes between old and new data and examining specific legal documents or announcements to ensure the acquired update information is valid and compliant. For some regions, updated tariff information may contain multiple sub-items or appendices; therefore, a pre-defined data analysis model is needed to extract these contents and integrate them into existing data tables. Using the newly acquired updated tariff information, new three-level mapping relationships are extracted and updated accordingly with the existing cross-border knowledge graph. First, the system needs to parse the updated tariff data, identify and extract the required three-level mapping relationships. Simultaneously, the system must also consider the existing knowledge graph to determine which mappings need updating. For example, the HS code of a product may be replaced by a new classification in the new tariff schedule, a change that directly affects the tariff map structure. After these changes are completed, the tariff map will be updated to reflect the new tariff information. By systematically monitoring the tariff update status of each region, obtaining the new tariff schedule, and extracting the corresponding three-level mapping relationships, the entire verification process can remain up-to-date and highly compliant, thereby improving the accuracy of cross-border document processing and reducing compliance risks.
[0094] In one embodiment, after step S103 of constructing the preset cross-border knowledge graph based on the preset three-level mapping relationship, the method further includes:
[0095] S1141: Obtain the mandatory authentication request for each of the preset regions;
[0096] S1142: Based on the mandatory authentication request, obtain the three-level mapping relationship to be authenticated corresponding to the preset cross-border knowledge graph;
[0097] S1143: Determine whether the three-level mapping relationship to be authenticated satisfies the mandatory authentication request;
[0098] S1144: If the three-level mapping relationship to be authenticated does not meet the mandatory authentication request, a change prompt is generated based on the three-level mapping relationship to be authenticated and the mandatory authentication request;
[0099] S1145: Obtain the changed content based on the change prompt;
[0100] S1146: Modify the three-level mapping relationship to be certified based on the changes to form a compliant three-level mapping relationship;
[0101] S1147: Input the compliance three-level mapping relationship into the preset cross-border knowledge graph to update the corresponding three-level mapping relationship to be certified.
[0102] As described in step S1141 above, it is necessary to identify and collect mandatory certification requests for specific commodities in each preset region. This involves collecting relevant regulations and requirements from official channels of regulatory agencies or trade management bodies in each region. It should be noted that tariff data refers to relevant information from official data agencies or trade databases in each region, including the latest HS codes, commodity classifications, and tariff rates. Mandatory certification requests, on the other hand, are mandatory rules set separately by each region and are not reflected in the tariff data. For example, Region A requires imported electronic products to meet rule a, while Region B requires the product to meet rule b. Such certification requests typically include detailed regulations and required document types, such as product manuals and safety test reports. Obtaining this information requires a dynamic update mechanism to reflect policy changes in each region in a timely manner, ensuring that the certification requests used by the system are up-to-date. Since these mandatory certification requests are generally published through regional websites, industry associations, or international trade organizations, the system can be configured with an API interface or periodically access relevant websites to automatically extract relevant data. Furthermore, the acquired data needs to be standardized to ensure that certification request information from different regions can enter subsequent processing stages in the same format. This process ensures that the products to be certified can be effectively aligned with the relevant certification requirements in subsequent steps, reducing compliance risks caused by incomplete or inaccurate information.
[0103] As described in step S1142 above, the three-level mapping relationship to be certified corresponding to the preset cross-border knowledge graph is obtained based on the mandatory certification request. The obtained mandatory certification request is matched with the three-level mapping relationship to be certified in the preset cross-border knowledge graph, which contains three-level mapping information such as product description, harmonization regime, and local extension code. Information related to the mandatory certification request is extracted. For example, the system needs to identify which product categories meet the mandatory certification request of a specific region, and then obtain the corresponding three-level mapping relationship. Specifically, the knowledge graph can be queried, and necessary information can be extracted according to the description of the certification request. This process may also involve the use of natural language processing (NLP) technology to ensure that the system can understand the complexity of the mandatory certification request and the accuracy of the matching relationship information. For example, the system may refine the broad category of "electronic devices" to confirm which of all devices need to meet specific certification conditions. In this way, the three-level mapping relationship to be certified can be clearly identified, providing a basis for subsequent compliance judgment and potential changes.
[0104] As described in step S1143 above, it is determined whether the three-level mapping relationship to be certified meets the mandatory certification request. The evaluation of whether the three-level mapping relationship to be certified conforms to the corresponding mandatory certification request includes comparing the product description, HS code, and local extension code to be certified with the requirements in the mandatory certification request. For example, if the HS code of a product requires specific certification, but the mapping relationship to be certified for that product does not contain this certification information, the system will determine that the relationship to be certified does not meet the requirements. Generally, the system uses conditional judgment logic to check each field in the relationship to be certified one by one to confirm its completeness and compliance. To improve the accuracy of the judgment, the system can also integrate a dynamic rule engine to analyze and apply the compliance logic and certification requirements of different regions, ensuring a comprehensive evaluation of the relationship to be certified. If some requirements are found to be missing or mismatched during this process, the system will issue a corresponding alarm and initiate subsequent change prompt generation steps to ensure that compliance issues can be identified and handled in a timely manner.
[0105] As described in step S1144 above, if the three-level mapping relationship to be certified does not meet the mandatory certification request, a change prompt is generated based on the three-level mapping relationship to be certified and the mandatory certification request. When the system identifies that certain conditions in the certification request are not met, it automatically generates a change prompt. These prompts will indicate which specific requirements are not met, thus reminding relevant personnel to make necessary adjustments and additions. The process of generating change prompts involves comparing the relationship to be certified with the mandatory certification request and identifying missing or incorrect fields. For example, if the certification request requires a specific compliance document, but the document type is not listed in the relationship to be certified, the system will automatically generate a prompt asking the user to add "missing M code". These change prompts will be presented in a user-friendly format, such as displaying each unmet compliance item in a list, ensuring that relevant personnel can quickly understand and take appropriate measures. This process not only improves the intelligence and response speed of the system but also reduces the burden of manual verification, ensuring that compliance issues can be handled quickly, thereby guaranteeing the smooth operation of the entire cross-border transaction.
[0106] As described in step S1145 above, the change content is obtained based on the change prompt. According to the generated change prompt, the specific change content is further obtained. This change content will include all the information and data that needs to be supplemented, clearly showing the specific items that need to be corrected in the three-level mapping relationship to be certified. For example, if the system finds that the relationship to be certified lacks a specific certification document, such as "missing C field," this step will require the user to provide the corresponding certification information. To obtain this change content, the system may first perform a detailed analysis of each prompt to ensure that the system understands the type of missing information and its importance. This process may include calling the APIs of relevant databases or regulatory agencies to review the required document list or requirements. Alternatively, the system may prompt the user to obtain necessary documents from official channels, such as sending a prompt requiring the user to upload new certification codes, compliance documents, or other necessary supporting materials. The accuracy and comprehensiveness of this process are also necessary conditions to ensure the smooth completion of subsequent correction operations, avoiding correction failures and subsequent verification delays due to incomplete information.
[0107] As described in step S1146 above, the three-level mapping relationship to be certified is modified based on the changed content to form a compliant three-level mapping relationship. Using the acquired changed content, the three-level mapping relationship to be certified is modified as necessary to ensure it complies with all relevant mandatory certification requests. This process involves not only simple data replacement but may also require regenerating the entire mapping relationship based on new information. For example, if the M-code is missing from the relationship to be certified, it is added to the relevant field. After this change is completed, the system must ensure that the format and structure of all updated data still conform to the preset standards so that it can continue to be used effectively in subsequent steps. Furthermore, the system can trigger an automatic review mechanism to ensure that the converted compliant three-level mapping relationship complies with the latest legal and regulatory requirements, further reducing compliance risks. Through this process, the final compliant three-level mapping relationship will be considered a verified state, ensuring that the same compliance issues will not be encountered in future trade activities.
[0108] As described in step S1147 above, the compliance three-level mapping relationship is input into the preset cross-border knowledge graph to update the corresponding three-level mapping relationship to be certified. The modified compliance three-level mapping relationship will be re-input into the preset cross-border knowledge graph to replace the original three-level mapping relationship to be certified. This update process maintains the corrected data state in the knowledge graph, enabling the entire system to operate effectively based on the latest information. The update process typically needs to ensure data compatibility and correctness so that the resulting mapping relationship can adapt to future query and verification needs. The verification process may include a series of automatically triggered data integrity checks to ensure that the new data meets format and logical requirements. To ensure continuous monitoring of compliance, this process can also record all change operations for future audits and compliance checks. The new compliance three-level mapping relationship can provide real-time compliance support for subsequent document processing, ensuring that relevant goods always meet the legal and regulatory requirements of each region in cross-border transactions, thereby reducing customs clearance risks, improving overall transaction efficiency, and providing comprehensive information support for intelligent verification and compliance management of cross-border documents.
[0109] In one embodiment, after step S1146, which modifies the three-level mapping relationship to be authenticated based on the changed content to form a compliant three-level mapping relationship, the method further includes:
[0110] S11471: Obtain the identity of the operator who made the change to the three-level mapping relationship to be authenticated based on the changed content, and the time of the change;
[0111] S11472: Generate a change log based on the change time, operator identity, the three-level mapping relationship to be authenticated, the three-level mapping relationship for compliance, and the change prompt;
[0112] S11473: Upload the change log to a preset blockchain for storage.
[0113] As described in steps S11471-S11473 above, record the identity of the operator performing the change operation and the time of the operation. Obtaining this information is not only for auditing and compliance needs, but also helps track change records, ensuring that any data adjustments or corrections can be audited and investigated. Operator identity typically includes the user's unique identifier (such as user ID) and relevant role (such as auditor, administrator), which helps identify the responsible party when performing the operation in subsequent queries. The change time ensures that all modifications have a clear timestamp, forming a timeline of change records, which helps analyze the operation and decision-making process in subsequent compliance audits. This information may be automatically extracted through the system's user authentication module (such as SSO single sign-on) and stored in the data change log. Integrate the collected information to generate a detailed change log, which should include: change time, operator identity, original three-level mapping relationship to be authenticated, updated compliance three-level mapping relationship, and change notification content. The generation of the change log is to create an audit-friendly and transparent record that details the process and background of data modifications, making subsequent compliance verification and issue tracing easier. When generating change logs, this information needs to be formatted into standardized text or structured data for easy storage and retrieval. Simultaneously, this log must adhere to certain security standards to ensure it cannot be altered or deleted by unauthorized users, thereby reducing the risk of data tampering. Effective management of change logs is not only crucial for ensuring data transparency but also key to meeting relevant compliance requirements. The generated change logs are uploaded to a pre-defined blockchain for persistent storage. Leveraging the immutability of blockchain technology, all change records, once added to the blockchain, cannot be modified or deleted, ensuring data integrity and reliability. Uploading change logs to the blockchain typically involves encrypting the log data to ensure that only authorized personnel can access the relevant information, while also allowing for access control through blockchain technology. The upload process can be implemented through smart contracts, ensuring automatic execution of the upload operation when conditions are met. Each record in the blockchain carries a timestamp and operator identification information, ensuring that every change can be traced back to the responsible person and specific time. This process not only enhances data security but also meets relevant requirements, and is significant for the legitimacy of audit roles and the transparency of data processing. In summary, uploading change logs to the blockchain is a crucial step in ensuring the compliance and security of cross-border document data processing, forming a trusted audit chain for the entire system.
[0114] In one embodiment, prior to step S1 of extracting key paragraphs from a specified cross-border document, the method further includes:
[0115] S001: Collect multiple abnormal scenarios in advance to build an abnormal pattern library;
[0116] S002: Construct a structured rule template based on the aforementioned anomaly pattern library;
[0117] S003: Based on the structured rule template, the specified cross-border documents are inspected to obtain non-compliant data and corresponding anomaly levels;
[0118] S004: Based on the correspondence table between the anomaly level and the preset level and the data acquisition method, obtain the target data acquisition method;
[0119] S005: Obtain the compliant data corresponding to the non-compliant data based on the target data acquisition method, and replace the non-compliant data in the cross-border document with the compliant data.
[0120] As described in steps S001-S003 above, extensive data analysis and research are conducted to collect data on common anomaly scenarios in cross-border documents, such as over 300 common anomaly scenarios. This process can be achieved by analyzing historical data, audit records, and business feedback, with the aim of identifying frequently occurring errors, compliance issues, and legal risks during document processing. By studying the characteristics of non-compliant documents, the system can effectively summarize key anomaly scenarios, such as "HS code errors" or "inconsistencies between invoice amount and customs declaration content." A comprehensive anomaly pattern library is formed, storing various types of anomaly samples and providing various available attributes for subsequent analysis and classification. This data can be constructed from historical business data and information collection systems, or by referencing regional or industry standards to ensure the comprehensiveness and relevance of the anomaly pattern library. The purpose of establishing the anomaly pattern library is to enhance the system's intelligent identification capabilities, enabling it to capture potential problems in data processing through clearly defined anomaly patterns, thereby achieving more automated compliance verification and intelligent decision-making.
[0121] The collected anomaly scenarios are transformed into structured rule templates. This aims to represent complex anomaly patterns in a standardized format, facilitating subsequent system detection and verification. The design of these structured rule templates employs a unified format and syntax, ensuring high readability and maintainability. Simultaneously, they must be compatible with the mapping relationships in the knowledge graph to effectively support subsequent data detection and verification in practical applications. Based on the previously established structured rule templates, automated compliance checks are performed on specified cross-border documents. First, the system extracts relevant field information from the documents (such as product description, quantity, HS code, etc.). Then, using the rule engine, it applies each preset rule template to determine whether the extracted data complies with relevant laws and regulations and the company's internal policies.
[0122] During the detection process, the system records data items corresponding to defects and assigns an anomaly level to each non-compliance issue. Anomaly levels are typically categorized based on the severity of the violation, such as "minor violation," "moderate violation," and "serious violation." This assessment process may consider multiple factors, such as the impact of the violation on the compliance of the relevant data, the legal risks involved, and the frequency of the error. Recording these non-compliant data and anomaly levels is crucial for subsequent appropriate corrective actions and risk monitoring. Automated detection not only improves audit efficiency but also ensures consistency in the screening process and reduces subjective biases inherent in human operations. By identifying and recording non-compliant data and their anomaly levels, the system lays the foundation for subsequent response measures and data correction.
[0123] As described in steps S004-S005 above, the optimal method for obtaining compliant data is determined based on the anomaly level of the acquired non-compliant data, combined with a preset correspondence table between anomaly levels and data acquisition methods. This preset correspondence table represents a rule-based mapping relationship, defining the data correction strategy and data source required for rectification at a given anomaly level. For example, when an anomaly level is identified as a simple anomaly (such as a format error), it is automatically corrected by the system; complex issues (such as cross-border HS code conflicts) are transferred to manual review and recorded in the case library. In this way, the system can flexibly select appropriate data acquisition methods (such as querying the business database, contacting external partners to obtain certificates or reports, etc.) based on the specific anomaly type and severity, promoting efficient implementation of subsequent rectification measures. Based on the determined target data acquisition method, the corresponding compliant data for the non-compliant data is obtained. For example, if an HS code in a cross-border document is marked as "incorrect," but a correct HS code is obtained through an external query as replacement data, the system will perform a replacement operation for the non-compliant item. Specifically, the system interacts with relevant data sources to obtain compliance information, such as contacting relevant certification bodies to obtain updated compliance documents or retrieving verified data from internal databases. Simultaneously, the system performs data validation to ensure that the obtained compliant data can effectively replace the original non-compliant data. After data replacement, the system updates the document content and records all operations that replace non-compliant data with compliant data for subsequent auditing and verification. Through this step, the system not only ensures the compliance of documents but also lays a solid foundation for the effectiveness of subsequent processing. This enables cross-border documents to achieve largely intelligent and automated management while complying with international trade regulations and requirements, enhancing the stability and transparency of the entire process.
[0124] In one embodiment, step S1 of extracting key paragraphs from a specified cross-border document includes:
[0125] S111: Divide the specified cross-border document into multilingual regions using a preset YOLOv5 model;
[0126] S112: Obtain the corresponding OCR processing model according to the language type of each of the aforementioned language regions;
[0127] S113: Process the corresponding language regions through each of the OCR processing models to obtain sub-text data corresponding to each language region;
[0128] S114: Concatenate the individual sub-text data to obtain the cross-border document text data;
[0129] S115: Extract key paragraphs from the cross-border document text data using a preset large language model.
[0130] As described in steps S111-S113 above, a preset YOLOv5 model is used to perform object detection on the specified cross-border documents, identifying multilingual regions. YOLOv5 is a highly efficient object detection algorithm that can quickly locate different objects in an image through training. The model is trained using a dataset containing 100,000 labeled images of multilingual documents. In the scenario of cross-border documents, targets include text blocks in different languages and various graphic elements (such as stamps, barcodes, icons, etc.). First, the system passes the input document image to the YOLOv5 model for processing. The model extracts features through its convolutional neural network (CNN) and identifies text regions in the image based on the existing training data. For each identified region, the model not only returns the location (boundary box) but also provides a confidence score for each region to determine whether it is a target region. After detection, the system will label the document based on the recognition results, dividing it into various language regions. The system will analyze the language type of each identified text region and select an appropriate optical character recognition (OCR) processing model for each region. A multilingual OCR model library can be predefined, with each model specifically optimized for processing text in a particular language. For F-documents, language is automatically separated and OCR is performed by region, integrating variant recognition models. Generally, languages written from left to right may use specialized OCR models, while mainstream languages can use common OCR tools (such as Tesseract). The system selects and loads the appropriate OCR processing model based on the characteristics of the identified text region and language identifiers (e.g., Unicode range, characteristic character set). The main task of OCR technology is to convert character information in images into editable text. During processing, the system passes the image of each language region to the corresponding OCR model for parsing. This process generates text blocks, such as extracting key information like invoice numbers, product names, and prices. During execution, the system also performs post-processing steps, such as text correction and formatting, to clean up errors that may occur during the OCR process (e.g., character recognition confusion, spelling errors). The resulting subtext data will be stored in a structured format for easy subsequent analysis and processing.
[0131] As described in S114-S115 above, the previously extracted sub-text data from various language regions are integrated and concatenated into a complete cross-border document text data. The system needs to rationally arrange the concatenation order of each sub-text based on the text differentiation of different languages and the layout of the original document. During concatenation, texts can be connected based on preset templates or rules, or directly according to the identified regional order, ensuring the final text maintains a natural order. To consider the contextual structure of the text, the system should also check the grammatical and logical integrity of the concatenation during this process, ensuring the final result conforms to language conventions. After concatenation, the synthesized cross-border document text data will serve as the basis for subsequent key paragraph extraction. A preset Large Language Model (LLM) will be used to process the integrated cross-border document text data, extracting key paragraphs and information. The key to this process is effectively utilizing the natural language understanding capabilities of LLM, enabling it to identify key information from complex, structured text. Large Language Models, such as XLM-Roberta or mT5, exhibit good characteristics when processing multilingual data and can adapt to the contextual semantics of different languages. Given text input, the LLM will return a deep analysis of the text and be able to extract specific key information fragments (such as invoice information, product numbers, certification requirements, etc.) according to the set task. This process can usually be precisely controlled by defining specific extraction rules or task objectives. For example, the model can be instructed to identify and extract all paragraphs related to customs clearance and return these paragraphs and their content. This ensures the accurate extraction of important information and provides a reliable basis for subsequent compliance checks, data verification, and processing, guaranteeing the efficiency and compliance of cross-border document processing. Overall, the above steps complement each other, from the recognition of different languages to the extraction of key information, greatly improving the processing capabilities of cross-border documents.
[0132] In one embodiment, before step S114 of concatenating the various sub-text data to obtain the cross-border document text data, the method further includes:
[0133] S1131: Input each of the sub-text data into a preset context-aware model to correct each part of the sub-text data and obtain the corrected sub-text data.
[0134] As described in step S1131 above, each extracted sub-text data is input into a preset context-aware model to correct various parts of the sub-text data. First, the text data entering the context-aware model is the original text extracted during OCR processing. This text may contain spelling errors or incorrect formatting due to character recognition errors or blurry image quality during conversion. By analyzing these individual text fragments and utilizing pre-trained language understanding capabilities, the text is corrected. For example, when processing specific language regions, the model can understand the grammatical rules of that language, thereby identifying and correcting erroneous words. Similarly, for other languages with specific grammatical logic, the model can provide appropriate correction suggestions based on its grammar and context. This processing is not limited to word-level correction but also includes optimization of the entire sentence structure. For example, if a sub-text paragraph has logical inconsistencies in its context, the model will propose reasonable correction suggestions based on its understanding of the surrounding text to ensure the semantic coherence of the entire paragraph. After correction, the corrected sub-text data is stored to replace the original sub-text, providing a more reliable text data foundation for the subsequent splicing process. The automation and intelligentization of this process improves the efficiency of cross-border document processing and reduces data compliance risks caused by textual errors.
[0135] Reference Figure 3 The present invention also provides a data verification device for cross-border documents, the device comprising:
[0136] The first extraction module 902 is used to extract key paragraphs from specified cross-border documents;
[0137] The mapping module 904 is used to map the key paragraphs into a preset cross-border knowledge graph to obtain the corresponding three-level mapping relationship; wherein the three-level mapping relationship is product description-harmonized system code-local extension code;
[0138] Verification module 906 is used to verify the cross-border documents based on the three-level mapping relationship;
[0139] Module 908 is used to acquire tariff data for multiple preset regions;
[0140] The second extraction module 910 is used to extract the preset three-level mapping relationship from the tariff data;
[0141] The construction module 912 is used to construct the preset cross-border knowledge graph based on the preset three-level mapping relationship.
[0142] In one embodiment, the data verification device for cross-border documents further includes:
[0143] The tariff update status monitoring module is used to monitor the tariff update status of each of the preset regions in real time.
[0144] The updated tariff information acquisition module is used to acquire the updated tariff information of the target preset region if the updated status of the tariff information of the target preset region is detected to be updated.
[0145] The preset three-level mapping relationship update module is used to extract the target three-level mapping relationship based on the updated tariff and update the preset three-level mapping relationship corresponding to the preset cross-border knowledge graph.
[0146] In one embodiment, the data verification device for cross-border documents further includes:
[0147] The mandatory authentication request acquisition module is used to acquire the mandatory authentication requests for each of the preset areas;
[0148] The module for obtaining the three-level mapping relationship to be certified is used to obtain the three-level mapping relationship to be certified corresponding to the preset cross-border knowledge graph based on the mandatory certification request.
[0149] The three-level mapping relationship determination module is used to determine whether the three-level mapping relationship to be authenticated satisfies the mandatory authentication request;
[0150] The change prompt generation module is used to generate a change prompt based on the three-level mapping relationship to be authenticated and the mandatory authentication request if the three-level mapping relationship to be authenticated does not meet the mandatory authentication request.
[0151] The change content acquisition module is used to acquire the change content based on the change prompt;
[0152] The module for changing the three-level mapping relationship to be certified is used to change the three-level mapping relationship to be certified based on the changed content, so as to form a compliant three-level mapping relationship;
[0153] The compliance three-level mapping relationship input module is used to input the compliance three-level mapping relationship into the preset cross-border knowledge graph to update the corresponding three-level mapping relationship to be certified.
[0154] In one embodiment, the data verification device for cross-border documents further includes:
[0155] The change time acquisition module is used to acquire the identity of the operator who made the change to the three-level mapping relationship to be authenticated based on the change content, as well as the change time.
[0156] The change log generation module is used to generate a change log based on the change time, operator identity, the three-level mapping relationship to be authenticated, the three-level mapping relationship for compliance, and the change prompt.
[0157] The change log saving module is used to upload the change log to a preset blockchain for saving.
[0158] In one embodiment, the data verification device for cross-border documents further includes:
[0159] An exception pattern library building module is used to pre-collect multiple exception scenarios to build an exception pattern library;
[0160] A structured rule template construction module is used to construct structured rule templates based on the exception pattern library;
[0161] A designated cross-border document detection module is used to detect the designated cross-border documents based on the structured rule template in order to obtain non-compliant data and the corresponding anomaly level;
[0162] The target data acquisition method acquisition module is used to acquire the target data acquisition method based on the anomaly level and the preset level and the data acquisition method correspondence table.
[0163] The non-compliant data replacement module is used to obtain the compliant data corresponding to the non-compliant data based on the target data acquisition method, and replace the non-compliant data in the cross-border document with the compliant data.
[0164] In one embodiment, the first extraction module 902 includes:
[0165] The multilingual region division submodule is used to divide the specified cross-border document into multilingual regions using a preset YOLOv5 model;
[0166] The OCR processing model acquisition submodule is used to acquire the corresponding OCR processing model according to the language type of each of the aforementioned language regions.
[0167] The language region processing submodule is used to process the corresponding language region through each of the OCR processing models to obtain the sub-text data corresponding to each language region.
[0168] The sub-text data splicing submodule is used to splice together various sub-text data to obtain cross-border document text data;
[0169] The key paragraph extraction submodule is used to extract key paragraphs from the cross-border document text data using a preset large language model.
[0170] In one embodiment, the first extraction module 902 further includes:
[0171] The subtext data correction submodule is used to input each of the subtext data into a preset context-aware model to correct each part of the subtext data and obtain the corrected subtext data.
[0172] Figure 4 An internal structural diagram of an electronic device in one embodiment is shown. This electronic device can specifically be a terminal or a server, and more specifically, a computer device. Figure 4 As shown, the electronic device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement a data verification method for cross-border documents. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to implement the data verification method for cross-border documents. Those skilled in the art will understand that… Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0173] In one embodiment, an electronic device is provided, including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps:
[0174] Extract key paragraphs from specified cross-border documents;
[0175] The key paragraphs are mapped into a pre-defined cross-border knowledge graph to obtain the corresponding three-level mapping relationship; wherein, the three-level mapping relationship is product description - harmonized system code - local extension code;
[0176] The cross-border documents are verified based on the aforementioned three-level mapping relationship;
[0177] Before the step of obtaining the corresponding three-level mapping relationship in the preset cross-border knowledge graph through the key paragraph mapping, the method further includes:
[0178] Obtain tariff data from multiple preset regions;
[0179] Extract the preset three-level mapping relationship from the tariff data;
[0180] The preset cross-border knowledge graph is constructed based on the preset three-level mapping relationship.
[0181] By extracting key paragraphs from specified cross-border documents and mapping them to a pre-defined cross-border knowledge graph, the verification process can quickly obtain the three-level mapping relationship between product descriptions, harmonized system codes, and local extension codes, improving the accuracy of document processing and making compliance verification more efficient. This reduces omissions and misjudgments caused by manual review, achieving intelligent, automated, and standardized cross-border document processing, and giving global trade higher efficiency and security.
[0182] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the following steps:
[0183] Extract key paragraphs from specified cross-border documents;
[0184] The key paragraphs are mapped into a pre-defined cross-border knowledge graph to obtain the corresponding three-level mapping relationship; wherein, the three-level mapping relationship is product description - harmonized system code - local extension code;
[0185] The cross-border documents are verified based on the aforementioned three-level mapping relationship;
[0186] Before the step of obtaining the corresponding three-level mapping relationship in the preset cross-border knowledge graph through the key paragraph mapping, the method further includes:
[0187] Obtain tariff data from multiple preset regions;
[0188] Extract the preset three-level mapping relationship from the tariff data;
[0189] The preset cross-border knowledge graph is constructed based on the preset three-level mapping relationship.
[0190] By extracting key paragraphs from specified cross-border documents and mapping them to a pre-defined cross-border knowledge graph, the verification process can quickly obtain the three-level mapping relationship between product descriptions, harmonized system codes, and local extension codes, improving the accuracy of document processing and making compliance verification more efficient. This reduces omissions and misjudgments caused by manual review, achieving intelligent, automated, and standardized cross-border document processing, and giving global trade higher efficiency and security.
[0191] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0192] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0193] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A data verification method for cross-border documents, characterized in that, The method includes: Extract key paragraphs from specified cross-border documents; The key paragraphs are mapped into a pre-defined cross-border knowledge graph to obtain the corresponding three-level mapping relationship; wherein, the three-level mapping relationship is product description - harmonized system code - local extension code; The cross-border documents are verified based on the aforementioned three-level mapping relationship; Before the step of obtaining the corresponding three-level mapping relationship in the preset cross-border knowledge graph through the key paragraph mapping, the method further includes: Obtain tariff data from multiple preset regions; Extract the preset three-level mapping relationship from the tariff data; The preset cross-border knowledge graph is constructed based on the preset three-level mapping relationship; Following the step of constructing the preset cross-border knowledge graph based on the preset three-level mapping relationship, the method further includes: Obtain the mandatory authentication requests for each of the preset regions; Based on the mandatory authentication request, obtain the three-level mapping relationship to be authenticated corresponding to the preset cross-border knowledge graph; Determine whether the three-level mapping relationship to be authenticated satisfies the mandatory authentication request; If the three-level mapping relationship to be authenticated does not meet the mandatory authentication request, a change prompt is generated based on the three-level mapping relationship to be authenticated and the mandatory authentication request; Obtain the changed content based on the change prompt; Based on the changes, the three-level mapping relationship to be certified is modified to form a compliant three-level mapping relationship; The compliance three-level mapping relationship is input into the preset cross-border knowledge graph to update the corresponding three-level mapping relationship to be certified; After the step of modifying the three-level mapping relationship to be authenticated based on the changed content to form a compliant three-level mapping relationship, the method further includes: Obtain the identity of the operator who modified the three-level mapping relationship to be authenticated based on the modified content, as well as the time of the modification. A change log is generated based on the change time, operator identity, the three-level mapping relationship to be authenticated, the three-level mapping relationship for compliance, and the change prompt. The change log is uploaded to a pre-defined blockchain for storage.
2. The data verification method for cross-border documents according to claim 1, characterized in that, Following the step of constructing the preset cross-border knowledge graph based on the preset three-level mapping relationship, the method further includes: Real-time monitoring of the tariff update status of each of the preset regions; If the tariff update status of the target preset region is detected to be updated, then the updated tariff of the target preset region is obtained; Based on the updated tariff, the target three-level mapping relationship is extracted and the preset three-level mapping relationship corresponding to the preset cross-border knowledge graph is updated.
3. The data verification method for cross-border documents according to claim 1, characterized in that, Before the step of extracting key paragraphs from a specified cross-border document, the method further includes: Collect multiple abnormal scenarios in advance to build an abnormal pattern library; Construct structured rule templates based on the aforementioned anomaly pattern library; The specified cross-border documents are inspected based on the structured rule template to obtain non-compliant data and corresponding anomaly levels; Based on the correspondence table between the anomaly level and the preset level and the data acquisition method, the target data acquisition method is obtained; Obtain the compliant data corresponding to the non-compliant data based on the target data acquisition method, and replace the non-compliant data in the cross-border document with the compliant data.
4. The data verification method for cross-border documents according to claim 1, characterized in that, The step of extracting key paragraphs from a specified cross-border document includes: The specified cross-border documents are divided into multilingual regions using a preset YOLOv5 model; Obtain the corresponding OCR processing model based on the language type of each of the aforementioned language regions; The corresponding language regions are processed by each of the OCR processing models to obtain sub-text data corresponding to each language region; The individual sub-text data are concatenated to obtain the cross-border document text data; Key paragraphs are extracted from the cross-border document text data using a pre-defined large language model.
5. The data verification method for cross-border documents according to claim 4, characterized in that, Before the step of concatenating the various sub-text data to obtain the cross-border document text data, the method further includes: Each of the sub-text data is input into a preset context-aware model to correct each part of the sub-text data, resulting in corrected sub-text data.
6. A data verification device for cross-border documents, characterized in that, The device includes: The first extraction module is used to extract key paragraphs from specified cross-border documents; The mapping module is used to map the key paragraphs into a preset cross-border knowledge graph to obtain the corresponding three-level mapping relationship; wherein the three-level mapping relationship is product description - harmonized system code - local extension code; The verification module is used to verify the cross-border documents based on the three-level mapping relationship; The acquisition module is used to acquire tariff data from multiple preset regions; The second extraction module is used to extract the preset three-level mapping relationship from the tariff data; The construction module is used to construct the preset cross-border knowledge graph based on the preset three-level mapping relationship; The mandatory authentication request acquisition module is used to acquire the mandatory authentication requests for each of the preset areas; The module for obtaining the three-level mapping relationship to be certified is used to obtain the three-level mapping relationship to be certified corresponding to the preset cross-border knowledge graph based on the mandatory certification request. The three-level mapping relationship determination module is used to determine whether the three-level mapping relationship to be authenticated satisfies the mandatory authentication request; The change prompt generation module is used to generate a change prompt based on the three-level mapping relationship to be authenticated and the mandatory authentication request if the three-level mapping relationship to be authenticated does not meet the mandatory authentication request. The change content acquisition module is used to acquire the change content based on the change prompt; The module for changing the three-level mapping relationship to be certified is used to change the three-level mapping relationship to be certified based on the changed content, so as to form a compliant three-level mapping relationship; The compliance three-level mapping relationship input module is used to input the compliance three-level mapping relationship into the preset cross-border knowledge graph to update the corresponding three-level mapping relationship to be certified; The change time acquisition module is used to acquire the identity of the operator who made the change to the three-level mapping relationship to be authenticated based on the change content, as well as the change time. The change log generation module is used to generate a change log based on the change time, operator identity, the three-level mapping relationship to be authenticated, the three-level mapping relationship for compliance, and the change prompt. The change log saving module is used to upload the change log to a preset blockchain for saving.
7. A computer-readable storage medium, characterized in that, The document contains a computer program that, when executed by a processor, causes the processor to perform the steps of the data verification method for cross-border documents as described in any one of claims 1 to 5.
8. An electronic device, characterized in that, The device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the data verification method for cross-border documents as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Tax knowledge base system based on knowledge graph
CN112613611A
Consistent set of interfaces derived from a business object model
WO2006117680A2