Intelligent customs declaration compliance verification method and system based on multi-source data

By synchronizing and credible assessing multi-source data in a timely manner, and combining customs regulations database and semantic analysis, the contradictions in verification basis caused by differences in the generation time of multi-source data are resolved. This enables intelligent compliance verification of customs declarations, improves accuracy and efficiency, and reduces the risk of customs clearance delays.

CN121788053APending Publication Date: 2026-04-03广州市昊链信息科技股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, inconsistencies caused by differences in the generation time of customs declarations from multiple sources lead to conflicting verification criteria, requiring manual investigation, which is inefficient, error-prone, and increases the risk of customs clearance delays.

Method used

By connecting with the Single Window, ERP system and document documents to obtain data and generate timestamps, multi-source data time-series synchronization is performed, credibility weights are calculated, high credibility datasets are selected, and a dynamically updated customs law database is called for multi-level verification. Combined with the inherent risk level of commodities and trade agreement rules, intelligent compliance verification is performed, and semantic analysis is conducted on unstructured data to generate a visualized risk control report.

Benefits of technology

It effectively resolves verification discrepancies caused by time-series deviations in multi-source data, improves the accuracy and efficiency of compliance verification, and reduces the risk of customs clearance delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121788053A_ABST
    Figure CN121788053A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent customs declaration compliance verification method and system based on multi-source data, and relates to the technical field of customs declaration compliance verification. According to the method, through multi-source data time sequence synchronization, credibility evaluation, multi-level rule verification, multi-protocol conflict judgment and AI semantic cross validation, the problem of verification contradiction caused by multi-source data time sequence deviation in the prior art is solved. And the system is in butt joint with a single window, an ERP system and a document file, calculates a data credibility weight, generates an initial risk score and a conflict judgment score, and outputs a grading risk report. According to the scheme, the compliance verification accuracy and efficiency are improved, the customs clearance delay risk is reduced, and the method is suitable for an intelligent declaration scene of enterprise customs declaration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of customs declaration compliance verification technology, and to a method and system for intelligent compliance verification of customs declarations based on multi-source data. Background Technology

[0002] In the process of customs declaration, customs declaration data typically comes from multiple channels, including documents exported from the Customs Single Window, business data stored in the enterprise's ERP system, and externally provided Excel or PDF document files. Current technologies often use a simple integration method to process this multi-source data, directly extracting and merging field data from each channel without considering the differences in generation time between the different sources. For example, documents exported from the Single Window may have been generated the previous day, while the corresponding data in the ERP system may have been modified on the same day due to business updates. Furthermore, the document file's generation time may differ from the actual business time due to delays in manual entry. This time-series discrepancy can lead to inconsistencies in the same customs declaration field across different data sources, such as differences in key information like unit price and quantity, resulting in conflicting data for compliance verification. Verification personnel must manually investigate the causes of data discrepancies, which is not only time-consuming but also prone to errors in judgment, leading to incorrect data selection as verification basis, resulting in incorrect customs declarations, triggering customs reviews, and increasing the risk of customs clearance delays. Currently, there is no effective solution to address the time-series discrepancies of multi-source data, failing to resolve the problem of conflicting verification data at its source.

[0003] Based on the above problems, there is an urgent need for a technical solution that can combine data timeliness and credibility for intelligent verification, so as to improve the accuracy and efficiency of customs declaration compliance verification. Summary of the Invention

[0004] This invention provides an intelligent compliance verification method for customs declarations based on multi-source data. The method includes steps such as collecting relevant data from customs declarations and performing preliminary processing; it also includes: connecting to the Single Window system, ERP system, and document files to obtain data and corresponding timestamps, achieving time-series synchronization of multi-source data; calculating the credibility weights of data from each source and selecting high-credibility datasets as verification benchmarks; calling a dynamically updated customs regulatory database to perform multi-level verification of basic fields, business logic, and policy compliance on the high-credibility datasets, and statistically analyzing the number of violation rules and the logical deviation of declaration elements; generating an initial compliance risk score based on the inherent risk level of the commodity, credibility weights, and the above statistical results; if there are conflicts in rules from multiple trade agreements, calculating the final risk score by combining agreement priority, rule matching degree, and historical ruling similarity; performing semantic analysis on unstructured data to extract implicit regulatory elements and cross-validating the accuracy of the logical deviation of declaration elements; classifying risk levels according to the final risk score and outputting a visual risk control report with rectification suggestions; and supporting the re-execution of the above verification steps after customs declarations are modified.

[0005] Preferably, in the steps of connecting to the Single Window, ERP system, and document documents to obtain data and generate corresponding timestamps, the Single Window data is obtained through the official export function, and the exported file formats include Excel and PDF, with the file containing a data generation time field; the ERP system data is obtained in real time through the RestAPI interface, and the data returned by the interface includes the data entry timestamp; the document documents are parsed through a document recognition tool, and the file reception time is recorded during the parsing process as a substitute for the data generation timestamp.

[0006] In a further preferred embodiment, the step of calling the dynamically updated customs regulatory database to perform multi-level verification includes the following: the customs regulatory database contains HS code classification rules, origin determination standards, regulatory document requirements, business compliance rules, and risk characteristic rules; basic field verification includes non-empty verification, format verification, and numerical range verification; business logic verification includes classification consistency verification, price reasonableness verification, and unit conversion compliance verification; policy compliance verification includes trade agreement applicability verification and license requirement verification; the customs regulatory database is automatically updated by connecting to the official channels of the General Administration of Customs and the Single Window, while also supporting manual rule maintenance.

[0007] In a further preferred embodiment, in the step of extracting implicit regulatory elements from unstructured data through semantic analysis, the unstructured data includes invoice remarks, contract terms, and packing list descriptions; the semantic analysis is implemented using a large model, which is trained through a regulatory element feature library, which includes hazardous materials keywords, 3C certification-related expressions, and energy efficiency labeling-related terms; when cross-validating the accuracy of the logical deviation degree of the declaration elements, if the regulatory elements extracted by semantic analysis are not declared in the customs declaration, the logical deviation degree value of the declaration elements is corrected.

[0008] More preferably, in the step of calculating the credibility weight of each source data, the credibility weight is calculated using a formula, which is: ; in, The reliability coefficient for single-window data ranges from 0.8 to 1.0. It is dimensionless and is used to characterize the official authority and accuracy of single-window data. The reliability coefficient of ERP system data, with a value ranging from 0.7 to 0.95, is dimensionless and is used to characterize the standardization and completeness of data entry in an enterprise's internal ERP system. The reliability coefficient of the document data, with a value ranging from 0.6 to 0.9, is dimensionless and is used to characterize the authenticity and standardization of the external documents. The data generation time is in minutes, specifically the export time of single window data, the entry time of ERP data, or the receipt time of document files. The benchmark time for verification is in minutes, specifically the current time of the system performing the compliance verification. This is the maximum allowable time series deviation threshold, in minutes, with a value of 1440, corresponding to 24 hours, used to limit the upper limit of data timeliness; The ERP data integrity coefficient has a value range of 0 to 1 and is dimensionless. It is calculated by statistically analyzing the complete filling ratio of required fields in ERP data. The document integrity coefficient ranges from 0 to 1 and is dimensionless. It is calculated by statistically analyzing the percentage of missing key information in the document documents.

[0009] In a further preferred embodiment, in the step of generating the initial compliance risk score, the initial risk score is calculated using a formula, which is: ; in, The confidence weight is calculated using the formula described in claim 5 and is dimensionless. The inherent risk level of the product is represented by a value ranging from 0.5 to 5, which is dimensionless. High-risk products, such as those involving taxes or licenses, are represented by a value of 3 to 5, ordinary products by a value of 1 to 2, and low-risk products, such as ordinary daily necessities, by a value of 0.5 to 1. The number of rule violations triggered, in units of "rules", is a non-negative integer, and is obtained by counting the number of entries that do not conform to the rules in the multi-level validation. The logical deviation of the declared elements is dimensionless, ranging from 0 to 1, and is expressed by the formula... The calculations are all in RMB (Chinese Yuan). When the actual total price is 0, The value is 1.

[0010] More preferably, in the step of calculating the final risk score, the final risk score is calculated using a formula, which is: ; in, The initial risk score calculated using the formula described in claim 6 is dimensionless. This represents the weight of high-priority trade agreements, with a value of 0.7. It is dimensionless and includes RCEP. This represents the weight of low-priority trade agreements, with a value of 0.3. It is dimensionless and includes the CPTPP. The degree of matching between high-priority trade agreement rules and commodity characteristics is a dimensionless value ranging from 0 to 1, calculated by comparing the degree of fit between the commodity's HS code, country of origin, and agreement rules. This represents the degree of matching between low-priority trade agreement rules and commodity characteristics, with a value ranging from 0 to 1, dimensionless, and calculated in the same way as... Consistent; The similarity weight of historical rulings ranges from 0.1 to 0.3, is dimensionless, and is used to adjust the degree of influence of historical cases on the final score; The similarity between the current conflict and the adjudication of similar historical conflicts is a dimensionless value ranging from 0 to 1. It is obtained by calculating the similarity between the feature vector of the current conflict and the feature vector of historical cases using the cosine similarity algorithm.

[0011] More preferably, in the step of classifying risk levels, the risk levels include three categories: error, warning, and compliance; when the final risk score is... When this is determined to be an error level, the customs declaration data must be forcibly corrected; when When the situation is deemed to be at the warning level, manual review of the customs declaration data is required; when When the system is deemed compliant, a customs declaration can be exported for filing. The visual risk control report includes the risk level, the name of the violation field, the triggered rule content, rectification suggestions, and verification timestamp.

[0012] In a further preferred embodiment, during the process of supporting the re-execution of the verification step after the customs declaration is modified, the system saves logs and audit records for each verification. The log content includes the verification time, the version of the legal library used, the credibility weight of each source data, the risk score, and the modification record. The audit record is used to trace the verification process, supports manual querying of historical verification results, and provides traceable evidence of the verification process for customs supervision.

[0013] A customs declaration intelligent compliance verification system based on multi-source data, used to execute the method described in any one of the above, characterized in that it includes a multi-source data time-series synchronization module, a data credibility assessment module, a rule engine and conflict adjudication module, an AI semantic enhancement verification module, and a risk report generation module; the multi-source data time-series synchronization module is used to connect with the single window, ERP system, and documents, acquire data, generate timestamps, and complete time-series alignment; the data credibility assessment module is used to calculate credibility weights according to the formula described in claim 5 and filter high-credibility datasets; the rule engine and conflict adjudication module is used to call the customs regulation library to perform multi-level verification and handle rule conflicts according to the formula described in claim 7; the AI ​​semantic enhancement verification module is used to analyze unstructured data to extract implicit regulatory elements and correct logical deviations in declaration elements; the risk report generation module is used to classify risk levels according to the final risk score, output a visual risk control report, and support re-verification after customs declaration modification.

[0014] The technical effects achieved by the above embodiments include: The core inventive technology of this invention lies in the collaborative synchronization of multi-source data time series and credibility assessment. It solves the problem of data generation time difference through timestamp alignment, quantifies data credibility by combining reliability and integrity coefficients, and selects high-credibility data as the verification benchmark. Simultaneously, it integrates multi-level rule verification, multi-agreement conflict adjudication, and AI semantic cross-validation to form a full-process intelligent verification system. This solution effectively solves the verification contradiction problem caused by multi-source data time series deviations in the background technology, avoids errors in manual inspection, improves the accuracy and efficiency of compliance verification, and reduces the risk of customs clearance delays. Attached Figure Description

[0015] Figure 1 This is a flowchart of a smart compliance verification method for customs declarations based on multi-source data, as described in this application. Figure 2 This is a connection block diagram of a customs declaration intelligent compliance verification system based on multi-source data, as proposed in this application. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0017] Traditional technical solutions have the following technical problems: time series deviations of multi-source data lead to contradictions in verification basis; existing technologies directly integrate data from the single window ERP system and document files without considering differences in data generation time; when inconsistencies occur in the same field, manual investigation is required, which is inefficient, error-prone, and increases the risk of customs clearance delays.

[0018] Based on this, please refer to Figure 1 This embodiment provides an intelligent compliance verification method for customs declarations based on multi-source data. The method includes steps such as collecting relevant customs declaration data and performing preliminary processing, and further includes: connecting to the Single Window ERP system and document files to obtain data and corresponding timestamps, completing multi-source data time-series synchronization; calculating the credibility weight of each source data and selecting high-credibility datasets as verification benchmarks; calling the dynamically updated customs regulatory database to perform multi-level verification of basic field business logic and policy compliance on the high-credibility datasets, and statistically analyzing the number of violation rules and the logical deviation of declaration elements; generating an initial compliance risk score based on the inherent risk level credibility weight of the commodity and the above statistical results; if there are conflicts in rules from multiple trade agreements, calculating the final risk score by combining the agreement priority rule matching degree and historical ruling similarity; performing semantic analysis on unstructured data to extract implicit regulatory elements, and cross-validating the accuracy of the logical deviation of declaration elements; classifying risk levels according to the final risk score and outputting a visual risk control report with rectification suggestions; and supporting the re-execution of the above verification steps after the customs declaration is modified.

[0019] It's worth noting that: Firstly, in the multi-source data time-series synchronization process, the system connects to the Customs Single Window through a pre-defined interface, calling the official export API of the Single Window to obtain relevant data from customs declarations in Excel or PDF format. The exported data includes a data generation time field, which the system extracts as... For ERP systems, the system uses the RestAPI interface to establish real-time communication with the enterprise ERP system. The interface request parameters include the data query time period, and the returned data carries the entry timestamp, which is converted to minutes and used as the input data. For external documents, users upload Excel or PDF files through the system. The system records the file reception time and converts it to minutes. Simultaneously, the POI tool is used to parse the Excel file, and the Aspose tool is used to convert the PDF file to Excel before parsing it with POI to extract key fields such as product name, HS code, unit price, and quantity. During time-series alignment, the system parses all data sources... With reference time Compare and calculate the time difference. To ensure data timeliness Data outside the specified range is marked as timed out and prompts for manual confirmation.

[0020] Secondly, in the data credibility assessment stage, the system sets parameters based on the data source type. The initial value, the single window data comes from official customs channels, The initial value is set to 0.9. If the data is complete and formatted correctly, it can be increased to 1.0. Since the data in the ERP system is entered internally by the company, The initial value is set to 0.8. If required fields are filled completely, When the value is 1.0, and the field missing rate is 10%. The value is 0.9, and so on; the data in the documents may contain errors due to manual entry. The initial value is set to 0.7. If no key information is missing, The value is 1.0, indicating that a key piece of information is missing. Decrease by 0.1. The system substitutes the above parameters into the credibility weight formula. The calculations obtained from each data source Values, Filtering The dataset is used as a high-confidence verification benchmark. The dataset was marked as low-confidence data, and the fields requiring review were listed. Then, in the multi-level rule verification stage, the system called the customs regulation database, which is automatically updated daily through the rule update interface of the General Administration of Customs' Internet + Customs platform and Single Window, while also providing a visual interface to support manual addition of temporary rules. During basic field verification, the system checked whether required fields in high-confidence data were not empty, whether the HS code format conformed to the 10-digit standard, and whether the unit price and quantity were positive. During business logic verification, the system verified the consistency between the HS code and the product name, calculated the difference between the unit price × quantity and the actual total price, and passed the verification. Obtain the logical deviation degree of the application elements , The time stamp is marked as a logical anomaly; during policy compliance verification, the system determines whether regulatory certificates are required based on the commodity's HS code and whether trade agreements such as RCEP apply based on the country of origin. Next, in the risk scoring and conflict resolution phase, the system queries the inherent risk level based on the commodity's HS code. For example, HS code 30049090 is set to 4.0, and HS code 62052000 is set to 1.0; count the number of violation rules triggered in multi-level verification. If both HS encoding format error and negative unit price are violated simultaneously, then... ;Will Substitute into the initial risk scoring formula ,get Value. If there are conflicts between multiple trade agreement rules, the system will set... The degree of conformity between the product characteristics and the agreement rules was compared and obtained. and The system queries the historical adjudication database and uses a cosine similarity algorithm to calculate the correlation between the current conflict and historical cases. Value, setting Substitute into the final risk scoring formula ,get The value is then determined. In the subsequent AI semantic cross-validation stage, the system performs semantic analysis on unstructured data such as invoice remarks and contract terms. A large model trained on a regulatory element feature library is used. This feature library includes keywords such as "flammable lithium battery" and "3C certification," along with corresponding regulatory requirements. After extracting implicit regulatory elements, the large model checks whether the customs declaration includes relevant information. If the remarks mention lithium batteries but the customs declaration does not declare battery information, the error is corrected. A value of 0.8 ensures The value can reflect the risk of omissions in implicit regulatory elements; the revised value... Substitute the value back into the initial risk scoring formula and update. The value, and thus the final risk score. .

[0021] Finally, in the risk report output and re-verification phase, the system... Values ​​are used to classify risk levels. Generate error reports in real time, listing the fields that need to be forcibly corrected; Early warning reports are generated in real time, prompting manual review; The system generates compliance reports in real time and supports exporting customs declarations for use in declarations. After a user modifies the customs declaration data, the system re-executes all the above verification steps and saves logs for each verification. Audit records can be exported through the system's query interface for customs supervision and verification. This solution resolves inconsistencies in verification basis through multi-source data time-series synchronization and credibility assessment, and improves verification accuracy by combining multi-level verification conflict adjudication and AI semantic analysis, effectively reducing the risk of customs clearance delays.

[0022] Traditional technical solutions have the following technical problems: the lack of a unified timestamp recording method when acquiring multi-source data; the inconsistent storage format of time information in single-window ERP systems and document documents; which makes time synchronization difficult and affects the judgment of data credibility.

[0023] Based on this, in the steps of connecting to the Single Window ERP system and document documents to obtain data and generate corresponding timestamps, Single Window data is obtained through the official export function, and the exported file formats include Excel and PDF, with the file containing a data generation time field; ERP system data is obtained in real time through the RestAPI interface, and the interface returns data containing the data entry timestamp; document documents are parsed through a document recognition tool, and the file reception time is recorded during the parsing process as a substitute for the data generation timestamp.

[0024] It's worth noting that regarding single-window data acquisition, the system triggers the official export function by simulating user operations or calling the single-window open API. The export request specifies the file format as Excel or PDF. In the file returned by the single window, the Excel format header contains a data generation time field in the format YYYY-MM-DDHH:MM:SS. The system reads this field using the POI tool and converts it to minutes. In PDF files, the data generation time is usually located in the header or footer of the first page. The system uses Aspose to convert the PDF to text, then extracts the time information using regular expressions (--::), and converts it to minutes. Regarding ERP system data acquisition, the system establishes a connection with the enterprise ERP system's API gateway. The URL format for the REST API request is http: / / ERP server IP:port / api / data / query, and the request parameters include startTime, endTime, and dataType. In the JSON format data returned by the ERP system, each data entry contains a createTimestamp field, which the system converts into minutes. Simultaneously, the system verifies the signature information of the returned data to ensure that the data transmission process has not been tampered with. Regarding the acquisition of document documents, users select Excel or PDF format documents through the system's file upload interface. The system uses JavaScript to display the file upload progress and records the file reception time, converting it to minutes. For Excel files, the system uses POI version 5.2.4 to parse them, reading xlsx files using the XSSFWorkbook class and xls files using the HSSFWorkbook class, extracting the rows and columns containing fields such as product name, HS code, unit price, and quantity to ensure accurate field extraction. For PDF files, the system uses Aspose.PDFforJava version 23.5, loading PDF files using the Document class, calling the save method to convert them to Excel format, setting the ExcelSaveOptions parameter during the conversion process to ensure the integrity of the table structure, and then using POI to parse the converted Excel file and extract fields. This solution solves the problem of time series synchronization difficulties by unifying the timestamp conversion standard and clarifying the acquisition and processing methods of time information from various data sources, providing accurate time parameters for subsequent credibility assessment.

[0025] Traditional technical solutions have the following technical problems: the customs regulatory database is not updated in a timely manner, the verification rules are simple and only cover basic field verification, lack business logic and policy compliance verification, and cannot cope with changes in trade agreements and complex business scenarios, resulting in missed judgments of compliance risks.

[0026] Based on this, the steps of performing multi-level verification by calling the dynamically updated customs regulatory database include: HS code classification rules, origin determination standards, regulatory document requirements, business compliance rules, and risk characteristic rules; basic field verification includes non-empty verification, format verification, and numerical range verification; business logic verification includes classification consistency verification, price reasonableness verification, and unit conversion compliance verification; policy compliance verification includes trade agreement applicability verification and license requirement verification; the customs regulatory database is automatically updated by connecting to the official channels of the General Administration of Customs and the Single Window, and also supports manual rule maintenance.

[0027] It is worth mentioning that, in terms of the construction of the customs regulations database, the system uses a relational database to store various rules. The HS code classification rule table includes fields such as HS code, commodity name, range, classification basis, effective date, and expiration date. For example, HS code 95030089 corresponds to the commodity name range of "Other Toys," and the classification basis is General Administration of Customs Announcement No. 56 of 2023. The origin determination standard table includes fields such as trade agreement name, origin standard, and applicable commodity range. For example, the RCEP agreement's origin standards include the fully obtained standard, regional value component standard, and other relevant standards. The regulatory document requirements table includes fields such as HS code, required document name, document number, and rule. For example, HS code 85176239 corresponds to the required document, a telecommunications equipment network access license. The business compliance rule table includes fields such as rule name, rule expression, and violation handling method. For example, the expression for the rule "Unit Price × Quantity = Total Price" is... The risk characteristic rule table includes risk type, characteristic keywords, and risk level fields. For example, the characteristic keywords for high-risk commodities are ivory and red coral, with a risk level of "high." Regarding regulatory database updates, the system connects daily at 2:00 AM via a scheduled task to the regulatory standard interface of the General Administration of Customs' "Internet + Customs" platform and the rule update interface of the Single Window. The interface returns update data in JSON format. After parsing, the system compares the effective and expiration times of existing rules in the database, marks expired rules as invalid, inserts new rules into the corresponding table, and generates an update log recording the update time, the number of updated rules, and a list of failed rules. Simultaneously, the system provides a visual rule maintenance interface. Administrators can add temporary rules through this interface, requiring them to fill in relevant rule fields and select the rule's effective time. Temporary rules have higher priority than automatically updated rules. In terms of multi-level verification, during basic field verification, the system iterates through the required fields of high-confidence data. Non-empty verification checks whether the field is an empty string or null; if it is empty, the non-empty verification is marked as failed. Format verification checks whether the HS code is a 10-digit number, whether the country of origin is a standard country or region name, and whether the unit price and quantity are positive numbers. Numerical range verification checks whether the unit price exceeds ±50% of the historical declared price of the product and whether the quantity is a reasonable integer. During business logic validation, the classification consistency check compares the product name with the HS code using the HS code classification rule table. For example, if the product name is "Children's Electric Toy Car" but the HS code is 95063000, the classification is considered inconsistent. The price reasonableness check calculates the difference between the product's unit price and the average unit price of products with the same HS code and country of origin. If the difference exceeds 30%, the price is marked as abnormal. The unit conversion compliance check verifies the correctness of conversions between different units. For example, 1 kilogram = 1000 grams, 1 meter = 100 centimeters. If the declared quantity is 5 kilograms but the corresponding unit of measurement is grams, the converted quantity should be 5000 grams. If this does not match, the unit conversion is marked as incorrect. During policy compliance validation, the trade agreement applicability check determines whether agreements such as RCE, EPC, and PTP are applicable based on the product origin and origin determination criteria table. For example, if the product's country of origin is Japan and the regional value component... If the HS code is not specified, the RCEP agreement applies. License requirement verification involves consulting the regulatory document requirements table based on the HS code to determine if a regulatory document is required. For example, if the HS code is 05071000, an import permit for endangered species is required; failure to provide this permit will result in a missing license. This solution utilizes a dynamically updated, multi-dimensional regulatory database to achieve multi-level verification, covering both basic business and policy aspects, addressing compliance risk omissions, and adapting to changes in trade agreements and complex business scenarios.

[0028] Traditional technical solutions have the following technical problems: they only verify structured data and ignore implicit regulatory elements in unstructured data such as invoice remarks and contract terms, which leads to the omission of special regulatory requirements such as 3C certification for dangerous goods, and there is a risk of concealment and misreporting.

[0029] Based on this, in the step of extracting implicit regulatory elements from unstructured data through semantic analysis, the unstructured data includes invoice remarks, contract terms, and packing list descriptions; the semantic analysis is implemented using a large model, which is trained through a regulatory element feature library, which includes dangerous goods keywords, 3C certification-related expressions, and energy efficiency label-related terms; when cross-validating the accuracy of the logical deviation degree of the declaration elements, if the regulatory elements extracted by semantic analysis are not declared in the customs declaration, the logical deviation degree value of the declaration elements is corrected.

[0030] It is worth mentioning that, in terms of unstructured data collection, when the system obtains customs declaration-related data, it simultaneously collects associated unstructured data. Invoice remarks are usually located in the remarks column or additional instructions section of the invoice document, contract terms are located in the technical specifications and packaging requirements section of the contract document, and packing list descriptions are located in the goods description field of the packing list. When the system parses these documents using the POIAspose tool mentioned above, it specifically extracts the above-mentioned unstructured text fields and stores them in an unstructured data table. The fields include file ID, text content, file type, and associated customs declaration number, which facilitates subsequent correlation analysis. In terms of constructing the regulatory element feature database, the system uses a structured database to store feature data. The hazardous materials keyword table includes fields for hazardous material category keywords and regulatory requirements. For example, lithium batteries correspond to Category 9 Miscellaneous Hazardous Substances and Articles, and the regulatory requirement is to provide a hazardous materials transportation certificate. The 3C certification related description table includes fields for product type description keywords and certification standards. For example, smartphones correspond to the description keywords "mobile phone," and the certification standard is GB4943.1-2011. The energy efficiency labeling related terminology table includes fields for product type terms and energy efficiency level requirements. For example, refrigerators correspond to the terms "energy efficiency" and "energy consumption," and the energy efficiency level requirement is Level 2 or above. The feature database is updated regularly by connecting to official announcements from the State Administration for Market Regulation and the General Administration of Customs, and also supports administrators to manually add new keywords and terms.

[0031] For large-scale model training and semantic analysis, the system uses the BERT-base model as the basic model. A training dataset is constructed using keywords and terms from the regulatory element feature library. Each training data point contains unstructured text labeled with regulatory elements and requirements. For example, if the text "goods are laptops containing lithium batteries" is labeled with the regulatory element "lithium battery laptops," the regulatory requirement is to provide a lithium battery transportation certificate and a 3C certification for the laptops. The training process uses the PyTorch framework, with a batch size of 32, a learning rate of 2e-5, and 10 training epochs. Model parameters are adjusted through cross-validation to ensure the model's accuracy in recognizing regulatory elements is ≥95%. During semantic analysis, the system inputs unstructured text into the trained large-scale model. The model outputs extracted regulatory elements and corresponding regulatory requirements, which are stored in the regulatory element extraction result table and associated with the corresponding customs declaration number. Regarding cross-validation and deviation correction, the system queries the corresponding customs declaration data based on the regulatory element extraction result table to check whether the extracted regulatory elements are included in the declaration data. For example, if the large model extracts the regulatory element for lithium batteries, but the commodity name field of the customs declaration does not specify lithium batteries and the dangerous goods transportation certificate is not declared, then it is determined that the declaration element is missing. At this time, the system corrects the logical deviation of the declaration element. ,Original When the value is 0.02, it is corrected to 0.8. When the value is 0.05, it is corrected to 0.9 to ensure... The value can reflect the risk of omissions in implicit regulatory elements; the revised value... Substitute the value back into the initial risk scoring formula ,renew The value, and thus the final risk score. This solution extracts implicit regulatory elements through semantic analysis of unstructured data, cross-validates structured declaration data, addresses the issue of missed judgments under special regulatory requirements, and reduces the risk of concealment and misreporting.

[0032] Traditional technical solutions have the following technical problems: the credibility assessment of multi-source data lacks quantitative methods, relies solely on subjective human judgment to select data, and the judgment standards are inconsistent, leading to incorrect selection of verification benchmarks and affecting the accuracy of compliance verification.

[0033] Based on this, the step of calculating the credibility weight of each source data uses a formula to calculate the credibility weight, which is: ; in, The reliability coefficient for single-window data ranges from 0.8 to 1.0. It is dimensionless and is used to characterize the official authority and accuracy of single-window data. The reliability coefficient of ERP system data, with a value ranging from 0.7 to 0.95, is dimensionless and is used to characterize the standardization and completeness of data entry in an enterprise's internal ERP system. The reliability coefficient of the document data, with a value ranging from 0.6 to 0.9, is dimensionless and is used to characterize the authenticity and standardization of the external documents. This refers to the data generation time, in minutes, specifically the export time of single window data, the entry time of ERP data, or the receipt time of document files. The benchmark time for verification is in minutes, specifically the current time of the system performing the compliance verification. This is the maximum allowable time series deviation threshold, in minutes, with a value of 1440, corresponding to 24 hours, used to limit the upper limit of data timeliness; The ERP data integrity coefficient has a value range of 0 to 1 and is dimensionless. It is calculated by statistically analyzing the complete filling ratio of required fields in ERP data. The document integrity coefficient ranges from 0 to 1 and is dimensionless. It is calculated by statistically analyzing the percentage of missing key information in the document documents.

[0034] It is worth mentioning that, regarding the parameter definition and the basis for its value, The value should be between 0.8 and 1.0. Since the Single Window data is generated by the customs system, it has high official authority and guaranteed accuracy. The initial value is set to 0.9. If the data format fully conforms to customs specifications and there are no missing fields, Upgraded to version 1.0. If there are minor formatting anomalies in a few non-critical fields, Lowered to 0.85; The value should range from 0.7 to 0.95. Since ERP data is entered by internal personnel, its accuracy is affected by the company's management level. The initial value is set to 0.8. If the company's ERP system has a field mandatory validation function... Increased to 0.95; if any required fields are missing, Lowered to 0.7; The value should be between 0.6 and 0.9. Due to the potential for errors in document completion or forgery, the initial value is set to 0.7. If the document bears an official seal and the key information is complete... Adjusted to 0.9; if there are signs of alteration to key information, Lowered to 0.6. The calculation method is as follows: Single window data export time and ERP data entry time are converted to a total number of minutes calculated from 00:00:00 on January 1, 1970. For example, 14:30 on May 20, 2024 is converted to... Minutes; the document reception time is the number of minutes converted from the system's current time, and... The calculation method is consistent to ensure the time difference. The unit is standardized to minutes. The value is 1440 minutes, which corresponds to 24 hours. Since customs declaration data usually needs to be submitted within 24 hours after the business occurs, data beyond this time is not timely enough and needs to be manually confirmed. The calculation method is as follows: count the number of required fields in the ERP data, and calculate the ratio of the number of fields that are fully filled to the total number of required fields. For example, if 5 fields are fully filled and the total number of required fields is 6, then... If a field has an incorrect format, the field will be considered incompletely filled. The corresponding reduction. The calculation method is as follows: count the number of key information items in the document, and calculate the ratio of the number of key information items without missing items to the total number of key information items. For example, if there are 3 key information items without missing items and the total number of key information items is 5, then... If key information is unclear, that information is considered missing. Regarding formula calculation and application, taking single-window data as an example, assuming... , minute, minute, minute, minute, ERP data , , Document data , , ;but ; This dataset serves as a high-reliability verification benchmark. The scheme calculates data reliability weights using quantified parameters and formulas, unifies judgment standards, addresses the uncertainty of subjective human judgment, and ensures the accuracy of the selected verification benchmark.

[0035] Traditional technical solutions have the following technical problems: compliance risk scoring only considers the number of violations, without taking into account data credibility and the inherent risks of the product. The scoring results are one-sided and cannot accurately reflect the risk differences of different data quality and product types, leading to incorrect risk level determination.

[0036] Based on this, the initial risk score is calculated using a formula in the step of generating the initial compliance risk score. The formula is as follows: ;

[0037] in, The confidence weight is calculated using the formula described in claim 5 and is dimensionless. The inherent risk level of the product is represented by a value ranging from 0.5 to 5, which is dimensionless. High-risk products, such as those involving taxes or licenses, are represented by a value of 3 to 5, ordinary products by a value of 1 to 2, and low-risk products, such as ordinary daily necessities, by a value of 0.5 to 1. The number of rule violations triggered, in units of "rules", is a non-negative integer, and is obtained by counting the number of entries that do not conform to the rules in the multi-level validation. The logical deviation of the declared elements is dimensionless, ranging from 0 to 1, and is expressed by the formula... The calculations are in RMB yuan for both the actual total price, unit price, and quantity. When the actual total price is 0, The value is 1. It's worth noting that regarding the parameter definition and the basis for its value... The result calculated according to claim 5, with values ​​ranging from 0 to 1, is dimensionless. The higher the value, the higher the data credibility and the better the initial risk score. The impact tends to be more biased towards lower risk. The lower the value, the lower the data credibility, and the more the score is biased towards high risk, reflecting the impact of data quality on risk assessment. Based on HS codes and regulatory requirements, high-risk goods include those involving taxation, certification, or inspection. Values ​​range from 3 to 5, such as HS code 30049090 having a value of 4.0 and HS code 05071000 having a value of 5.0; general merchandise includes clothing and electronic products. Values ​​range from 1 to 2, such as 1.0 for HS code 62052000 and 1.5 for HS code 85287210; low-risk goods include ordinary daily necessities. The value ranges from 0.5 to 1, such as HS code 96032100 which has a value of 0.5 and HS code 63029300 which has a value of 0.8; The value is stored in the product risk level table, and the system automatically retrieves it based on the product's HS code. The results of multi-level validation were analyzed, and each violation of a rule in the basic field validation, business logic validation, and policy compliance validation was counted as one violation. For example, violations such as HS encoding format errors, business logic violations (unit price × quantity ≠ total price), and policy compliance violations (missing license) were counted. If the same violation triggers multiple rules, they will be counted separately to ensure... It can reflect the severity of the violation. Through formula The calculations are all in RMB (Chinese Yuan) to ensure the results are dimensionless. For example, if the total price is 1000 yuan, the unit price is 100 yuan, and the quantity is 10 pieces, then the unit price multiplied by the quantity is 1000 yuan. There is no logical deviation; for example, if the actual total price is 1000 yuan, the unit price is 100 yuan, the quantity is 9 pieces, and the unit price × quantity = 900 yuan. There is a 10% logical deviation; when the actual total price is 0, the deviation cannot be calculated. A value of 1 is marked as high logic deviation and requires manual confirmation. Regarding formula calculation and application, it is assumed that... , , , ,but , , ; , , , ; The initial risk score was approximately 3.51, providing a basis for the subsequent final risk score calculation.

[0038] This scheme calculates an initial risk score using a multi-parameter fusion formula, taking into account data credibility, inherent product risks, number of violations, and logical biases, thus addressing the issue of one-sided scoring and accurately reflecting the level of compliance risk.

[0039] Traditional technical solutions have the following technical problems: they lack an intelligent adjudication mechanism when multiple trade agreement rules conflict, and the rules are selected based on a single agreement or manual judgment, resulting in low adjudication efficiency and easy application errors, and they cannot adapt to a trade environment with multiple agreements running in parallel.

[0040] Therefore, in the step of calculating the final risk score, the formula is used to calculate the final risk score, which is: ; in, The initial risk score calculated using the formula described in claim 6 is dimensionless. This represents the weight of high-priority trade agreements, with a value of 0.7. It is dimensionless and includes RCEP. This represents the weight of low-priority trade agreements, with a value of 0.3. It is dimensionless and includes the CPTPP. The degree of matching between high-priority trade agreement rules and commodity characteristics, with a value ranging from 0 to 1, is dimensionless and is calculated by comparing the degree of fit between the commodity's HS code country of origin and the agreement rules. This represents the degree of matching between low-priority trade agreement rules and commodity characteristics, with a value ranging from 0 to 1, dimensionless, and calculated in the same way as... Consistent; The similarity weight of historical rulings ranges from 0.1 to 0.3, is dimensionless, and is used to adjust the degree of influence of historical cases on the final score; The similarity between the current conflict and the adjudication of similar historical conflicts is a dimensionless value ranging from 0 to 1. It is obtained by calculating the similarity between the feature vector of the current conflict and the feature vector of historical cases using the cosine similarity algorithm.

[0041] It is worth mentioning that, regarding the parameter definition and the basis for its value, The result calculated according to claim 6 is dimensionless and reflects the basic risk level without considering rule conflicts; The value is 0.7. The value is set to 0.3 because the RCEP agreement covers more countries and accounts for a larger share of trade, and has a higher priority than the CPTPP. The weight setting reflects the difference in importance between the agreements. If other trade agreements are added in the future, the weight parameters can be expanded to ensure that the total weight is 1.0. The calculation method is as follows: For the RCEP agreement, the system obtains the rules from the origin determination criteria table, compares the characteristics of the commodity, such as its HS code, origin, region of origin, and value composition, with the rules, and determines the origin of the commodity. If the commodity's HS code falls within the scope of the RCEP agreement, the country of origin is an RCEP member country. ,but ;like ,but If the HS code is not applicable, then . Calculation method and Consistent, if the goods comply with the CPTPP agreement rules, then Otherwise, the degree of fit will be reduced. The value ranges from 0.1 to 0.3, and is adjusted according to the size of the historical case library. When the case library is large... Take 0.3, case library size in hours We set the value to 0.1 to ensure the impact of historical cases is reasonable. The calculation method is as follows: Convert the features of the current rule conflict into feature vectors; query the feature vectors of similar conflicts from the historical adjudication database; and calculate the similarity using a cosine similarity algorithm. For example, if the cosine value between the current vector and the historical vector is 0.8, then... Regarding formula calculation and application, it is assumed that... , , , , , , ; , , ; The final risk score was approximately 2.76, which was determined to be at the compliance level.

[0042] This solution calculates the final risk score using a formula that combines agreement priority rule matching and historical case collaboration, intelligently adjudicates rule conflicts, and solves the problems of low efficiency and incorrect application of manual adjudication, thus adapting to a multi-agreement trade environment.

[0043] Traditional technical solutions have the following technical problems: the risk level classification standards are vague, simply dividing them into two categories: compliant and non-compliant, without distinguishing between errors and warnings, which makes it impossible for users to quickly determine the processing priority and delays the modification and declaration of customs declarations.

[0044] Based on this, the risk level classification process includes three categories: error warning and compliance; when the final risk score is determined... When this is determined to be an error level, the customs declaration data must be forcibly corrected; when When the situation is deemed to be at the warning level, manual review of the customs declaration data is required; when When the system is deemed compliant, a customs declaration can be exported for filing. The visual risk control report includes the rule content, rectification suggestions, and verification timestamps triggered by the risk level violation field name.

[0045] It is worth mentioning that, in terms of setting risk level classification standards, the system determines the scoring threshold based on a large amount of historical verification data and customs supervision requirements. This is an error level, corresponding to a serious violation. If such a violation is not corrected, it will inevitably trigger a customs review, leading to customs clearance delays. Therefore, it must be corrected. This is the warning level, corresponding to minor violations or situations requiring confirmation. In such cases, manual judgment is needed to determine whether correction is necessary, in order to avoid over-verification. The compliance level corresponds to no violations or very minor violations. These situations do not affect customs declarations and can be directly exported. Regarding the level determination and processing procedures, the system calculates... After the value is calculated, it is automatically compared with the threshold to determine the risk level and trigger the corresponding processing flow: When the error level is high, the system marks the non-compliant field in the customs declaration editing interface, pops up a forced correction prompt box, and prevents the user from exporting the customs declaration until all erroneous fields are corrected and re-verified. Only after these conditions are met can the process continue. At the warning level, the system highlights the violation fields in yellow and displays a warning prompt box, suggesting manual review. Users can choose to review and modify immediately or temporarily ignore and continue exporting, but the system records a warning log for later traceability. At the compliance level, the system marks compliance in green and displays an export declaration button, supporting users to export customs declarations in Excel format. The exported file includes a verification timestamp and compliance identifier. Regarding the generation of visual risk control reports, the reports are in HTML format, supporting online viewing in a browser and downloading in PDF format. The report structure includes five parts: verification overview, risk level details, violation details, rectification suggestions, and verification log. The verification overview includes the customs declaration number, verification time, data source, and final risk score. Risk level; Risk level details: Criteria for determining the current level and processing requirements; Violation details: List the names of the violating fields, the rules triggered, and the severity of the violation; Rectification suggestions: Provide specific modification directions for each violating field; Verification logs include the credibility weights of each data source. Initial risk score Information such as the rule conflict adjudication process is included. Regarding report interaction, users can click on the name of the violation field in the report to directly jump to the corresponding field in the customs declaration editing interface for quick modification; it supports filtering violation details by severity; a re-verification button is provided, allowing users to trigger a re-verification with one click after modifying data, and the system updates the report content; the verification timestamp format in the report is YYYY-MM-DDHH:MM:SS, consistent with the data generation timestamp format, ensuring accurate time traceability. This solution, through a three-level risk classification and detailed visual reports, clarifies processing priorities, solves users' difficulties in judgment, and improves the efficiency of customs declaration modification and declaration.

[0046] Traditional technical solutions have the following technical problems: the verification process lacks logs and audit records, making it impossible to trace historical verification results, and it cannot provide verification basis during customs supervision. In addition, users cannot compare the verification status of different versions after modifying the customs declaration, which affects the troubleshooting of problems.

[0047] Based on this, during the process of re-executing the verification step after the customs declaration is modified, the system saves logs and audit records for each verification. The log content includes the verification time, the version of the regulatory library used, and the credibility weights of the data from each source. Risk Score And modification records; audit records are used to trace the verification process, support manual query of historical verification results, and provide traceable evidence of the verification process for customs supervision.

[0048] It's worth noting that, regarding log content and format, the system uses a combination of text files and a database to store logs. Text files are used for long-term archiving, while the database is used for fast retrieval. Log content includes verification ID, customs declaration number, verification time, regulatory database version, data source information, risk score, risk level, and modification records. The verification ID format is BG+YYYYMMDD+6-digit random number. Data source information includes the name, credibility weight, and other relevant data sources. Value; risk score includes initial risk score. Final risk score The modification record includes the field values ​​before and after the modification, the modification time, and the person who modified it. Regarding log storage and management, text-formatted log files are stored in folders by date, with each file not exceeding 100MB; files exceeding this size are automatically split into new files. Log data in the database is stored in a verification log table, with fields corresponding to log content. Queries are supported by conditions such as customs declaration number, verification time, and risk level. The system sets the log retention period to 3 years; text logs exceeding this period are automatically compressed and archived, and database log data is transferred to a historical data table to ensure reasonable utilization of storage resources. Regarding audit record generation and application, audit records are generated based on log data and include fields such as audit ID, verification ID, audit time, audit content, and audit result. Audit content includes compliance checks during the verification process. Audit results are categorized as pass or fail; if the regulatory database is an older version, the audit result is fail, prompting an update to the regulatory database. Audit records support both manual generation by administrators and automatic generation by the system. Automatically generated audit records are for all verification logs from the previous day, while manually generated audit records are for logs with a specified verification ID. For historical verification result queries, the system provides a historical verification query interface. Users can enter the customs declaration number or verification ID to query all verification records for that customs declaration, displaying the time, risk level, and risk score of each verification. The system modifies records and supports comparing verification results from different versions, displaying changes in fields before and after modification and corresponding risk score changes in a table format. Query results can be exported to Excel for user archiving and troubleshooting. Regarding customs supervision integration, the system provides a customs supervision interface. Customs officials can input the customs declaration number through this interface to query the corresponding verification logs and audit records, verifying the compliance of the verification process and the credibility of the data source. The interface returns data in JSON format, including the verification ID and the legal database version. value value This solution preserves key information such as modification records to ensure that customs supervision has a basis. Through complete logs and audit records, the verification process is traceable, resolving issues such as the inability to query historical results and the lack of a basis for supervision. It also assists users in identifying changes in risk before and after modifications.

[0049] Traditional technical solutions have the following technical problems: the division of verification system modules is chaotic, the functions of each module overlap, the data interaction logic is unclear, which makes system maintenance difficult. At the same time, there is a lack of coordination between modules, making it impossible to achieve full-process automated verification and relying on manual intervention.

[0050] Based on this, please refer to Figure 2This embodiment provides a customs declaration intelligent compliance verification system based on multi-source data, used to execute any of the methods described above. Its features include a multi-source data time-series synchronization module, a data credibility assessment module, a rule engine and conflict resolution module, an AI semantic enhancement verification module, and a risk report generation module. The multi-source data time-series synchronization module is used to interface with the single-window ERP system and documentation documents to acquire data and generate timestamps. And complete the time alignment; the data credibility assessment module is used to calculate the credibility weight according to the formula described in claim 5. The system filters high-confidence datasets; the rule engine and conflict adjudication module are used to call the customs regulation library to perform multi-level verification, based on... Handling rule conflicts; the AI ​​semantic enhancement verification module is used to analyze unstructured data to extract implicit regulatory elements and correct logical deviations in declaration elements. The risk report generation module is used to generate reports based on the final risk score. It classifies risk levels, outputs visual risk control reports, and supports re-verification after customs declarations are modified.

[0051] It's worth noting that in terms of module hardware deployment, the system adopts a distributed architecture. The multi-source data time-series synchronization module is deployed on the enterprise's intranet server, configured with an Intel Xeon E5-2680v4 CPU, 64GB RAM, and 1TB SSD, ensuring secure intranet communication with the ERP system. The data credibility assessment module's rule engine and conflict resolution module are deployed on a cloud server, configured with an AMD EPYC7551 CPU, 128GB RAM, and 2TB SSD, leveraging the cloud server's high-performance computing capabilities to handle complex formula calculations and multi-level verifications. The AI ​​semantic enhancement verification module is deployed on a GPU server, configured with an NVIDIA Tesla V100 GPU, 128GB RAM, and 4TB SSD, improving the speed and accuracy of semantic analysis. The risk report generation module is deployed on a cloud server, communicating with other modules via RESTful APIs to ensure real-time data interaction. Regarding module functionality and data interaction, the multi-source data time-series synchronization module connects to external data sources through preset interfaces to obtain data and generate timestamps. Next, the data is encapsulated in JSON format, including the data source type, data content, and generation timestamp. The file ID field is transmitted to the data credibility assessment module via HTTPS protocol; the data credibility assessment module receives the JSON data, parses it, extracts parameters, and calculates the credibility weight. ,filter The dataset is packaged into a high-confidence dataset JSON, containing the data content. The integrity field of the value field is transmitted to the rule engine and conflict resolution module. After receiving the high-confidence dataset, the rule engine and conflict resolution module call the locally deployed customs regulation library to perform multi-level verification and statistics. and Query the product risk level table to obtain Calculate the initial risk score If there are rule conflicts, consult the trade agreement rule table for details. Query the historical ruling database to obtain Calculate the final risk score ,Will The data is packaged into a risk score JSON and transmitted to the AI ​​semantic enhancement verification module and the risk report generation module. The AI ​​semantic enhancement verification module receives the risk score JSON and unstructured JSON text transmitted by the multi-source data time-series synchronization module, including file ID, text content, and file type fields. It then extracts implicit regulatory elements using a trained model, compares this data with the customs declaration data, and corrects any missing elements. The value will be corrected. The extracted values ​​and regulatory elements are encapsulated into semantically validated JSON and fed back to the rule engine and conflict resolution module; the rule engine and conflict resolution module then recalculate the results. and The risk score JSON is updated and transmitted again to the risk report generation module. The risk report generation module receives the final risk score JSON, violation detail JSON, and semantic verification JSON. The violation detail JSON includes the severity of the rules triggered by the violation fields. It calls an HTML template to generate a visual risk control report, supporting online viewing and PDF download. It also provides a customs declaration editing interface. After the user modifies the data and clicks the re-verification button, the multi-source data time-series synchronization module is triggered to re-acquire the modified data, repeating the entire verification process until a compliant risk report is generated. Regarding module collaboration and maintenance, each module achieves dynamic communication through a service registration and discovery mechanism. When a module fails, the system automatically triggers a backup module to take over. For example, if the AI ​​semantic enhancement verification module fails, a backup semantic analysis module based on keyword matching is activated to ensure system stability. During module maintenance, online upgrades are supported without stopping system operation. Upgrade packages are transmitted through an encrypted channel, and functional testing is automatically performed after the upgrade. The upgrade takes effect upon passing the test; otherwise, it rolls back to the original version. This solution automates the entire process of data collection credibility assessment rule verification and semantic analysis report generation through clear module division and collaborative logic. It solves the problems of traditional system module confusion, maintenance difficulties and reliance on manual intervention, and improves system operation stability and verification efficiency.

[0052] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A method for intelligent compliance verification of customs declarations based on multi-source data, comprising the steps of collecting relevant data from customs declarations and performing preliminary processing, characterized in that, Also includes: It integrates with the Single Window system, ERP system, and document files to obtain data and generate corresponding timestamps, thereby achieving time-series synchronization of multi-source data. Calculate the credibility weight of each source of data and select high-credibility datasets as the verification benchmark; The system calls upon the dynamically updated customs regulations database to perform multi-level checks on basic fields, business logic, and policy compliance of high-credibility datasets, and to count the number of violation rules and the logical deviation between the declared elements. An initial compliance risk score is generated based on the inherent risk level of the product, the credibility weight, and the above statistical results. If there are conflicts between the rules of multiple trade agreements, the final risk score is calculated by combining the priority of the agreements, the degree of rule matching, and the similarity of historical rulings. Semantic analysis is performed on unstructured data to extract implicit regulatory elements, and the accuracy of logical deviation of declaration elements is cross-validated. The risk level is determined based on the final risk score, and a visual risk control report with rectification suggestions is output; the above verification steps can be re-executed after the customs declaration is modified.

2. The method according to claim 1, characterized in that, In the steps of connecting to the Single Window, ERP system, and document documents to obtain data and generate corresponding timestamps, Single Window data is obtained through the official export function, and the exported file formats include Excel and PDF, with the file containing a data generation time field; ERP system data is obtained in real time through the RestAPI interface, and the interface returns data containing a data entry timestamp; document documents are parsed through a document recognition tool, and the file reception time is recorded during the parsing process as a substitute for the data generation timestamp.

3. The method according to claim 1, characterized in that, The step of calling the dynamically updated customs regulatory database to perform multi-level verification includes: HS code classification rules, origin determination standards, regulatory document requirements, business compliance rules, and risk characteristic rules; basic field verification includes non-empty verification, format verification, and numerical range verification; business logic verification includes classification consistency verification, price reasonableness verification, and unit conversion compliance verification; policy compliance verification includes trade agreement applicability verification and license requirement verification; the customs regulatory database is automatically updated by connecting to the official channels of the General Administration of Customs and the Single Window, and also supports manual rule maintenance.

4. The method according to claim 1, characterized in that, In the step of extracting implicit regulatory elements from unstructured data through semantic analysis, the unstructured data includes invoice remarks, contract terms, and packing list descriptions. The semantic analysis is implemented using a large model, which is trained through a regulatory element feature library. The regulatory element feature library includes hazardous materials keywords, 3C certification-related expressions, and energy efficiency labeling-related terms. When cross-validating the accuracy of the logical deviation of the declaration elements, if the regulatory elements extracted by the semantic analysis are not declared in the customs declaration, the logical deviation value of the declaration elements is corrected.

5. The method according to claim 1, characterized in that, In the step of calculating the credibility weight of each source data, the credibility weight is calculated using a formula, which is: ; in, The reliability coefficient for single-window data ranges from 0.8 to 1.

0. It is dimensionless and is used to characterize the official authority and accuracy of single-window data. The reliability coefficient of ERP system data, with a value ranging from 0.7 to 0.95, is dimensionless and is used to characterize the standardization and completeness of data entry in an enterprise's internal ERP system. The reliability coefficient of the document data, with a value ranging from 0.6 to 0.9, is dimensionless and is used to characterize the authenticity and standardization of the external documents. The data generation time is in minutes, specifically the export time of single window data, the entry time of ERP data, or the receipt time of document files. The benchmark time for verification is in minutes, specifically the current time of the system performing the compliance verification. This is the maximum allowable time series deviation threshold, in minutes, with a value of 1440, corresponding to 24 hours, used to limit the upper limit of data timeliness; The ERP data integrity coefficient has a value range of 0 to 1 and is dimensionless. It is calculated by statistically analyzing the complete filling ratio of required fields in ERP data. The document integrity coefficient ranges from 0 to 1 and is dimensionless. It is calculated by statistically analyzing the percentage of missing key information in the document documents.

6. The method according to claim 1, characterized in that, In the step of generating the initial compliance risk score, the initial risk score is calculated using a formula, which is: ; in, The confidence weight is calculated using the formula described in claim 5 and is dimensionless. The inherent risk level of the product is represented by a value ranging from 0.5 to 5, which is dimensionless. High-risk products, such as those involving taxes or licenses, are represented by a value of 3 to 5, ordinary products by a value of 1 to 2, and low-risk products, such as ordinary daily necessities, by a value of 0.5 to 1. The number of rule violations triggered, expressed in units of "rules", is a non-negative integer and is obtained by counting the number of entries that do not conform to the rules in the multi-level validation process. The logical deviation of the declared elements is dimensionless, ranging from 0 to 1, and is expressed by the formula... The calculations are all in RMB (Chinese Yuan). When the actual total price is 0, The value is 1.

7. The method according to claim 1, characterized in that, In the step of calculating the final risk score, the final risk score is calculated using a formula, which is: ; in, The initial risk score calculated using the formula described in claim 6 is dimensionless. This represents the weight of high-priority trade agreements, with a value of 0.

7. It is dimensionless and includes RCEP. This represents the weight of low-priority trade agreements, with a value of 0.

3. It is dimensionless and includes the CPTPP. The degree of matching between high-priority trade agreement rules and commodity characteristics is a dimensionless value ranging from 0 to 1, calculated by comparing the degree of fit between the commodity's HS code, country of origin, and agreement rules. This represents the degree of matching between low-priority trade agreement rules and commodity characteristics, with a value ranging from 0 to 1, dimensionless, and calculated in the same way as... Consistent; The similarity weight of historical rulings ranges from 0.1 to 0.3, is dimensionless, and is used to adjust the degree of influence of historical cases on the final score; The similarity between the current conflict and the adjudication of similar historical conflicts is a dimensionless value ranging from 0 to 1. It is obtained by calculating the similarity between the feature vector of the current conflict and the feature vector of historical cases using the cosine similarity algorithm.

8. The method according to claim 1, characterized in that, In the process of classifying risk levels, the risk levels include three categories: error, warning, and compliance; when the final risk score is... When this is determined to be an error level, the customs declaration data must be forcibly corrected; when When the situation is deemed to be at the warning level, manual review of the customs declaration data is required; when When the system is deemed compliant, a customs declaration can be exported for filing. The visual risk control report includes the risk level, the name of the violation field, the triggered rule content, rectification suggestions, and verification timestamp.

9. The method according to claim 1, characterized in that, During the process of re-executing the verification step after the customs declaration is modified, the system saves logs and audit records for each verification. The log content includes the verification time, the version of the legal library used, the credibility weight of each source data, the risk score, and the modification record. The audit record is used to trace the verification process, supports manual query of historical verification results, and provides traceable evidence of the verification process for customs supervision.

10. A customs declaration intelligent compliance verification system based on multi-source data, used to execute the method described in any one of claims 1 to 9, characterized in that, It includes a multi-source data time-series synchronization module, a data credibility assessment module, a rule engine and conflict adjudication module, an AI semantic enhancement verification module, and a risk report generation module; the multi-source data time-series synchronization module is used to connect with the single window, ERP system and documents, acquire data and generate timestamps and complete time-series alignment; The data credibility assessment module is used to calculate credibility weights according to the formula described in claim 5 and to screen high-credibility datasets; the rule engine and conflict adjudication module is used to call the customs regulation library to perform multi-level verification and to handle rule conflicts according to the formula described in claim 7; the AI ​​semantic enhancement verification module is used to analyze unstructured data to extract implicit regulatory elements and correct the logical deviation of declaration elements; the risk report generation module is used to classify risk levels according to the final risk score, output a visual risk control report, and support re-verification after the customs declaration is modified.