A method for mining customs risk assessment rules based on interpretable deep learning

Through the method based on interpretability deep learning, global and local interpretable rules are generated, which solves the problems of low efficiency and insufficient interpretation of manual experience design in customs risk analysis and analysis, and improves the accuracy of risk assessment and business expert acceptance.

CN119624143BActive Publication Date: 2025-05-02TRS INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510161594.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-02
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The existing customs risk analysis methods rely on manual experience design rules, which have problems of inefficiency and insufficient interpretation, making it difficult to fully consider the rich information in the customs declaration form.

Method used

Using an interpretability deep learning method, through data preprocessing, feature acquisition, global and local rule generation, combined with a decision tree model, we generate customs risk assessment rules that are easy to understand and interpret.

Benefits of technology

It improves the efficiency and accuracy of customs risk assessment rules, ensures comprehensiveness of feature selection, reduces the rate of misjudgment, and provides explanatory rules that are easy to understand and accept by business experts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624143B_ABST
    Figure CN119624143B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of customs risk monitoring technology, and proposes a customs risk assessment rule mining method based on interpretable deep learning. The feature items of the customs declaration form are expanded by cleaning the specification model column and introducing Internet public information from the customs risk knowledge base. The interpretability of the tabnet model is used to locate the key feature items from the expanded many customs declaration feature items. Then, combined with the decision tree model, the feature attribution method is used to generate global risk rules and local risk rules respectively, which provides a reference for the design of customs risk assessment rules. The present invention can improve the efficiency of the design of customs risk assessment rules, ensure the comprehensiveness of feature selection, and make the judgment results more accurate. By generating global risk rules and local risk rules, the present invention comprehensively solves the difficulties faced by customs business experts when designing rules for the customs risk judgment rule engine, that is, the inability to fully consider the many feature items of the customs declaration form and the difficulty in using public information on the Internet.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of customs risk monitoring, and in particular to a customs risk assessment rule mining method based on explainable deep learning. Background Art

[0002] The existing customs risk assessment method mainly relies on manual experience to design if-then rules to screen risk documents from massive customs declarations. However, this method is limited by limited human energy and experience, and secondly, customs declarations contain rich and diverse information, and it is difficult to discover hidden risks based on manual experience alone. At present, customs are also using computer technology to continuously improve the automation and intelligence of risk assessment. With the continuous development and improvement of deep learning technology, deep learning models have been widely used in computer vision and natural language processing, but they are still rarely used in tabular data such as customs declarations. The tabnet deep learning model proposed by Google can effectively process tabular data, and its network structure has both built-in feature selection and interpretability. As an end-to-end technology, tabnet can comprehensively consider all the information of the customs declaration and fully explore the potential connections between data. Summary of the invention

[0003] In order to solve the technical problems that the existing technology is not comprehensive enough in considering the customs risk assessment rules based on manual design, and the customs risk assessment rules based on technologies such as deep learning are not interpretable enough, the present invention proposes a customs risk assessment rule mining method based on interpretable deep learning, which improves the efficiency of customs risk assessment rules in screening customs risk commodities and has a certain degree of interpretability, providing support for customs commodity risk assessment work.

[0004] The specific plan is as follows:

[0005] A customs risk assessment rule mining method based on explainable deep learning.

[0006] S1, data preprocessing: obtain the target type commodity declaration form of the same type as the target type commodity through field screening; split and clean the data in the specification column of the target type commodity declaration form to obtain the "declaration element name" and "declaration element content" fields, and query the customs risk knowledge base to obtain the fusion knowledge data as a supplementary field, and store the customs risk feature field list;

[0007] S2: Obtain features: Take the customs risk feature field list as a sample, input it into the tabnet model, adjust the hyperparameters to obtain the best performance, and then generate the global feature importance and local feature importance mask; the global feature importance is the frequency or weight of each commodity feature composed of the field being selected in all samples for inference whether there is a customs risk;

[0008] S3: Generate globally interpretable rules: Select a set of customs risk feature items based on the importance of global features and combine them with prior high-frequency feature items to form model input, and generate globally interpretable rules for customs risks based on the decision tree model;

[0009] S4: Generate local interpretable rules: obtain black samples in the customs risk feature field list, and generate local interpretable rules using a bottom-up rule learning method based on the importance of local features of the black samples;

[0010] S5: Merge the local interpretable rules and the global interpretable rules and input them into the customs rule engine to obtain customs risk assessment rules.

[0011] Preferably, step S1 specifically includes:

[0012] S11: Obtaining declaration element content: The declaration element content is obtained by cleaning and splitting the data in the specification model column in the customs declaration form information, including: extracting the contract signing date, brand, model, purpose, material, function, ingredient, content, packaging specification and applicable vehicle model;

[0013] S12: Obtain the name of the declaration element: the name of the declaration element is obtained by querying the import and export trademark standard declaration catalog based on the tax number;

[0014] S13: Obtain output result: match the declaration element content with the declaration element name in order to obtain "declaration element name": "declaration element content" as the output result;

[0015] S14: Obtain supplementary fields: introduce fused knowledge data related to the declared commodity from the customs risk knowledge base as supplementary fields to be added to the output result, wherein the fused knowledge data includes: relevant commodity prices, safety precautions and relevant commodity customs risks.

[0016] Preferably, the step of obtaining the importance of local features in the tabnet model in S2 is:

[0017] S21: Feature selection: Through the attention mechanism, a feature subset is selected from the customs risk feature field list at each step to form a feature selection vector Mask;

[0018] S22: Feature transformation: performing feature transformation based on the selected feature subset through the ReLU activation function, that is, multiplying the feature selection vector Mask by the ReLU activation function;

[0019] S23: Obtaining local feature importance: Obtain the multiplication result of each step as the local feature importance.

[0020] Preferably, in S3, the step of generating a global interpretable rule based on the global feature importance is:

[0021] S31: According to the ranking of customs risk features in the global feature importance, the top m% features with the highest scores are extracted from the customs risk feature field list to generate a customs risk feature item set;

[0022] S32: extracting n customs risk feature items with the highest number of uses from all the rules currently being used in the customs rule engine as prior high-frequency feature items;

[0023] S33: adding the two customs risk feature items with the highest frequency of occurrence among the a priori high-frequency feature items in step S32 and not included in step S31 to the customs risk feature item set in step S31;

[0024] S34: Based on the customs risk feature item set sorted out in step S33, an efficient feature data set is extracted from the customs risk feature field list, and input into a decision tree model for training;

[0025] S35: From the final training result of the decision tree model in step S34, extract all qualified customs risk assessment rules and calculate the confidence, output the q customs risk assessment rules with the highest confidence as global interpretable rules, and output them.

[0026] Preferably, step S4 specifically includes:

[0027] S41: Obtain the feature items corresponding to the first w values ​​with the largest values ​​in the feature selection vector Mask of all black samples in each step of the validation set, which are called step-by-step target feature items;

[0028] S42: For feature items of numerical type, the intervals are divided according to the frequency of occurrence of the numerical values, so that each interval contains the same number of data points, that is, the frequency of occurrence of the numerical values ​​is consistent;

[0029] S43: traverse all black samples in the validation set of the tabnet model, and generate rules for judging whether there is customs risk based on each sample in the black samples as the initial special rules;

[0030] S44: based on the generalization strategy, the initial special rule obtained in S43 is gradually generalized to obtain an extended rule, and a likelihood ratio statistic LRS after each step of rule generalization is calculated;

[0031] S45: sorting the likelihood ratio statistic LRS of the expansion rule of each black sample in S44, and selecting the top 10 rules as candidate rules;

[0032] S46: Summarize the candidate rules of all black samples, and select them by likelihood ratio statistic LRS sorting to obtain local interpretable rules, wherein the local interpretable rules are the top 10 candidate rules ranked in the likelihood ratio statistic LRS sorting and the top 10 rules containing the most prior high-frequency feature items.

[0033] Preferably, in step S41, the number w of the step-by-step target feature items is calculated as follows:

[0034] (1)

[0035] Among them, w is the number of step-by-step target feature items, N is the number of feature items in the feature subset in each step; step is the number of steps of the tabnet model.

[0036] Preferably, in step S44, the formula for calculating the likelihood ratio statistic LRS is:

[0037] (2)

[0038] : The number of classes included in the rule coverage samples;

[0039] :kind The number of samples contained in ;

[0040] : The expected frequency with which the rule makes random guesses in the dataset.

[0041] Preferably, in step S44, the generalization strategy is:

[0042] A1: Generalize the target feature item of the last step obtained from the tabnet model to the target feature item of the first step;

[0043] A2: In the target feature item of each step, the value of each target feature item in the local feature importance is obtained, and the coverage of the feature value in the target feature item is expanded horizontally from small to large, and then the feature items are deleted in turn, that is, the restrictions of the feature items on the customs risk assessment rules are lifted, and the coverage of the customs risk assessment rules is expanded vertically.

[0044] Beneficial effects:

[0045] The present invention discloses a method for mining customs risk assessment rules based on interpretable deep learning. The method uses the interpretability of tabnet to locate key feature items from a large number of customs declaration feature items, and then combines the decision tree model to generate global risk rules and local risk rules using a feature attribution method, providing a reference for the design of customs risk assessment rules. The present invention can improve the efficiency of the design of customs risk assessment rules, ensure the comprehensiveness of feature selection, and make the assessment results more accurate. By designing global interpretable rules and local interpretable rules, the present invention comprehensively solves the difficulties faced by customs business experts when designing rules acting on the customs risk assessment rule engine, namely, the inability to fully consider the many feature items of the customs declaration form and the difficulty in using public information on the Internet.

[0046] First, in the process of data preprocessing, the present invention designs a cleaning method based on the commodity specification model column, and adds supplementary fields through the customs risk knowledge base, solving the problem of difficulty in using public information on the Internet. The present invention cleans the data of the target type commodity declaration form, especially the commodity specification model column that is difficult to use directly, and the customs risk feature field generated after cleaning can be directly input into the tabnet model. At the same time, the customs risk knowledge base constructed based on Internet data is used to generate supplementary fields, solving the problem of difficulty in using public information on the Internet.

[0047] Secondly, the present invention constructs globally interpretable rules that are easily accepted and recognized by business experts by designing the global feature importance. If the pre-processed field features are directly used to generate customs risk assessment rules, due to the complexity of customs business and too many feature fields, even if the time complexity of the rule generation process is ignored, these rules are still difficult for business experts to read and understand, and cannot provide them with effective reference value. However, by applying the global feature importance evaluation method proposed in the present invention, the features can be ranked by importance, and a few key features can be selected based on the insights of business experts. Based on these highly important features, globally interpretable rules that are easily accepted and recognized by business experts can be constructed.

[0048] Third, the present invention further optimizes the global interpretable rules by constructing local interpretable rules to make them more in line with the current business logic, thereby reducing the misjudgment rate. In view of the complexity and variability of customs business, those global interpretable rules generated based on overall historical business data may not be able to fully capture the latest business logic. Local interpretable rules are derived by generalizing the decision-making process and feature contribution of the model under specific black sample instance inputs. They can effectively correct the inaccurate predictions that may be generated by global interpretable rules under the current business logic, thereby reducing misjudgments and deviations. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 Flowchart of a customs risk assessment rule mining method based on explainable deep learning.

[0050] Figure 2 Data flow diagram of a customs risk assessment rule mining method based on explainable deep learning.

[0051] Figure 3 A flowchart of generating a customs risk feature field list through preprocessing in an embodiment.

[0052] Figure 4 A partial display of the results of the importance of local features in the embodiments. DETAILED DESCRIPTION

[0053] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0054] like Figure 1 and Figure 2 As shown in the figure, a method for mining customs risk assessment rules based on interpretable deep learning is proposed.

[0055] S1, data preprocessing: obtain the target type commodity declaration form of the same type as the target type commodity through field screening; split and clean the data of the target type commodity declaration form to obtain the "declaration element name" and "declaration element content" fields, and query the customs risk knowledge base to obtain the fusion knowledge data as a supplementary field, and store it to obtain the customs risk feature field list;

[0056] S2: Obtain features: Take the customs risk feature field list as a sample, input it into the tabnet model, adjust the hyperparameters to obtain the best performance, and then generate the global feature importance and local feature importance mask; the global feature importance is the frequency or weight of each commodity feature composed of the field being selected in all samples for inference whether there is a customs risk;

[0057] S3: Generate globally interpretable rules: Select a set of customs risk feature items based on the importance of global features and combine them with prior high-frequency feature items to form model input, and generate globally interpretable rules for customs risks based on the decision tree model;

[0058] S4: Generate local interpretable rules: obtain black samples in the customs risk feature field list, and generate local interpretable rules using a bottom-up rule learning method based on the importance of local features of the black samples;

[0059] S5: Merge the local interpretable rules and the global interpretable rules and input them into the customs rule engine to obtain customs risk assessment rules.

[0060] Preferably, if Figure 3 As shown, step S1 specifically includes:

[0061] S11: Obtaining declaration element content: The declaration element content is obtained by cleaning and splitting the data in the specification model column in the customs declaration form information, including: extracting the contract signing date, brand, model, purpose, material, function, ingredient, content, packaging specification and applicable vehicle model, etc.;

[0062] S12: Obtain the name of the declaration element: the name of the declaration element is obtained by querying the import and export trademark standard declaration catalog based on the tax number;

[0063] S13: Obtain output result: match the declaration element content with the declaration element name in order to obtain "declaration element name": "declaration element content" as the output result;

[0064] S14: Obtain supplementary fields: introduce fused knowledge data related to the declared commodity from the customs risk knowledge base as supplementary fields to be added to the output result, wherein the fused knowledge data includes: relevant commodity prices, safety precautions and relevant commodity customs risks.

[0065] In this embodiment, the customs risk knowledge base further covers multiple fields including Internet public information, such as commodity prices (for different sales markets), origin, brand, packaging type, and production and operation dynamics of the companies involved, etc. This information can be added to the output results as supplementary fields according to specific business needs.

[0066] The customs big data resource pool contains the import and export declaration data of all commodities. Different commodities have different import and export requirements. Therefore, in order to extract effective customs risk rules from business data more accurately, it is necessary to provide business experts with an operation interface for data screening during the data preprocessing stage. In this interface, business experts can filter out the declaration forms of target type commodities through different fields.

[0067] In the embodiment, the specification and model column of the customs declaration form is a field in the customs declaration form, which contains the specification and model information of the commodity: 4|3|for laboratory use|silicone rubber|NITTO|PQA-43|00000000|00-00-0. The purpose of cleaning this field is to split the specification and model text, extract the content corresponding to the declaration elements such as the signing date, brand and model, and fill in the corresponding content of each data in its corresponding declaration element field to prepare for subsequent model processing.

[0068] Preferably, the step of obtaining the importance of local features in the tabnet model in S2 is:

[0069] S21: Feature selection: Through the attention mechanism, a feature subset is selected from the customs risk feature field list at each step to form a feature selection vector Mask;

[0070] S22: Feature transformation: performing feature transformation based on the selected feature subset through the ReLU activation function, that is, multiplying the feature selection vector Mask by the ReLU activation function;

[0071] S23: Obtaining the importance of local features: Obtain the multiplication result of each step as the importance of local features. Figure 4 As shown, some results of the importance of local features of the tabnet model are presented.

[0072] In the embodiment, the declaration form fields plus the fields cleaned from the specification and model column plus the Internet public information imported from the customs risk knowledge base query have a total of nearly 200 feature items. Generating rules based on so many feature items is difficult for business experts to read and understand even if the time complexity is not considered, and it cannot provide effective reference for business experts. With the global feature importance, we can sort based on the global feature importance, and then select a few features based on the opinions of business experts, and generate global interpretable rules that are easily accepted and recognized by business experts based on these highly important features.

[0073] Preferably, in S3, the step of generating a global interpretable rule based on the global feature importance is:

[0074] S31: According to the ranking of customs risk features in the global feature importance, the top m% features with the highest scores are extracted from the customs risk feature field list to generate a customs risk feature item set;

[0075] S32: extracting n customs risk feature items with the highest number of uses from all the rules currently being used in the customs rule engine as prior high-frequency feature items;

[0076] S33: adding the two customs risk feature items with the highest frequency of occurrence among the a priori high-frequency feature items in step S32 and not included in step S31 to the customs risk feature item set in step S31;

[0077] S34: Based on the customs risk feature item set sorted out in step S33, an efficient feature data set is extracted from the customs risk feature field list, and input into a decision tree model for training;

[0078] S35: From the final training result of the decision tree model in step S34, extract all qualified customs risk assessment rules and calculate the confidence, output the q customs risk assessment rules with the highest confidence as global interpretable rules, and output them.

[0079] In the embodiment, the top 10% features with the highest scores can be extracted from the customs risk feature field list, and then the 10 most frequently used feature items are extracted from all the rules currently being used in the customs rule engine, which are called prior high-frequency feature items. The two most frequently used feature items that did not appear in S32 are added to the feature items in S31. All rules are extracted from the final training results of the decision tree model and the confidence is calculated. The 10 rules with the highest confidence can be selected as the global interpretable rule output. The number of rules and the ratio of feature extraction can be set manually.

[0080] In an embodiment, the division of the training set and the test set of the decision tree model may be consistent with that of the tabnet model.

[0081] Preferably, step S4 specifically includes:

[0082] S41: Obtain the feature items corresponding to the first w values ​​with the largest values ​​in the feature selection vector Mask of all black samples in each step of the validation set, which are called step-by-step target feature items;

[0083] S42: For feature items of numerical type, the intervals are divided according to the frequency of occurrence of the numerical values, so that each interval contains the same number of data points, that is, the frequency of occurrence of the numerical values ​​is consistent;

[0084] S43: traverse all black samples in the validation set of the tabnet model, and generate rules for judging whether there is customs risk based on each sample in the black samples as the initial special rules;

[0085] S44: based on the generalization strategy, the initial special rule obtained in S43 is gradually generalized to obtain an extended rule, and a likelihood ratio statistic LRS after each step of rule generalization is calculated;

[0086] S45: sorting the likelihood ratio statistic LRS of the expansion rule of each black sample in S44, and selecting the top 10 rules as candidate rules;

[0087] S46: Summarize the candidate rules of all black samples, and select them by likelihood ratio statistic LRS sorting to obtain local interpretable rules, wherein the local interpretable rules are the top 10 candidate rules ranked in the likelihood ratio statistic LRS sorting and the top 10 rules containing the most prior high-frequency feature items.

[0088] Preferably, in step S41, the number w of the step-by-step target feature items is calculated as follows:

[0089] (1)

[0090] Among them, w is the number of step-by-step target feature items, N is the number of feature items in the feature subset in each step; step is the number of steps of the tabnet model.

[0091] Preferably, in step S44, the formula for calculating the likelihood ratio statistic LRS is:

[0092] (2)

[0093] : The number of classes included in the rule coverage samples;

[0094] :kind The number of samples contained in ;

[0095] : The expected frequency with which the rule makes random guesses in the dataset.

[0096] Preferably, in step S44, the generalization strategy is:

[0097] A1: Generalize the target feature item of the last step obtained from the tabnet model to the target feature item of the first step;

[0098] A2: In the target feature item of each step, the value of each target feature item in the local feature importance is obtained, and the coverage of the feature value in the target feature item is expanded horizontally from small to large, and then the feature items are deleted in turn, that is, the restrictions of the feature items on the customs risk assessment rules are lifted, and the coverage of the customs risk assessment rules is expanded vertically.

[0099] The important features of each step of the tabnet model are different. For example, the process of generalizing a black sample is:

[0100] The top five local features, i.e., target features, ranked in importance among the local features obtained in the first step of the tabnet model are: import and export type, country of production and sales, declared quantity, total transaction price of materials, and number of pieces;

[0101] The top five local features, i.e., target feature items, ranked in importance of local features obtained in the second step of the tabnet model are: declaration date, brand, model, purpose, and material;

[0102] First, we expand the target feature items of the second step, and then expand the target feature items of the first step. When executing the expansion rules of the second step, we first expand the coverage of the feature values. For example, for the two feature values ​​with declaration dates of January 20, 2025 and January 24, 2025, we can summarize them as the time interval of January 2025, and generate new expansion rules based on this. At the same time, we compare the likelihood ratio statistics LRS before and after the rule expansion. Secondly, we expand each target feature item. For example, in the second step, we only consider the four feature items of declaration date, model, purpose and material, and compare the likelihood ratio statistics LRS before and after the rule expansion.

[0103] Compared with directly inputting the customs commodity risk characteristics into the tabnet model, the preprocessing method of the present invention is used to deeply clean the specification model column of the customs declaration form and extract the relevant content of the customs knowledge base as a supplementary field, and then input it into the tabnet model. The model can more effectively identify key features. In addition, the accuracy of the training set of the model is stable at about 94%, and the accuracy of the verification set is stable at about 92%. At the same time, the global interpretable rules and local interpretable rules generated by the present invention are effective after review and feedback from customs business experts and testing of the customs risk assessment rule engine. This greatly improves the efficiency of customs business experts in designing rules, and significantly improves the coverage of rules in a simple and efficient manner through the algorithm of the present invention, which is several times higher than before. In particular, the introduction of Internet public information from the customs risk knowledge base, such as commodity price and other fields, significantly improves the real hit effect of the rules.

[0104] It should be noted that the above-described specific implementations can enable those skilled in the art to more fully understand the invention, but do not limit the invention in any way. Therefore, although this specification has described the invention in detail with reference to the drawings and embodiments, those skilled in the art should understand that the invention can still be modified or replaced by equivalents. In short, all technical solutions and improvements that do not deviate from the spirit and scope of the invention should be included in the protection scope of the patent for the invention.

Claims

1. A customs risk assessment rule mining method based on explainable deep learning, characterized in that: S1, data preprocessing: obtain the target type commodity declaration form of the same type as the target type commodity through field screening; split and clean the data of the specification column of the target type commodity declaration form to obtain the "declaration element name" and "declaration element content" fields, and query the customs risk knowledge base to obtain the fusion knowledge data as a supplementary field, and store it to obtain the customs risk feature field list; S2: Get features: Take the list of customs risk feature fields as samples, input them into the tabnet model, adjust the hyperparameters to get the best performance, and then generate the global feature importance and local feature importance mask; The global feature importance is the frequency or weight of each commodity feature composed of fields being selected in all samples for inferring whether there is a customs risk; S3: Generate globally interpretable rules: Select a set of customs risk feature items based on the importance of global features and combine them with prior high-frequency feature items to form model input, and generate globally interpretable rules for customs risks based on the decision tree model; S4: Generate local interpretable rules: obtain black samples in the customs risk feature field list, and generate local interpretable rules using a bottom-up rule learning method based on the importance of local features of the black samples; S5: Merge the local interpretable rules and the global interpretable rules and input them into the customs rule engine to obtain customs risk assessment rules.

2. A method for mining customs risk assessment rules based on explainable deep learning according to claim 1, characterized in that: The S1 step specifically includes: S11: Obtaining declaration element content: The declaration element content is obtained by cleaning and splitting the data in the specification model column in the customs declaration form information, including: extracting the contract signing date, brand, model, purpose, material, function, ingredient, content, packaging specification and applicable vehicle model; S12: Obtain the name of the declaration element: the name of the declaration element is obtained by querying the import and export trademark standard declaration catalog based on the tax number; S13: Obtain output result: match the declaration element content with the declaration element name in order to obtain "declaration element name": "declaration element content" as the output result; S14: Obtain supplementary fields: Introduce fused knowledge data related to the declared commodity from the customs risk knowledge base as supplementary fields to the output result, wherein the fused knowledge data includes: relevant commodity prices, safety precautions and relevant commodity customs risks. The customs risk knowledge base is constructed by local data and Internet public data.

3. A method for mining customs risk assessment rules based on explainable deep learning according to claim 1, characterized in that: Steps to obtain the importance of local features in the tabnet model in S2: S21: Feature selection: Through the attention mechanism, a feature subset is selected from the customs risk feature field list at each step to form a feature selection vector Mask; S22: Feature transformation: performing feature transformation based on the selected feature subset through the ReLU activation function, that is, multiplying the feature selection vector Mask by the ReLU activation function; S23: Obtaining local feature importance: Obtain the multiplication result of each step as the local feature importance.

4. A method for mining customs risk assessment rules based on explainable deep learning according to claim 1, characterized in that: In S3, the steps to generate global interpretable rules based on global feature importance are: S31: According to the ranking of customs risk features in the global feature importance, the top m% features with the highest scores are extracted from the customs risk feature field list to generate a customs risk feature item set; S32: extracting n customs risk feature items with the highest number of uses from all the rules currently being used in the customs rule engine as prior high-frequency feature items; S33: adding the two customs risk feature items with the highest frequency of occurrence among the a priori high-frequency feature items in step S32 and not included in step S31 to the customs risk feature item set in step S31; S34: Based on the customs risk feature item set sorted out in step S33, an efficient feature data set is extracted from the customs risk feature field list, and input into a decision tree model for training; S35: From the final training result of the decision tree model in step S34, extract all qualified customs risk assessment rules and calculate the confidence, output the q customs risk assessment rules with the highest confidence as global interpretable rules, and output them.

5. A method for mining customs risk assessment rules based on explainable deep learning according to claim 1, characterized in that: The S4 step specifically includes: S41: Obtain the feature items corresponding to the first w values ​​with the largest values ​​in the feature selection vector Mask of all black samples in each step of the validation set, which are called step-by-step target feature items; S42: For feature items of numerical type, the intervals are divided according to the frequency of occurrence of the numerical values, so that each interval contains the same number of data points, that is, the frequency of occurrence of the numerical values ​​is consistent; S43: traverse all black samples in the validation set of the tabnet model, and generate rules for judging whether there is customs risk based on each sample in the black samples as the initial special rules; S44: based on the generalization strategy, the initial special rule obtained in S43 is gradually generalized to obtain an extended rule, and a likelihood ratio statistic LRS after each step of rule generalization is calculated; S45: sorting the likelihood ratio statistic LRS of the expansion rule of each black sample in S44, and selecting the top 10 rules as candidate rules; S46: Summarize the candidate rules of all black samples, and select them by likelihood ratio statistic LRS sorting to obtain local interpretable rules, wherein the local interpretable rules are the top 10 candidate rules ranked in the likelihood ratio statistic LRS sorting and the top 10 rules containing the most prior high-frequency feature items.

6. A method for mining customs risk assessment rules based on explainable deep learning according to claim 5, characterized in that: In step S41, the number w of the step-by-step target feature items is calculated as follows: (1) Among them, w is the number of step-by-step target feature items, N is the number of feature items in the feature subset in each step; step is the number of steps of the tabnet model.

7. A method for mining customs risk assessment rules based on explainable deep learning according to claim 5, characterized in that: In step S44, the formula for calculating the likelihood ratio statistic LRS is: (2) : The number of classes included in the rule coverage samples; :kind The number of samples contained in ; : The expected frequency with which the rule makes random guesses in the dataset.

8. A method for mining customs risk assessment rules based on explainable deep learning according to claim 5, characterized in that: In step S44, the generalization strategy is: A1: Generalize the target feature item of the last step obtained from the tabnet model to the target feature item of the first step; A2: In the target feature item of each step, the value of each target feature item in the local feature importance is obtained, and the coverage of the feature value in the target feature item is expanded horizontally from small to large, and then the feature items are deleted in turn, that is, the restrictions of the feature items on the customs risk assessment rules are lifted, and the coverage of the customs risk assessment rules is expanded vertically.

Citation Information

Patent Citations

  • Customs business data structured management method and device, computer equipment group and storage medium

    CN116450753A

  • Customs declaration document information risk rule generation method and system

    CN116542523A

  • Automatic generation method and device for security rule of intrusion detection system, and storage medium

    CN117113337A