Customs risk indicator screening method, system and medium based on random forest algorithm
By establishing and optimizing customs risk stochastic forests based on stochastic forest algorithm, the problem of indicator deviation in customs risk indicator screening is solved, and more accurate data screening and judgment are achieved.
Patent Information
- Application Number
- CN202410885970.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-03
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-07-03
AI Technical Summary
The existing technology has index deviations in data screening in customs risk indicator screening, which makes it impossible to accurately determine whether the data meets the standards.
The random forest algorithm is used to obtain multiple customs risk types, establish and optimize the random forest for each customs risk, and use random forests to screen customs risk indicators to avoid deviations caused by feature selection.
It improves the accuracy of customs risk data screening, ensures that the screening results can correctly reflect whether the data meets the standards, and provide a reliable basis for subsequent processing.
Smart Images

Figure CN119066528B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of customs risk management, and in particular to a method, system and medium for screening customs risk indicators based on a random forest algorithm. Background Art
[0002] Customs risk refers to potential non-compliance with customs laws. Existing improvements for screening customs risk indicators generally involve screening data, evaluating the screened data, and issuing returns or warnings based on the evaluation results. For example, Chinese patent application publication number CN108959648A discloses a customs business declaration risk verification system. This solution uses the parameter rules to filter data from the parameter table, evaluates the screened data, and issues returns or warnings for customs business that does not meet the requirements.
[0003] Improvements in the existing technology for screening customs risk indicators usually focus on accurate inspection of goods, but there is a lack of effective improvement methods for screening data in customs risks. When screening data in customs risks using existing data screening methods, the need to perform feature selection on the data will cause deviations in the screening indicators, resulting in the screening results being unable to accurately judge whether the data in customs risks meet the standards. In view of this, it is necessary to improve the existing customs risk indicator screening methods. Summary of the Invention
[0004] The present invention aims to solve, at least to a certain extent, one of the technical problems in the prior art. By proposing a customs risk indicator screening method, system and medium based on a random forest algorithm, the present invention is used to solve the problem in the prior art of screening data in customs risks. When screening data in customs risks using existing data screening methods, the need to perform feature selection on the data may cause deviations in the screening indicators, resulting in the screening results being unable to accurately reflect whether the data in the customs risks meet the standards.
[0005] To achieve the above objectives, in a first aspect, the present invention provides a method for screening customs risk indicators based on a random forest algorithm, comprising:
[0006] Get multiple customs risk types, which are recorded as customs risk FX1 to customs risk FX N ;
[0007] Analyze all customs risks separately and obtain the random forest corresponding to each customs risk;
[0008] When new customs risk indicators are obtained, the customs risk indicators are screened based on the random forest corresponding to each customs risk, and corresponding processing is performed based on the screening results.
[0009] Furthermore, all customs risks are analyzed separately, and the random forest corresponding to each customs risk is obtained, including:
[0010] For Customs Risk FX1 to Customs Risk FX N Any customs risk FX N1 , using the random forest construction method to obtain customs risk FX N1 Corresponding data subtable ZB N1 Based on multiple decision trees, the customs risk FX N1 Initial random forest of
[0011] FX for customs risk N1 The initial random forest is optimized, and the customs risk FX is obtained based on the optimization results. N1 Random Forest;
[0012] Use the random forest construction method to obtain multiple decision trees for all customs risks and the initial random forests corresponding to the multiple decision trees. Optimize all the initial random forests to obtain the random forests corresponding to all customs risks.
[0013] Furthermore, the random forest construction method includes:
[0014] Obtaining Customs Risk FX N1 The data names of the multiple data contained in the data are sequentially recorded as data name SC1 to data name SC M ;
[0015] For data name SC1 to data name SC M Any data name in SC M1 ;
[0016] Get data name SC M1 The most recently obtained standard historical quantity data in the corresponding data is recorded as data name SC M1 Reference history database;
[0017] Set the data name to SC M1 The maximum value, minimum value, average value and variance in the reference history database are recorded as U1, U2, U3 and U4 respectively, and U1 to U4 are recorded as data names SC M1 Data characteristics;
[0018] Get data name SC M1 The data measurement standard is SC M1 Data characteristics and data name SC M1 The data measurement standard is compared, when the data name SC M1 Any of the data characteristics does not match the data name SCM1 When measuring the data standard, the data name SC M1 Recorded as failed; when the data name SC M1 All data characteristics of the data name SC M1 When measuring the data standard, the data name SC M1 Recorded as passed;
[0019] Get the data characteristics and pass status of all data names, where the pass status includes failed and passed.
[0020] Furthermore, the random forest construction method also includes:
[0021] Create a table with T rows and Y columns, record it as the data table SB N1 , of which the data table SB N1 The first row of the table SB is filled with the maximum value, minimum value, average value, variance and pass status from left to right except the first cell. N1 Except for the first cell in the leftmost column, fill in the data names SC1 to SC from top to bottom. M ;
[0022] Based on data name SC1 to data name SC M The data characteristics and passing status are shown in the data summary table SB N1 Fill in the corresponding data.
[0023] Furthermore, the random forest construction method also includes:
[0024] Create multiple tables with the same format, each of which is recorded as a data sub-table ZB N11 To data subtable ZB N1K ;
[0025] For data subtable ZB N11 To data subtable ZB N1K Any data subtable ZB in N1K1 , data subtable ZB N1K1 It is a table with T rows and Y columns, and the data subtable ZB N1K1 Fill in the maximum value, minimum value, average value, variance and pass status in the first row except the first cell from left to right;
[0026] In data name SC1 to data name SC M Extract M data names with replacement and record them as sub-table data names ZM K1 , in the data subtable ZB N1K1 Except for the first cell in the leftmost column, fill in the subtable data name ZM from top to bottom K1 ; Based on the subtable data name ZM K1The data characteristics and passing status of each data name are in the data sub-table ZB N1K1 Fill in the corresponding data;
[0027] Randomly obtain J data features from the maximum value, minimum value, average value and variance and record them as decision points, and establish a table of T rows × J1 columns, which is recorded as data subtable ZB N1K1 The first row of the analysis subtable is filled with J decision points from left to right except the first cell, and the leftmost column of the analysis subtable is filled with the subtable data name ZM from top to bottom except the first cell. K1 , and in the analysis sub-table based on the sub-table data name ZM K1 Fill in the corresponding data for each data name in the data feature;
[0028] A decision tree is established based on the data in the analysis subtable, wherein the decision nodes of the decision tree are determined by the values of the decision points, and the end points of the decision tree are the passing conditions of the data types;
[0029] Get all data subtable ZB N1 The decision tree will be used to store all data in the sub-table ZB N1 The decision tree is denoted as customs risk FX N1 The initial random forest.
[0030] Furthermore, the customs risk FX N1 The initial random forest is optimized, and the customs risk FX is obtained based on the optimization results. N1 The random forest consists of:
[0031] Randomly obtain customs risk FX N1 Any set of data names SC1 to SC in the historical data M The corresponding data is recorded as test data, the data features of the test data are obtained, and the passing conditions of all data names in the test data are recorded as standard passing conditions BZ1 to standard passing conditions BZ M ;
[0032] Based on customs risk FX N1 All decision trees in the initial random forest use the decision tree processing method to process all test data and their data features, and obtain the passing status of the data names of all test data in the initial random forest, which are recorded as forest passing status SL1 to forest passing status SL M ;
[0033] The decision tree processing method includes: obtaining the decision point of the decision node in the decision tree, and obtaining the numerical name of the same as the decision point in the data feature of the test data, recording it as the decision name, substituting all the data corresponding to the decision name in the test data into the decision tree containing the decision name, and obtaining the terminal point to which the data corresponding to the decision name leads in the decision tree; for any data name SC in the test data M1 , for data name SC M1 Any data feature TZ A , get the data feature TZ A The number of decision trees inserted after the decision name is recorded as L, and the data feature TZ is obtained. A The decision name is recorded and substituted into the decision tree. The endpoint is the number of passed points, recorded as L1. When L1 is greater than L / 2, the data feature TZ A The processing result is recorded as passed. When L1 is less than or equal to L / 2, the data feature TZ A The result of the processing is recorded as failed; when the data name SC M1 When all data characteristics of the data are passed, the data name SC M1 The forest passing situation is recorded as passed, when the data name SC M1 If any of the data features is not passed, the data name SC M1 The forest pass situation is recorded as failed.
[0034] Furthermore, the customs risk FX N1 The initial random forest is optimized, and the customs risk FX is obtained based on the optimization results. N1 The random forest also includes:
[0035] When the standard pass situation BZ1 and forest pass situation SL1 to the standard pass situation BZ M and forest through the situation SL M When the values are equal, the test data is recorded as successful test data; when the standard passes the case BZ1 and the forest passes the case SL1 to the standard passes the case BZ M and forest through the situation SL M Any one of the criteria in BZ M1 and forest through the situation SL M1 If they are not equal, the test data is recorded as failed test data;
[0036] When test data is recorded as failed test data, add the test data to the Customs Risk FX N1 The corresponding reference history database is retrieved and the customs risk FX N1 The corresponding initial random forest;
[0037] When the test data is recorded as successful test data, the initial random forest is recorded as customs risk FX N1 Random Forest.
[0038] Furthermore, when new customs risk indicators are obtained, the customs risk indicators are screened based on the random forest corresponding to each customs risk, and corresponding processing based on the screening results includes:
[0039] When a new customs risk index is obtained, the customs risks in the customs risk index are classified based on the different customs risk types corresponding to the customs risk index. Based on the classification results, the customs risks obtained are recorded as risks to be analyzed DFX1 to risks to be analyzed DFX G ;
[0040] For the risks to be analyzed DFX1 to DFX G Any of the risks to be analyzed DFX G1 , using the risk to be analyzed DFX G1 Random Forest for Risk Analysis DFX G1 Analyze the data in the analysis and obtain the risk DFX to be analyzed based on the analysis results. G1 The customs risk indicators are screened based on the passing status of each data.
[0041] In a second aspect, the present invention also provides a customs risk indicator screening system based on a random forest algorithm, comprising a risk acquisition module, a random forest generation module, and a random forest application module:
[0042] The risk acquisition module is used to obtain multiple customs risk types, which are recorded as customs risk FX1 to customs risk FX N ;
[0043] The random forest generation module is used to analyze all customs risks separately and obtain the random forest corresponding to each customs risk;
[0044] The random forest application module is used to screen the customs risk indicators based on the random forest corresponding to each customs risk when new customs risk indicators are obtained, and perform corresponding processing based on the screening results.
[0045] In a third aspect, the present invention provides a storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps in the above method are executed.
[0046] Beneficial effects of the present invention: The present invention obtains multiple customs risk types and records them in sequence as customs risk FX1 to customs risk FX N, and then analyze all customs risks separately to obtain the random forest corresponding to each customs risk. Finally, when a new customs risk indicator is obtained, the customs risk indicator is screened based on the random forest corresponding to each customs risk, and corresponding processing is performed based on the screening results. The advantage of this is that when analyzing customs risks, by establishing a random forest and using the random forest to screen the data corresponding to the customs risks, the feature subset corresponding to the data corresponding to the customs risks can be randomly extracted, thereby avoiding feature selection when screening the data in the customs risks, and then preventing the deviation of the indicators during screening, so that the overall data screening can more correctly judge whether the customs risks meet the standards, which is helpful for further processing of the screened data.
[0047] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or understood by practicing the present application. The purposes and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a principle block diagram of the system of the present invention;
[0049] Figure 2 is a flow chart of the steps of the method of the present invention;
[0050] Figure 3 is a schematic diagram of a decision tree of the present invention;
[0051] Figure 4 This is a partial processing diagram of the decision tree processing method of the present invention. DETAILED DESCRIPTION
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0053] Example 1, please refer to Figure 1 As shown, in the first aspect, this application provides a customs risk indicator screening system based on the random forest algorithm, including a risk acquisition module, a random forest generation module and a random forest application module:
[0054] The risk acquisition module is used to obtain multiple customs risk types, which are recorded as customs risk FX1 to customs risk FX N ;
[0055] The random forest generation module is used to analyze all customs risks separately and obtain the random forest corresponding to each customs risk;
[0056] The random forest generation module is configured with a forest building strategy. The forest building strategy includes the first forest building sub-strategy and the second forest building sub-strategy. The first forest building sub-strategy includes:
[0057] For Customs Risk FX1 to Customs Risk FX N Any customs risk FX N1 , using the random forest construction method to obtain customs risk FX N1 Corresponding data subtable ZB N1 Based on multiple decision trees, the customs risk FX N1 Initial random forest of
[0058] Random forest construction methods include:
[0059] Obtaining Customs Risk FX N1 The data names of the multiple data contained in the data are sequentially recorded as data name SC1 to data name SC M ;
[0060] For data name SC1 to data name SC M Any data name in SC M1 ;
[0061] Get data name SC M1 The most recently obtained standard historical quantity data in the corresponding data is recorded as data name SC M1 Reference history database;
[0062] Set the data name to SC M1 The maximum value, minimum value, average value and variance in the reference history database are recorded as U1, U2, U3 and U4 respectively, and U1 to U4 are recorded as data names SC M1 Data characteristics;
[0063] In the specific implementation process, the standard historical data volume can be determined according to the number of data that can be analyzed in the historical data corresponding to the data name. In this embodiment, the standard historical volume is recorded as 5. For example, in one test, the data name SC M1 is the number of import and export transactions, and the reference historical database corresponding to the import and export transaction numbers is 5000, 6000, 6550, 5800, and 5300. Then, by calculation, the data characteristics of the import and export transaction numbers are: U1 equals 6550, U2 equals 5000, U3 equals 5730, and U4 equals 293600;
[0064] Get data name SC M1 The data measurement standard is SC M1 Data characteristics and data name SC M1 The data measurement standard is compared, when the data name SC M1 Any of the data characteristics does not match the data name SC M1 When measuring the data standard, the data name SC M1 Recorded as failed; when the data name SC M1 All data characteristics of the data name SC M1 When measuring the data standard, the data name SC M1 Recorded as passed;
[0065] In the specific implementation process, the data measurement standard can be changed according to the standard corresponding to the data name. The purpose of the data measurement standard is to determine whether the data corresponding to the data name meets the standard. For example, in a test, the data name SC M1 The number of import and export transactions, and when the data measurement criteria for the number of import and export transactions are: U1 is greater than or equal to 6000, U2 is greater than or equal to 4500, U3 is greater than or equal to 5500, and U4 is less than or equal to 300000, the data characteristics of the number of import and export transactions obtained by calculation are: U1 equals 6550, U2 equals 5000, U3 equals 5730, and U4 equals 293600, then all the data characteristics of the number of import and export transactions obtained by comparison meet the data measurement criteria for the number of import and export transactions, and the number of import and export transactions is recorded as passed;
[0066] Get the data characteristics and pass status of all data names, where the pass status includes failed and passed.
[0067] Random forest construction methods also include:
[0068] Create a table with T rows and Y columns, record it as the data table SB N1 , of which the data table SB N1 The first row of the table SB is filled with the maximum value, minimum value, average value, variance and pass status from left to right except the first cell. N1 Except for the first cell in the leftmost column, fill in the data names SC1 to SC from top to bottom. M In the specific implementation process, T is set to M+1, Y is set to 6, and the data table SB N1 Please refer to Table 1 for the filling format. Table 1 is the data summary table SB N1 Fill out the form;
[0069] Table 1: Data Summary Table SB N1 Fill out the form
[0070]
[0071] In the specific implementation process, during a data analysis, the data names SC1 to SC M They are import and export transaction quantity, import and export transaction price, import and export tax payment quantity, import and export tax unpaid quantity and import and export transportation cost, and the data characteristics corresponding to each data name and the passing situation are shown in Table 2. Table 2 is the data summary table SB for this data analysis. N1 The data entry table:
[0072] Table 2: Data Summary Table SB N1 Data entry table
[0073]
[0074] Based on data name SC1 to data name SC M The data characteristics and passing status are shown in the data summary table SB N1 Fill in the corresponding data.
[0075] Random forest construction methods also include:
[0076] Create multiple tables with the same format, each of which is recorded as a data sub-table ZB N11 To data subtable ZB N1K ;
[0077] For data subtable ZB N11 To data subtable ZB N1K Any data subtable ZB in N1K1 , data subtable ZB N1K1 It is a table with T rows and Y columns, and the data subtable ZB N1K1 Fill in the maximum value, minimum value, average value, variance and pass status in the first row except the first cell from left to right;
[0078] In data name SC1 to data name SC M Extract M data names with replacement and record them as sub-table data names ZM K1 , in the data subtable ZB N1K1 Except for the first cell in the leftmost column, fill in the subtable data name ZM from top to bottom K1 ; Based on the subtable data name ZM K1 The data characteristics and passing status of each data name are in the data sub-table ZB N1K1 Fill in the corresponding data;
[0079] In the specific implementation process, in one test, the data sub-table ZB N1K1As shown in Table 3, Table 3 is the data sub-table ZB N1K1 ;
[0080] Table 3: Data subtable ZB N1K1
[0081]
[0082] Randomly obtain J data features from the maximum value, minimum value, average value and variance and record them as decision points, and establish a table of T rows × J1 columns, which is recorded as data subtable ZB N1K1 The first row of the analysis subtable is filled with J decision points from left to right except the first cell, and the leftmost column of the analysis subtable is filled with the subtable data name ZM from top to bottom except the first cell. K1 , and in the analysis sub-table based on the sub-table data name ZM K1 Fill in the corresponding data for each data name in the data feature;
[0083] In the specific implementation process, J can be determined according to the number of decision points that need to be extracted and analyzed in the actual detection. The decision point is a part of the data feature, which is used for random extraction and random analysis to prevent the deviation of the analysis result. In this embodiment, the value of J is set to 2. In one detection, two data features are randomly obtained from the maximum value, minimum value, average value and variance and recorded as decision points. The extracted decision points are the minimum value and variance respectively. The data subtable ZB is established. N1K1 The analysis sub-table is shown in Table 4, which is the data sub-table ZB N1K1 Analytical subtable of
[0084] Table 4: Data subtable ZB N1K1 Analysis subtable
[0085]
[0086] A decision tree is established based on the data in the analysis subtable, wherein the decision nodes of the decision tree are determined by the values of the decision points, and the end points of the decision tree are the passing conditions of the data types;
[0087] In the specific implementation process, the decision tree established based on Table 4 can be found in Figure 3 As shown;
[0088] Get all data subtable ZB N1 The decision tree will be used to store all data in the sub-table ZB N1 The decision tree is denoted as customs risk FX N1 Initial random forest of
[0089] The second sub-strategy of the forest is used to analyze the customs risk FX N1The initial random forest is optimized, and the customs risk FX is obtained based on the optimization results. N1 Random Forest;
[0090] The second sub-strategy of forest construction includes:
[0091] Randomly obtain customs risk FX N1 Any set of data names SC1 to SC in the historical data M The corresponding data is recorded as test data, the data features of the test data are obtained, and the passing conditions of all data names in the test data are recorded as standard passing conditions BZ1 to standard passing conditions BZ M ;
[0092] Based on customs risk FX N1 All decision trees in the initial random forest use the decision tree processing method to process all test data and their data features, and obtain the passing status of the data names of all test data in the initial random forest, which are recorded as forest passing status SL1 to forest passing status SL M ;
[0093] Decision tree processing methods include: See Figure 4 As shown, get the decision point of the decision node in the decision tree, and get the numerical name of the same as the decision point in the data feature of the test data, record it as the decision name, substitute all the data corresponding to the decision name in the test data into the decision tree containing the decision name, and get the terminal point that the data corresponding to the decision name leads to in the decision tree; for any data name SC in the test data M1 , for data name SC M1 Any data feature TZ A , get the data feature TZ A The number of decision trees inserted after the decision name is recorded as L, and the data feature TZ is obtained. A The decision name is recorded and substituted into the decision tree. The endpoint is the number of passed points, recorded as L1. When L1 is greater than L / 2, the data feature TZ A The processing result is recorded as passed. When L1 is less than or equal to L / 2, the data feature TZ A The result of the processing is recorded as failed; when the data name SC M1 When all data characteristics of the data are passed, the data name SC M1 The forest passing situation is recorded as passed, when the data name SC M1 If any of the data features is not passed, the data name SC M1 The forest pass situation is recorded as failed;
[0094] In the specific implementation process, in one test, the data feature TZA The corresponding L is 100, and the data feature TZ A The corresponding L1 is 90, then by comparison, the data feature TZ A The processing result is recorded as passed;
[0095] When the standard pass situation BZ1 and forest pass situation SL1 to the standard pass situation BZ M and forest through the situation SL M When the values are equal, the test data is recorded as successful test data; when the standard passes the case BZ1 and the forest passes the case SL1 to the standard passes the case BZ M and forest through the situation SL M Any one of the criteria in BZ M1 and forest through the situation SL M1 If they are not equal, the test data is recorded as failed test data;
[0096] In the specific implementation process, in one test, the standard passing situation BZ1 to the standard passing situation BZ M They are: passed, failed, passed, passed and failed, and the forest pass situation SL1 to the forest pass situation SL are obtained through the decision tree processing method. M They are: passed, failed, failed, passed, and passed. Then the standard pass case 3 and the forest pass case 3 as well as the standard pass case 5 and the forest pass case 5 are not equal, so the test data is recorded as failed test data;
[0097] When test data is recorded as failed test data, add the test data to the Customs Risk FX N1 The corresponding reference history database is retrieved and the customs risk FX N1 The corresponding initial random forest;
[0098] In the specific implementation process, by adding failed test data to the customs risk FX N1 The corresponding reference history database can be used to analyze the customs risk FX N1 The corresponding initial random forest is further optimized and trained to make the test results of the initial random forest more accurate;
[0099] When the test data is recorded as successful test data, the initial random forest is recorded as customs risk FX N1 Random Forest;
[0100] Use the random forest construction method to obtain multiple decision trees for all customs risks and the initial random forests corresponding to the multiple decision trees. Optimize all the initial random forests to obtain the random forests corresponding to all customs risks.
[0101] The random forest application module is used to screen new customs risk indicators based on the random forest corresponding to each customs risk when new customs risk indicators are obtained, and perform corresponding processing based on the screening results;
[0102] The random forest application module is configured with a forest application processing strategy, which includes:
[0103] When a new customs risk index is obtained, the customs risks in the customs risk index are classified based on the different customs risk types corresponding to the customs risk index. Based on the classification results, the customs risks obtained are recorded as risks to be analyzed DFX1 to risks to be analyzed DFX G ;
[0104] For the risks to be analyzed DFX1 to DFX G Any of the risks to be analyzed DFX G1 , using the risk to be analyzed DFX G1 Random Forest for Risk Analysis DFX G1 Analyze the data in the analysis and obtain the risk DFX to be analyzed based on the analysis results. G1 The customs risk indicators are screened based on the passing status of each data;
[0105] In the specific implementation process, when using the risk to be analyzed DFX G1 Random Forest for Risk Analysis DFX G1 When analyzing the data in the forest, you can refer to the analysis process of the test data in the second sub-strategy of forest construction to obtain the risk DFX to be analyzed. G1 The data in the forest passes the situation, when the risk DFX to be analyzed G1 When all the data in the forest pass situation meet the standards, the risk DFX to be analyzed can be G1 Recorded as qualified risk and further screened; for example, in a test process, the risk to be analyzed DFX G1 The data in the data include data name Q1, data name Q2, data name Q3 and data name Q4, and the risk DFX to be analyzed G1 The passing criteria for the data in the data are: Data Name Q1 has passed, Data Name Q2 has passed, Data Name Q3 has failed, and Data Name Q4 has failed; by putting Data Name Q1, Data Name Q2, Data Name Q3, and Data Name Q4 into the risk DFX to be analyzed G1 After analyzing the random forest of the data, the forest pass status is: data name Q1 has passed, data name Q2 has passed, data name Q3 has failed, and data name Q4 has failed. The forest pass status of all data in the risk G1 meets the standard, and the risk to be analyzed DFXG1 Recorded as qualified risk and further screened.
[0106] Example 2, please refer to Figure 2 As shown, in a second aspect, the present invention provides a customs risk indicator screening method based on a random forest algorithm, comprising:
[0107] Step S1: Obtain multiple customs risk types, which are recorded as customs risk FX1 to customs risk FX N .
[0108] Step S2: Analyze all customs risks separately and obtain the random forest corresponding to each customs risk;
[0109] Step S2 includes the following sub-steps:
[0110] Step S201: For customs risk FX1 to customs risk FX N Any customs risk FX N1 , using the random forest construction method to obtain customs risk FX N1 Corresponding data subtable ZB N1 Based on multiple decision trees, the customs risk FX N1 Initial random forest of
[0111] Random forest construction methods include:
[0112] Step S2011, obtain customs risk FX N1 The data names of the multiple data contained in the data are sequentially recorded as data name SC1 to data name SC M ;
[0113] For data name SC1 to data name SC M Any data name in SC M1 ;
[0114] Step S2012, obtain the data name SC M1 The most recently obtained standard historical quantity data in the corresponding data is recorded as data name SC M1 Reference history database;
[0115] Set the data name to SC M1 The maximum value, minimum value, average value and variance in the reference history database are recorded as U1, U2, U3 and U4 respectively, and U1 to U4 are recorded as data names SC M1 Data characteristics;
[0116] Step S2013, obtain the data name SC M1 The data measurement standard is SC M1Data characteristics and data name SC M1 The data measurement standard is compared, when the data name SC M1 Any of the data characteristics does not match the data name SC M1 When measuring the data standard, the data name SC M1 Recorded as failed; when the data name SC M1 All data characteristics of the data name SC M1 When measuring the data standard, the data name SC M1 Recorded as passed;
[0117] Step S2014: Obtain data features and pass status of all data names, wherein the pass status includes failed and passed.
[0118] The random forest construction method further includes: step S2015, creating a table of T rows × Y columns, recorded as the data summary table SB N1 , of which the data table SB N1 The first row of the table SB is filled with the maximum value, minimum value, average value, variance and pass status from left to right except the first cell. N1 Except for the first cell in the leftmost column, fill in the data names SC1 to SC from top to bottom. M ;
[0119] Step S2016, based on the data name SC1 to the data name SC M The data characteristics and passing status are shown in the data summary table SB N1 Fill in the corresponding data.
[0120] Step S2017: Create multiple tables with the same format, each of which is recorded as a data sub-table ZB N11 To data subtable ZB N1K ;
[0121] For data subtable ZB N11 To data subtable ZB N1K Any data subtable ZB in N1K1 , data subtable ZB N1K1 It is a table with T rows and Y columns, and the data subtable ZB N1K1 Fill in the maximum value, minimum value, average value, variance and pass status in the first row except the first cell from left to right;
[0122] In data name SC1 to data name SC M Extract M data names with replacement and record them as sub-table data names ZM K1 , in the data subtable ZB N1K1 Except for the first cell in the leftmost column, fill in the subtable data name ZM from top to bottomK1 ; Based on the subtable data name ZM K1 The data characteristics and passing status of each data name are in the data sub-table ZB N1K1 Fill in the corresponding data;
[0123] Step S2018: randomly obtain J data features from the maximum value, minimum value, average value and variance and record them as decision points, and create a table with T rows and J1 columns, which is recorded as data sub-table ZB N1K1 The first row of the analysis subtable is filled with J decision points from left to right except the first cell, and the leftmost column of the analysis subtable is filled with the subtable data name ZM from top to bottom except the first cell. K1 , and in the analysis sub-table based on the sub-table data name ZM K1 Fill in the corresponding data for each data name in the data feature;
[0124] Step S2019: Building a decision tree based on the data in the analysis subtable, wherein the decision nodes of the decision tree are determined by the values of the decision points, and the endpoints of the decision tree are the passing status of the data types;
[0125] Get all data subtable ZB N1 The decision tree will be used to store all data in the sub-table ZB N1 The decision tree is denoted as customs risk FX N1 The initial random forest.
[0126] Step S202: Customs risk FX N1 The initial random forest is optimized, and the customs risk FX is obtained based on the optimization results. N1 Random Forest;
[0127] Step S202 includes: Step S2021, randomly obtain customs risk FX N1 Any set of data names SC1 to SC in the historical data M The corresponding data is recorded as test data, the data features of the test data are obtained, and the passing conditions of all data names in the test data are recorded as standard passing conditions BZ1 to standard passing conditions BZ M ;
[0128] Step S2022, based on customs risk FX N1 All decision trees in the initial random forest use the decision tree processing method to process all test data and their data features, and obtain the passing status of the data names of all test data in the initial random forest, which are recorded as forest passing status SL1 to forest passing status SL M .
[0129] The decision tree processing method includes: obtaining the decision point of the decision node in the decision tree, and obtaining the numerical name of the same as the decision point in the data feature of the test data, recording it as the decision name, substituting all the data corresponding to the decision name in the test data into the decision tree containing the decision name, and obtaining the terminal point to which the data corresponding to the decision name leads in the decision tree; for any data name SC in the test data M1 , for data name SC M1 Any data feature TZ A , get the data feature TZ A The number of decision trees inserted after the decision name is recorded as L, and the data feature TZ is obtained. A The decision name is recorded and substituted into the decision tree. The endpoint is the number of passed points, recorded as L1. When L1 is greater than L / 2, the data feature TZ A The processing result is recorded as passed. When L1 is less than or equal to L / 2, the data feature TZ A The result of the processing is recorded as failed; when the data name SC M1 When all data characteristics of the data are passed, the data name SC M1 The forest passing situation is recorded as passed, when the data name SC M1 If any of the data features is not passed, the data name SC M1 The forest pass situation is recorded as failed;
[0130] When the standard pass situation BZ1 and forest pass situation SL1 to the standard pass situation BZ M and forest through the situation SL M When the values are equal, the test data is recorded as successful test data; when the standard passes the case BZ1 and the forest passes the case SL1 to the standard passes the case BZ M and forest through the situation SL M Any one of the criteria in BZ M1 and forest through the situation SL M1 If they are not equal, the test data is recorded as failed test data;
[0131] When test data is recorded as failed test data, add the test data to the Customs Risk FX N1 The corresponding reference history database is retrieved and the customs risk FX N1 The corresponding initial random forest;
[0132] When the test data is recorded as successful test data, the initial random forest is recorded as customs risk FX N1 Random Forest;
[0133] Step S203: Use a random forest construction method to obtain multiple decision trees for all customs risks and initial random forests corresponding to the multiple decision trees, optimize all the initial random forests, and obtain random forests corresponding to all customs risks.
[0134] Step S3: When new customs risk indicators are obtained, the customs risk indicators are screened based on the random forest corresponding to each customs risk, and corresponding processing is performed based on the screening results;
[0135] Step S3 includes: when a new customs risk index is obtained, classifying the customs risks in the customs risk index based on the different customs risk types corresponding to the customs risk index, and recording the customs risks obtained as risks to be analyzed DFX1 to risks to be analyzed DFX2 based on the classification results. G ;
[0136] For the risks to be analyzed DFX1 to DFX G Any of the risks to be analyzed DFX G1 , using the risk to be analyzed DFX G1 Random Forest for Risk Analysis DFX G1 Analyze the data in the analysis and obtain the risk DFX to be analyzed based on the analysis results. G1 The customs risk indicators are screened based on the passing status of each data.
[0137] Embodiment 3, in a third aspect, the present application provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the above method are executed. Through the above technical solution, when the computer program is executed by the processor, the method in any optional implementation of the above embodiment is executed to achieve the following functions: monitoring and updating the stored data and production data in the oil and gas production area by establishing a digital model of the oil and gas production area, and then presetting the production parameters according to the basic information of the oil and gas field by establishing an oil and gas production cycle model, and finally sending the processing results of the digital model of the oil and gas production area and the oil and gas production cycle model to the staff, reminding the staff to adjust the production process of the oil and gas production in a timely manner based on the processing results.
[0138] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0139] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0140] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
Claims
1. A method for screening customs risk indicators based on the random forest algorithm, characterized by: include: Get multiple customs risk types, which are recorded as customs risk FX1 to customs risk FX N ; Analyze all customs risks separately and obtain the random forest corresponding to each customs risk; When new customs risk indicators are obtained, the customs risk indicators are screened based on the random forest corresponding to each customs risk, and corresponding processing is performed based on the screening results; All customs risks are analyzed separately, and the random forest corresponding to each customs risk is obtained, including: For Customs Risk FX1 to Customs Risk FX N Any customs risk FX N1 , using the random forest construction method to obtain customs risk FX N1 Corresponding data subtable ZB N1 Based on multiple decision trees, the customs risk FX N1 Initial random forest of FX for customs risk N1 The initial random forest is optimized, and the customs risk FX is obtained based on the optimization results. N1 Random Forest; Use the random forest construction method to obtain multiple decision trees for all customs risks and the initial random forests corresponding to the multiple decision trees. Optimize all the initial random forests to obtain the random forests corresponding to all customs risks. FX for customs risk N1 The initial random forest is optimized, and the customs risk FX is obtained based on the optimization results. N1 The random forest consists of: Randomly obtain customs risk FX N1 Any set of data names SC1 to SC in the historical data M The corresponding data is recorded as test data, the data features of the test data are obtained, and the passing conditions of all data names in the test data are recorded as standard passing conditions BZ1 to standard passing conditions BZ M ; Based on customs risk FX N1 All decision trees in the initial random forest use the decision tree processing method to process all test data and their data features, and obtain the passing status of the data names of all test data in the initial random forest, which are recorded as forest passing status SL1 to forest passing status SL M ; The decision tree processing method includes: obtaining the decision point of the decision node in the decision tree, and obtaining the numerical name of the same as the decision point in the data feature of the test data, recording it as the decision name, substituting all the data corresponding to the decision name in the test data into the decision tree containing the decision name, and obtaining the terminal point to which the data corresponding to the decision name leads in the decision tree; for any data name SC in the test data M1 , for data name SC M1 Any data feature TZ A , get the data feature TZ A The number of decision trees inserted after the decision name is recorded as L, and the data feature TZ is obtained. A The decision name is recorded and substituted into the decision tree. The endpoint is the number of passed points, recorded as L1. When L1 is greater than L / 2, the data feature TZ A The processing result is recorded as passed. When L1 is less than or equal to L / 2, the data feature TZ A The result of the processing is recorded as failed; when the data name SC M1 When all data characteristics of the data are passed, the data name SC M1 The forest passing situation is recorded as passed, when the data name SC M1 If any of the data features is not passed, the data name SC M1 The forest pass situation is recorded as failed.
2. The method for screening customs risk indicators based on random forest algorithm according to claim 1, characterized in that: Random forest construction methods include: Obtaining Customs Risk FX N1 The data names of the multiple data contained in the data are sequentially recorded as data name SC1 to data name SC M ; For data name SC1 to data name SC M Any data name in SC M1 ; Get data name SC M1 The most recently obtained standard historical quantity data in the corresponding data is recorded as data name SC M1 Reference history database; Set the data name to SC M1 The maximum value, minimum value, average value and variance in the reference history database are recorded as U1, U2, U3 and U4 respectively, and U1 to U4 are recorded as data names SC M1 Data characteristics; Get data name SC M1 The data measurement standard is SC M1 Data characteristics and data name SC M1 The data measurement standard is compared, when the data name SC M1 Any of the data characteristics does not match the data name SC M1 When measuring the data standard, the data name SC M1 Recorded as failed; when the data name SC M1 All data characteristics of the data name SC M1 When measuring the data standard, the data name SC M1 Recorded as passed; Get the data characteristics and pass status of all data names, where the pass status includes failed and passed.
3. The method for screening customs risk indicators based on the random forest algorithm according to claim 2, characterized in that: Random forest construction methods also include: Create a table with T rows and Y columns, record it as the data table SB N1 , of which the data table SB N1 The first row of the table SB is filled with the maximum value, minimum value, average value, variance and pass status from left to right except the first cell. N1 Except for the first cell in the leftmost column, fill in the data names SC1 to SC from top to bottom. M ; Based on data name SC1 to data name SC M The data characteristics and passing status are shown in the data summary table SB N1 Fill in the corresponding data.
4. The method for screening customs risk indicators based on random forest algorithm according to claim 3, characterized in that: Random forest construction methods also include: Create multiple tables with the same format, each of which is recorded as a data sub-table ZB N11 To data subtable ZB N1K ; For data subtable ZB N11 To data subtable ZB N1K Any data subtable ZB in N1K1 , data subtable ZB N1K1 It is a table with T rows and Y columns, and the data subtable ZB N1K1 Fill in the maximum value, minimum value, average value, variance and pass status in the first row except the first cell from left to right; In data name SC1 to data name SC M Extract M data names with replacement and record them as sub-table data names ZM K1 , in the data subtable ZB N1K1 Except for the first cell in the leftmost column, fill in the subtable data name ZM from top to bottom K1 ; Based on the subtable data name ZM K1 The data characteristics and passing status of each data name are in the data sub-table ZB N1K1 Fill in the corresponding data; Randomly obtain J data features from the maximum value, minimum value, average value and variance and record them as decision points, and establish a table of T rows × J1 columns, which is recorded as data subtable ZB N1K1 The first row of the analysis subtable is filled with J decision points from left to right except the first cell, and the leftmost column of the analysis subtable is filled with the subtable data name ZM from top to bottom except the first cell. K1 , and in the analysis sub-table based on the sub-table data name ZM K1 Fill in the corresponding data for each data name in the data feature; A decision tree is established based on the data in the analysis subtable, wherein the decision nodes of the decision tree are determined by the values of the decision points, and the end points of the decision tree are the passing conditions of the data types; Get all data subtable ZB N1 The decision tree will be used to store all data in the sub-table ZB N1 The decision tree is denoted as customs risk FX N1 The initial random forest.
5. The method for screening customs risk indicators based on random forest algorithm according to claim 4, characterized in that: FX for customs risk N1 The initial random forest is optimized, and the customs risk FX is obtained based on the optimization results. N1 The random forest also includes: When the standard pass situation BZ1 and forest pass situation SL1 to the standard pass situation BZ M and forest through the situation SL M When the values are equal, the test data is recorded as successful test data; when the standard passes the case BZ1 and the forest passes the case SL1 to the standard passes the case BZ M and forest through the situation SL M Any one of the criteria in BZ M1 and forest through the situation SL M1 If they are not equal, the test data is recorded as failed test data; When test data is recorded as failed test data, add the test data to the Customs Risk FX N1 The corresponding reference history database is retrieved and the customs risk FX N1 The corresponding initial random forest; When the test data is recorded as successful test data, the initial random forest is recorded as customs risk FX N1 Random Forest.
6. The method for screening customs risk indicators based on random forest algorithm according to claim 5, characterized in that: When new customs risk indicators are obtained, the customs risk indicators are screened based on the random forest corresponding to each customs risk. The corresponding processing based on the screening results includes: When a new customs risk index is obtained, the customs risks in the customs risk index are classified based on the different customs risk types corresponding to the customs risk index. Based on the classification results, the customs risks obtained are recorded as risks to be analyzed DFX1 to risks to be analyzed DFX G ; For the risks to be analyzed DFX1 to DFX G Any of the risks to be analyzed DFX G1 , using the risk to be analyzed DFX G1 Random Forest for Risk Analysis DFX G1 Analyze the data in the analysis and obtain the risk DFX to be analyzed based on the analysis results. G1 The customs risk indicators are screened based on the passing status of each data.
7. A customs risk indicator screening system based on a random forest algorithm, implemented based on the customs risk indicator screening method based on a random forest algorithm according to any one of claims 1 to 6, characterized in that: Including risk acquisition module, random forest generation module and random forest application module: The risk acquisition module is used to obtain multiple customs risk types, which are recorded as customs risk FX1 to customs risk FX N ; The random forest generation module is used to analyze all customs risks separately and obtain the random forest corresponding to each customs risk; The random forest application module is used to screen the customs risk indicators based on the random forest corresponding to each customs risk when new customs risk indicators are obtained, and perform corresponding processing based on the screening results.
8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are executed.
Citation Information
Patent Citations
Customs service declaration risk checking system
CN108959648A
Resource allocation method based on random forest algorithm
CN112416588A
Method for risk rating prediction based on big data and computing device
CN116385151A