A method and system for analyzing a customs declaration based on hierarchical density clustering and frame of reference
Through the hierarchical density clustering and reference system customs declaration analysis method, using intelligent unsupervised clustering algorithm and sorting model, the number of customs declarations that need to be inspected can be quickly and accurately determined, solving the problem of low analysis efficiency in existing technologies and significantly improving the inspection effect.
Patent Information
- Application Number
- CN202210108682.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-01-28
AI Technical Summary
How to quickly and effectively analyze batches of customs declarations to determine the number of declarations that need to be inspected to meet customs supervision requirements.
A customs declaration analysis method based on hierarchical density clustering and reference system is adopted. The historical blacklist dataset is clustered and analyzed using an intelligent unsupervised clustering algorithm to obtain the coordinates of the cluster center points. The minimum Euclidean distance from the customs declaration sample to each cluster center point is calculated. A reference model is constructed through a sorting algorithm to determine the proportion of customs declarations that need to be analyzed.
It has achieved rapid and accurate analysis of customs declarations, reduced the control rate by 50%, and increased the seizure rate by 200%-300%.
Smart Images

Figure CN116561606B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of customs declaration analysis, in particular to a customs declaration analysis method and system based on hierarchical density clustering and reference system. BACKGROUND
[0002] According to the provisions of the Customs Law of the People's Republic of China and the Customs Supervision Measures of the Bonded Area, the goods entering and leaving the bonded area must be supervised, whether the goods are from abroad or from the non-bonded area in the territory. At the same time, the Customs Law also provides that the customs declaration and tax payment procedures shall be handled by the customs-registered customs declaration enterprises or enterprises authorized to operate import and export business.
[0003] In recent years, with the continuous and rapid growth of China's import and export trade, how to effectively and quickly analyze and check the batch of customs declarations is a problem to be solved. SUMMARY
[0004] The main purpose of the present application is to overcome the above-mentioned defects in the prior art, and to provide a customs declaration analysis method based on hierarchical density clustering and reference system, which can accurately analyze the number of customs declarations that need to be analyzed and checked in the batch of customs declarations, and is fast and effective.
[0005] The present application adopts the following technical solutions:
[0006] A customs declaration analysis method based on hierarchical density clustering and reference system, comprising:
[0007] The historical black list data set is analyzed by using an intelligent unsupervised clustering algorithm to obtain the coordinate sequence of each clustering center point after clustering;
[0008] The reference system data set is obtained, the minimum Euclidean distance of each sample in the reference system data set to each clustering center point coordinate is calculated, and a reference model is constructed according to a sorting algorithm;
[0009] The customs declaration sample is received in real time, the minimum Euclidean distance to each clustering center point coordinate is calculated, and the minimum Euclidean distance sequence of the customs declaration sample is obtained;
[0010] The minimum Euclidean distance sequence is input into the reference model to obtain a sorting score;
[0011] According to the sorting score, the contribution degree corresponding to the customs declaration sample, and the business hit threshold, the proportion of the customs declaration that needs to be analyzed in the received customs declaration sample is determined.
[0012] Specifically, the intelligent unsupervised clustering algorithm is used, which is specifically:
[0013] The historical blacklist dataset is clustered by using HDBSCAN, i.e., a hierarchical clustering algorithm, including:
[0014] Output the serial number of the center node;
[0015] Remove noise points;
[0016] Group the center nodes and count the number of data falling in each center node, when the center node data is uniform, the sample data cluster centroids are selected, and the hyperparameter values of the HDBSCAN constructed class can be determined, and the cluster centroids are the clustering center points.
[0017] Specifically, a reference frame dataset is obtained, and the minimum Euclidean distance of each sample in the reference frame dataset to the coordinates of the clustering center points is calculated, wherein the Euclidean distance of each sample to the coordinates of the clustering center points is calculated as:
[0018]
[0019] Wherein, n=256, the dimension of the sequence of coordinates of the clustering center points, x1x2…x n The characteristic values of each dimension of the reference frame declaration form, y1y2…y n The dimension characteristic values of the clustering center points.
[0020] Specifically, the minimum Euclidean distance sequence is input into the reference model to obtain a ranking score, specifically:
[0021]
[0022] Wherein, f(r) is the ranking score, M is the total ranking number in the reference model, and r is the average ranking of the received declaration form sample.
[0023] Specifically, according to the ranking score, the contribution degree corresponding to the declaration form sample, and the business hit threshold, the proportion of declaration forms to be analyzed in the received declaration form sample is determined, specifically:
[0024] s=f(r)*α*K
[0025] Wherein, α is the contribution degree corresponding to the declaration form sample, which means the risk contribution degree of the declaration company itself / the median contribution degree of all declaration companies; K is the business hit threshold, which is 5 / 1000.
[0026] Another aspect of the embodiment of the application provides a declaration form analysis system based on hierarchical density clustering and reference frame, including:
[0027] The historical blacklist clustering unit: using intelligent unsupervised clustering algorithm, the historical blacklist dataset is clustered and analyzed to obtain the sequence of coordinates of the clustering center points after clustering;
[0028] Referential model construction unit: obtain a reference system dataset, calculate the minimum Euclidean distance of each sample in the reference system dataset to the coordinates of each cluster center point, and construct a referential model according to a sorting algorithm;
[0029] Customs declaration form Euclidean distance calculation unit: real-time receive a customs declaration form sample, calculate the minimum Euclidean distance to each cluster center point coordinate, and obtain the minimum Euclidean distance sequence of the customs declaration form sample;
[0030] Sorting score calculation unit: input the minimum Euclidean distance sequence into the referential model to obtain a sorting score;
[0031] Analysis of the proportion of the customs declaration form calculation unit: according to the sorting score and the contribution degree corresponding to the customs declaration form sample and the business hit threshold, determine the proportion of the customs declaration form sample to be analyzed in the received customs declaration form sample.
[0032] Specifically, in the historical blacklist clustering unit, an intelligent unsupervised clustering algorithm is used, specifically:
[0033] The historical blacklist dataset is clustered and analyzed by using an HDBSCAN, i.e., a hierarchical clustering algorithm, including:
[0034] Output the center node serial number;
[0035] Remove noise points;
[0036] Group the center nodes and count the number of data falling in each center node. When the center point data is uniform, the sample data cluster center points are selected, and the hyperparameter values of the HDBSCAN constructed class can be determined. The cluster center points are the cluster center points.
[0037] Specifically, in the referential model construction unit, a reference system dataset is obtained, and the minimum Euclidean distance of each sample in the reference system dataset to the coordinates of each cluster center point is calculated, wherein the Euclidean distance of each sample to the coordinates of each cluster center point is calculated as:
[0038]
[0039] Wherein, n = 256, which is the dimension of the cluster center point coordinate sequence, x1x2…x n The characteristic value of each dimension of the reference system customs declaration form is y1y2…y n The dimension characteristic value of each cluster center point is.
[0040] Specifically, in the sorting score calculation unit, the minimum Euclidean distance sequence is input into the referential model to obtain a sorting score, specifically:
[0041]
[0042] Wherein, f(r) is the ranking score, M is the total number of ranking in the reference model, and r is the average ranking of the received declaration sample.
[0043] Specifically, in the analysis declaration proportion calculation unit, the proportion of the received declaration sample that needs to be analyzed is determined according to the ranking score, the contribution degree corresponding to the declaration sample, and the business hit threshold, specifically:
[0044] s = f(r) * a * K
[0045] Wherein, a is the contribution degree corresponding to the declaration sample, which means the self-risk contribution degree of the declaration company of the declaration sample / the median contribution degree of all declaration company risks, and K is the business hit threshold, which is 5 / 1000.
[0046] From the above description of the present application, compared with the prior art, the present application has the following beneficial effects:
[0047] (1) The present application proposes a declaration analysis method based on hierarchical density clustering and reference system, which uses an intelligent unsupervised clustering algorithm to perform clustering analysis on historical black list data sets and obtain the coordinate sequence of each clustering center point after clustering; obtains a reference system data set, calculates the minimum Euclidean distance from each sample in the reference system data set to each clustering center point coordinate, and constructs a reference model according to a ranking algorithm; receives declaration samples in real time, calculates the minimum Euclidean distance to each clustering center point coordinate, and obtains the minimum Euclidean distance sequence of the declaration sample; inputs the minimum Euclidean distance sequence into the reference model to obtain a ranking score; determines the proportion of the received declaration sample that needs to be analyzed according to the ranking score, the contribution degree corresponding to the declaration sample, and the business hit threshold; the method proposed by the present application can accurately analyze the number of declaration samples that need to be analyzed and inspected in batches, and is fast and effective. Practical application shows that the historical control rate is reduced by 50% compared with the use of the method of the present application, and the seizure rate is increased by 200%-300%. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 A flowchart of a declaration analysis method based on hierarchical density clustering and reference system is provided for the present application.
[0049] Figure 2 A declaration analysis architecture diagram based on hierarchical density clustering and reference system is provided for the embodiment of the present application.
[0050] Figure 3 An electronic device schematic diagram is provided for the embodiment of the present application.
[0051] Figure 4 An embodiment schematic diagram of a computer readable storage medium is provided for the embodiment of the present application.
[0052] The application will be further described in detail below in combination with the drawings and specific embodiments. DETAILED DESCRIPTION
[0053] As Figure 1 A flow chart of a customs declaration analysis method based on hierarchical density clustering and reference system provided for an embodiment of the application, comprising:
[0054] S101: using an intelligent unsupervised clustering algorithm, clustering analysis is performed on the historical black list data set to obtain a sequence of clustering center point coordinates after clustering;
[0055] Collect national historical data of express delivery business and divide it into two types of data, namely black list and white list. The white list data is the customs declaration form identified as normal by the customs inspection personnel after opening the package, and the black list is the customs declaration form identified as abnormal after opening the package. Abnormal customs declaration forms are generally used by declarers to evade taxes or conduct illegal purchases through false declaration. Before importing data into the system, dirty data needs to be cleaned, standardized in format and classified, i.e. labeled in artificial intelligence. 400,000 white lists and 400,000 black lists from 2019 to 2020 are labeled as white lists and black lists respectively.
[0056] The intelligent unsupervised clustering algorithm uses the HDBSCAN hierarchical clustering algorithm to extend the DBSCAN algorithm to cluster structured data. Based on clustering stability, hierarchical clustering technology is used, and it can handle clustering problems with different densities. The cluster hierarchy is constructed, and the optimal hyperparameter configuration is determined by grid search on the HDBSCAN construction class parameters. The settings of the core parameters are as follows:
[0057] The min_cluster_size parameter represents the minimum cluster size; the min_samples model measures the degree of cluster conservation. The larger the min_samples value, the more conservative the cluster, and more points will be declared as noise, and the cluster will be limited to more and more dense areas.
[0058] The alpha model measures the degree of cluster conservation parameter, and the default is 1.0. Increasing alpha will make the clustering more conservative, but on a tighter scale, we can set alpha to 1.3;
[0059] The gen_min_span_tree parameter represents whether to use the minimum spanning tree method;
[0060] Sub-step 1, output the center node number
[0061] center_num = db.labels_.astype(np.int)
[0062] Sub-step 2, remove noise points
[0063] x_pca_centers = x_pca_centers[x_pca_centers["cluster_db"]!= -1]
[0064] pd.DataFrame(np.hstack((x_pca1, center_num.reshape(-1, 1))), columns=self.center_cols)
[0065] x_pca_centers = x_pca_centers[x_pca_centers["cluster_db"]!= -1]
[0066] Sub-step 3, group the center nodes and count the number of data falling in each center node. When the data of most center nodes is relatively uniform, the sample data cluster centroids are selected, and the parameter values of the HDBSCAN construction class can be determined. The centroid point set is used as the basis point for similarity calculation.
[0067] Cluster Centroid Point: x_pca_centers =
[0068] x_pca_centers.groupby('cluster_db').mean()
[0069] S102: Obtain a reference system data set, calculate the minimum Euclidean distance from each sample in the reference system data set to each cluster center point coordinate, and construct a reference model according to a sorting algorithm;
[0070] Collect 7-day independent historical data of express business, which is not included in the national historical training data. After cleaning, standardization and labeling, it is used as an abnormal class similarity reference data set. There are 1.5 million * 5 days + 5000 (Saturday and Sunday) = 80,000 customs data in total. Calculate the minimum Euclidean distance from each sample in the reference system data set to each cluster center point coordinate,
[0071] Specifically, the reference system data set is obtained, and the minimum Euclidean distance from each sample in the reference system data set to each cluster center point coordinate is calculated, wherein the Euclidean distance from each sample to each cluster center point coordinate is calculated as:
[0072]
[0073] Wherein, n = 256, is the dimension of the cluster center point coordinate sequence, x1x2…xn y1, y2, …, y are the characteristic values of each dimension of the customs declaration form as the reference system. n y1, y2, …, y are the characteristic values of each dimension of the cluster center point.
[0074] The minimum Euclidean distances corresponding to the 80,000 customs declaration data are sorted in descending order to construct a reference sequence table, i.e., a reference model.
[0075] S103: Real-time receive customs declaration form samples, calculate the minimum Euclidean distance to each cluster center point coordinate, and obtain the minimum Euclidean distance sequence of the customs declaration form samples;
[0076] Real-time receive customs declaration form samples, and calculate the minimum Euclidean distance from the samples to each cluster center point coordinate to obtain the minimum Euclidean distance sequence of the customs declaration form samples.
[0077] S104: Input the minimum Euclidean distance sequence into the reference model to obtain a ranking score.
[0078] Specifically, input the minimum Euclidean distance sequence into the reference model to obtain a ranking score, which is specifically:
[0079]
[0080] Where f(r) is the ranking score, M is the total ranking number in the reference model, and r is the average ranking of the received customs declaration form samples.
[0081] S105: According to the ranking score, the contribution degree corresponding to the customs declaration form sample, and the business hit threshold, determine the proportion of customs declaration forms that need to be analyzed in the received customs declaration form samples.
[0082] Specifically, according to the ranking score, the contribution degree corresponding to the customs declaration form sample, and the business hit threshold, determine the proportion of customs declaration forms that need to be analyzed in the received customs declaration form samples, which is specifically:
[0083] s = f(r) * a * K
[0084] Where a is the contribution degree corresponding to the customs declaration form sample, which means the risk contribution degree of the declaration company itself / the median contribution degree of all declaration company risks; K is the business hit threshold, which is 5 / 1000.
[0085] The customs declaration risk contribution degree of each declaration company and the median of the customs declaration risk contribution degree are derived from xgb_model.feature_importances after training by the xgboost algorithm.
[0086] From the model effect, its effect is outstanding. In terms of express service, the data is characterized by large daily customs declaration quantity. According to the data access condition of customs declaration, the customs declaration data quantity of weekends and holidays is less, and the daily quantity of customs declaration is about 20,000 per day from Monday to Friday. The abnormal proportion of customs declaration is only 0.0013-0.004 of the total. From the data distribution, it is an extremely unbalanced data distribution, and it is very difficult to hit effectively through an intelligent model. By using the method provided by the application, the overall control rate and the acquisition rate of the express risk intelligent model from January to February, 2021 are counted, the control rate is about 0.5%, and the acquisition efficiency is between 2-3%, which is 50% lower than the historical control rate, and is 200%-300% higher than the traditional expert rule acquisition rate.
[0087] As Figure 2 For another aspect of the embodiment of the application, a customs declaration analysis system structure diagram based on hierarchical density clustering and reference system is provided, comprising:
[0088] The historical blacklist clustering unit 201: using an intelligent unsupervised clustering algorithm, the historical blacklist data set is clustered and analyzed to obtain the clustered center point coordinate sequence;
[0089] The historical data of express service nationwide is collected and divided into two types of data, namely blacklist and whitelist. The whitelist data is the customs declaration identified as normal by the customs inspection personnel after opening the package, and the blacklist is the customs declaration identified as abnormal after opening the package. Abnormal customs declaration is generally to evade taxes or conduct illegal purchase through false declaration. Before importing data into the system, dirty data needs to be cleaned, standardized and classified, that is, labeled in artificial intelligence. 400,000 white lists and 400,000 blacklists from 2019 to 2020 are labeled as white lists and blacklists respectively.
[0090] The intelligent unsupervised clustering algorithm uses HDBSCAN, i.e. hierarchical clustering algorithm, to extend DBSCAN algorithm to cluster structured data. Based on clustering stability, hierarchical clustering technology is used, and different density clustering problems can be handled. The cluster hierarchy is constructed, and the optimal hyperparameter configuration is determined by grid search of HDBSCAN construction class parameters. The settings of the core parameters are as follows:
[0091] The min_cluster_size parameter represents the minimum cluster size; the min_samples model cluster conservativeness measures the conservativeness of the cluster. The larger the min_samples value, the more conservative the cluster, and more points will be declared as noise, and the cluster will be limited to an increasingly dense area.
[0092] The alpha model cluster conservativeness parameter is 1.0 by default, and increasing alpha will make the cluster more conservative, but on a tighter scale, we can set alpha to 1.3;
[0093] The gen_min_span_tree parameter represents whether to use the minimum spanning tree method;
[0094] Sub-step 1, output the center node serial number
[0095] center_num = db.labels_.astype(np.int)
[0096] Sub-step 2, remove noise points
[0097] x_pca_centers = x_pca1[db.labels_!= -1]
[0098] pd.DataFrame(np.hstack((x_pca1, center_num.reshape(-1, 1))), columns=self.center_cols)
[0099] x_pca_centers = x_pca_centers[x_pca_centers["cluster_db"]!= -1]
[0100] Sub-step 3, group the center nodes and count the number of data falling on each center node. When most of the center point data is relatively uniform, the sample data cluster center points are selected, and the HDBSCAN construction class parameter values can be determined. The center point set is used as the basis for similarity calculation.
[0101] Cluster center points: x_pca_centers =
[0102] x_pca_centers.groupby('cluster_db').mean().
[0103] The reference model construction unit 202 obtains the reference system data set, calculates the minimum Euclidean distance from each sample in the reference system data set to the coordinates of each cluster center point, and constructs a reference model according to the sorting algorithm;
[0104] The 7-day independent historical data of express collection service is collected, which is not included in the national historical training data. After cleaning, standardization and labeling, it is used as an abnormal class similarity reference data set. There are 1.5 million*5 days+5000 (Saturdays and Sundays)=800,000 customs declaration data in total. The minimum Euclidean distance of each sample in the reference system data set to the coordinates of each cluster center is calculated.
[0105] Specifically, the reference system data set is obtained, and the minimum Euclidean distance of each sample in the reference system data set to the coordinates of each cluster center is calculated, wherein the Euclidean distance of each sample to the coordinates of each cluster center is calculated as:
[0106]
[0107] Wherein, n=256, is the dimension of the sequence of coordinates of each cluster center, x1x2…x n is the feature value of each dimension of the reference system declaration, y1y2…y n is the dimension feature value of each cluster center.
[0108] And the minimum Euclidean distance of the 800,000 customs declaration data corresponding to the minimum Euclidean distance is sorted from large to small, and the reference sequence table is constructed, that is, the reference model.
[0109] The customs declaration Euclidean distance calculation unit 203: real-time receives the customs declaration sample, calculates the minimum Euclidean distance to the coordinates of each cluster center, and obtains the minimum Euclidean distance sequence of the customs declaration sample;
[0110] Real-time receives the customs declaration sample, and calculates the minimum Euclidean distance of the sample to the coordinates of each cluster center, and obtains the minimum Euclidean distance sequence of the customs declaration sample.
[0111] The sorting score calculation unit 204: the minimum Euclidean distance sequence is input into the reference model to obtain the sorting score;
[0112] Specifically, the minimum Euclidean distance sequence is input into the reference model to obtain the sorting score, which is specifically:
[0113]
[0114] Wherein, f(r) is the sorting score, M is the total ranking number in the reference model, and r is the average ranking of the received customs declaration sample.
[0115] The analysis of the customs declaration proportion calculation unit 205: according to the sorting score and the contribution degree corresponding to the customs declaration sample and the business hit threshold, the proportion of the customs declaration sample to be analyzed in the received customs declaration sample is determined
[0116] Specifically, according to the ranking score, the contribution degree corresponding to the declaration sample, and the business hit threshold, the proportion of declaration samples to be analyzed in the received declaration samples is determined, and the proportion is specifically:
[0117] s=f(r)*a*K
[0118] Wherein, a is the contribution degree corresponding to the declaration sample, which means the risk contribution degree of the declaration company itself / the median contribution degree of all declaration company risks; K is the business hit threshold, which is 5 / 1000.
[0119] The risk contribution degree of each declaration company and the median risk contribution degree are derived from xgb_model.feature_importances after xgboost algorithm training.
[0120] As shown in Figure 3 , the embodiment of the present application provides an electronic device 300, which comprises a memory 310, a processor 320, and a computer program 311 stored in the memory 320 and executable on the processor 320, and the processor 320 implements the method for analyzing declaration provided by the embodiment of the present application when executing the computer program 311.
[0121] Since the electronic device introduced in the embodiment is the device used in the embodiment of the present application, the specific implementation of the electronic device of the embodiment and its various forms can be understood by those skilled in the art based on the method introduced in the embodiment of the present application, so the method of the electronic device how to implement the method in the embodiment of the present application will not be described in detail, as long as the device used by those skilled in the art to implement the method in the embodiment of the present application belongs to the scope of the present application.
[0122] Please refer to Figure 4 , Figure 4 for an embodiment of a computer readable storage medium provided by the embodiment of the present application.
[0123] As shown in Figure 4 , the embodiment provides a computer readable storage medium 400, which stores a computer program 411, and the computer program 411 is executed by a processor to implement the method for analyzing declaration based on hierarchical density clustering and reference system provided by the embodiment of the present application.
[0124] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0125] Those skilled in the art will appreciate that embodiments of the application can be devised for a method, a system, or a computer program product. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer readable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.
[0126] The application proposes a customs declaration analysis method based on hierarchical density clustering and reference system, using an intelligent unsupervised clustering algorithm to perform clustering analysis on historical black list data sets, obtaining a coordinate sequence of each clustering center point after clustering; obtaining a reference system data set, calculating the minimum Euclidean distance of each sample in the reference system data set to each clustering center point coordinate, and constructing a reference model according to a sorting algorithm; receiving a customs declaration sample in real time, calculating the minimum Euclidean distance to each clustering center point coordinate, obtaining a minimum Euclidean distance sequence of the customs declaration sample; inputting the minimum Euclidean distance sequence into the reference model to obtain a sorting score; according to the sorting score and the contribution degree corresponding to the customs declaration sample and the business hit threshold, determining the proportion of customs declarations that need to be analyzed in the received customs declaration sample; the method proposed by the application can accurately analyze the number of customs declarations that need to be analyzed and inspected in batches of customs declarations, and is fast and effective, and practical application shows that the historical control rate is reduced by 50% compared with the use of the method, and the seizure rate is increased by 200%-300%.
[0127] It should be noted that, in this document, relational terms such as "first" and "second", and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element. The terms "an embodiment", "one embodiment", or the like, in this document, do not necessarily all refer to the same embodiment, although they can. Such terms should be interpreted in the context of this document with the full scope of the present application being set out in the claims. The above specification, examples and data provide essential information for a person of ordinary skill in the art to make and use the application. The present application can be practiced according to the claims without resorting to the details of the above specification. The above specification, examples and data also exhibit the best mode for practicing the present application. While the application has been described in terms of embodiments, it will be apparent to those of ordinary skill in the art that many modifications, substitutions, and deletions can be made to the embodiments without departing from the spirit and scope of the application. Accordingly, many modifications, substitutions, and deletions can be made by those of ordinary skill in the art without departing from the spirit and scope of the application in its broadest form. The specification and examples given herein are the most complete and specific examples available to the inventor at the time of the filing of the application. However, in light of developments in the art, the inventor expects further modifications will be made by those of ordinary skill in the art and encourages others to make their own changes and additions.
[0128] The above merely illustrates the specific embodiments of the present application, but the design concept of the present application is not limited thereto, and any non-essential modification of the present application using the concept shall belong to the infringement of the protection scope of the present application.
Claims
1. A customs declaration analysis method based on hierarchical density clustering and reference system, characterized by: include: Using intelligent unsupervised clustering algorithms, the historical blacklist dataset is clustered and analyzed to obtain the coordinate sequence of each cluster center point after clustering; Obtain the reference data set, calculate the minimum Euclidean distance between each sample in the reference data set and the coordinates of each cluster center, and construct a reference model based on the sorting algorithm; Receive customs declaration samples in real time, calculate the minimum Euclidean distance to the coordinates of each cluster center, and obtain the minimum Euclidean distance sequence of customs declaration samples; Input the minimum Euclidean distance sequence into the reference model to obtain the ranking score; Based on the ranking scores, the contribution corresponding to the customs declaration samples, and the business hit threshold, the proportion of customs declarations that need to be analyzed in the received customs declaration samples is determined.
2. The method for analyzing customs declaration forms based on hierarchical density clustering and reference system according to claim 1, characterized in that: The intelligent unsupervised clustering algorithm is specifically: The HDBSCAN hierarchical clustering algorithm is used to perform cluster analysis on the historical blacklist dataset, including: Output the central node number; Remove noise points; The central nodes are grouped and the number of data falling on each central node is counted. When the central point data is uniform, the centroid points of each cluster of sample data are selected, and the hyperparameter values of the HDBSCAN construction class can be determined. The centroid points of each cluster are the center points of each cluster.
3. The method for analyzing customs declaration forms based on hierarchical density clustering and reference system according to claim 1, characterized in that: Obtain the reference data set and calculate the minimum Euclidean distance from each sample in the reference data set to the coordinates of each cluster center point. The Euclidean distance from each sample to the coordinates of each cluster center point is calculated as: Among them, n = 256, is the coordinate sequence dimension of each cluster center point, x1x2…x n is the characteristic value of each dimension of the reference declaration form, y1y2…y n is the dimensional feature value of each cluster center point.
4. The method for analyzing customs declaration forms based on hierarchical density clustering and reference system according to claim 1, characterized in that: Input the minimum Euclidean distance sequence into the reference model to obtain the ranking score, which is: Where f(r) is the ranking score, M is the total number of rankings in the reference model, and r is the average ranking of the samples of received customs declarations.
5. The method for analyzing customs declaration forms based on hierarchical density clustering and reference system according to claim 4, characterized in that: Based on the ranking scores, the contribution of the customs declaration samples, and the business hit threshold, the proportion of customs declarations to be analyzed in the received customs declaration samples is determined, specifically: s=f(r)*α*K Among them, α is the contribution rate corresponding to the customs declaration sample, which means the risk contribution rate of the declaring company of the customs declaration sample itself / the median contribution rate of the risks of all declaring companies; K is the business hit threshold, which is 0.5%.
6. A customs declaration analysis system based on hierarchical density clustering and reference system, characterized by: include: Historical blacklist clustering unit: uses an intelligent unsupervised clustering algorithm to perform cluster analysis on the historical blacklist dataset and obtain the coordinate sequence of each cluster center point after clustering; Reference model construction unit: obtains the reference system data set, calculates the minimum Euclidean distance between each sample in the reference system data set and the coordinates of each cluster center point, and constructs the reference model according to the sorting algorithm; Customs declaration Euclidean distance calculation unit: receives customs declaration samples in real time, calculates the minimum Euclidean distance to the coordinates of each cluster center, and obtains the minimum Euclidean distance sequence of customs declaration samples; Ranking score calculation unit: input the minimum Euclidean distance sequence into the reference model to obtain the ranking score; The customs declaration proportion calculation unit for analysis is used to determine the proportion of customs declarations that need to be analyzed in the received customs declaration samples based on the ranking scores, the contribution corresponding to the customs declaration samples, and the business hit threshold.
7. The customs declaration form analysis system based on hierarchical density clustering and reference system according to claim 6 is characterized in that: In the historical blacklist clustering unit, an intelligent unsupervised clustering algorithm is used, specifically: The HDBSCAN hierarchical clustering algorithm is used to perform cluster analysis on the historical blacklist dataset, including: Output the central node number; Remove noise points; The central nodes are grouped and the number of data falling on each central node is counted. When the central point data is uniform, the centroid points of each cluster of sample data are selected, and the hyperparameter values of the HDBSCAN construction class can be determined. The centroid points of each cluster are the center points of each cluster.
8. The customs declaration form analysis system based on hierarchical density clustering and reference system according to claim 6 is characterized in that: In the reference model construction unit, a reference system data set is obtained, and the minimum Euclidean distance from each sample in the reference system data set to the coordinates of each cluster center point is calculated, wherein the Euclidean distance from each sample to the coordinates of each cluster center point is calculated as: Among them, n = 256, is the coordinate sequence dimension of each cluster center point, x1x2…x n is the characteristic value of each dimension of the reference declaration form, y1y2…y n is the dimensional feature value of each cluster center point.
9. The customs declaration form analysis system based on hierarchical density clustering and reference system according to claim 6 is characterized in that: In the ranking score calculation unit, the minimum Euclidean distance sequence is input into the reference model to obtain the ranking score, specifically: Where f(r) is the ranking score, M is the total number of rankings in the reference model, and r is the average ranking of the samples of received customs declarations.
10. The customs declaration form analysis system based on hierarchical density clustering and reference system according to claim 9, characterized in that: The customs declaration proportion calculation unit determines the proportion of customs declarations to be analyzed in the received customs declaration samples based on the ranking scores, the contribution corresponding to the customs declaration samples, and the business hit threshold, specifically: s=f(r)*α*K Among them, α is the contribution rate corresponding to the customs declaration sample, which means the risk contribution rate of the declaring company of the customs declaration sample itself / the median contribution rate of the risks of all declaring companies. K is the business hit threshold, which is 0.5%.
Citation Information
Patent Citations
Method and system for label-free data classification and predication
CN107679734A
Method and device for sending text messages, computer device and storage medium
WO2020062702A1