Anomaly detection method and apparatus, and computer readable storage medium
By calculating the evaluation distance and association relationship of the item set in the customs declaration data, and using anomaly detection algorithm to process the evaluation indicators, the problem of the inability to accurately detect abnormal customs clearance trade behavior in the existing technology is solved, and the accurate identification of trade regional risks, regional dispersion risks and port drift risks is achieved.
Patent Information
- Application Number
- PCT/CN2024/142754
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-26
- Publication Date
- 2025-07-03
AI Technical Summary
There is a lack of solutions in the prior art that can accurately detect abnormal customs clearance trade behavior.
By obtaining the items and address association fields in the customs declaration statement data, calculate the evaluation distance and association relationship between the item sets, and use an exception detection algorithm to process the evaluation distance and indicators to determine whether there are abnormalities in the customs declaration statement data.
Accurate detection of abnormal customs clearance trade behaviors, especially the identification of trade regional risks, regional dispersed risks and port drift risks, which improves the accuracy of detection and reduces the false alarm rate.
Smart Images

Figure CN2024142754_03072025_PF_FP_ABST
Abstract
Description
Abnormality detection method, device and computer-readable storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure is based on and claims priority to an application filed in China with application number 202311835399.0, filed on December 28, 2023. The disclosure of this application is hereby incorporated into this disclosure as a whole. Technical Field
[0003] The present disclosure relates to the field of anomaly detection technology, and in particular to an anomaly detection method, device, and computer-readable storage medium. Background Art
[0004] At present, there are some abnormal customs clearance trade behaviors, which can be detected by analyzing the customs declaration data. Summary of the Invention
[0005] However, related technologies lack solutions that can accurately detect abnormal customs clearance trade behaviors.
[0006] In view of this, the embodiments of the present disclosure propose the following solution, which can accurately detect abnormal customs clearance trade behavior.
[0007] According to one aspect of an embodiment of the present disclosure, there is provided an anomaly detection method, comprising: obtaining first customs declaration data to be detected, the first customs declaration data comprising a plurality of fields in which corresponding items are associated with addresses, the plurality of items corresponding to the plurality of fields forming a plurality of item sets, each item set comprising two items of the plurality of items; calculating an evaluation distance between two addresses associated with the two items in each item set; performing association rule mining on a plurality of historical second customs declaration data to obtain an evaluation index representing an association relationship between the two items in each item set; processing the evaluation distance and the evaluation index using an anomaly detection algorithm to obtain a first result representing whether there is a risk between the two items in each item set; and determining whether the first customs declaration data has an anomaly based on the first result.
[0008] In some embodiments, the multiple item sets include a first item set, and the two fields corresponding to the two items in the first item set are both unit fields related to units; wherein, based on the first result, determining whether there is an anomaly in the first customs declaration data includes: when the first result indicates that there is a risk between the two items in the first item set, determining that the first customs declaration data has an anomaly related to trade regional risk.
[0009] In some embodiments, the multiple item sets include a second item set, and the two fields corresponding to the two items in the second item set are one and the other of a unit field related to a unit, a port field related to a port, and a destination field related to a destination; wherein, based on the first result, determining whether there is an anomaly in the first customs declaration data includes: when the first result indicates that there is a risk between the two items in the second item set, determining that there is an anomaly in the first customs declaration data related to a geographical dispersion risk.
[0010] In some embodiments, calculating the evaluation distance between the two addresses associated with the two items in each item set includes: resolving the two addresses associated with the two items in each item set to obtain two resolved addresses; obtaining the geographic coordinates of the two resolved addresses; and calculating the geographic distance between the two resolved addresses based on the geographic coordinates, wherein the evaluation distance includes the geographic distance.
[0011] In some embodiments, the multiple item sets include a third item set, and the two fields corresponding to the two items in the third item set do not include a port field related to the port; wherein, calculating the evaluation distance between the two addresses associated with the two items in each item set also includes: calculating a first text distance between two strings corresponding to the two addresses associated with the two items in the third item set; and obtaining a final text distance based on the first text distance, wherein the evaluation distance of the third item set also includes the final text distance.
[0012] In some embodiments, the two resolved addresses both include multiple levels; obtaining the final text distance based on the first text distance includes: when the strings corresponding to the multiple levels all contain space, using the first text distance as the final text distance; and when the string corresponding to at least one level in the multiple levels does not contain space, calculating the second text distance between the strings corresponding to each level in the at least one level; wherein the final text distance is obtained based on the first text distance and the second text distance.
[0013] In some embodiments, the evaluation metric includes at least one of frequency and confidence.
[0014] In some embodiments, the anomaly detection algorithm includes an isolation forest algorithm, a single-class support vector machine algorithm, and a statistical anomaly threshold algorithm.
[0015] In some embodiments, the multiple fields include a unit field related to a unit; the method further includes: filtering out a first group of second customs declaration data in which an item corresponding to the unit field appears from the multiple second customs declaration data; determining a commodity number whose number of occurrences is greater than a threshold in the first group of second customs declaration data; filtering out a second group of second customs declaration data in which the commodity number appears from the first group of second customs declaration data; and determining a second result indicating whether there is a risk for the item corresponding to the unit field based on whether an item corresponding to the port field related to the port appears in the second customs declaration data within the second time window in the second group of second customs declaration data, wherein the start time of the first time window is not earlier than the end time of the second time window; wherein whether there is an abnormality in the first customs declaration data is further determined based on the second result.
[0016] In some embodiments, when the second result indicates that there is a risk in the item corresponding to the unit field, it is determined that there is an anomaly related to the port drift risk in the first customs declaration data.
[0017] According to another aspect of an embodiment of the present disclosure, there is provided an anomaly detection apparatus, comprising: a module configured to execute the anomaly detection method described in any one of the above embodiments.
[0018] According to another aspect of an embodiment of the present disclosure, there is provided an anomaly detection device, comprising: a memory; and a processor coupled to the memory, configured to execute the anomaly detection method described in any one of the above embodiments based on instructions stored in the memory.
[0019] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, comprising computer program instructions, wherein when the computer program instructions are executed by a processor, the anomaly detection method described in any one of the above embodiments is implemented.
[0020] According to another aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the anomaly detection method described in any one of the above embodiments is implemented.
[0021] According to another aspect of the embodiments of the present disclosure, a computer program is provided, which, when executed by a processor, implements the anomaly detection method described in any one of the above embodiments.
[0022] In the disclosed embodiment, first, the first customs declaration data to be tested is obtained. Multiple items associated with addresses in the first customs declaration data may comprise multiple item sets. Then, the evaluation distance between two addresses associated with two items in each item set is calculated, and association rule mining is performed on multiple historical second customs declaration data to obtain an evaluation index representing the association relationship between the two items in each item set. In this case, the evaluation distance and evaluation index are processed using an anomaly detection algorithm to obtain a first result indicating whether there is a risk between the two items in each item set. Based on this first result, it is then possible to determine whether the first customs declaration data contains an anomaly. In this way, abnormal customs clearance trade behavior can be accurately detected.
[0023] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0025] FIG1 is a flow chart of an anomaly detection method according to some embodiments of the present disclosure;
[0026] FIG2 is a schematic diagram of a process for obtaining a final text distance according to some embodiments of the present disclosure;
[0027] FIG3 is a flow chart of an anomaly detection method according to other embodiments of the present disclosure;
[0028] FIG4 is a flow chart of an anomaly detection method according to yet other embodiments of the present disclosure;
[0029] FIG5 is a schematic structural diagram of an anomaly detection device according to some embodiments of the present disclosure;
[0030] FIG6 is a schematic structural diagram of an abnormality detection device according to other embodiments of the present disclosure;
[0031] FIG7 is a schematic structural diagram of an abnormality detection device according to yet other embodiments of the present disclosure. DETAILED DESCRIPTION
[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0033] Unless specifically stated otherwise, the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0034] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0035] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0036] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0037] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0038] In summary, the main anomalies in customs clearance trade behavior currently include those related to trade geographic risk between enterprises, those related to regional dispersion risk between enterprises and ports, and those related to port drift risk between enterprises and ports. These anomalies are described below.
[0039] Customs clearance trade between enterprises (also known as units or companies) with infrequent trade and long distances may have anomalies related to trade geographical risks.
[0040] Typically, due to factors such as geography, logistics, social relationships, and the industrial chain, it's more reasonable for a company to declare its customs trade activities in its home region, have it represented by a customs broker in that region, and clear customs at the port of entry in that region. With the development of a nationwide integrated customs clearance model, customs trade activities across customs districts may also be normal. However, frequent declarations outside the company's home region, frequent representation by customs brokers in different regions, or frequent customs clearance in customs districts outside the company's home region may indicate anomalies related to geographical dispersion risks.
[0041] If an enterprise has never imported or exported common commodities at a certain port within a certain time window, but has recently started to import or export common commodities at that port, there may be anomalies related to port drift risk.
[0042] By analyzing the above-mentioned major anomalies, we found that these anomalies are basically related to the historical customs clearance and trade behaviors of enterprises, and are related to the distance between enterprises, ports, destinations, etc.
[0043] In view of this, the embodiments of the present disclosure propose the following anomaly detection method, which can accurately detect abnormal customs clearance trade behavior based on the distance between addresses and frequent item sets.
[0044] FIG1 is a flowchart of an anomaly detection method according to some embodiments of the present disclosure.
[0045] As shown in FIG1 , the anomaly detection method includes steps 102 to 110 .
[0046] In step 102, first customs declaration data to be detected is obtained.
[0047] Here, the first customs declaration data includes a plurality of fields associated with corresponding items and addresses.
[0048] The multiple fields may include but are not limited to unit fields related to the unit (for example, including operating units, customs declaration units, receiving units, and sending units), port fields related to ports (for example, including import ports, export ports, and declaration ports), and destination fields related to destinations (for example, including domestic destinations).
[0049] The items corresponding to these fields are associated with addresses, that is, they have corresponding addresses. For example, the item corresponding to the business unit can be the identifier of a certain enterprise (such as enterprise name, enterprise customs code, enterprise credit code, etc.), and the enterprise corresponding to the enterprise identifier is associated with the corresponding address.
[0050] Here, the multiple items corresponding to the multiple fields may form multiple item sets, and each item set includes two items of the multiple items.
[0051] In step 104 , the evaluation distance between two addresses associated with two items in each item set is calculated.
[0052] For example, the geographical distance between two addresses associated with two items in each itemset can be calculated as the evaluation distance.
[0053] In step 106, association rule mining is performed on a plurality of historical second customs declaration data to obtain an evaluation index representing the association relationship between two items in each item set.
[0054] In some embodiments, the evaluation indicator may include at least one of frequency and confidence. For example, the evaluation indicator may include frequency. For another example, the evaluation indicator may include confidence. For another example, the evaluation indicator may include frequency and confidence.
[0055] In some embodiments, a frequent pattern (FP) algorithm may be used to perform association rule mining on multiple historical second customs declaration data to obtain evaluation indicators. The FP algorithm may be, for example, a frequent pattern growth (FP-growth) algorithm.
[0056] The FP-growth algorithm can construct a tree structure to scan multiple pieces of second customs declaration data to obtain an evaluation metric representing the association between two items in each itemset. In some implementations, a minimum support can be set to determine whether an itemset is a frequent itemset based on its frequency. In other implementations, a minimum confidence can be set to determine whether the relationship between items in an itemset is strong based on the confidence level of the itemset.
[0057] Taking the item set including items corresponding to the receiving unit and the operating unit as an example, by performing association rule mining, the evaluation index of the item set can be obtained as follows:
[0058] Among them, (Enterprise A, Enterprise B) is an item set, Enterprise A is the item corresponding to the receiving unit, and Enterprise B is the item corresponding to the operating unit. The example indicates that the frequency of trade between the two items in (Enterprise A, Enterprise B) is 3 and the confidence level is 1.
[0059] In step 108 , the evaluation distance and the evaluation index are processed using an anomaly detection algorithm to obtain a first result indicating whether there is a risk between two items in each item set.
[0060] In other words, by taking the evaluation distance and evaluation index between item sets as the input of the anomaly detection algorithm, an output (ie, the first result) can be obtained to indicate whether there is a risk in the relationship between two items in the item set.
[0061] Here is an example of the first result for three item sets of the first customs declaration data:
[0062] {
[0063] "Relationship between the recipient and the operating unit": {"(Enterprise A, Enterprise B)": 0},
[0064] "Relationship between the recipient and the customs declaration unit":{"(Company A, Company C)":0},
[0065] "Relationship between business unit and customs declaration unit":{"(Enterprise B, Enterprise C)":0},
[0066] }
[0067] Here, (Enterprise A, Enterprise B), (Enterprise A, Enterprise C), and (Enterprise B, Enterprise C) are all item sets, with Enterprise C being the item corresponding to the customs declaration entity. A 0 after an item set indicates that the relationship between the two items in the item set is risk-free, while a 1 after an item set indicates that the relationship between the two items in the item set is risky. In the above example, all three item sets are risk-free.
[0068] In some embodiments, the anomaly detection algorithm may include one or more of an isolation forest algorithm, a one-class support vector machine (One Class SVM) algorithm, and a statistical anomaly threshold algorithm. The isolation forest algorithm can isolate abnormal data by constructing a binary tree, thereby detecting abnormal data. The One Class SVM algorithm can determine normal data by determining a hyperplane, thereby separating abnormal data. The statistical anomaly threshold algorithm can mine abnormal data based on the quartiles and standard deviation of the data distribution.
[0069] As some implementations, anomaly detection algorithms can include the isolation forest algorithm, the one-class support vector machine (SVM) algorithm, and the statistical anomaly threshold algorithm. In this approach, the isolation forest algorithm, the one-class support vector machine (SVM), and the statistical anomaly threshold algorithm can be used to process the evaluation distance and evaluation index of an item set, respectively, to obtain sub-results indicating whether a risk exists between two items in the item set. A first result can then be determined based on the sub-results of each of the three algorithms. For example, if a sub-result from any of the three algorithms indicates a risk exists between two items in a particular item set, the first result indicates a risk exists between the two items in the item set. This improves the accuracy of the first result, facilitating more accurate detection of abnormal customs trade behavior and reducing false positive rates.
[0070] In step 110, based on the first result, it is determined whether there is any abnormality in the first customs declaration data.
[0071] In some embodiments, when the first result indicates that there is a risk between two items in a plurality of item sets exceeding a preset number of item sets, it is determined that there is an abnormality in the first customs declaration data (i.e., there is an abnormal customs clearance trade behavior); and when the first result indicates that there is a risk between two items in a plurality of item sets not exceeding a preset number of item sets, it is determined that there is no abnormality in the first customs declaration data (i.e., there is no abnormal customs clearance trade behavior).
[0072] For example, the preset number may be 0. In this case, as long as the first result indicates that there is a risk between two items in any item set among the multiple item sets, it is determined that the first customs declaration data has an anomaly.
[0073] In the above embodiment, first, the first customs declaration data to be tested is obtained. Multiple items associated with addresses in the first customs declaration data may form multiple item sets. Then, the evaluation distance between two addresses associated with two items in each item set is calculated, and association rule mining is performed on multiple historical second customs declaration data to obtain an evaluation index representing the association relationship between the two items in each item set. In this case, the evaluation distance and evaluation index are processed using an anomaly detection algorithm to obtain a first result indicating whether there is a risk between the two items in each item set. Based on this first result, it is then possible to determine whether the first customs declaration data contains an anomaly. In this way, anomalous customs clearance trade behavior can be accurately detected.
[0074] Next, the anomaly detection method shown in FIG1 is further described with reference to some embodiments.
[0075] In some embodiments, the plurality of item sets includes an item set (hereinafter referred to as a first item set) in which both fields corresponding to two items are unit fields.
[0076] It is understandable that the number of first item sets can be one or more. For example, the three items corresponding to the three unit fields of operating unit, customs declaration unit, and receiving (shipping) unit can be combined in pairs to obtain three first item sets: the item set corresponding to the receiving (shipping) unit and the customs declaration unit, the item set corresponding to the receiving (shipping) unit and the operating unit, and the item set corresponding to the operating unit and the customs declaration unit.
[0077] In these embodiments, when the first result indicates that there is a risk between two items in the first item set, it can be determined that the first customs declaration data has an anomaly related to the trade region risk.
[0078] For example, if there are multiple first item sets, as long as the first result indicates that there is a risk between two items in any item set, it can be determined that the first customs declaration data contains an anomaly related to trade region risk. Conversely, if the first result indicates that there is no risk between the two items in each first item set, it can be determined that there is no anomaly related to trade region risk in the first customs declaration data.
[0079] As explained above, anomalies related to trade geographical risks usually occur between companies with infrequent trade and long distances between them.
[0080] This disclosed embodiment determines whether the first customs declaration data contains anomalies related to trade regional risk based on whether the first item set, where both fields corresponding to the two items represented by the first result are unit fields, contains risk. This allows accurate detection of abnormal customs clearance trade behavior related to inter-enterprise trade regional risk.
[0081] In some embodiments, the plurality of item sets include an item set (hereinafter referred to as a second item set) in which two fields corresponding to two items are respectively one and the other of a unit field, a port field, and a destination field.
[0082] It is understood that the number of second item sets can be one or more. For example, there can be the following eight second item sets: the item set corresponding to the receiving (sending) unit and the domestic destination, the item set corresponding to the receiving (sending) unit and the import (export) port, the item set corresponding to the operating unit and the domestic destination, the item set corresponding to the operating unit and the import (export) port, the item set corresponding to the customs declaration unit and the domestic destination, the item set corresponding to the customs declaration unit and the import (export) port, the item set corresponding to the customs declaration unit and the declaration port, and the item set corresponding to the import (export) port and the domestic destination.
[0083] In these embodiments, when the first result indicates that there is a risk between two items in the second item set, it can be determined that the first customs declaration data has an anomaly related to the geographical dispersion risk.
[0084] For example, if there are multiple second item sets, as long as the first result indicates that there is a risk between two items in any second item set, it can be determined that the first customs declaration data contains an anomaly related to regional dispersion risk. Conversely, if the first result indicates that there is no risk between the two items in the second item set, it can be determined that there is no anomaly related to regional dispersion risk in the first customs declaration data.
[0085] As explained above, anomalies related to geographical dispersion risks usually include frequent declarations outside the enterprise's location, frequent representation by customs declaration agencies in different regions, or frequent customs clearance in customs areas outside the enterprise's location.
[0086] The disclosed embodiment determines whether the first customs declaration data contains anomalies related to regional dispersion risk based on whether a risk exists in a second item set where two fields corresponding to two items represented by the first result are one unit field, one port field, and one destination field. In this way, abnormal customs clearance trade behavior related to regional dispersion risk can be accurately detected.
[0087] In combination with the above embodiments, it can be seen that by adjusting the fields corresponding to the two items in the item set in the anomaly detection method shown in Figure 1, abnormal customs clearance trade behaviors can be accurately detected, especially abnormal customs clearance trade behaviors related to trade regional risks and regional dispersion risks.
[0088] Next, some embodiments of step 104 are described.
[0089] In some embodiments, the evaluation distance between two addresses associated with two items in each item set may be calculated in the following manner.
[0090] First, the two addresses associated with the two items in each item set may be resolved to obtain two resolved addresses.
[0091] In some embodiments, the Python-based cpcA module can be used to parse two addresses to obtain two intermediate resolved addresses. The intermediate resolved addresses can be in the format of [province, city, district, sub-address], where "sub-address" can include street / town / village, community, and other information. Because the cpcA module can complete ambiguous or missing addresses, the intermediate resolved addresses are more accurate.
[0092] As some implementations, the intermediate resolution address may be directly used as the resolution address.
[0093] As another implementation, a regular expression algorithm can be used to further parse the "sub-address" in the intermediate resolved address. The format of the resolved sub-address can be [town, community]. The intermediate resolved address and the addresses at each level in the resolved sub-address are then concatenated to obtain the resolved address in the format of [province, city, county-level city / district / county, street / town / township, community|subdistrict|village|village|industrial park|industrial park|science and technology park|factory|company|office building|building]. This can further improve the accuracy of the resolved address, thereby more accurately detecting abnormal customs trade behavior.
[0094] Then, the geographic coordinates of the two resolved addresses may be obtained. The geographic coordinates may include, for example, the longitude and latitude corresponding to the resolved addresses.
[0095] In some implementations, when obtaining the geographic coordinates of a resolved address, the lowest-level geographic coordinates within the resolved address are obtained whenever possible. For example, an exact match is performed first on the lowest-level geographic coordinates. If the exact match fails, a fuzzy match is performed on the lowest-level geographic coordinates. If the fuzzy match also fails, the next-level geographic coordinates are matched, first with an exact match and then with a fuzzy match. This process is repeated until a match is successful.
[0096] Finally, the geographic distance between the two resolved addresses is calculated based on their geographic coordinates. In this case, the evaluated distance includes the geographic distance.
[0097] Taking geographic coordinates including longitude and latitude as an example, the geographic distance di,j between two resolved addresses can be calculated based on the following formula:
[0098] Where i represents one of the two resolved addresses, j represents the other of the two resolved addresses, W represents latitude, J represents longitude, and R represents the radius of the earth.
[0099] In some embodiments, the plurality of item sets include an item set (hereinafter referred to as a third item set) in which the two fields corresponding to the two items do not include a port field. For example, the third item set may include the first item set.
[0100] In this case, a first text distance between two character strings corresponding to two addresses associated with two items in the third item set may also be calculated, and a final text distance may be obtained based on the first text distance. In these embodiments, the evaluation distance of the third item set includes the final text distance in addition to the geographic distance.
[0101] Since the address of a port is usually in the form of "such and such port", the text distance (also called edit distance) between the port address and the addresses of other items is difficult to accurately reflect the distance between the addresses. Therefore, the embodiment of the present disclosure does not need to calculate the text distance of the item set including the items corresponding to the port field, but only calculates the text distance of the third item set that does not include the items corresponding to the port field.
[0102] In the above embodiment, for the third item set whose corresponding fields do not include the port field, text distance is also calculated as one of the evaluation distances. However, for the item set whose corresponding fields include the port field, text distance is not calculated as one of the evaluation distances. This allows for a more accurate first result based on the evaluation distance, thereby more accurately detecting abnormal customs clearance trade behavior, while also reducing computational burden.
[0103] Some implementations of obtaining the final text distance based on the first text distance are described below.
[0104] As some implementations, both resolved addresses include multiple levels.
[0105] When the character strings corresponding to multiple levels all contain empty (ie, None), the first text distance is used as the final text distance.
[0106] When the character strings corresponding to at least one of the multiple levels do not contain null, a second text distance between the character strings corresponding to each of the at least one level is calculated, and a final text distance is obtained based on the first text distance and the second text distance.
[0107] For example, if the resolved address is the intermediate resolved address in the previous example, each resolved address includes four levels: province, city, district, and sub-address. In this case, the two resolved addresses are [province 1, city 1, district 1, sub-address 1] and [province 2, city 2, district 2, sub-address 2].
[0108] If the strings corresponding to multiple levels all contain empty spaces, that is, any one of "Province 1" and "Province 2" is empty, any one of "City 1" and "City 2" is empty, any one of "District 1" and "District 2" is empty, and any one of "Subaddress 1" and "Subaddress 2" is empty, then the first text distance is taken as the final text distance.
[0109] If the character strings corresponding to at least one level do not contain empty characters, for example, "Province 1" and "Province 2" are not empty, "City 1" and "City 2" are not empty, "District 1" and "District 2" are not empty, or "Sub-address 1" and "Sub-address 2" are not empty, then the second text distance between the character strings corresponding to each level in at least one level is calculated, and the final text distance is obtained based on the first text distance and the second text distance.
[0110] In some implementations, each level in the multiple levels has a corresponding weight. The second text distances of the corresponding levels whose character strings do not contain spaces can be weighted summed, and a weighted average of the weighted summation result and the first text distance is calculated as the final text distance.
[0111] For ease of understanding, the following description is made in conjunction with Figure 2. Figure 2 is a schematic diagram of a process for obtaining a final text distance according to some embodiments of the present disclosure.
[0112] As shown in FIG2 , for two addresses associated with two items in an item set, we can first determine whether the two addresses are the same, and then use the cpca module to parse the two addresses to obtain two parsed addresses.
[0113] If the two unresolved addresses are the same, the first text distance is directly determined to be equal to 1 without further calculation. In this case, the final text distance is equal to the first text distance.
[0114] It can be understood that the closer the text distance is to 1, the higher the similarity between the two strings.
[0115] If the two unresolved addresses are different, the first text distance between the character strings corresponding to the two addresses is calculated, and it is determined whether the character strings corresponding to the multiple levels of the two resolved addresses all contain spaces.
[0116] If the strings corresponding to multiple levels all contain spaces, the first calculated text distance is directly used as the final text distance.
[0117] If the strings corresponding to multiple levels do not contain spaces, the second text distances between the strings corresponding to the levels that do not contain spaces are calculated. Then, a weighted sum of the second text distances between the strings corresponding to the levels that do not contain spaces is performed. In this case, the final text distance is the weighted average of the calculated first text distance and the weighted sum.
[0118] As some implementations, when performing a weighted summation of the second text distance, the smaller the level among the multiple levels, the greater the weight. Taking the format of the parsed address as [province, city, district, sub-address] as an example, the weight of the level "province" can be 0.1, the weight of the level "city" can be 0.2, the weight of the level "district" can be 0.3, and the weight of the level "sub-address" can be 0.4. In this way, the final text distance can be calculated more accurately, thereby more accurately detecting abnormal customs clearance trade behavior.
[0119] As some implementations, if only one of the multiple levels, except for the smallest level, contains a space, and the levels other than the space-containing level are identical, the second text distance of each level can be directly determined as 1 without calculating separately. For example, if the text of the level "province" is identical, the text of the level "district" is identical, and the text of the level "sub-address" is identical, and only the level "city" contains a space, the two resolved addresses can be considered identical, and the second text distance of each level can be directly determined as 1.
[0120] As some implementations, the text distance (e.g., the first text distance and the second text distance) can be calculated based on the Jaro-Winkler algorithm. In the Jaro-Winkler algorithm, the text distance D = d + L*p(1-d), where d represents the calculated score of the Jaro distance between the two strings, L is the length of the same prefix (maximum value is 4), and p represents a constant for adjusting the score (maximum value is 25, for example, can be 0.1).
[0121] A text distance D=0 indicates that the two character strings are completely different, and a text distance D=1 indicates that the two character strings are exactly the same. Moreover, the closer the text distance D is to 1, the higher the similarity between the two character strings.
[0122] The Jaro distance score d can be calculated using the following formula:
[0123] Where m is the number of characters to match, t is the number of character transitions, and |l1| and |l2| represent the lengths of the two strings to be compared.
[0124] So far, some embodiments of the anomaly detection method shown in Figure 1 have been described. Next, some other embodiments of the anomaly detection method of the present disclosure will be described.
[0125] FIG3 is a flowchart of an anomaly detection method according to other embodiments of the present disclosure.
[0126] As shown in Figure 3, the anomaly detection method includes steps 302 to 312. Steps 302 to 312 that are similar to those described above will not be described again, and the relevant embodiments can be found in the previous text.
[0127] In step 302, first customs declaration data to be detected is obtained.
[0128] In step 304, a first group of second customs declaration form data including items corresponding to the unit field is selected from a plurality of historical second customs declaration form data.
[0129] In step 306, the commodity numbers whose occurrence times in the first set of second customs declaration data are greater than a threshold are determined.
[0130] The product number may represent a common commodity of the item corresponding to the unit field. For example, the product numbers that appear most frequently (e.g., the top 20) in the first set of second customs declaration data may be determined as common commodities. The product number may be, for example, a commodity tax number.
[0131] In step 308, a second group of second customs declaration form data containing the commodity number is filtered out from the first group of second customs declaration form data.
[0132] In step 310, based on whether an item corresponding to the port field related to the port in the second customs declaration data within the first time window in the second group of second customs declaration data appears in the second customs declaration data within the second time window in the second group of second customs declaration data, a second result indicating whether the item corresponding to the unit field has a risk is determined.
[0133] Here, the start time of the first time window is not earlier than the end time of the second time window. That is, the first time window is later than the second time window. For example, the first time window is within the last 30 days, and the second time window is within the last 2 years to the last 30 days ago.
[0134] In some embodiments, if an item corresponding to the port field (i.e., a port identifier) in the second customs declaration data within the first time window of the second set of second customs declaration data appears in the second customs declaration data within the second time window of the second set of second customs declaration data, this indicates that the most recent customs clearance port for the enterprise's common commodities is a previous customs clearance port. In this case, the second result indicates that the item corresponding to the unit field (i.e., the enterprise) does not pose a risk.
[0135] Conversely, if the item corresponding to the port field in the second customs declaration data within the first time window of the second set of second customs declaration data does not appear in the second customs declaration data within the second time window of the second set of second customs declaration data, this indicates that the most recent customs clearance port for the company's common products is not a previous customs clearance port. In this case, the second result indicates that the item corresponding to the unit field is risky.
[0136] An example of the second result is as follows:
[0137] This example shows that in the first customs declaration data, the item corresponding to the receiving unit (i.e., enterprise A), the item corresponding to the operating unit (i.e., enterprise B), and the item corresponding to the customs declaration unit (i.e., enterprise C) are all enterprises with import drift risk and export drift risk. That is, the items corresponding to the unit field in the first customs declaration data all have risks.
[0138] In step 312, based on the second result, it is determined whether there is any abnormality in the first customs declaration data.
[0139] In some embodiments, if the second result indicates that the item corresponding to the unit field has a risk, the first customs declaration data is determined to have an anomaly related to the risk of port drift. Conversely, if the second result indicates that the item corresponding to the unit field does not have a risk, the first customs declaration data is determined to have no anomaly related to the risk of port drift.
[0140] According to the explanation in the previous article, the company has never imported or exported common goods at a certain port within a certain time window, but has recently started to import or export common goods at the port. This may be an anomaly related to the port drift risk.
[0141] In the disclosed embodiment, by screening out a first set of second customs declaration data from historical second customs declaration data that includes a company identifier in the current first customs declaration data, and determining the commodity numbers that appear more than a threshold number of times in the first set of second customs declaration data, common commodities of the company can be found. Then, by determining to find a second set of second customs declaration data from the first set of second customs declaration data that includes the common commodity numbers, it is possible to determine whether the port of entry in the second set of second customs declaration data within the first time window appears in the second set of second customs declaration data within the second time window and obtain a second result. Furthermore, based on the second result, it is possible to determine whether the first customs declaration data is abnormal. In this way, abnormal customs clearance trade behavior, particularly abnormal customs clearance trade behavior associated with port drift risk, can be accurately detected.
[0142] FIG1 and FIG3 above illustrate two different methods for detecting anomalies in customs clearance trade behaviors. However, the embodiments of FIG1 and FIG3 can also be combined with each other. This will be explained below in conjunction with FIG4.
[0143] FIG4 is a flowchart of an anomaly detection method according to yet other embodiments of the present disclosure.
[0144] As shown in FIG4 , the anomaly detection method includes steps 402 to 410 .
[0145] In step 402, first customs declaration data is obtained.
[0146] In step 404, a trade region risk test is performed to obtain a first result of the first part (also referred to as a trade region risk result).
[0147] Specifically, the geographic distance and final textual distance (i.e., evaluation distance) between the two addresses associated with the two items in the first itemset can be calculated, and association rule mining can be performed to obtain evaluation metrics for the first itemset. Subsequently, the isolation forest algorithm, one-class support vector machine algorithm, and statistical anomaly threshold algorithm can be used to process the data to obtain trade regional risk results. Related embodiments can be found in the previous description and will not be elaborated here.
[0148] In step 406, a regional dispersion risk test is performed to obtain a first result of the second part (also referred to as a regional dispersion risk result).
[0149] Specifically, the geographic distance (i.e., evaluation distance) between two addresses associated with two items in the second itemset can be calculated, and association rule mining can be performed to obtain evaluation metrics for the second itemset. Subsequently, the isolation forest algorithm, one-class support vector machine algorithm, and statistical anomaly threshold algorithm can be used to process the data to obtain a regionally dispersed risk result. Related embodiments can be found in the previous description and will not be elaborated on here.
[0150] In step 408, a port drift risk detection is performed to obtain a second result (also referred to as a port drift enterprise list).
[0151] Specifically, the common commodities (e.g., common import commodities and common export commodities) of the enterprise in the first customs declaration data can be first screened, and then it can be determined whether the customs clearance port for the enterprise's common commodities in the most recent time window (i.e., the first time window) is the customs clearance port in the earlier time window (i.e., the second time window), so as to obtain a list of enterprises with port drift. Relevant embodiments can be found in the previous description and will not be elaborated here.
[0152] It is understandable that FIG4 schematically shows that the first time window is within the last 30 days, and the second time window is from the last 2 years to the last 30 days ago, but the embodiments of the present disclosure are not limited thereto.
[0153] In step 410 , it may be determined whether there is an abnormality in the first customs declaration data based on the first result and the second result.
[0154] For example, a comprehensive risk value indicating whether the first customs declaration data has an anomaly can be determined based on the first result and the second result, and the first result, the second result, and the determined comprehensive risk value can be output. A user (e.g., customs staff) can assess the risk of the customs declaration based on the output first result, the second result, and the comprehensive risk value. A higher comprehensive risk value indicates a higher risk of the customs declaration and a more serious anomaly.
[0155] The following is an example of the output of the first result, second result, and overall risk value:
[0156] In the above example, a risk value of 1 indicates that the item set is at risk, while a risk value of 0 indicates that the item set is not at risk. Furthermore, the list of port drifting enterprises is empty, indicating that the item corresponding to the unit field in the first customs declaration data is not at risk. If the list of port drifting enterprises is not empty, the risk value corresponding to the list is the number of enterprises on the list.
[0157] It is understood that the risk value corresponding to the port drift enterprise list can be any integer in [0, 3]. It is also understood that the receiving unit in the above examples can be the sending unit in some cases, and the import port can be the export port in some cases.
[0158] In some embodiments, the comprehensive risk value can be obtained by summing the individual risk values. In this case, the comprehensive risk value can be any integer in [0, 14].
[0159] In other embodiments, the comprehensive risk value may be obtained by weighted summing of the individual risk values.
[0160] In some embodiments, it may be determined that the first customs declaration data does not have an anomaly when the comprehensive risk value is equal to 0. In other embodiments, it may be determined that the first customs declaration data does have an anomaly when the comprehensive risk value is not equal to 0.
[0161] If an anomaly is determined in the first customs declaration data, the type of anomaly can be further determined. Referring to Table 1 below, if the risk value of any of the objects numbered 1 to 3 is 1, the first customs declaration data is determined to have an anomaly related to trade region risk; if the risk value of any of the objects numbered 4 to 11 is 1, the first customs declaration data is determined to have an anomaly related to regional dispersion risk; and if the risk value of any of the objects numbered 12 to 14 is 1 (i.e., the risk value corresponding to the list of port drifting enterprises is not 0), the first customs declaration data is determined to have an anomaly related to port drift risk.
[0162] Table 1
[0163] Based on the anomaly detection method shown in Figure 4, abnormal customs clearance trade behaviors related to trade regional risk, regional dispersion risk, and port drift risk can be accurately detected simultaneously, thereby more comprehensively identifying possible anomalies in customs declaration data and helping users to more accurately assess the risks of customs declarations.
[0164] The embodiment of the present disclosure also provides an abnormality detection device.
[0165] In some embodiments, the anomaly detection apparatus includes a module configured to execute the anomaly detection method of any one of the above embodiments.
[0166] In other embodiments, the anomaly detection device includes a memory and a processor coupled to the memory, and the processor is configured to execute the anomaly detection method of any one of the above embodiments based on instructions stored in the memory.
[0167] FIG5 is a schematic structural diagram of an anomaly detection device according to some embodiments of the present disclosure.
[0168] As shown in FIG5 , the anomaly detection device 500 includes an acquisition module 501 , a calculation module 502 , a mining module 503 , a processing module 504 and a determination module 505 .
[0169] The acquisition module 501 is configured to acquire first customs declaration data to be detected. Here, the first customs declaration data includes multiple fields in which corresponding items are associated with addresses. The multiple items corresponding to the multiple fields can form multiple item sets, and each item set includes two items from the multiple items.
[0170] The calculation module 502 is configured to calculate an evaluation distance between two addresses associated with two items in each item set.
[0171] The mining module 503 is configured to perform association rule mining on multiple historical second customs declaration data to obtain an evaluation index representing the association relationship between two items in each item set. In some embodiments, the evaluation index includes at least one of frequency and confidence.
[0172] The processing module 504 is configured to process the evaluation distance and the evaluation index using an anomaly detection algorithm to obtain a first result indicating whether there is a risk between two items in each item set. In some embodiments, the anomaly detection algorithm includes an isolation forest algorithm, a single-class support vector machine algorithm, and a statistical anomaly threshold algorithm.
[0173] The determination module 505 is configured to determine whether there is any abnormality in the first customs declaration data based on the first result.
[0174] In some embodiments, the multiple item sets include a first item set, and both fields corresponding to the two items in the first item set are unit fields related to units. In this case, determining module 505 determines whether the first customs declaration data has an anomaly based on the first result, including: if the first result indicates that there is a risk between the two items in the first item set, determining that the first customs declaration data has an anomaly related to trade region risk.
[0175] In some embodiments, the multiple item sets include a second item set, and the two fields corresponding to the two items in the second item set are respectively one and the other of a unit field related to the unit, a port field related to the port, and a destination field related to the destination. In this case, determining module 505 determines whether the first customs declaration data has an anomaly based on the first result, including: if the first result indicates that there is a risk between the two items in the second item set, determining that the first customs declaration data has an anomaly related to the risk of geographical dispersion.
[0176] In some embodiments, the calculation module 502 calculates the estimated distance between two addresses associated with two items in each item set, including: resolving the two addresses associated with the two items in each item set to obtain two resolved addresses; obtaining geographic coordinates of the two resolved addresses; and calculating the geographic distance between the two resolved addresses based on the geographic coordinates. In this case, the estimated distance includes the geographic distance.
[0177] In some embodiments, the multiple item sets include a third item set, and the two fields corresponding to the two items in the third item set do not include a port field related to a port. In this case, calculating module 502 calculates the estimated distance between the two addresses associated with the two items in each item set, further comprising: calculating a first text distance between the two strings corresponding to the two addresses associated with the two items in the third item set; and obtaining a final text distance based on the first text distance. Here, the estimated distance of the third item set also includes the final text distance.
[0178] In some embodiments, both resolved addresses obtained by the calculation module 502 include multiple levels. In this case, the calculation module 502 obtains the final text distance based on the first text distance, including: when the strings corresponding to the multiple levels all contain spaces, using the first text distance as the final text distance; and when the strings corresponding to at least one level in the multiple levels do not contain spaces, calculating the second text distance between the strings corresponding to each level in the at least one level. Here, the calculation module 502 obtains the final text distance based on the first text distance and the second text distance.
[0179] FIG6 is a schematic structural diagram of an anomaly detection device according to other embodiments of the present disclosure.
[0180] As shown in FIG6 , the abnormality detection device 600 includes an acquisition module 601 , a screening module 602 and a determination module 603 .
[0181] The acquisition module 601 is configured to acquire the first customs declaration data to be detected.
[0182] The screening module 602 is configured to screen out a first group of second customs declaration form data in which an item corresponding to the unit field appears from the plurality of second customs declaration form data.
[0183] The determination module 603 is configured to determine the commodity numbers in the first set of second customs declaration data that appear more than a threshold number of times.
[0184] The screening module 602 is further configured to screen out a second group of second customs declaration form data in which the commodity number appears from the first group of second customs declaration form data.
[0185] Determination module 603 is further configured to determine a second result indicating whether the item corresponding to the unit field is risky based on whether the item corresponding to the port field related to the port in the second customs declaration data within the first time window in the second set of second customs declaration data appears in the second customs declaration data within the second time window in the second set of second customs declaration data. Here, the start time of the first time window is not earlier than the end time of the second time window.
[0186] The determination module 603 is further configured to determine whether there is any abnormality in the first customs declaration data based on the second result.
[0187] In some embodiments, when the second result indicates that the item corresponding to the unit field has a risk, the determination module 603 determines that the first customs declaration data has an anomaly related to the port drift risk.
[0188] It can be understood that the anomaly detection apparatus 500 / 600 may include other modules not shown to execute the anomaly detection method of any of the above embodiments.
[0189] FIG7 is a schematic structural diagram of an abnormality detection device according to yet other embodiments of the present disclosure.
[0190] As shown in FIG7 , an anomaly detection device 700 includes a memory 701 and a processor 702 coupled to the memory 701 . The processor 702 is configured to execute the anomaly detection method of any one of the above embodiments based on instructions stored in the memory 701 .
[0191] The memory 701 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory may store, for example, an operating system, an application program, a boot loader, and other programs.
[0192] The anomaly detection device 700 may also include an input / output interface 703, a network interface 704, a storage interface 705, and the like. These interfaces, such as the input / output interface 703, the network interface 704, and the storage interface 705, as well as the memory 701 and the processor 702, may be connected via, for example, a bus 706. The input / output interface 703 provides a connection interface for input / output devices such as a display, mouse, keyboard, and touch screen. The network interface 704 provides a connection interface for various networked devices. The storage interface 705 provides a connection interface for external storage devices such as SD cards and USB flash drives.
[0193] An embodiment of the present disclosure further provides a computer-readable storage medium, comprising computer program instructions, which, when executed by a processor, implement the anomaly detection method of any one of the above embodiments.
[0194] An embodiment of the present disclosure further provides a computer program product, including a computer program, which implements the anomaly detection method of any one of the above embodiments when executed by a processor.
[0195] The embodiments of the present disclosure further provide a computer program, which, when executed by a processor, implements the anomaly detection method of any one of the above embodiments.
[0196] Thus far, various embodiments of the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details known in the art have not been described. Based on the above description, those skilled in the art can fully understand how to implement the technical solutions disclosed herein.
[0197] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the device embodiments, since they are essentially identical to the method embodiments, their descriptions are relatively simple. For relevant parts, reference can be made to the descriptions of the method embodiments.
[0198] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transient storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0199] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0200] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0201] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0202] Although some specific embodiments of the present disclosure have been described in detail through examples, those skilled in the art will understand that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that the above embodiments may be modified or some technical features may be replaced with equivalents without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. An anomaly detection method, comprising: Obtaining first customs declaration data to be detected, where the first customs declaration data includes multiple fields with corresponding items associated with addresses, and multiple items corresponding to the multiple fields can form multiple item sets, and each item set includes two items among the multiple items; Calculating an evaluation distance between two addresses associated with the two items in each item set; Performing association rule mining on multiple historical second customs declaration data to obtain an evaluation index representing the association relationship between the two items in each item set; Processing the evaluation distance and the evaluation index by using an anomaly detection algorithm to obtain a first result indicating whether there is a risk between the two items in each item set; And Based on the first result, determining whether the first customs declaration data is abnormal.
2. The method according to claim 1, wherein The multiple item sets include a first item set, and the two fields corresponding to the two items in the first item set are both unit fields related to the unit; Wherein, determining whether the first customs declaration data is abnormal based on the first result includes: In the case where the first result indicates that there is a risk between the two items in the first item set, determining that the first customs declaration data has an anomaly related to the trade region risk.
3. The method according to claim 1, wherein The multiple item sets include a second item set, and the two fields corresponding to the two items in the second item set are respectively one and the other of a unit field related to the unit, a port field related to the port, and a destination field related to the destination; Wherein, determining whether the first customs declaration data is abnormal based on the first result includes: In the case where the first result indicates that there is a risk between the two items in the second item set, determining that the first customs declaration data has an anomaly related to the geographical dispersion risk.
4. The method according to any one of claims 1-3, wherein, Calculating the evaluation distance between two addresses associated with the two items in each item set includes: Parsing the two addresses associated with the two items in each item set to obtain two parsed addresses; Obtaining the geographical coordinates of the two parsed addresses; and Based on the geographical coordinates, calculating the geographical distance between the two parsed addresses, where the evaluation distance includes the geographical distance.
5. The method according to claim 4, wherein The multiple item sets include a third item set, and the two fields corresponding to the two items in the third item set do not include a port field related to the port; Wherein, calculating the evaluation distance between two addresses associated with the two items in each item set further includes: Calculating a first text distance between two strings corresponding to the two addresses associated with the two items in the third item set; and Obtaining a final text distance based on the first text distance, where the evaluation distance of the third item set further includes the final text distance.
6. The method according to claim 5, wherein, Both of the two parsed addresses include multiple levels; Obtaining the final text distance based on the first text distance includes: In the case where the strings corresponding to the multiple levels all contain null, taking the first text distance as the final text distance; and In the case where at least one of the strings corresponding to the multiple levels does not contain null, calculating a second text distance between the strings corresponding to each level in the at least one level; Wherein, the final text distance is obtained based on the first text distance and the second text distance.
7. The method according to any one of claims 1-6, wherein, The evaluation metrics include at least one of frequency and confidence.
8. The method according to any one of claims 1-6, wherein, The anomaly detection algorithms include the Isolation Forest algorithm, the one-class Support Vector Machine algorithm, and the statistical anomaly threshold algorithm.
9. The method according to any one of claims 1-6, wherein The multiple fields include a unit field related to a unit; The method further includes: screening out a first set of second customs declaration data in which items corresponding to the unit field appear from the multiple sets of second customs declaration data; determining product numbers that appear more than a threshold number of times in the first set of second customs declaration data; screening out a second set of second customs declaration data in which the product numbers appear from the first set of second customs declaration data; and determining a second result indicating whether there is a risk for the item corresponding to the unit field based on whether an item corresponding to a port field related to a port in the second customs declaration data within a first time window appears in the second customs declaration data within a second time window in the second set of second customs declaration data, wherein a start time of the first time window is not earlier than an end time of the second time window; Wherein, it is further determined whether the first customs declaration data is abnormal based on the second result.
10. The method according to claim 9, wherein, In a case where the second result indicates that there is a risk for the item corresponding to the unit field, it is determined that the first customs declaration data has an anomaly related to a port drift risk.
11. An anomaly detection device, comprising: a module configured to execute the anomaly detection method according to any one of claims 1-10.
12. An anomaly detection device, comprising: a memory; and a processor coupled to the memory and configured to execute the anomaly detection method according to any one of claims 1-10 based on instructions stored in the memory.
13. A computer-readable storage medium, comprising computer program instructions, wherein, When the computer program instructions are executed by the processor, the anomaly detection method according to any one of claims 1-10 is implemented.
14. A computer program product, comprising a computer program, wherein, When the computer program is executed by the processor, the anomaly detection method according to any one of claims 1-10 is implemented.
15. A computer program, which when executed by a processor, implements the anomaly detection method according to any one of claims 1-10.
Citation Information
Patent Citations
Customs declaration risk management and control method, device, compute device and storage medium
CN109146248A
Express receiving and sending organization discovery method based on relationship mining and related equipment
CN112581062A
Custom clearance risk identification method and device, equipment, medium and program product
CN116911591A
Abnormality detection method and device and computer readable storage medium
CN117744006A
A method for performing a customs procedure
EP1653402A1
Cited By
Rehabilitation glove evaluation method and system based on multi-source data analysis
CN120954754A
A rehabilitation glove evaluation method and system based on multi-source data analysis
CN120954754B
Data cleaning and repairing method and device
CN121579461A