Abnormal transaction identification method and system
By calculating the distance matrix between transaction features and updating, extracting dark blocks and adjusting cluster division, the accuracy and noise sensitivity of transaction anomaly data detection in the prior art are solved, and higher abnormal transaction recognition accuracy and cluster analysis accuracy are achieved.
Patent Information
- Application Number
- CN202510184986.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult to accurately identify abnormal transactions in the detection of transaction abnormal data, especially when facing changing transaction patterns, the preset number of clusters may lead to misjudgment and missed detection, and the noise data is highly sensitive, affecting the detection accuracy.
By calculating the distance matrix between transaction features and updating them with direct and indirect paths, reordering different similar images are obtained. Then, dark blocks are extracted, the number of transaction clusters is estimated based on the number of dark blocks and the characteristic importance, and cluster division is adjusted by re-identifying dark blocks to improve the automation and accuracy of cluster analysis.
It improves the accuracy of abnormal transaction identification, especially when processing long-chain clusters, reduces misjudgment and missed detection, enhances resistance to noise data, and improves the automation and accuracy of clustering analysis.
Smart Images

Figure CN120070055A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of transaction identification, and particularly to an abnormal transaction identification method and system. Background Art
[0002] Online shopping and online payment are becoming more and more common, and it is increasingly difficult to identify abnormalities in transactions. The detection of abnormal transaction data has become a key task that cannot be ignored in financial risk management. Abnormal transaction data usually refers to abnormal behaviors or data patterns that occur during transactions, and these abnormalities may be caused by various reasons, such as malicious attacks, system failures, user operation errors, etc. Timely identification of these abnormal transaction data can effectively prevent fraud, avoid financial risks, ensure system security, and provide a reliable monitoring basis for regulatory authorities; moreover, the detection of abnormal transactions can ensure the normal operation of financial transactions.
[0003] Currently, there are various detection methods for abnormal transaction data. Using clustering methods can effectively identify abnormal transactions that are significantly different from most transaction behaviors. However, the diversity and dynamics of transaction behaviors make it difficult to determine the appropriate number of clusters based on prior experience or assumptions. Especially when facing constantly changing transaction patterns, this preset number of clusters may lead to misjudgment and missed detection; moreover, existing clustering methods are highly sensitive to noise data, resulting in blurred boundaries of clustering results, thereby affecting the accuracy of abnormal detection. Summary of the Invention
[0004] In order to improve the accuracy of identifying abnormal transactions, first, the present invention provides an abnormal transaction identification method, and the method includes: Obtain the features of each transaction, calculate the distances between the same features of different transactions to obtain a distance matrix of features. For any two points in the distance matrix, obtain the direct path connecting the two points and all non-direct paths, and update the distance matrix based on the direct path and all non-direct paths; rearrange the updated distance matrix to obtain a reordered dissimilarity image of features; Extract the dark blocks in the reordered dissimilarity image of features, obtain the number of transaction clusters according to the number of dark blocks in the reordered dissimilarity image of features and the importance of features, re-identify the dark blocks of the reordered dissimilarity image of each feature according to the number of transaction clusters, and obtain the clusters of transactions by using the re-identified dark blocks; judge whether the transaction to be identified is abnormal according to the clusters of transactions.
[0005] Secondly, the present invention provides an abnormal transaction identification system, and the system includes: A clustering trend calculation module, which is used to obtain the features of each transaction, calculate the distances between the same features of different transactions to obtain a distance matrix of features. For any two points in the distance matrix, obtain the direct path and all non-direct paths connecting the two points, and update the distance matrix based on the direct path and all non-direct paths; rearrange the updated distance matrix to obtain a reordered dissimilarity image of the features; An abnormal transaction identification module, which is used to extract the dark blocks in the reordered dissimilarity image of the features, obtain the number of transaction clusters according to the number of dark blocks in the reordered dissimilarity image of the features and the importance of the features, re-identify the dark blocks of the reordered dissimilarity image of each feature according to the number of transaction clusters, and obtain the clusters of transactions by using the re-identified dark blocks; judge whether the transaction to be identified is abnormal according to the clusters of transactions.
[0006] Finally, the present invention provides an abnormal transaction identification program, and the identification program includes: A clustering trend calculation module, which is used to obtain the features of each transaction, calculate the distances between the same features of different transactions to obtain a distance matrix of features. For any two points in the distance matrix, obtain the direct path and all non-direct paths connecting the two points, and update the distance matrix based on the direct path and all non-direct paths; rearrange the updated distance matrix to obtain a reordered dissimilarity image of the features; An abnormal transaction identification module, which is used to extract the dark blocks in the reordered dissimilarity image of the features, obtain the number of transaction clusters according to the number of dark blocks in the reordered dissimilarity image of the features and the importance of the features, re-identify the dark blocks of the reordered dissimilarity image of each feature according to the number of transaction clusters, and obtain the clusters of transactions by using the re-identified dark blocks; judge whether the transaction to be identified is abnormal according to the clusters of transactions.
[0007] Preferably, the updating of the distance matrix based on the direct path and all non-direct paths is specifically as follows: Obtain the maximum distance between two adjacent nodes in each non-direct path, calculate the minimum value among the maximum distances of all non-direct paths. If the distance of the direct path is greater than the minimum value, replace the distance between the two points in the distance matrix with the minimum value.
[0008] Preferably, the updating of the distance matrix based on the direct path and all non-direct paths is specifically as follows: Calculate the length of each non-direct path, and calculate the minimum value or average value of the lengths of all non-direct paths whose distances are less than the distance of the direct path. If the distance of the direct path is greater than the minimum value or the average value, replace the distance between the two points in the distance matrix with the minimum value or the average value.
[0009] Preferably, the extraction of the dark blocks in the reordered dissimilarity image is specifically as follows: Obtain the average value of the re-ordered dissimilar images, obtain a threshold based on the average value, move the center point along the diagonal of the re-ordered dissimilar images by moving one pixel at a time, each center point corresponds to a square, the square is centered on the center point, the average value of the square in the re-ordered dissimilar images is less than the threshold, and the area of the square is the largest; If there is an intersection between the squares of two adjacent center points on the diagonal, merge the squares of the two center points into one square, and iterate continuously until there is no intersection between the squares of all adjacent center points on the diagonal; Take the square as the dark block of the re-ordered dissimilar image.
[0010] Preferably, obtain the number of trading clusters according to the number of dark blocks in the re-ordered dissimilar image of the feature and the importance of the feature, specifically: Calculate the average value of the product of the importance of the feature and the number of dark blocks corresponding to the feature, and round the average value to obtain the number of trading clusters.
[0011] Preferably, re-identify the dark blocks of the re-ordered dissimilar image of each feature according to the number of trading clusters, specifically: If the number of dark blocks in the re-ordered dissimilar image is greater than the number of trading clusters, merge the two closest dark blocks, and repeat continuously until the number of dark blocks in the re-ordered dissimilar image is equal to the number of trading clusters; If the number of dark blocks in the re-ordered dissimilar image is less than the number of trading clusters, split the dark block with the largest area or the largest average value into two dark blocks, and repeat continuously until the number of dark blocks in the re-ordered dissimilar image is equal to the number of trading clusters.
[0012] Preferably, obtain the clusters of transactions by using the re-identified dark blocks, specifically: Take the dark blocks of the re-ordered dissimilar images with the same feature as a cluster of the feature; Obtain all cluster combinations composed of different clusters of different features, the cluster combination contains one cluster of each feature; obtain the intersection of the clusters in the cluster combination, and establish the corresponding relationship between the cluster combination and the intersection; Sort the combinations according to the number of transactions in the intersection, add the first cluster combination to the combination set, judge whether the next cluster combination is the same as a certain cluster in the combination set, if the same, skip it, otherwise, add the next cluster combination to the combination set, and repeat continuously until there are the number of trading clusters of cluster combinations in the combination set; Obtain the intersection corresponding to the cluster combination in the combination set, and obtain the cluster of transactions according to the intersection.
[0013] Preferably, judge whether the transaction to be identified is abnormal according to the cluster of transactions, specifically: If the transaction to be recognized belongs to a cluster of transactions and the cluster of transactions to which it belongs is an abnormal cluster, or the transaction to be recognized is a discrete point, then the transaction to be recognized is determined to be abnormal.
[0014] Regarding the problem that the boundaries between clusters are not obvious in abnormal transaction recognition, resulting in recognition errors, the present invention calculates the distance matrix between transaction features and updates it using direct paths and indirect paths, making the cluster division of features more accurate, especially for long-chain clusters. In addition, the number of transaction clusters is estimated by combining feature importance and the number of feature clusters, and the clusters are adjusted by re-identifying dark blocks, optimizing the clustering effect and improving the automation and accuracy of clustering analysis. Brief Description of the Drawings
[0015] Figure 1 Is a flowchart of the first embodiment; Figure 2 Is a visualization result diagram of the distance matrix; Figure 3 Is a schematic diagram of dark block merging; Figure 4 Is a flowchart of obtaining the cluster of transactions using the re-identified dark blocks; Figure 5 Is a schematic diagram of a feature clustering. Detailed Embodiment
[0016] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0018] Figure 1 The first embodiment of the present invention is shown. The first embodiment provides a method for recognizing abnormal transactions, and the method includes: S1. Obtain the features of each transaction, calculate the distances between the same features of different transactions to obtain a distance matrix of features. For any two points in the distance matrix, obtain the direct path connecting the two points and all non-direct paths, and update the distance matrix based on the direct path and all non-direct paths; Rearrange the updated distance matrix to obtain a feature re-ordered dissimilarity image; In transaction data, each transaction has multiple features (attributes), such as transaction amount, transaction time, transaction type, user behavior, etc. Extract these features from each transaction to form a data set, and each transaction corresponds to a set of feature values. For example, for a transaction data set, assume that each transaction contains three features: transaction amount, transaction time, and transaction frequency. Then each transaction will have these three features. For the sake of illustration, the present invention will be described below using these three features, but it should be noted that the features of the transaction are not limited to these three features.
[0019] For each pair of transactions in the transaction data set, calculate the distances on the same features to obtain a distance matrix, the visualization of which is as Figure 2 shown. The distance matrix records the distances or dissimilarities between the same features of all transactions. For example, if there are 5 transactions and each transaction has 3 features, the size of the distance matrix is 5×5, where each element D(i, j) represents the distance between the i-th transaction and the j-th transaction on the same feature. Since there are 3 features, 3 distance matrices will be obtained. The first distance matrix is the distance matrix of the transaction amounts of different transactions in the transaction data set, the second distance matrix is the distance matrix of the transaction times of different transactions in the transaction data set, and the third distance matrix is the distance matrix of the transaction frequencies of different transactions in the transaction data set.
[0020] For each pair of transaction points in the distance matrix, that is, each pair of elements in the distance matrix, in addition to directly calculating the distance between the two points, i.e., the direct path, non-direct paths also need to be considered. The non-direct path refers to connecting these two points through other points; the direct path is that the two points are directly connected without passing through other points. The directly connected distance between the first transaction and the second transaction in the distance matrix of transaction frequency, that is, the corresponding value between the two in the distance matrix, is the direct distance. If the first transaction connects the third transaction and the third transaction connects the second transaction, this path is a non-direct path; the length of the non-direct path can be calculated from the direct path: D(T1, T3, T2) = D(T1, T3) + D(T3, T2), where D(T1, T3) is the direct path length or distance from transaction 1 to transaction 3.
[0021] If, under the same feature, the direct distance between two transactions is very far, but the distance connecting them through other transactions is very close, the actual distance between the two should be very close, and they should also be assigned to the same cluster. Therefore, after calculating the direct path and all non-direct paths, update the original distance matrix based on this path information. By rearranging the updated distance matrix, that is, by reordering the matrix, similar data points are arranged together, and dissimilar data points are arranged farther apart. The rearranged distance matrix is the rearranged dissimilarity image. Among them, the rearrangement method preferably uses the rearrangement method in iVAT or VAT clustering for rearrangement. In the rearranged dissimilarity image, the pixel points represent the distance between two transactions. The smaller the distance, the smaller the value of the pixel points in the rearranged dissimilarity image, and the darker it is. Figure 2 Shows the rearranged dissimilarity graph of a feature.
[0022] S2. Extract the dark blocks in the rearranged dissimilarity image. Obtain the number of transaction clusters according to the number of dark blocks in the rearranged dissimilarity image of the feature and the importance of the feature. Re-identify the dark blocks of the rearranged dissimilarity image of each feature according to the number of transaction clusters, and obtain the clusters of transactions by using the re-identified dark blocks; judge whether the transaction to be identified is abnormal according to the clusters of transactions.
[0023] In the rearranged dissimilarity image, similar data points are arranged more closely, while dissimilar data points are separated. The dark blocks in the image represent the dense regions within the clusters. Identify all the dark blocks from the rearranged dissimilarity image. Each dark block represents a potential clustering. Among them, the extraction method of the dark blocks includes but is not limited to the aVAT algorithm or image recognition, and the image recognition method includes but is not limited to threshold segmentation, deep learning, etc.
[0024] Different features have different impacts on transactions. Some features are more conducive to distinguishing whether a transaction is abnormal. For example, the transaction frequency is more helpful for distinguishing whether a transaction is abnormal than the transaction time. In one embodiment, the importance of the feature is specified manually. In another embodiment, the importance of the feature is determined by statistical methods, information gain, variance, etc. Furthermore, according to the number of dark blocks in the rearranged dissimilarity image and the importance of each feature, infer the number of transaction clusters of the data. In an alternative embodiment, the number of transaction clusters is obtained according to the average value of the number of dark blocks in the rearranged dissimilarity image of the feature.
[0025] Since the number of dark blocks of different features, that is, the number of clusters of different features may be different, and may also be different from the number of transaction clusters, it is necessary to further re-identify the dark blocks so that the number of dark blocks of the feature, that is, the number of clusters, is the same as the number of transaction clusters. For example, adjust the division of the clusters of the feature by merging or splitting the dark blocks to make it consistent with the number of transaction clusters.
[0026] The dark block of the feature, i.e., the cluster, includes different transactions, and the number of transactions and transactions in different feature clusters are different. The re-identified dark block is further used to obtain the cluster of transactions. In one embodiment, the center of the dark block of the feature with the largest weight, i.e., the cluster, is taken as the center of the cluster corresponding to the transaction, and then the transactions are clustered according to the center of the cluster of transactions. If the transaction to be identified does not belong to any transaction cluster, or is far away from the center of the cluster to which it belongs, or has a low similarity with other transaction points in the cluster, the transaction can be considered abnormal. For example, the transactions are divided into 3 clusters: Cluster 1 contains a large number of small transactions, and the transaction type is online shopping, Cluster 2 contains a small number of large transactions, and the transaction type is transfer, and Cluster 3 contains a medium number of transactions, and the transaction type is credit card consumption. There is a transaction amount of 10,000 yuan for a transaction to be identified, and the transaction type is online shopping. If this transaction does not belong to any cluster, it is considered an abnormal transaction.
[0027] A second embodiment provides an abnormal transaction identification system, the system comprising: A clustering trend calculation module is used to obtain the features of each transaction, calculate the distances between the same features of different transactions to obtain a distance matrix of the features, obtain a direct path and all indirect paths connecting the two points for any two points in the distance matrix, and update the distance matrix based on the direct path and all indirect paths; rearrange the updated distance matrix to obtain a reordered dissimilar image of the features; The abnormal transaction identification module is used to extract dark blocks in the reordered dissimilar images, obtain the number of transaction clusters according to the number of dark blocks in the reordered dissimilar images of the features and the importance of the features, re-identify the dark blocks of the reordered dissimilar images of each feature according to the number of transaction clusters, and obtain the transaction cluster using the re-identified dark blocks; and determine whether the transaction to be identified is abnormal based on the transaction cluster.
[0028] The third embodiment provides an abnormal transaction identification program, the identification program comprising: A clustering trend calculation module is used to obtain the features of each transaction, calculate the distances between the same features of different transactions to obtain a distance matrix of the features, obtain a direct path and all indirect paths connecting the two points for any two points in the distance matrix, and update the distance matrix based on the direct path and all indirect paths; rearrange the updated distance matrix to obtain a reordered dissimilar image of the features; The abnormal transaction identification module is used to extract dark blocks in the reordered dissimilar images, obtain the number of transaction clusters according to the number of dark blocks in the reordered dissimilar images of the features and the importance of the features, re-identify the dark blocks of the reordered dissimilar images of each feature according to the number of transaction clusters, and obtain the transaction cluster using the re-identified dark blocks; and determine whether the transaction to be identified is abnormal based on the transaction cluster.
[0029] There may be some slender clusters or chain-like clusters. The characteristics of such clusters are that the elements within the cluster are distributed along a certain linear path. The direct distance between points may be very far, but the path through intermediate points is relatively short. Such scattered data points are connected into a cluster, rather than being wrongly divided into different clusters because of the far direct distance. In one embodiment, the distance matrix is updated based on the direct path and all non-direct paths, specifically as follows: Obtain the maximum distance between two adjacent nodes in each non-direct path, calculate the minimum value among the maximum distances of all non-direct paths. If the distance of the direct path is greater than the minimum value, then replace the distance between the two points in the distance matrix with the minimum value.
[0030] Obtain the non-direct path between two points, and then calculate the maximum distance for each pair of points on the path. For example, a non-direct path between two points m1 and m2 is m1 - m4 - m2. Assume the distance between m1 and m4 is 1, and the distance between m4 and m2 is 2, then the maximum distance between two adjacent nodes in this non-direct path is 2. Another non-direct path between m1 and m2 is m1 - m3 - m4 - m2, and the maximum distance of this non-direct path is 4. Similarly, the maximum values {2, 4,...} for each non-direct path between m1 and m2 can be obtained. Take the minimum value among the maximum distances {2, 4,...} of all non-direct paths, assume it is 2. If the direct path distance between m1 and m2 is 8, then use 2 to replace 8 in the distance matrix.
[0031] In another alternative embodiment, the distance matrix is updated based on the direct path and all non-direct paths, specifically as follows: Calculate the length of each non-direct path, and calculate the minimum value or average value of the lengths of all non-direct paths whose distances are less than the distance of the direct path. If the distance of the direct path is greater than the minimum value or the average value, then replace the distance between the two points in the distance matrix with the minimum value or the average value.
[0032] The length of each non-direct path refers to the sum of the distances between all adjacent data points on the path. For each pair of data points mi and mj, if there is no direct connection but they can be connected through other intermediate points such as mk, then calculate the sum of the distances between all adjacent points from mi to mk and then to mj. Screen out those non-direct paths whose path lengths are less than the distance of the direct path. From all non-direct paths whose lengths are less than the direct path, take the minimum path length as the minimum value, or calculate the average value of the path lengths from all non-direct paths whose lengths are less than the direct path as the average value. If the distance of the direct path is greater than the minimum value or the average value, then update the values of points i and j in the distance matrix to the calculated minimum value or average value.
[0033] In addition, the present invention also proposes a method for extracting dark blocks from re-ordered dissimilar images. Specifically, Obtain the average value of the re-ordered dissimilar images, obtain a threshold according to the average value, move the center point along the diagonal of the re-ordered dissimilar images one pixel at a time. Each center point corresponds to a square. The square is centered on the center point, and the average value of the square in the re-ordered dissimilar images is less than the threshold, and the area of the square is the largest; If there is an intersection between the squares of two adjacent center points on the diagonal, then merge the squares of the two center points into one square, and continuously iterate until there is no intersection between the squares of all adjacent center points on the diagonal; Take the square as the dark block of the re-ordered dissimilar image.
[0034] Calculate the average value of the entire re-ordered dissimilar image. The average value represents the average similarity level or distance of the entire data set. Determine a threshold through the average value. This threshold is used to screen those regions with higher similarity, that is, dark blocks. Preferably, multiply the average value by an adjustment factor to obtain the threshold. For example, the adjustment factor is 1.1 or 1.3, etc. In the re-ordered image, the points on the diagonal are the centers of the dark blocks. Take the points on the diagonal as the center points. Each center point corresponds to a square. When moving the center point to a new position, generate a square centered on that position. The size of the square is determined according to the threshold, that is, the average value within the square needs to be less than the calculated threshold and the area of the square is the largest. When there is an intersection between the squares of two adjacent center points, it means that these two squares may belong to the same cluster. Merge the two intersecting squares into a larger square. As Figure 3 shown, in one embodiment, obtain the union of the squares of two adjacent center points, calculate the ratio of the intersection to the union. When the ratio meets a preset condition, such as being greater than a preset value, the two squares will be merged. Perform through iteration until the intersections of all adjacent squares are eliminated, that is, each cluster no longer overlaps with its adjacent clusters; the merged squares and the remaining unmerged squares are used as the dark blocks of the re-ordered dissimilar image.
[0035] A transaction includes multiple features, and the importance of each feature for determining whether it is an abnormal transaction is different, and the role is also different when clustering transactions. In one embodiment, the number of transaction clusters is obtained according to the number of dark blocks in the re-ordered dissimilar image of the feature and the importance of the feature. Specifically: Calculate the average value of the product of the importance of the feature and the number of dark blocks corresponding to the feature, and round the average value to obtain the number of transaction clusters.
[0036] For each feature, calculate the product of its importance and the number of dark blocks corresponding to the feature, and then calculate the average value of the product results for all features. This average value reflects the comprehensive influence of all features in the dataset cluster division. Round the calculated average value to obtain the estimated number of transaction clusters in the dataset.
[0037] Due to different features and different data distribution situations of features, the number of transaction clusters may be different from the number of dark blocks of features. In an alternative embodiment, the re-identification of the dark blocks of non-similar images sorted by the number of transaction clusters is specifically as follows: If the number of dark blocks of non-similar images sorted is greater than the number of transaction clusters, merge the two closest dark blocks, and repeat continuously until the number of dark blocks of non-similar images sorted is equal to the number of transaction clusters; If the number of dark blocks of non-similar images sorted is less than the number of transaction clusters, split the dark block with the largest area or the largest average value into two dark blocks, and repeat continuously until the number of dark blocks of non-similar images sorted is equal to the number of transaction clusters.
[0038] If the number of dark blocks of a feature is greater than the number of transaction clusters, choose to merge the two closest dark blocks into one cluster. The "distance" refers to the boundary distance between the two dark blocks or the distance between the centers of the two dark blocks. If two dark blocks are close to each other, then they are very likely to belong to the same cluster, and merge them into one dark block. After merging, update the image and recalculate the number of dark blocks, and then continue to merge the two closest dark blocks until the final number of dark blocks is equal to the estimated number of clusters.
[0039] If the number of dark blocks of a feature is less than the number of transaction clusters, some large clusters need to be subdivided to reach the number of transaction clusters. Dark blocks with large areas or large average values usually represent loose regions. Splitting them can better capture more sub-clusters in the data. When splitting, choose the dark block with the largest area or the largest average value for splitting. After splitting, check the new number of dark blocks, and continue to split the dark block with the largest area or the largest average value until the number of dark blocks reaches the estimated number of clusters.
[0040] If the number of dark blocks of a feature is equal to the number of transaction clusters, no further splitting or merging is performed.
[0041] In an alternative embodiment, obtaining the transaction clusters using the re-identified dark blocks is as Figure 4 shown, specifically as follows: S21, regard the dark blocks of non-similar images with the same feature as one cluster of the feature; For each feature, the dark blocks that reorder dissimilar images form a cluster, that is, each dark block is a cluster of this feature. At the same time, the transactions included in this cluster can be obtained. For example, the first dark block of the feature "transaction amount", that is, the first cluster, includes transaction 1 and transaction 3. Then, transaction 1 and transaction 3 are the transactions in the first cluster.
[0042] S22. Obtain all cluster combinations composed of different clusters of different features, where the cluster combination contains one cluster of each feature; obtain the intersections of the clusters in the cluster combination, and establish the corresponding relationship between the cluster combination and the intersection. In the way of permutation and combination, combine different clusters of different features. For example, there are three features, and each feature has three clusters. The clusters of one feature are as Figure 5 shown, then 27 combination methods will be obtained. An exemplary combination method is A1, B1, C3. Among them, A1 is the first cluster of the first feature, B1 is the first cluster of the second feature, and C3 is the third cluster of the third feature. For each combination, calculate the intersection of the clusters in the combination. For example, A1 includes transaction 1, B1 includes transaction 1 and transaction 3, and C3 includes transaction 1 and transaction 4. Then the intersection between different clusters of different features is transaction 1.
[0043] Establish the corresponding relationship between the combination and the intersection. The above example obtains the corresponding relationship as {(A1, B1, C3), (transaction 1)}.
[0044] S23. Sort the combinations according to the number of transactions in the intersection, add the first cluster combination to the combination set, and judge whether the next cluster combination is the same as a certain cluster in the combination set. If they are the same, skip it. Otherwise, add the next cluster combination to the combination set and repeat continuously until there are as many cluster combinations as the number of transaction clusters in the combination set. Sort these cluster combinations according to the number of transactions in the intersection, and the cluster combination with more transactions is preferred. Add the first sorted cluster combination to the combination set, select the next cluster combination. If the next cluster combination has an intersection with any combination in the combination set, skip it. For example, if the next cluster combination is A1, B2, C1, and any combination in the combination set includes A1 or B2 or C1, then skip the next cluster combination. If the next cluster combination has no intersection with any cluster combination in the combination set, add the next cluster combination to the combination set. Repeat this process until the number of cluster combinations in the combination set is equal to the number of transaction clusters.
[0045] S24. Obtain the intersections corresponding to the cluster combinations in the combination set, and obtain the transaction clusters according to the intersections.
[0046] Determine the transactions in each cluster combination. Since the correspondence between the cluster combination and the intersection is established in step S22, the intersection corresponding to the cluster combination can be directly obtained, and each intersection is used as a cluster of the transaction.
[0047] In an alternative embodiment, the determination of whether the transaction to be recognized is abnormal based on the clusters of the transaction is specifically as follows: If the transaction to be recognized belongs to a cluster of the transaction and the cluster of the transaction to which it belongs is an abnormal cluster, or the transaction to be recognized is a discrete point, then the transaction to be recognized is determined to be abnormal.
[0048] If the transaction to be recognized belongs to a known abnormal cluster, or the transaction itself is a discrete point, then the transaction to be recognized is abnormal. For example, there are three types of transaction clusters, namely high-frequency transaction clusters, low-frequency transaction clusters, and abnormal clusters. If a transaction to be recognized belongs to an abnormal cluster, or it is an isolated point and does not belong to any cluster, then this transaction is an abnormal transaction.
[0049] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0050] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Other embodiments can also be adopted; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying abnormal transactions, characterized in that: The method comprises: Acquire the features of each transaction, calculate the distances between the same features of different transactions to obtain a distance matrix of the features, obtain a direct path and all indirect paths connecting the two points for any two points in the distance matrix, and update the distance matrix based on the direct path and all indirect paths; rearrange the updated distance matrix to obtain a reordered dissimilar image of the features; The dark blocks in the reordered dissimilar images are extracted, and the number of transaction clusters is obtained according to the number of dark blocks in the reordered dissimilar images of the features and the importance of the features. The dark blocks of the reordered dissimilar images of each feature are re-identified according to the number of transaction clusters, and the clusters of the transactions are obtained using the re-identified dark blocks. Whether the transaction to be identified is abnormal is determined based on the clusters of the transactions.
2. The method according to claim 1, characterized in that The updating of the distance matrix based on the direct path and all the indirect paths is specifically as follows: The maximum distance between two adjacent nodes in each indirect path is obtained, and the minimum value among the maximum distances of all indirect paths is calculated. If the distance of the direct path is greater than the minimum value, the distance between the two points in the distance matrix is replaced by the minimum value.
3. The method according to claim 1, characterized in that The updating of the distance matrix based on the direct path and all the indirect paths is specifically as follows: Calculate the length of each indirect path, and calculate the minimum or average value of the lengths of all indirect paths whose distances are less than the direct path; if the distance of the direct path is greater than the minimum or average value, replace the distance between the two points in the distance matrix with the minimum or average value.
4. The method according to claim 1, characterized in that The extraction and reordering of dark blocks in dissimilar images is specifically as follows: Obtaining an average value of the reordered dissimilar images, obtaining a threshold value according to the average value, and moving the center point along the diagonal of the reordered dissimilar images by one pixel at a time, wherein each center point corresponds to a square, the square is centered on the center point, the average value of the square in the reordered dissimilar images is less than the threshold value, and the area of the square is the largest; If two blocks of adjacent center points on the diagonal line have an intersection, the blocks of the two center points are merged into one block, and the process is repeated until the blocks of all adjacent center points on the diagonal line have no intersection; Treat the squares as dark patches in reordered dissimilar images.
5. The method according to claim 1, characterized in that The number of transaction clusters is obtained by reordering the features according to the number of dark blocks in the dissimilar images and the importance of the features, specifically: The average value of the product of the importance of the feature and the number of dark blocks corresponding to the feature is calculated, and the average value is rounded to the integer value as the number of transaction clusters.
6. The method according to claim 1, characterized in that The dark blocks of the reordered dissimilar images of each feature are re-identified according to the number of transaction clusters, specifically: If the number of dark blocks in the reordered dissimilar images is greater than the number of transaction clusters, the two dark blocks closest to each other are merged, and the process is repeated until the number of dark blocks in the reordered dissimilar images is equal to the number of transaction clusters; If the number of dark blocks in the reordered dissimilar images is less than the number of transaction clusters, the dark block with the largest area or the largest average value is split into two dark blocks, and the process is repeated until the number of dark blocks in the reordered dissimilar images is equal to the number of transaction clusters.
7. The method according to claim 1, characterized in that The cluster of transactions obtained by using the re-identified dark blocks is specifically: The dark blocks of reordered dissimilar images with the same features are regarded as a cluster of features; Obtain all cluster combinations consisting of different clusters of different features, wherein the cluster combination contains a cluster for each feature; obtain the intersection of clusters in the cluster combination, and establish a corresponding relationship between the cluster combination and the intersection; Sort the combinations according to the number of transactions in the intersection, add the first cluster combination to the combination set, and determine whether the next cluster combination is the same as a cluster in the combination set. If so, skip it. Otherwise, add the next cluster combination to the combination set, and repeat until there are cluster combinations equal to the number of transaction clusters in the combination set. The intersection corresponding to the cluster combination in the combination set is obtained, and the transaction cluster is obtained according to the intersection.
8. The method according to claim 1, characterized in that The step of judging whether the transaction to be identified is abnormal based on the transaction cluster is as follows: If the transaction to be identified belongs to a cluster of transactions and the cluster of transactions to which it belongs is an abnormal cluster, or the transaction to be identified is a discrete point, the transaction to be identified is judged to be abnormal.
9. An abnormal transaction identification system, characterized in that: The system comprises: A clustering trend calculation module is used to obtain the features of each transaction, calculate the distances between the same features of different transactions to obtain a distance matrix of the features, obtain a direct path and all indirect paths connecting the two points for any two points in the distance matrix, and update the distance matrix based on the direct path and all indirect paths; rearrange the updated distance matrix to obtain a reordered dissimilar image of the features; The abnormal transaction identification module is used to extract dark blocks in the reordered dissimilar images, obtain the number of transaction clusters according to the number of dark blocks in the reordered dissimilar images of the features and the importance of the features, re-identify the dark blocks of the reordered dissimilar images of each feature according to the number of transaction clusters, and obtain the transaction cluster using the re-identified dark blocks; and determine whether the transaction to be identified is abnormal based on the transaction cluster.
10. An abnormal transaction identification program, characterized in that: The identification procedure includes: A clustering trend calculation module is used to obtain the features of each transaction, calculate the distances between the same features of different transactions to obtain a distance matrix of the features, obtain a direct path and all indirect paths connecting the two points for any two points in the distance matrix, and update the distance matrix based on the direct path and all indirect paths; rearrange the updated distance matrix to obtain a reordered dissimilar image of the features; The abnormal transaction identification module is used to extract dark blocks in the reordered dissimilar images, obtain the number of transaction clusters according to the number of dark blocks in the reordered dissimilar images of the features and the importance of the features, re-identify the dark blocks of the reordered dissimilar images of each feature according to the number of transaction clusters, and obtain the transaction cluster using the re-identified dark blocks; and determine whether the transaction to be identified is abnormal based on the transaction cluster.