Commercial tenant information anomaly identification method and device

By clustering and similarity distance calculation of the time series data of merchant nodes, abnormal merchants are identified, and the problem of low recognition accuracy in the existing technology is solved, and more efficient abnormal merchant identification is achieved.

CN120047163APending Publication Date: 2025-05-27CHINA UNIONPAY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510072527.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When identifying abnormal merchants, the prior art lacks targeted mining of merchant characteristics, resulting in a low recognition accuracy rate.

Method used

By obtaining the time series data of the sample merchant nodes, the similarity distance between each two sample merchant nodes is calculated, and the merchant nodes are clustered to determine the neighborhood of each merchant node. Then, based on the number of merchant nodes in the neighborhood and whether they are located in the neighborhood of the core merchant node, it is determined whether the merchant node is an abnormal merchant.

Benefits of technology

It improves the accuracy of abnormal merchant identification, can efficiently and quickly locate abnormal merchants, and fully considers the inherent characteristics of the merchant industry to which the merchant belongs and the merchant clustering characteristics of the same industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047163A_ABST
    Figure CN120047163A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a merchant information anomaly identification method and device, which can be applied to the technical field of computers, in the method, respective time sequence data of N sample merchant nodes are acquired, and N is greater than 1; obtaining a similarity distance between every two sample merchant nodes based on the respective time sequence data of every two sample merchant nodes in the N sample merchant nodes; based on the similarity distance between every two sample merchant nodes, clustering the N sample merchant nodes to obtain a neighborhood of each sample merchant node and other sample merchant nodes located in the neighborhood; for each sample merchant node, if the quantity of the merchant nodes in the neighborhood of the sample merchant node is smaller than a first preset threshold value and the sample merchant node is not located in the neighborhood of the core merchant node, the sample merchant node is an abnormal merchant, and the quantity of the merchant nodes in the neighborhood corresponding to the core merchant node is not smaller than the first preset threshold value; and the accuracy of abnormal merchant identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technologies, and in particular, to a method and device for abnormal identification of merchant information. Background Art

[0002] In the process of analyzing the changing trends of industry consumption, it is necessary to pre-eliminate abnormal merchants that do not belong to this industry. In related technologies, merchants or non-genuine merchants that do not belong to this industry, such as merchants engaged in code swapping and cash laundering, are identified based on the transaction behaviors of merchants. Specifically, taking the merchant characteristics of all merchants as input, machine learning algorithms are used to determine abnormal merchants.

[0003] However, this method starts from all merchants to find abnormal characteristics, lacking pertinence in the excavation of merchant characteristics, resulting in the difficulty of comprehensively representing the characteristics of merchants by merchant characteristics, thereby reducing the accuracy of abnormal merchant identification. Summary of the Invention

[0004] The embodiments of the present invention provide a method and device for abnormal identification of merchant information, which are used to identify abnormal merchant information based on time series data, and improve the accuracy of abnormal merchant identification.

[0005] On the one hand, the embodiments of the present application provide a method for abnormal identification of merchant information, and the method includes:

[0006] Obtain the time series data of each of the N sample merchant nodes, where N is greater than 1;

[0007] For every two of the N sample merchant nodes, based on the time series data of the two sample merchant nodes, obtain the similarity distance between the two sample merchant nodes;

[0008] Cluster the N sample merchant nodes based on the similarity distances between every two of the N sample merchant nodes, and obtain the neighborhood of each sample merchant node and other sample merchant nodes located in the neighborhood;

[0009] For each sample merchant node, if the number of merchant nodes in the neighborhood of the sample merchant node is less than a first preset threshold, and the sample merchant node is not located in the neighborhood of a core merchant node, then the sample merchant node is an abnormal merchant, where the number of merchant nodes in the neighborhood corresponding to the core merchant node is not less than the first preset threshold.

[0010] Optionally, the obtaining the time series data of each of the N sample merchant nodes includes:

[0011] For each sample merchant node, obtain the original data of the sample merchant node in M index dimensions respectively, where M is greater than 1;

[0012] Slice the original data of each metric dimension from L time dimensions to obtain the metric time series of each said metric dimension in the L time dimensions, where L is greater than 1;

[0013] Based on the metric time series of each said metric dimension in the L time dimensions obtained, determine the time series data of the sample merchant node.

[0014] Optionally, the obtaining the similarity distance between the two sample merchant nodes based on the time series data of the two sample merchant nodes respectively includes:

[0015] For each sample merchant node, normalize the time series data of the sample merchant node to obtain the normalized time series data of the sample merchant node;

[0016] Based on the target change value between every two time points in the normalized time series data, convert the normalized time series data into time series change data;

[0017] Based on the time series change data of the two sample merchant nodes respectively, calculate the similarity distance between the two sample merchant nodes.

[0018] Optionally, before converting the normalized time series data into time series change data based on the target change value between every two time points in the normalized time series data, it further includes:

[0019] Set multiple value ranges, and each value range corresponds to a fixed threshold;

[0020] For every two time points in the normalized time series data, when the metric change value of the two time points is within a target range among the multiple value ranges, use the fixed threshold corresponding to the target range as the target change value between the two time points.

[0021] Optionally, the normalized time series data of the sample merchant node includes: sub-time series data of multiple dimensions; the calculating the similarity distance between the two sample merchant nodes based on the time series change data of the two sample merchant nodes respectively includes:

[0022] For the sub-time series data of the multiple dimensions, perform the following operations respectively: Based on the sub-time series data of the two sample merchant nodes in one dimension, obtain the sub-similarity distance of the two sample merchant nodes in the one dimension;

[0023] Based on the obtained sub-similarity distances of multiple dimensions, obtain the similarity distance between the two sample merchant nodes.

[0024] Optionally, clustering the N sample merchant nodes based on the similarity distances between every two of the N sample merchant nodes to obtain the neighborhood of each sample merchant node and other sample merchant nodes located in the neighborhood, including:

[0025] For each sample merchant node, taking the sample merchant node as the center and the second preset threshold as the radius to obtain the neighborhood of the sample merchant node;

[0026] For each other sample merchant node, when the similarity distance between the other sample merchant node and the sample merchant node is less than or equal to the second preset threshold, the other sample merchant node is located in the neighborhood of the sample merchant node.

[0027] Optionally, it further includes:

[0028] If the number of merchant nodes in the neighborhood of the sample merchant node is less than the first preset threshold and the sample merchant node is located in the neighborhood of the core merchant node, then the sample merchant node is a normal merchant.

[0029] On the one hand, an embodiment of the present application provides an abnormal identification device for merchant information, and the device includes:

[0030] An acquisition module, configured to acquire the time series data of each of the N sample merchant nodes, where N is greater than 1;

[0031] A calculation module, configured to obtain the similarity distance between every two of the N sample merchant nodes based on the time series data of the two sample merchant nodes;

[0032] A clustering module, configured to cluster the N sample merchant nodes based on the similarity distances between every two of the N sample merchant nodes to obtain the neighborhood of each sample merchant node and other sample merchant nodes located in the neighborhood;

[0033] For each sample merchant node, if the number of merchant nodes in the neighborhood of the sample merchant node is less than the first preset threshold and the sample merchant node is not located in the neighborhood of the core merchant node, then the sample merchant node is an abnormal merchant, where the number of merchant nodes in the neighborhood corresponding to the core merchant node is not less than the first preset threshold.

[0034] Optionally, the acquisition module is specifically configured to:

[0035] For each sample merchant node, acquire the original data of the sample merchant node in M index dimensions respectively, where M is greater than 1;

[0036] Slice the original data of each metric dimension from L time dimensions to obtain the metric time series of each said metric dimension in the L time dimensions, where L is greater than 1;

[0037] Based on the metric time series of each said metric dimension in the L time dimensions obtained, determine the time series data of the sample merchant node.

[0038] Optionally, the calculation module is specifically configured to:

[0039] For each sample merchant node, perform normalization processing on the time series data of the sample merchant node to obtain the normalized time series data of the sample merchant node;

[0040] Based on the target change value between every two time points in the normalized time series data, convert the normalized time series data into time series change data;

[0041] Based on the time series change data of the two sample merchant nodes respectively, calculate the similarity distance between the two sample merchant nodes.

[0042] Optionally, the calculation module is further configured to:

[0043] Set multiple value ranges, and each value range corresponds to a fixed threshold;

[0044] For every two time points in the normalized time series data, when the metric change values of the two time points are within a target range among the multiple value ranges, use the fixed threshold corresponding to the target range as the target change value between the two time points.

[0045] Optionally, the normalized time series data of the sample merchant node includes: sub-time series data of multiple dimensions, and the calculation module is specifically configured to:

[0046] For the sub-time series data of the multiple dimensions, perform the following operations respectively: Based on the sub-time series data of one dimension of the two sample merchant nodes respectively, obtain the sub-similarity distance of the two sample merchant nodes in the one dimension;

[0047] Based on the obtained sub-similarity distances of multiple dimensions, obtain the similarity distance between the two sample merchant nodes.

[0048] Optionally, the clustering module is specifically configured to:

[0049] For each sample merchant node, with the sample merchant node as the center and the second preset threshold as the radius, obtain the neighborhood of the sample merchant node;

[0050] For each other sample merchant node, when the similarity distance between the other sample merchant node and the sample merchant node is less than or equal to the second preset threshold, the other sample merchant node is within the neighborhood of the sample merchant node.

[0051] Optionally, the clustering module is further configured to:

[0052] If the number of merchant nodes in the neighborhood of the sample merchant node is less than the first preset threshold and the sample merchant node is within the neighborhood of the core merchant node, then the sample merchant node is a normal merchant.

[0053] On the one hand, an embodiment of the present application provides a computer device, including:

[0054] A memory for storing program instructions;

[0055] A processor for calling the program instructions stored in the memory and executing the steps of the above-mentioned abnormal identification method of merchant information according to the obtained program.

[0056] On the one hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program executable by a computer device. When the program runs on the computer device, the computer is caused to execute the steps of the above-mentioned abnormal identification method of merchant information.

[0057] On the one hand, an embodiment of the present application provides a computer program product, including a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the above-mentioned abnormal identification method of merchant information.

[0058] In the embodiment of the present application, based on the time series data of each sample merchant node among N sample merchant nodes, the similarity distance between every two sample merchant nodes is calculated. Based on the similarity distances between pairs of the N sample merchant nodes, each sample merchant node is clustered to obtain the neighborhood of each sample merchant node and other sample merchant nodes located within the neighborhood. For each sample merchant node, if the number of merchant nodes in the neighborhood of the sample merchant node is less than the first preset threshold and the sample merchant node is not within the neighborhood of the core merchant node, then the sample merchant node is determined to be an abnormal merchant. The method of the embodiment of the present application fully considers the inherent characteristics of the industry to which the merchant belongs and the merchant aggregation characteristics of the same industry, strengthens the pertinence of merchant information mining, and can efficiently and quickly locate abnormal merchants by clustering the sample merchant nodes and determining abnormal merchants within the scope of the neighborhood, while improving the accuracy of abnormal merchant identification. Description of the Drawings

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0060] Figure 1 It is a schematic structural diagram of a system architecture provided by an embodiment of the present application;

[0061] Figure 2 It is a schematic flowchart of a method for abnormal identification of merchant information provided by an embodiment of the present application;

[0062] Figure 3 It is a schematic flowchart of a method for abnormal identification of merchant information provided by an embodiment of the present application;

[0063] Figure 4 It is a schematic structural diagram of a device for abnormal identification of merchant information provided by an embodiment of the present application;

[0064] Figure 5 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0065] In order to make the purpose, technical solutions and beneficial effects of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0066] The terms "first", "second", etc. in the description, claims and above-mentioned accompanying drawings of the present application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.

[0067] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware or / and software code that can perform functions related to the element.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0069] The following briefly introduces the system architecture diagram applicable to the technical solutions of the embodiments of the present application. It should be noted that the following introduced processes are only used to illustrate the embodiments of the present application rather than to limit them.

[0070] Refer to Figure 1 , which is a system architecture diagram applicable to the embodiments of the present application. The system architecture at least includes a terminal device 101 and a server 102. The number of terminal devices 101 can be one or more, and the number of servers 102 can also be one or more. The present application does not specifically limit the number of terminal devices 101 and servers 102.

[0071] The terminal device 101 is pre-installed with an application for identifying abnormal merchants. This application can be a client application, a web version application, a mini-program application, etc. The terminal device 101 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart home appliance, a smart voice interaction device, a smart vehicle-mounted device, etc., but is not limited thereto.

[0072] The server 102 is the background server of the application. The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, but is not limited thereto.

[0073] It should be noted that the method in the embodiments of the present application can be executed independently by the terminal device 101 or the server 102, or jointly executed by the terminal device 101 and the server 102.

[0074] In the embodiments of the present application, the terminal device 101 and the server 102 can be directly or indirectly communicatively connected through one or more networks. The network can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network, or it can be a Wireless-Fidelity (WIFI) network. Of course, it can also be other possible networks, and the embodiments of the present application do not limit this.

[0075] Based on the following Figure 1 system architecture diagram shown, an embodiment of the present application provides a process for an abnormal recognition method of merchant information. The process of this method can be executed by the Figure 1 terminal device 101 shown, or can be executed by the server 102, or can also be executed by the interaction between the terminal device 101 and the server 102. As Figure 2 shown, it includes the following steps:

[0076] Step 201, obtain the time series data of each of the N sample merchant nodes, where N is greater than 1.

[0077] Specifically, the time series data of the sample merchant node refers to the data of the merchant's transaction indicators changing in chronological order.

[0078] In some embodiments, for each sample merchant node, obtain the original data of the sample merchant node in M metric dimensions respectively, where M is greater than 1; slice the original data of each metric dimension from L time dimensions to obtain the metric time series of each metric dimension in L time dimensions respectively, where L is greater than 1; based on the obtained metric time series of each metric dimension in L time dimensions respectively, determine the time series data of the sample merchant node.

[0079] Specifically, the M metric dimensions include: the number of transactions, the transaction amount, the number of membership cards, the number of users, etc.; the L time dimensions include: hour, date, week, month, etc. Slice the metrics of each metric dimension according to the L time dimensions respectively, so that each metric corresponds to L metric time series respectively. The time series data of each sample merchant node includes M * L metric time series.

[0080] Specifically, for a combination of one time dimension and one metric dimension. First, obtain multiple time slices corresponding to this time dimension. For example, in the hour dimension, obtain 24 time slices such as 0 - 1 hour, 1 - 2 hours, …, 23 - 24 hours. Then count the value of one metric dimension in each time slice. This value can be the mean value of the metric dimension in the past period of time in this time slice. For example, the value of the number of transactions in the 0 - 1 hour in the hour dimension is the mean value of the number of transactions collected every day in the past month at 0 - 1 hour.

[0081] For example, for the sample merchant node A, set M = 2, L = 4. The 2 metric dimensions include: the number of transactions, the number of users; the 4 time dimensions include: hour, date, week, month.

[0082] Taking the number of transactions as the metric dimension, for the hourly dimension, first obtain 24 time slices from 0-1 hour, 1-2 hour, …, 23-24 hour, and then obtain the number of transactions within each time slice. Based on the number of transactions of each of the 24 time slices, obtain the metric time series of the number of transactions in the hourly dimension.

[0083] For the weekly dimension, first obtain 7 time slices from Monday, Tuesday, …, Sunday, and obtain the number of transactions within each time slice. Based on the number of transactions of each of the 7 time slices, obtain the metric time series of the number of transactions in the weekly dimension.

[0084] For the date dimension, first obtain 30 time slices from the first day, the second day, …, the 30th day, and obtain the number of transactions within each time slice. Based on the number of transactions of each of the 30 time slices, obtain the metric time series of the number of transactions in the date dimension.

[0085] For the monthly dimension, first obtain 12 time slices from January, February, …, December, and obtain the number of transactions within each time slice. Based on the number of transactions of each of the 12 time slices, obtain the metric time series of the number of transactions in the monthly dimension.

[0086] The method of obtaining the 4 metric time series corresponding to the number of users is the same as the method of obtaining the 4 metric time series corresponding to the number of transactions, which will not be elaborated here. The 4 metric time series corresponding to the number of users and the 4 metric time series corresponding to the number of transactions are the time series data of the sample merchant node A.

[0087] In the embodiments of the present application, the metrics of each sample merchant node are sliced according to different time dimensions to determine the time series data of the sample merchant node, fully considering the off-peak and peak trading periods of different industries, which is beneficial to improving the accuracy of identifying abnormal merchants in the industry.

[0088] Step 202, for every two sample merchant nodes among the N sample merchant nodes, based on the time series data of the two sample merchant nodes respectively, obtain the similarity distance between the two sample merchant nodes.

[0089] In some embodiments, for each sample merchant node, normalize the time series data of the sample merchant node to obtain the normalized time series data of the sample merchant node; based on the target change value between every two time points in the normalized time series data, convert the normalized time series data into time series change data; based on the time series change data of the two sample merchant nodes respectively, calculate the similarity distance between the two sample merchant nodes.

[0090] Specifically, the normalized time series data of the sample merchant node includes multiple sub-time series data.

[0091] For each sample merchant node, normalize the M*L indicator time series of the sample merchant node to eliminate the influence of dimension, and obtain M*L sub-time series data of the sample merchant node. Each sub-time series data is obtained by the following formula (1):

[0092]

[0093] where, x i represents the i-th data of any one of the M*L indicator time series; x i ′ represents x i after normalization; x min represents the minimum value data of any one of the M*L indicator time series; x max represents the maximum value data of any one of the M*L indicator time series.

[0094] In some embodiments, set multiple value ranges, and each value range corresponds to a fixed threshold; for every two time points in the normalized time series data, when the indicator change value between the two time points is within a target range among the multiple value ranges, use the fixed threshold corresponding to the target range as the target change value between the two time points.

[0095] Specifically, for the sub-time series data of each sample merchant data, determine the indicator change value or slope value between each time point and the next time point in the sub-time series data. If the indicator change value or slope value falls within a target range among the multiple pre-set value ranges, use the fixed threshold corresponding to the target range as the target change value or target slope value between this time point and the next time point. In this way, convert the n-dimensional sub-time series data of the sample merchant node into (n - 1)-dimensional sub-time series data, and based on the M*L (n - 1)-dimensional sub-time series data of the sample merchant node, obtain the time series change data of the sample merchant node.

[0096] For example, taking the indicator dimension of the number of transactions of the sample merchant node as an example, set 5 value ranges, which are [-1, -0.2), [-0.2, 0), 0, (0, 0.2], (0.2, 1] respectively. Among them, the change pattern corresponding to the value range [-1, -0.2) is a large decrease, and the corresponding fixed threshold is -2; the change pattern corresponding to the value range [-0.2, 0) is a small decrease, and the corresponding fixed threshold is -1; the value 0 corresponds to the change pattern of no change, and the corresponding fixed threshold is 0; the change pattern corresponding to the value range (0, 0.2] is a small increase, and the corresponding fixed threshold is 1; the change pattern corresponding to the value range (0.2, 1] is a large increase, and the corresponding fixed threshold is 2.

[0097] If the change value of the number of transactions between two time points in the sub-time series data of the sample merchant node is -0.4, then -2 is used as the target change value of the number of transactions between the two time points.

[0098] In the embodiments of the present application, the n-dimensional sub-time series data of the sample merchant node is converted into (n - 1)-dimensional sub-time series data, reducing the computational complexity and the probability of errors, and further improving the accuracy of abnormal merchant identification.

[0099] In some embodiments, for the sub-time series data of multiple dimensions, the following operations are respectively performed: based on the sub-time series data of each of two sample merchant nodes in one dimension, obtain the sub-similarity distance between the two sample merchant nodes in one dimension; based on the obtained sub-similarity distances of multiple dimensions, obtain the similarity distance between the two sample merchant nodes.

[0100] Specifically, each of the above dimensions is a combination of an index dimension and a time dimension. For example, one dimension is composed of the transaction volume dimension and the week dimension. The sub-time series data of the sample merchant node in this dimension is: the change value of the transaction volume of the sample merchant node during the period from Monday to Sunday. The specific calculation formula for calculating the sub-similarity distance of two sub-time series data is shown in the following formula (2):

[0101]

[0102] where d represents the sub-similarity distance between two sub-time series data of two sample merchant nodes in one dimension (including the index dimension and the time dimension); n represents the number of data in one sub-time series data; k represents the k-th data in one sub-time series data; x ak represents the k-th data in the sub-time series data of the sample merchant node; x bk represents the k-th data in the sub-time series data of the sample merchant node.

[0103] Based on the sub-similarity distances of two sample merchant data in M * L dimensions, obtain the similarity distance between the two sample merchant nodes. The specific calculation formula is shown in the following formula (3):

[0104]

[0105] where represents the similarity distance between two sample merchant nodes; M represents the number of index dimensions, L represents the number of time dimensions, M * L represents the number of sub-time series data of one sample merchant node; d j represents the sub-similarity distance between two sample merchant data in the j-th dimension.

[0106] Step 203: Based on the similarity distances between every two of the N sample merchant nodes, cluster the N sample merchant nodes to obtain the neighborhood of each sample merchant node and the other sample merchant nodes located in the neighborhood.

[0107] Specifically, for each sample merchant node, the density-based clustering algorithm (DBSCAN) idea is adopted to cluster the sample merchant nodes. Its core idea is to divide nodes by density and mark the areas with low density as noise points or outliers. Specifically, a neighborhood is set for each node, and core stores, border points, and noise points (also called outliers) are divided according to the number of nodes in the neighborhood of each node. Categories are divided through the continuous diffusion of core points, and finally, the nodes not assigned to any category are noise points.

[0108] In some embodiments, for each sample merchant node, with the sample merchant node as the center and the second preset threshold as the radius, the neighborhood of the sample merchant node is obtained; for each other sample merchant node, when the similarity distance between the other sample merchant node and the sample merchant node is less than or equal to the second preset threshold, the other sample merchant node is located within the neighborhood of the sample merchant node.

[0109] Specifically, the second preset threshold can be set according to actual situations such as industry characteristics and the similarity distances between each sample merchant node, and this application does not make any limitations.

[0110] Step 204: For each sample merchant node, if the number of merchant nodes in the neighborhood of the sample merchant node is less than the first preset threshold and the sample merchant node is not located within the neighborhood of the core merchant node, then the sample merchant node is an abnormal merchant, where the number of merchant nodes in the neighborhood corresponding to the core merchant node is not less than the first preset threshold.

[0111] In some embodiments, if the number of merchant nodes in the neighborhood of the sample merchant node is less than the first preset threshold and the sample merchant node is located within the neighborhood of the core merchant node, then the sample merchant node is a normal merchant.

[0112] Specifically, if the number of merchant nodes in the neighborhood of the sample merchant node is greater than or equal to the first preset threshold, then the sample merchant node is a high-density node, called a core point; if the number of merchant nodes in the neighborhood of the sample merchant node is less than the first preset threshold, then the sample merchant node is a low-density node. If the low-density node is within the neighborhood of other core points, then the low-density node is a border point; if the low-density node is not within the neighborhood of any other core points, then the low-density node is a noise point, and the sample merchant node represented by the noise point is the abnormal merchant in the industry. The first preset threshold is set according to actual situations such as an industry category and the number of merchants, and this application does not make any limitations.

[0113] In the embodiments of the present application, based on the time series data of each sample merchant node among N sample merchant nodes, the similarity distance between every two sample merchant nodes is calculated. Based on the similarity distances between pairs of the N sample merchant nodes, each sample merchant node is clustered to obtain the neighborhood of each sample merchant node and other sample merchant nodes located within that neighborhood. For each sample merchant node, if the number of merchant nodes in the neighborhood of the sample merchant node is less than a first preset threshold and the sample merchant node is not located within the neighborhood of a core merchant node, then the sample merchant node is determined to be an abnormal merchant. The method of the embodiments of the present application fully considers the inherent characteristics of the industry to which the merchant belongs and the clustering characteristics of merchants in the same industry, strengthens the pertinence of merchant information mining, clusters the sample merchant nodes and determines abnormal merchants within the scope of the neighborhood, efficiently and quickly locates abnormal merchants, and at the same time improves the accuracy of abnormal merchant identification.

[0114] To better explain the embodiments of the present application, the following introduces an abnormal identification method for merchant information provided by the embodiments of the present application in combination with an actual scenario. The process of this method is executed by the Figure 1 server shown in Figure 3 and includes the following steps, as shown in

[0115] Step 301, construct time series data for each sample merchant node.

[0116] Step 302, convert the time series data of each sample merchant node into time series change data.

[0117] Step 303, for each sample merchant node, based on the time series change data, calculate the similarity distance between every two sample merchant nodes.

[0118] Step 304, set a threshold to divide the sample merchant nodes to be clustered into core points, boundary points, and noise points.

[0119] Specifically, for each sample merchant node, with each sample merchant node as the center and the second preset threshold as the radius, determine the neighborhood of the sample merchant node. If the number of sample nodes in the neighborhood of the sample merchant node is greater than or equal to the first preset threshold, then the sample merchant node is a core point; if the number of sample nodes in the neighborhood of the sample merchant node is less than the first preset threshold, but the sample merchant node is located within the neighborhood of a core point, then the sample merchant node is a boundary point; if the number of sample nodes in the neighborhood of the sample merchant node is less than the first preset threshold and the sample merchant node is not located within the neighborhood of a core point, then the sample merchant node is a noise point.

[0120] Step 305, complete the clustering of the sample merchant nodes, and take the merchant corresponding to the noise point as the abnormal merchant.

[0121] In the embodiments of the present application, time series data is constructed for each sample merchant node based on the metric dimension and the time dimension, and the time series data is converted into time series change data. The similarity distance between every two sample merchant nodes is calculated based on the time series change data of each sample merchant node. The neighborhood of each sample merchant node and other sample merchant nodes located within the neighborhood are determined based on a first preset threshold and a second preset threshold. Finally, noise points are obtained, and the merchants corresponding to the noise points are used as abnormal merchants. By measuring the similarity of the time series data between sample merchant nodes, the change trend of the metric dimension is focused on, the difference in dimensions between different sample merchant nodes is weakened, and noise points are identified by combining a density-based clustering algorithm, thereby improving the accuracy of identifying abnormal merchants.

[0122] Based on the same technical concept, the embodiments of the present application provide a structural schematic diagram of an abnormal identification device for merchant information, as Figure 4 shown. The abnormal identification device 400 for merchant information includes:

[0123] An acquisition module 401, configured to acquire the time series data of each of the N sample merchant nodes, where N is greater than 1;

[0124] A calculation module 402, configured to obtain the similarity distance between every two of the N sample merchant nodes based on the time series data of the two sample merchant nodes for each pair of the N sample merchant nodes;

[0125] A clustering module 403, configured to cluster the N sample merchant nodes based on the similarity distance between every two of the N sample merchant nodes, and obtain the neighborhood of each sample merchant node and other sample merchant nodes located in the neighborhood;

[0126] For each sample merchant node, if the number of merchant nodes in the neighborhood of the sample merchant node is less than the first preset threshold, and the sample merchant node is not located in the neighborhood of the core merchant node, then the sample merchant node is an abnormal merchant, where the number of merchant nodes in the neighborhood corresponding to the core merchant node is not less than the first preset threshold.

[0127] Optionally, the acquisition module 401 is specifically configured to:

[0128] For each sample merchant node, acquire the original data of the sample merchant node in M metric dimensions respectively, where M is greater than 1;

[0129] Slice the original data of each metric dimension from L time dimensions to obtain the metric time series of each metric dimension in the L time dimensions respectively, where L is greater than 1;

[0130] Determine the time series data of the sample merchant node based on the index time series of each of the obtained metric dimensions in the L time dimensions.

[0131] Optionally, the calculation module 402 is specifically configured to:

[0132] For each sample merchant node, normalize the time series data of the sample merchant node to obtain the normalized time series data of the sample merchant node;

[0133] Based on the target change value between every two time points in the normalized time series data, convert the normalized time series data into time series change data;

[0134] Based on the time series change data of the two sample merchant nodes respectively, calculate the similarity distance between the two sample merchant nodes.

[0135] Optionally, the calculation module 402 is further configured to:

[0136] Set multiple value ranges, each value range corresponding to a fixed threshold;

[0137] For every two time points in the normalized time series data, when the index change value of the two time points is within a target range among the multiple value ranges, use the fixed threshold corresponding to the target range as the target change value between the two time points.

[0138] Optionally, the normalized time series data of the sample merchant node includes sub-time series data of multiple dimensions, and the calculation module 402 is specifically configured to:

[0139] For the sub-time series data of the multiple dimensions, perform the following operations respectively: Based on the sub-time series data of one dimension of the two sample merchant nodes respectively, obtain the sub-similarity distance of the two sample merchant nodes in the one dimension;

[0140] Based on the obtained sub-similarity distances of multiple dimensions, obtain the similarity distance between the two sample merchant nodes.

[0141] Optionally, the clustering module 403 is specifically configured to:

[0142] For each sample merchant node, with the sample merchant node as the center and the second preset threshold as the radius, obtain the neighborhood of the sample merchant node;

[0143] For each other sample merchant node, when the similarity distance between the other sample merchant node and the sample merchant node is less than or equal to the second preset threshold, the other sample merchant node is within the neighborhood of the sample merchant node.

[0144] Optionally, the clustering module 403 is further configured to:

[0145] If the number of merchant nodes in the neighborhood of the sample merchant node is less than a first preset threshold, and the sample merchant node is located in the neighborhood of the core merchant node, then the sample merchant node is a normal merchant.

[0146] In the embodiment of the present application, based on the time series data of each sample merchant node among the N sample merchant nodes, the similarity distance between every two sample merchant nodes is calculated. Based on the similarity distances between pairs of the N sample merchant nodes, each sample merchant node is clustered to obtain the neighborhood of each sample merchant node and other sample merchant nodes located in the neighborhood. For each sample merchant node, if the number of merchant nodes in the neighborhood of the sample merchant node is less than the first preset threshold, and the sample merchant node is not located in the neighborhood of the core merchant node, then it is determined that the sample merchant node is an abnormal merchant. The method of the embodiment of the present application fully considers the inherent characteristics of the merchant's industry and the merchant aggregation characteristics of the same industry, strengthens the pertinence of merchant information mining, clusters the sample merchant nodes and determines abnormal merchants within the scope of the neighborhood, efficiently and quickly locates abnormal merchants, and at the same time improves the accuracy of abnormal merchant identification.

[0147] Based on the same technical concept, the embodiment of the present application provides a computer device, which may be Figure 1 the server shown in Figure 5 shown, including at least one processor 501 and a memory 502 connected to the at least one processor. In the embodiment of the present application, the specific connection medium between the processor 501 and the memory 502 is not limited. Figure 5 Taking the example that the processor 501 and the memory 502 are connected by a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0148] In the embodiment of the present application, the memory 502 stores instructions executed by the at least one processor 501. The at least one processor 501 can execute the steps of the abnormal identification method of the above merchant information by executing the instructions stored in the memory 502.

[0149] Among them, the processor 501 is the control center of the computer device. It can connect various parts of the computer device through various interfaces and lines. By running or executing the instructions stored in the memory 502 and calling the data stored in the memory 502, the abnormal identification of merchant information can be realized. Optionally, the processor 501 may include one or more processing modules. The processor 501 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 501. In some embodiments, the processor 501 and the memory 502 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.

[0150] The processor 501 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0151] The memory 502 serves as a non-volatile computer-readable storage medium and can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 502 can include at least one type of storage medium. For example, it can include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical disks, and so on. The memory 502 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited thereto. The memory 502 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0152] Based on the same inventive concept, embodiments of the present application provide a computer-readable storage medium storing a computer program executable by a computer device. When the program runs on the computer device, the computer device is caused to execute the steps of the above-mentioned method for identifying anomalies in merchant information.

[0153] Based on the same inventive concept, embodiments of the present application provide a computer program product including a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer device, cause the computer device to execute the steps of the above-mentioned method for identifying anomalies in merchant information.

[0154] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0155] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in a process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0156] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in a process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in a process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0158] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A method for identifying abnormalities in merchant information, characterized in that: include: Obtain the time series data of each of N sample merchant nodes, where N is greater than 1; For every two sample merchant nodes among the N sample merchant nodes, based on the respective time series data of the two sample merchant nodes, obtain a similarity distance between the two sample merchant nodes; Based on the similarity distance between every two sample merchant nodes in the N sample merchant nodes, clustering the N sample merchant nodes to obtain a neighborhood of each sample merchant node and other sample merchant nodes located in the neighborhood; For each sample merchant node, if the number of merchant nodes in the neighborhood of the sample merchant node is less than a first preset threshold, and the sample merchant node is not located in the neighborhood of the core merchant node, then the sample merchant node is an abnormal merchant, wherein the number of merchant nodes in the neighborhood corresponding to the core merchant node is not less than the first preset threshold.

2. The method according to claim 1, characterized in that The obtaining of the time series data of each of the N sample merchant nodes includes: For each sample merchant node, obtain the original data of the sample merchant node in M ​​indicator dimensions, where M is greater than 1; The original data of each indicator dimension is segmented from L time dimensions to obtain the indicator time series of each indicator dimension in the L time dimensions, where L is greater than 1; Based on the obtained indicator time series of each indicator dimension in the L time dimensions respectively, the time series data of the sample merchant node is determined.

3. The method according to claim 1, characterized in that The obtaining the similarity distance between the two sample merchant nodes based on the respective time series data of the two sample merchant nodes includes: For each sample merchant node, normalize the time series data of the sample merchant node to obtain the normalized time series data of the sample merchant node; Based on the target change value between every two time points in the normalized time series data, converting the normalized time series data into time series change data; Based on the respective time series change data of the two sample merchant nodes, the similarity distance between the two sample merchant nodes is calculated.

4. The method according to claim 3, characterized in that Before converting the normalized time series data into time series change data based on the target change value between every two time points in the normalized time series data, the method further includes: Set multiple value ranges, each value range corresponds to a fixed threshold; For every two time points in the normalized time series data, when the indicator change values ​​of the two time points are within a target range among the multiple value ranges, a fixed threshold corresponding to the target range is used as the target change value between the two time points.

5. The method according to claim 3, characterized in that The normalized time series data of the sample merchant node includes: sub-time series data of multiple dimensions; The calculating the similarity distance between the two sample merchant nodes based on the respective time series change data of the two sample merchant nodes includes: For the sub-time series data of the multiple dimensions, the following operations are performed respectively: based on the sub-time series data of the two sample merchant nodes in one dimension respectively, a sub-similarity distance of the two sample merchant nodes in the one dimension is obtained; Based on the obtained sub-similarity distances of the multiple dimensions, the similarity distance between the two sample merchant nodes is obtained.

6. The method according to claim 1, characterized in that The clustering of the N sample merchant nodes based on the similarity distance between every two sample merchant nodes in the N sample merchant nodes to obtain the neighborhood of each sample merchant node and other sample merchant nodes located in the neighborhood includes: For each sample merchant node, taking the sample merchant node as the center and the second preset threshold as the radius, obtaining the neighborhood of the sample merchant node; For each other sample merchant node, when the similarity distance between the other sample merchant node and the sample merchant node is less than or equal to the second preset threshold, the other sample merchant node is located in the neighborhood of the sample merchant node.

7. The method according to claim 6, characterized in that Also includes: If the number of merchant nodes in the neighborhood of the sample merchant node is less than a first preset threshold, and the sample merchant node is located in the neighborhood of the core merchant node, then the sample merchant node is a normal merchant.

8. A device for identifying abnormalities in merchant information, characterized in that: include: An acquisition module is used to acquire the time series data of each of N sample merchant nodes, where N is greater than 1; A calculation module, configured to obtain, for each two sample merchant nodes among the N sample merchant nodes, a similarity distance between the two sample merchant nodes based on respective time series data of the two sample merchant nodes; A clustering module, configured to cluster the N sample merchant nodes based on the similarity distance between every two sample merchant nodes in the N sample merchant nodes, and obtain a neighborhood of each sample merchant node and other sample merchant nodes located in the neighborhood; For each sample merchant node, if the number of merchant nodes in the neighborhood of the sample merchant node is less than a first preset threshold, and the sample merchant node is not located in the neighborhood of the core merchant node, then the sample merchant node is an abnormal merchant, wherein the number of merchant nodes in the neighborhood corresponding to the core merchant node is not less than the first preset threshold.

9. A computer device, characterized in that: include: A memory for storing program instructions; A processor is used to call the program instructions stored in the memory and execute the steps of any one of the methods of claims 1 to 7 according to the obtained program.

10. A computer-readable storage medium, characterized in that: It stores a computer program executable by a computer device. When the program is run on the computer device, the computer device executes the steps of any method described in claims 1 to 7.

11. A computer program product, characterized in that The computer program product comprises a computer program stored on a computer-readable storage medium, wherein the computer program comprises program instructions, and when the program instructions are executed by a computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 7.