Cluster identification method and apparatus, computing device, storage medium, and program product

By establishing a multi-directed graph for resource transfer, and identifying abnormal clusters based on the resource transfer level and hierarchy of nodes, the problem of ignoring subject association and resource system transfer process in existing technologies is solved, and higher identification accuracy is achieved.

CN116776259BActive Publication Date: 2026-01-02TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210215971.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2026-01-02
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

Existing technologies, when identifying abnormal clusters, neglect the interrelationships among the various entities in the cluster participating in abnormal resource transfer activities and the resource transfer process between different resource systems, resulting in a high rate of missed detections and false detections.

Method used

By acquiring resource transfer data from multiple entities, a multi-part directed graph of resource transfer is established. Based on the resource transfer level and hierarchy of nodes, target clusters are identified. Considering the resource transfer relationships and flow characteristics between entities, the attribute values ​​in the multi-part directed graph of resource transfer are used to identify target clusters.

Benefits of technology

It improves the accuracy of abnormal cluster identification, eliminates interference from irrelevant entities, and ensures that the identified clusters are those of entities participating in the complete resource transfer process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116776259B_ABST
    Figure CN116776259B_ABST
Patent Text Reader

Abstract

The application discloses a cluster identification method, comprising: obtaining resource transfer data of a plurality of subjects in a first resource system, wherein the resource transfer data comprises resource receiving data and resource expenditure data of each subject in the plurality of subjects; determining candidate subjects from the plurality of subjects based on the resource transfer data of the plurality of subjects, wherein the resource receiving data of the candidate subjects is related to a second resource system different from the first resource system, and the resource expenditure data of the candidate subjects is related to a third resource system different from the first resource system, wherein the second resource system is used to expend resources to the candidate subjects in the first resource system, and the third resource system is used to receive resources from the candidate subjects in the first resource system; determining a resource transfer level of each subject in the candidate subjects according to the resource receiving data and the resource expenditure data of the candidate subjects; and identifying a target cluster from the candidate subjects according to at least the resource transfer level of each subject in the candidate subjects, wherein the target cluster comprises at least one candidate subject involved in a resource transfer process from the second resource system to the third resource system via the first resource system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of Internet, and particularly relates to a cluster identification method and device, a computing device, a computer readable storage medium and a computer program product. BACKGROUND

[0002] The rapid development of Internet technology provides a new way for resource transfer and circulation. In daily resource transactions, there may be abnormal or abnormal resource transfer activities, and the abnormal resource transfer activities are often composed of multiple subjects, which can be called abnormal clusters, and a single subject can be called an abnormal subject. In the related art, the method for identifying abnormal clusters includes determining the transaction characteristics of a single subject engaged in abnormal resource transfer activities in advance, developing an identification model or training a machine learning model according to the transaction characteristics, so as to determine whether a to-be-identified resource transfer subject is a suspicious subject engaged in abnormal resource transfer activities by using the above model; in addition, it also includes using existing algorithms or models to classify or cluster subjects, so as to determine a suspicious subject cluster.

[0003] The abnormal cluster identification method in the related art often takes the transaction characteristics of a single subject as the core, and ignores the mutual correlation of each subject in the cluster participating in abnormal resource transfer activities and the resource transfer process between different resource systems, which leads to the one-sidedness and limitation of the cluster identification idea, and cannot accurately mine potential abnormal clusters, resulting in a large false negative rate and / or false positive rate. SUMMARY

[0004] In view of this, the present application provides a cluster identification method and device, a computing device, a computer readable storage medium and a computer program product, which are expected to alleviate or overcome some or all of the defects mentioned above and other possible defects.

[0005] According to an aspect of the present application, a cluster identification method is provided, which comprises: obtaining resource transfer data of a plurality of subjects in a first resource system, wherein the resource transfer data comprises resource receiving data and resource spending data of each subject in the plurality of subjects; determining candidate subjects from the plurality of subjects based on the resource transfer data of the plurality of subjects, wherein the resource receiving data of the candidate subjects is related to a second resource system different from the first resource system, and the resource spending data of the candidate subjects is related to a third resource system different from the first resource system, wherein the second resource system is used to spend resources to the candidate subjects in the first resource system, and the third resource system is used to receive resources from the candidate subjects in the first resource system; determining resource transfer levels of each subject in the candidate subjects according to the resource receiving data and the resource spending data of the candidate subjects; and identifying a target cluster from the candidate subjects according to at least the resource transfer levels of each subject in the candidate subjects, wherein the target cluster comprises at least one candidate subject involved in a resource transfer process from the second resource system via the first resource system to the third resource system.

[0006] In the cluster identification method according to some embodiments of the present application, identifying the target cluster from the candidate subjects according to at least the resource transfer levels of each subject in the candidate subjects comprises: taking each subject in the candidate subjects as a node, defining directed edges between different nodes based on resource transfer relationships between the nodes, and establishing a resource transfer multi-partite directed graph according to the nodes and the directed edges between the nodes, wherein the resource transfer multi-partite directed graph comprises a plurality of partite graphs, the nodes in a same partite graph have a same level, and the level of the node indicates a degree of relevance of the node to the second resource system or the third resource system; determining attribute values of each node in the resource transfer multi-partite directed graph according to at least the resource transfer level and the level of each node in the resource transfer multi-partite directed graph, wherein the attribute value of each node indicates a possibility of the node belonging to the target cluster; and identifying the target cluster from the nodes of the resource transfer multi-partite directed graph based on the attribute values of each node in the resource transfer multi-partite directed graph.

[0007] In the cluster identification method according to some embodiments of the present application, the resource transfer level of each node comprises a resource transfer retention value of each node, wherein the resource transfer retention value of each node represents an absolute value of a difference between a total amount of resource spending and a total amount of resource receiving of the node.

[0008] In the cluster identification method according to some embodiments of the present application, the resource transfer level of each node further comprises a resource spending-receiving minimum value of each node, wherein the resource spending-receiving minimum value of each node represents a smaller one of the total amount of resource spending and the total amount of resource receiving of the node.

[0009] In the cluster identification method according to some embodiments of the present application, the hierarchy of a node includes a first type of hierarchy indicating the degree of relevance of the node to a second resource system and a second type of hierarchy indicating the degree of relevance of the node to a third resource system, and the attribute value of each node in the resource transfer multi-partite graph satisfies at least one of the following conditions: negatively correlated with the resource transfer retention value of the node; negatively correlated with the smaller one of the first type of hierarchy and the second type of hierarchy of the node; and positively correlated with the resource balance minimum value of the node.

[0010] In the cluster identification method according to some embodiments of the present application, the hierarchy of a node includes a first type of hierarchy indicating the degree of relevance of the node to a second resource system and a second type of hierarchy indicating the degree of relevance of the node to a third resource system, and the attribute value of each node in the resource transfer multi-partite graph is determined according to at least the resource transfer level and the hierarchy of each node in the resource transfer multi-partite graph, including: for each node in the resource transfer multi-partite graph, determining the attribute value of the node using the following formula:

[0011]

[0012] wherein represents the attribute value of node i in the resource transfer multi-partite graph S, represents the smaller one of the total resource expenditure and the total resource reception of node i, represents the larger one of the total resource expenditure and the total resource reception of node i, represents the smaller one of the first type of hierarchy and the second type of hierarchy of node i, and a is a preset parameter greater than or equal to 0 and less than 1.

[0013] In the cluster identification method according to some embodiments of the present application, the target cluster is identified from the nodes of the resource transfer multi-partite graph based on the attribute value of each node in the resource transfer multi-partite graph, including:

[0014] The set of nodes in the resource transfer multi-partite graph is taken as a current candidate cluster, and the following steps are sequentially executed in an iterative manner to obtain a candidate cluster set: an iteration end determination step: in response to the existence of a node-empty subgraph in the resource transfer multi-partite graph corresponding to the current candidate cluster, ending the iteration; a cluster characteristic value calculation step: calculating the cluster characteristic value of the current candidate cluster based on the attribute values of the nodes in the current candidate cluster, and taking the current candidate cluster as a candidate cluster in the candidate cluster set; a current candidate cluster updating step: removing the node with the lowest attribute value from the current candidate cluster and updating the attribute values of the adjacent nodes of the removed node to update the current candidate cluster, and going to the iteration end determination step; and identifying a target cluster from the candidate cluster set based on the cluster characteristic values of the candidate clusters in the candidate cluster set.

[0015] In the cluster identification method according to some embodiments of the present application, obtaining the resource transfer data of the plurality of subjects in the first resource system comprises: obtaining the resource transfer data of the plurality of subjects within a preset time period.

[0016] In the cluster identification method according to some embodiments of the present application, the length of the preset time period is greater than or equal to 3 hours.

[0017] In the cluster identification method according to some embodiments of the present application, determining candidate subjects from the plurality of subjects based on the resource transfer data of the plurality of subjects comprises: filtering a first subject set having a resource transfer relationship with the second resource system from the plurality of subjects based on the resource transfer data of the plurality of subjects; filtering a second subject set having a resource transfer relationship with the third resource system from the plurality of subjects based on the resource transfer data of the plurality of subjects; and determining candidate subjects according to the intersection of the first subject set and the second subject set.

[0018] In the cluster identification method according to some embodiments of the present application, the first subject set having a resource transfer relationship with the second resource system is screened from the plurality of subjects based on the resource receiving data of the plurality of subjects, including: extracting, from the plurality of subjects, direct receiving subjects having a direct resource transfer relationship with the second resource system based on the resource receiving data of the plurality of subjects; extracting, from the plurality of subjects, indirect receiving subjects having an indirect resource transfer relationship with the second resource system according to the resource receiving data of the direct receiving subjects; determining the first subject set based on the direct receiving subjects and the indirect receiving subjects, and / or the second subject set having a resource transfer relationship with the third resource system is screened from the plurality of subjects based on the resource expenditure data of the plurality of subjects, including: extracting, from the plurality of subjects, direct expenditure subjects having a direct resource transfer relationship with the third resource system based on the resource expenditure data of the plurality of subjects; extracting, from the plurality of subjects, indirect expenditure subjects having an indirect resource transfer relationship with the third resource system according to the resource expenditure data of the direct expenditure subjects; determining the second subject set based on the direct expenditure subjects and the indirect expenditure subjects.

[0019] In the cluster identification method according to some embodiments of the present application, the target cluster is identified from the candidate subjects according to at least the resource transfer level of each subject in the candidate subjects, including: calculating the total amount of resource transfer between subject pairs having a direct resource transfer relationship in the candidate subjects based on the resource receiving data and the resource expenditure data of the candidate subjects; in response to the total amount of resource transfer between the subject pairs being less than a resource transfer threshold, deleting the subject pairs from the candidate subjects to obtain updated candidate subjects; and identifying the target cluster from the updated candidate subjects according to at least the resource transfer level of each subject in the updated candidate subjects.

[0020] According to another aspect of the present application, there is provided a cluster identification apparatus, comprising: an obtaining module configured to obtain resource transfer data of a plurality of subjects in a first resource system, wherein the resource transfer data comprises resource receiving data and resource spending data of each of the plurality of subjects; a first determining module configured to determine candidate subjects from the plurality of subjects based on the resource transfer data of the plurality of subjects, wherein the resource receiving data of the candidate subjects is related to a second resource system different from the first resource system and the resource spending data of the candidate subjects is related to a third resource system different from the first resource system, wherein the second resource system is used to spend resources to the candidate subjects in the first resource system and the third resource system is used to receive resources from the candidate subjects in the first resource system; a second determining module configured to determine resource transfer levels of each of the candidate subjects according to the resource receiving data and the resource spending data of the candidate subjects; and an identifying module configured to identify a target cluster from the candidate subjects according to at least the resource transfer levels of each of the candidate subjects, wherein the target cluster comprises at least one candidate subject involved in a resource transfer process from the second resource system via the first resource system to the third resource system.

[0021] According to yet another aspect of the present application, there is provided a computing device comprising a memory and a processor, wherein the memory has stored therein a computer program which, when executed by the processor, causes the processor to perform the steps of the cluster method according to some embodiments of the present application.

[0022] According to still another aspect of the present application, there is provided a computer-readable storage medium having stored thereon computer-readable instructions which, when executed, implement the cluster identification method according to some embodiments of the present application.

[0023] According to another aspect of the present application, there is provided a computer program product comprising computer instructions which, when executed by a processor, implement the steps of the cluster identification method according to some embodiments of the present application.

[0024] In the cluster identification method and device according to some embodiments of the present application, from the plurality of subjects in the first resource system, a subject cluster participating in an abnormal resource transfer process (i.e., a resource transfer process in which resources received from another resource system (i.e., a second resource system) flow through the first resource system or are spent to yet another resource system (i.e., a third resource system)) is identified as a target cluster (i.e., an abnormal cluster) by using the resource receiving data and resource spending data of each subject, so that the complete resource transfer process or resource flow in the first resource system of abnormal resources (e.g., certain resources from the second resource system transferred into the third resource system through the first resource system) is considered; moreover, the resource transfer level of each subject is calculated based on the complete resource transfer process, and the resource transfer level of each subject is used as the identification basis for identifying the target cluster. Such an identification method not only considers the transaction characteristics of each subject itself, but also takes into account the resource transfer relationship and resource flow characteristics between different subjects, overcoming the one-sidedness and limitations of related technologies; at the same time, since this identification method only involves a subject cluster participating in the complete resource transfer process in the first resource system, it can effectively exclude the interference of irrelevant subjects (e.g., normal subjects engaged in normal resource transfer activities) or the interference of normal resource transfer processes, thereby significantly improving the accuracy of identifying the cluster.

[0025] These and other advantages of the application will become apparent from the following description, taken in conjunction with the accompanying drawings, illustrating the principles of the application. BRIEF DESCRIPTION OF DRAWINGS

[0026] Embodiments of the application will now be described in greater detail, by way of example, and with reference to the accompanying drawings, in which:

[0027] Figure 1 An exemplary application scenario of the cluster identification method according to some embodiments of the present application is shown;

[0028] Figure 2 An exemplary flowchart of the cluster identification method according to some embodiments of the present application is shown;

[0029] Figure 3 A schematic principle diagram for determining candidate subjects from a plurality of subjects according to some embodiments of the present application is shown;

[0030] Figure 4 An exemplary flowchart of the cluster identification method according to some embodiments of the present application is shown;

[0031] Figure 5A A schematic diagram for establishing a resource transfer multi-partite directed graph according to some embodiments of the present application is shown;

[0032] Figures 5B-5HA diagram showing a target cluster determined based on resource transfer multi-partite graph according to some embodiments of the present application is shown.

[0033] Figure 6 An exemplary block diagram of a cluster identification apparatus according to some embodiments of the present application is shown; and

[0034] Figure 7 An example system includes an example computing device that is representative of one or more systems and / or devices that can implement the various methods described herein. DETAILED DESCRIPTION

[0035] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, however, can be implemented in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example embodiments to those skilled in the art. Like reference numerals refer to like elements throughout the figures, and descriptions of the same or similar elements can be incorporated throughout this specification by reference.

[0036] Moreover, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the

[0037] The block diagrams in the drawings show only the functionality of the features and can not imply a physical or architectural arrangement of the features. That is, the functionality can be implemented in software, hardware, or a combination thereof. The functionality can be implemented in one or more modules or components, which can collect data and / or perform one or more operations.

[0038] The flow diagrams depicted herein are merely exemplary sequences of operations and are not necessarily implemented in the order as shown. For example, some operations can be performed in parallel, some operations can be omitted, and some operations can be combined, and some operations can be performed in an order other than that which is described here. Furthermore, the flow diagrams can be implemented by hardware, software, firmware, or a combination thereof.

[0039] It should be understood that although the terms first, second, third, etc. can be used herein to describe various components, these components should not be limited by these terms. These terms are used only to distinguish one component from another. Thus, a first component discussed below could be termed a second component without departing from the teachings of the present application. As used herein, the term "and / or" and similar terms include any and all combinations of one or more of the associated listed items.

[0040] Those skilled in the art can understand that the modules or flows in the drawings are not necessarily required for implementing the present application, and therefore cannot be used to limit the protection scope of the present application.

[0041] Before the embodiments of the present application are described in detail, some related concepts are first explained for the sake of clarity.

[0042] Multigraph: also known as multi-part graph, is a special model in graph theory. Let G=(V, E) be an undirected graph, if the vertices V can be partitioned into n mutually disjoint subsets (A1, A2, …, An), n is a positive integer greater than 1, and each edge (i, j) in the graph is associated with two vertices i and j belonging to two different vertex sets (A1, A2, …, An) in the n vertex sets, then the graph G is called a multi-part graph, where A1, A2, …, An are called the multi-part graph G. n n ), n is a positive integer greater than 1, and each edge (i, j) in the graph is associated with two vertices i and j belonging to two different vertex sets (A1, A2, …, An) in the n vertex sets, then the graph G is called a multi-part graph, where A1, A2, …, An are called the multi-part graph G.

[0043] Resource: The resources referred to herein include electronic assets that can be transferred between different subjects through a network, including but not limited to funds, monetary assets for network transactions, cash assets in financial accounts (including bank accounts), financial product resources (such as funds, stocks, bonds, futures), and other assets.

[0044] Subject: The subject referred to herein is a machine device used by an institution, organization or individual for the purpose of providing a certain service or commodity or engaging in certain operational activities or without any commercial purpose, including but not limited to computers, servers or mobile devices such as mobile phones of merchants, individuals and enterprises.

[0045] Cluster: The cluster referred to herein refers to a collection of one or more subjects.

[0046] Target cluster: The target cluster referred to herein refers to a collection of subjects (also called abnormal subjects) participating in or engaging in abnormal resource transfer activities, also called abnormal cluster.

[0047] Figure 1 ​An exemplary application scenario 100 of the cluster identification method according to some embodiments of the present application is shown. The application scenario 100 can include a database 101, a server 102, a network 103, and a terminal device 104, the server 102 and the terminal device 104 are communicatively coupled together through the network 103.

[0048] In this embodiment, the database 101 can store resource transfer data of a plurality of subjects, the resource transfer data of the plurality of subjects including resource receiving data and resource spending data of each subject in the plurality of subjects. The database 101 can transmit the resource transfer data of the plurality of subjects to the server 102 in a wired or wireless manner according to the requirement of the server 102. Optionally, the database 101 can also store the receiving time of each piece of resource receiving data and the spending time of each piece of resource spending data of each subject in the plurality of subjects.

[0049] The server 102 obtains resource transfer data of a plurality of subjects in a first resource system from the database 101, wherein the resource transfer data includes resource receiving data and resource spending data of each subject in the plurality of subjects.

[0050] Then the server 102 determines candidate subjects from the plurality of subjects based on the resource transfer data of the plurality of subjects, the resource receiving data of the candidate subjects being related to a second resource system different from the first resource system, and the resource spending data of the candidate subjects being related to a third resource system different from the first resource system, wherein the second resource system is used to spend resources to the candidate subjects in the first resource system, and the third resource system is used to receive resources from the candidate subjects in the first resource system.

[0051] Next, the server 102 determines resource transfer levels of each subject in the candidate subjects according to the resource receiving data and the resource spending data of the candidate subjects.

[0052] Finally, the server 102 identifies a target cluster from the candidate subjects according to at least the resource transfer levels of each subject in the candidate subjects, the target cluster including at least one candidate subject involved in a resource transfer process from the second resource system via the first resource system to the third resource system. After obtaining the target cluster, the server 102 can send the target cluster to the terminal device 104 for presentation for use by an operator.

[0053] Additionally, the operator can initiate the request to identify the cluster from the terminal device 104, for example, by clicking a button to initiate the identification of the cluster or by inputting an instruction to initiate the identification of the cluster. The request can be transmitted to the server 102 via the network 103 or directly input to the server 102 through an input device of the server 102. The network 103 can be, for example, a wide area network (WAN), a local area network (LAN), a wireless network, a public telephone network, an intranet, and any other type of network known to those skilled in the art.

[0054] It should be noted that the database 101 can be a medium and / or device capable of persistently storing information, and / or a tangible storage device. Therefore, the computer-readable storage medium refers to a non-signal bearing medium. The computer-readable storage medium includes hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented in a method or technology suitable for storing information (such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data). As understood by those of ordinary skill in the art, the instance of the server 102 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited in the present application. The terminal device 104 can receive the target cluster from the server 102 and present the subjects in the cluster to the user one by one.

[0055] The terminal device 104 can be any type of mobile computing device, including a mobile computer (e.g., a personal digital assistant (PDA), a laptop computer, a notebook computer, a tablet computer, a netbook, etc.), a mobile phone (e.g., a cellular phone, a smartphone, etc.), a wearable computing device (e.g., a smartwatch, a head-mounted device, including smart glasses, etc.), or other types of mobile devices. In some embodiments, the terminal device 104 can also be a stationary computing device, such as a desktop computer, a game console, a smart television, etc. Furthermore, in the case where the application scenario 100 includes multiple terminal devices 104, the multiple terminal devices 104 can be the same or different types of computing devices.

[0056] As Figure 1As shown, the terminal device 104 can include a display screen and a terminal application that can interact with the end user via the display screen. The terminal application can be a native application program, a web (Web) application program, or a Little App (e.g., a mobile applet, a WeChat applet) as a lightweight application. In the case where the terminal application is a native application program that needs to be installed, the terminal application can be installed in the terminal device 104. In the case where the terminal application is a Web application program, the terminal application can be accessed through a browser. In the case where the terminal application is a Little App, the terminal application can be directly opened on the terminal device 104 without installation of the terminal application by searching for relevant information of the terminal application (e.g., the name of the terminal application), scanning a graphic code (e.g., a bar code, a QR code) of the terminal application, and the like.

[0057] Figure 2 An exemplary flowchart of a cluster identification method according to some embodiments of the present application is shown. The method 200 shown can be implemented on the server side (e.g., can be implemented on the server 102 shown). Figure 1 Alternatively, in some embodiments, the resource management method according to the present application can be directly executed on the terminal device 104 in the case where the terminal device 104 has sufficient computing resources and computing capabilities. In other embodiments, the resource management method according to the present application can also be executed by the server and the terminal device in combination. As shown, Figure 2 As shown, the cluster identification method according to some embodiments of the present application can include steps S201-S204.

[0058] In step S201, resource transfer data of a plurality of subjects in a first resource system is obtained, wherein the resource transfer data includes resource receiving data and resource spending data of each subject in the plurality of subjects.

[0059] According to the concept of the present application, in order to identify a cluster of subjects participating in abnormal resource transfer activities based on a complete resource transfer process, data of a plurality of subjects that can participate in abnormal resource transfer activities and resource transfer data thereof must be obtained first. The resource receiving data of a subject includes an identification of the subject, an amount of the resource received by the subject, and an identification of another subject that spends the resource. The resource spending data of a subject includes an identification of the subject, an amount of the resource spent by the subject, and an identification of another subject that receives the resource. The first resource system is a system in which resource transfer activities of the plurality of subjects occur, for example, a WeChat payment system.

[0060] Since the transfer activities of abnormal resources are usually completed within two hours, the time period covered by the resource transfer data of the plurality of subjects has an important impact. Too short a time period mostly reflects immediate transaction behavior, resulting in resource transfer data that is not valuable for identification. For example, the resource transfer data of the plurality of subjects within the past 10 minutes mostly reflects normal transaction behavior. Therefore, the operator can pre-set the time period in which the resource transfer data of the plurality of subjects occurs as needed. In some embodiments, obtaining the resource transfer data of the plurality of subjects in the first resource system comprises: obtaining the resource transfer data of the plurality of subjects within a preset time period. Optionally, in some embodiments, the length of the preset time period is greater than or equal to 3 hours. A time period greater than or equal to 3 hours generally covers the complete transfer process of most abnormal resources, which can effectively improve the accuracy of identifying the target cluster.

[0061] In step S202, based on the resource transfer data of the plurality of subjects, candidate subjects are determined from the plurality of subjects, the resource receiving data of the candidate subjects is related to a second resource system different from the first resource system, and the resource expenditure data of the candidate subjects is related to a third resource system different from the first resource system, wherein the second resource system is used to expend resources to the candidate subjects in the first resource system, and the third resource system is used to receive resources from the candidate subjects in the first resource system.

[0062] According to the concept of the present application, after obtaining the resource transfer data of the plurality of subjects, the cluster can be identified. In order to improve the accuracy of identifying the cluster, a preliminary identification can be performed first to eliminate, for example, normal subjects engaged in normal resource transfer activities, and to retain subjects more likely to participate in abnormal resource transfer activities, i.e., candidate subjects, while ensuring a complete resource transfer process or resource flow in the first resource system. The second resource system is a resource system different from the first resource transfer system, for example, the second resource system is a bank card system and the first resource system is a WeChat payment system. The third resource system is also a resource system different from the first resource transfer system, for example, the third resource system is a bank card system, other network payment system, etc. It should be noted that the second resource system and the third resource system can be the same or different.

[0063] The resource receiving data of the candidate subject is related to the second resource system, and the relation includes direct relation and indirect relation. The resource receiving data of the candidate subject is directly related to the second resource system, which means that the candidate subject receives resources from the second resource system. For example, in the process that the WeChat user A receives currency resources from the bank card 1 by the way of WeChat recharge, the resource receiving data of the WeChat user A is directly related to the bank card 1. The resource receiving data of the candidate subject is indirectly related to the second resource system, which means that the candidate subject receives resources from the second resource system through at least one intermediate subject. Similarly, the resource spending data of the candidate subject is related to the third resource system, and the relation includes direct relation and indirect relation. The resource spending data of the candidate subject is directly related to the third resource system, which means that the candidate subject spends resources to the third resource system. For example, in the process that the WeChat user A withdraws currency resources from the bank card 2 by the way of WeChat withdrawal, the resource spending data of the WeChat user A is directly related to the bank card 2. The resource spending data of the candidate subject is indirectly related to the third resource system, which means that the candidate subject spends resources to the third resource system through at least one intermediate subject. By determining the candidate subject, the subsequent identification method can be applied to the complete resource transfer process, instead of only the legal resource transfer process in the first resource system, which improves the identification accuracy. Those subjects not related to the second resource system and those subjects not related to the third resource system will not interfere with the identification of the target cluster, which also improves the identification accuracy.

[0064] Generally, the subjects participating in the abnormal resource transfer activity transfer currency resources from the abnormal bank card to the legal bank card through the network payment system (such as the WeChat payment system), so as to realize the legalization of the currency resources. The cluster identification method of the present application can be used to identify the target cluster of the subjects participating in the abnormal resource transfer activity.

[0065] Figure 3 A schematic diagram illustrating the determination of a candidate subject from a plurality of subjects according to some embodiments of the present application is shown. As shown in Figure 3 The obtained resource transfer data of the subjects A, B, C and D is as follows: the subject A receives 90000 yuan from the second resource system (for example, by the way of WeChat recharge), the subject A spends 90000 yuan to the subject B, accordingly, the subject B receives 90000 yuan from the subject A, further, the subject B spends 90000 yuan to the third resource system (for example, by the way of WeChat withdrawal); moreover, the subject A spends 10000 yuan to the subject C, accordingly, the subject C receives 10000 yuan from the subject A; the subject C further spends 6000 yuan to the subject D, accordingly, the subject D receives 6000 yuan from the subject C; the subject D has no resource spending data, or the resource spending data is zero.

[0066] For subject A, subject A receives resources from the second resource system, so the resource receiving data of subject A is directly related to the second resource system; subject A pays out resources to subject B and subject C, and subject B pays out resources to the third resource system, so the resource payout data of subject A is indirectly related to the third resource system. Thus, subject A is a candidate subject.

[0067] For subject B, subject B pays out resources to the third resource system, so the resource payout data of subject B is directly related to the third resource system; subject B only receives data from subject A, and subject A receives resources from the second resource system, so the resource receiving data of subject B is indirectly related to the second resource system. Thus, subject B is a candidate subject.

[0068] For subject C, subject C receives resources from subject A, and subject A receives resources from the second resource system, so the resource receiving data of subject C is indirectly related to the second resource system; subject C pays out resources to subject D, and the resource payout data of subject D is zero, so the resource payout data of subject C is not related to the third resource system, and thus C is not a candidate subject.

[0069] For subject D, the resource payout data of subject D is zero, so the resource payout data of subject D is not related to the third resource system, and thus D is not a candidate subject.

[0070] Through the foregoing method of determining candidate subjects, subject A and subject B are candidate subjects, and subjects C and D are not candidate subjects.

[0071] In step S203, the resource transfer level of each subject in the candidate subjects is determined according to the resource receiving data and the resource payout data of the candidate subjects. The resource transfer level of a subject is a result calculated according to the resource receiving data and the resource payout data of the subject, and reflects the possibility of the subject participating in or engaging in abnormal resource transfer activities. A subject participating in or engaging in abnormal resource transfer activities often has a high resource transfer level in order to achieve rapid transfer of abnormal resources, so the resource transfer level of a subject can be used as an important identification basis for determining the possibility of the subject participating in or engaging in abnormal resource transfer activities.

[0072] There are various ways to calculate the resource transfer level of a candidate subject using the resource receiving data and resource expenditure data of the candidate subject, which are not limited in the present application. As an example, according to the resource receiving data of a candidate subject, the total amount of resource receiving of the candidate subject is f, and according to the resource expenditure data of the candidate subject, the total amount of resource expenditure of the candidate subject is q, the resource transfer level of the candidate subject can be calculated according to the formula f-a*q, where a is any constant, the total amount of resource receiving f of the candidate subject is a numerical representation of the resource receiving data of the candidate subject, and the total amount of resource expenditure q of the candidate subject is a numerical representation of the resource expenditure data of the candidate subject. For example, the resource transfer level of the candidate subject can also be calculated according to the formula f-q*q. More detailed and specific ways to determine the resource transfer level of each subject in the candidate subject will be given in subsequent embodiments.

[0073] According to the concept of the present application, since the resource transfer level of a subject plays an important role in the judgment of abnormal resource transfer process, in order to play the role of the resource transfer level of the subject, the resource transfer level of each subject should be calculated respectively, and the resource transfer level of each subject should be taken as the identification basis for identifying the target cluster, so as to significantly improve the accuracy of identifying the cluster.

[0074] In step S204, at least according to the resource transfer level of each subject in the candidate subject, a target cluster is identified from the candidate subject, and the target cluster includes at least one candidate subject involved in the resource transfer process from the second resource system to the third resource system via the first resource system.

[0075] According to the concept of the present application, after completing the preliminary identification and calculating the resource transfer level of each candidate subject, the target cluster, i.e. the cluster of subjects that may participate in abnormal resource transfer activities, should be identified from the candidate subject using the identification basis. Generally, one or more subjects participate in abnormal resource transfer activities, so the final identification result should be a cluster including at least one candidate subject. In this step S204, the resource transfer data between each candidate subject reflects not only the transaction characteristics of each subject itself, but also the resource transfer relationship and resource flow characteristics between different subjects, overcoming the one-sidedness and limitations of related technologies, and since the target cluster includes at least one candidate subject involved in the resource transfer process from the second resource system to the third resource system via the first resource system, it is ensured that the target cluster involves the complete resource transfer process or resource flow in the first resource system, and the role of the resource transfer level of the subject in determining the possibility of participating in or engaging in abnormal resource transfer activities is adopted, so as to significantly improve the accuracy of identifying the cluster.

[0076] In the cluster identification method according to some embodiments of the present application, from the plurality of subjects in the first resource system, a subject cluster participating in an abnormal resource transfer process (i.e., a resource transfer process in which resources received from another resource system (i.e., a second resource system) flow through the first resource system or are spent to yet another resource system (i.e., a third resource system)) is identified as a target cluster (i.e., an abnormal cluster) by using resource receiving data and resource spending data of each subject, so that the complete resource transfer process or resource flow direction of abnormal resources (e.g., certain resources from the second resource system transferred into the third resource system via the first resource system) in the first resource system is considered; moreover, the resource transfer level of each subject is calculated based on the complete resource transfer process, and the resource transfer level of each subject is used as the identification basis for identifying the target cluster. Such an identification method not only considers the transaction characteristics of each subject itself, but also takes into account the resource transfer relationship and resource flow characteristics between different subjects, overcoming the one-sidedness and limitations of related technologies; at the same time, since this identification method only involves a subject cluster participating in the complete resource transfer process in the first resource system, it can effectively exclude the interference of irrelevant subjects (e.g., normal subjects engaged in normal resource transfer activities) or the interference of normal resource transfer processes, thereby significantly improving the accuracy of identifying the cluster; in addition, the method uses the resource transfer level of the subject to determine the possibility of participating in or engaging in abnormal resource transfer activities, thereby significantly improving the accuracy of identifying the cluster.

[0077] In step S204, the process of identifying the target cluster from the candidate subjects includes various implementation schemes. Figure 4 An exemplary flowchart of a cluster identification method according to some embodiments of the present application is shown. Figure 4 The steps S401-S403 shown are Figure 2 an extension of step S204, and steps S401-S403 are based on the premise that resource transfer data of a plurality of subjects is obtained and candidate subjects are determined from the plurality of subjects.

[0078] In step S401, each entity in the candidate entities is treated as a node. Directed edges between different nodes are defined based on the resource transfer relationships between them. A multi-part directed graph of resource transfer is then established based on these edges. This multi-part directed graph includes multiple subgraphs, where nodes within the same subgraph have the same level. The level of a node indicates its relevance to the second or third resource system. Resource transfer relationships include both resource expenditure and resource receipt relationships, and the number of transferred resources corresponds to the number of such relationships. For example, if node 1 spends 10,000 yuan on node 2, it indicates a resource transfer relationship between node 1 and node 2, with the direction of the transfer from node 1 to node 2, and the number of such relationships is 10,000. For node levels, a smaller level indicates a greater relevance between the node and the second or third resource system. In this embodiment, the obtained resource transfer data of the candidate entities is mapped to a multi-part directed graph of resource transfer. Based on this graph, the resource transfer level of each entity is determined, and the target cluster is identified. In later embodiments, the hierarchy indicating the relevance of a node to the second resource system is a first type hierarchy, and the hierarchy indicating the relevance of a node to the third resource system is a second type hierarchy.

[0079] Figure 5A A schematic diagram illustrating the establishment of a resource transfer multi-directed graph S1 according to some embodiments of this application is shown. Figure 5A In the illustrated embodiment, "resources" refers to monetary resources, and the resource transfer data of the identified candidate entities are as follows.

[0080] Candidate subject P 1,1 Received 400,000 yuan from account C1 in the second resource system and 100,000 yuan from account C2 in the second resource system, and transferred the funds to candidate entity P. 2,1 100,000 yuan was spent on candidate entity P. 2,2 200,000 yuan was spent on candidate entity P. 2,3 180,000 yuan was spent.

[0081] Candidate subject P 1,2 Received 200,000 yuan from account C1 in the second resource system and 300,000 yuan from account C2 in the second resource system, and transferred the funds to candidate entity P. 2,3 300,000 yuan was spent on candidate entity P. 2,4 190,000 yuan was spent.

[0082] Candidate subject P 2,1 To candidate subject P 3,1 80,000 yuan was spent, candidate entity P 2,1 The resource receiving data will not be repeated.

[0083] Candidate subject P2,2 To candidate subject P 3,1 20,000 yuan was spent, candidate entity P 2,2 The resource receiving data will not be repeated.

[0084] Candidate subject P 2,3 To candidate subject P 3,3 480,000 yuan was spent, candidate entity P 2,3 The resource receiving data will not be repeated.

[0085] Candidate subject P 2,4 To candidate subject P 3,2 190,000 yuan was spent, candidate entity P 2,4 The resource receiving data will not be repeated.

[0086] Candidate subject P 3,1 A payment of 50,000 yuan was made to account D3 of the third-party resource system, candidate entity P. 3,1 The resource receiving data will not be repeated.

[0087] Candidate subject P 3,2 180,000 yuan was spent on account D4 of the third resource system, candidate entity P. 3,2 The resource receiving data will not be repeated.

[0088] Candidate subject P 3,3 480,000 yuan was spent on account D4 of the third resource system, candidate entity P. 3,3 The resource receiving data will not be repeated.

[0089] like Figure 5A As shown, each candidate entity is mapped to a node, and directed edges between different nodes are defined based on the resource transfer relationships between them. For example, node P 2,1 To node P 3,1 The expenditure of 80,000 yuan can be identified as the starting point, node P. 2,1 The endpoint is node P. 3,1 The edges have a size of 80,000. Then, based on the above nodes and directed edges, a multi-part directed graph S1 for resource transfer is constructed. The multi-part directed graph S1 includes three sub-graphs. The first sub-graph includes node P. 1,1 and P 1,2 The first type level of these two nodes is 1, and the second type level is 3. The second subgraph includes node P. 2,1 P 2,2 P 2,3 and P 2,4 The first type level of these four nodes is 2, the second type level is 2, and the third subgraph includes node P. 3,1 P 3,2 and 3,3The first type level of the three nodes is 3, and the second type level is 1.

[0090] It is worth noting that the accounts C1 and C2 of the second resource system and the accounts D3 and D4 of the third resource system are not included in the resource transfer multi-partite graph S1, and the resource transfer multi-partite graph only includes the nodes and directed edges between the dashed lines 501 and 502, Figures 5B-5H The same is true.

[0091] In step S402, the attribute value of each node in the resource transfer multi-partite graph is determined according to at least the resource transfer level and the level of each node in the resource transfer multi-partite graph, where the attribute value of each node indicates the possibility of the node belonging to the target cluster. After obtaining the resource transfer multi-partite graph, the resource transfer level and the level of each node can be easily calculated, and the two parameters help to accurately determine the possibility of the node belonging to the target cluster. Specifically, the subject corresponding to the node with a larger resource transfer level is more likely to participate in abnormal resource transfer, and the subject corresponding to the node with a smaller resource transfer level is less likely to participate in abnormal resource transfer; the subject corresponding to the node with a smaller level is more related to the second or third resource system, and is more likely to participate in abnormal resource transfer, and the subject corresponding to the node with a larger level is less related to the second or third resource system, and is less likely to participate in abnormal resource transfer. Therefore, the attribute value of the node is determined according to the two parameters, so that the attribute value of the node can accurately represent the possibility of the subject corresponding to the node being an abnormal resource transfer subject, the larger the attribute value of the node, the more likely the subject corresponding to the node is to participate in abnormal resource transfer activities, and the smaller the attribute value of the node, the less likely the subject corresponding to the node is to participate in abnormal resource transfer activities.

[0092] In step S403, the target cluster is identified from the nodes of the resource transfer multi-partite graph based on the attribute value of each node in the resource transfer multi-partite graph. In this embodiment, the target cluster participating in abnormal resource transfer activities is identified by calculating the possibility of each subject in the candidate subject participating in abnormal resource transfer activities one by one. The specific identification method can include various embodiments, which are not limited here. For example, a node attribute value threshold can be preset, and in the resource transfer multi-partite graph, the nodes with attribute values greater than the attribute value threshold are identified as nodes belonging to the target cluster. For another example, the 10% of the nodes with the largest attribute values can be identified as nodes belonging to the target cluster.

[0093] In some embodiments, identifying the target cluster from the candidate subjects according to at least the resource transfer level of each subject in the candidate subjects comprises: calculating, based on the resource receiving data and the resource expenditure data of the candidate subjects, the total amount of resource transfer between the subject pairs in the candidate subjects that have direct resource transfer relationship; in response to the total amount of resource transfer between the subject pairs being less than a resource transfer threshold, deleting the subject pairs from the candidate subjects to obtain updated candidate subjects; and identifying the target cluster from the updated candidate subjects according to at least the resource transfer level of each subject in the updated candidate subjects. The total amount of resource transfer indicates the total amount of receiving / expenditure resources between the corresponding subject pairs. In fact, the total amount of resource transfer of a subject being less than the resource transfer threshold indicates that the subject has not enough intensity of correlation with the second resource system or the third resource system, and can be determined as irrelevant to the second resource system or the third resource system, and not belonging to the candidate subjects. Therefore, deleting the subject pairs with the total amount of resource transfer less than the resource transfer threshold can effectively exclude the interference of irrelevant subjects (e.g. normal subjects engaged in normal resource transfer activities) or the interference of normal resource transfer process, thereby significantly improving the accuracy of identifying the cluster. The resource transfer threshold may, for example, be set to 1000 yuan or 2000 yuan according to experience.

[0094] In the following, the way of calculating the attribute value of the node is described in detail through other embodiments.

[0095] In some embodiments, the resource transfer level of each node comprises a resource transfer retention value of each node, and the resource transfer retention value of each node represents the absolute value of the difference between the total amount of resource expenditure and the total amount of resource receiving of the node.

[0096] In the subjects participating in abnormal resource transfer activities, in order to achieve the purpose of resource transfer, the resource transfer retention value of the subject is usually small, and these subjects act as resource transfer media and have a strong resource transfer level. Therefore, the resource transfer retention value of the node can be used to determine the possibility of the node participating in abnormal resource transfer activities. For example, in the resource transfer multi-directed graph S1 in Figure 5A , the resource transfer retention value of node P 2,1 is 20,000 yuan (100,000-80,000=20,000 yuan). In this embodiment, there are various ways to calculate the attribute value of the node according to at least the resource transfer retention value and the level in the resource transfer level of the node, which are not limited here. For example, the attribute value of the node can be equal to the reciprocal of the product of the resource transfer retention value of the node and the level of the node.

[0097] In some embodiments, the resource transfer level of each node further includes the minimum resource expenditure and receipt value of each node, where the minimum resource expenditure and receipt value represents the smaller of the node's total resource expenditure and total resource receipt. Among entities participating in abnormal resource transfer activities, to achieve the purpose of resource transfer, the entity's minimum resource expenditure and receipt value is often large, thereby achieving the goal of rapid resource transfer. Therefore, the minimum resource expenditure and receipt value of a node can also be used to determine the likelihood of a node participating in abnormal resource transfer activities. In this embodiment, there are various ways to determine the node's attribute value based on the resource transfer retention value and minimum resource expenditure and receipt value in the node's resource transfer level, as well as based on the hierarchy, and these are not limited here. For example, the result of the node's minimum resource expenditure and receipt value / (node's resource transfer retention value * node's hierarchy) can be used as the node's attribute value.

[0098] In some embodiments, the hierarchy of a node includes a first type hierarchy indicating the degree of relevance of the node to the second resource system and a second type hierarchy indicating the degree of relevance of the node to the third resource system, and the attribute value of each node in the resource transfer multi-directional graph satisfies at least one of the following conditions: (1) negatively correlated with the resource transfer retention value of the node; (2) negatively correlated with the smaller of the first type hierarchy and the second type hierarchy of the node; and (3) positively correlated with the minimum resource income and expenditure of the node. The smaller the first type hierarchy of the node, the stronger the relevance of the node to the second resource system, such as... Figure 3 As shown, node A's first type level is 1, meaning that node A is directly related to the second resource system. The higher the first type level of a node, the weaker the correlation between the node and the second resource system. Figure 3 As shown, node B's first type level is 2, meaning that node B is separated from the second resource system by an intermediate node A, indicating a weak correlation. The smaller the second type level of a node, the stronger its correlation with the third resource system. Figure 3 As shown, node B's second type level is 1, meaning that node B is directly related to the third resource system. The higher the second type level of a node, the weaker the correlation between the node and the third resource system. Figure 3 As shown, the second type level of node A is 2, which means that node A is separated from the third resource system by an intermediate node B, and the correlation is relatively weak.

[0099] As described above, for condition (1), as described above, in the subjects participating in the abnormal resource transfer activity, in order to achieve the purpose of resource transfer, the resource transfer retention value of the subject is often small, these subjects act as the function of resource transfer medium, and have a strong resource transfer level, therefore, the calculation method of the attribute value of the node is designed to be negatively correlated with the resource transfer retention value of the node, which can accurately identify the subjects participating in the abnormal resource transfer activity. For example, in the embodiment in which the attribute value of the node is equal to the reciprocal of the resource transfer retention value of the node, the attribute value of node P 2,1 is 20 millionths, and the attribute value of node P 2,2 is 180 millionths, which means that node P 2,1 is more likely to participate in the abnormal resource transfer activity than node P 2,2 . For condition (2), the first type level and the second type level of the node reflect the position of the node in the resource transfer process, because in the abnormal resource transfer activity, the first level and the last level close to the resource transfer process are more likely to appear subjects participating in the abnormal resource transfer activity, and the subjects in the middle level are more likely to appear subjects participating in the normal resource transfer activity, therefore, the attribute value of the node is negatively correlated with the smaller one of the first type level and the second type level of the node, which helps to accurately identify the subjects participating in the abnormal resource transfer activity. For condition (3), in the subjects participating in the abnormal resource transfer activity, in order to achieve the purpose of resource transfer, the minimum resource balance of the subject is often large, so as to achieve the purpose of fast resource transfer, therefore, the attribute value of the node is designed to be positively correlated with the minimum resource balance of the node, which can accurately identify the subjects participating in the abnormal resource transfer activity. Taking the resource transfer multi-partite directed graph S1 in Figure 5A as an example, when the attribute value w of the node is equal to one millionth of the minimum resource balance of the node, the attribute value of node P 2,1 is 8, and the attribute value of node P 2,2 is 2, which means that node P 2,1 is more likely to participate in the abnormal resource transfer activity than node P 2,2 . In summary, satisfying any one of the above conditions (1)-(3) helps to improve the identification accuracy.

[0100] In some more specific embodiments, the level of the node includes a first type level indicating the degree of relevance of the node to the second resource system and a second type level indicating the degree of relevance of the node to the third resource system, and the attribute value of each node in the resource transfer multi-partite directed graph is determined according to at least the resource transfer level and the level of each node in the resource transfer multi-partite directed graph, including: for each node in the resource transfer multi-partite directed graph, using formula (1) to determine the attribute value of the node:

[0101] Formula (1)

[0102] in This represents the attribute value of node i in the multi-directional graph S that represents resource transfer. This represents the smaller of the total resource expenditure and the total resource receipt of node i. This represents the larger of the total resource expenditure and the total resource receipt of node i. Let 'a' represent the smaller of the first type level and the second type level of node i, where 'a' is a preset parameter greater than or equal to 0 and less than 1. Here, the preset parameter 'a' acts as a harmonic parameter. The larger 'a' is, the greater the influence of the node's resource transfer retention value on its attribute value, and the smaller the influence of the node's minimum resource transfer amount on its attribute value. Conversely, the smaller 'a' is, the smaller the influence of the node's resource transfer retention value on its attribute value, and the greater the influence of the node's minimum resource transfer amount on its attribute value. As mentioned earlier, in abnormal resource transfer activities, entities closer to the first and last levels of the resource transfer process are more likely to participate in abnormal resource transfer activities, while entities in the middle levels tend to participate in normal resource transfer activities. Therefore, a negative correlation between a node's attribute value and the smaller of its first type level and the second type level helps to accurately identify entities participating in abnormal resource transfer activities.

[0103] Formula (1) combines conditions (1)-(3) from the aforementioned embodiments, w i (S) is negatively correlated with the resource transfer retention value of node i, negatively correlated with the smaller of the first type level and the second type level of node i, and positively correlated with the minimum resource transfer amount of node i. Therefore, using this formula to calculate the attribute value of a node can significantly improve the accuracy of node identification. The preset parameter 'a' can be set based on empirical values.

[0104] by Figure 5A Taking the resource transfer multi-directed graph S1 as an example, assuming that parameter a is preset to 0.5, node P 1,1 f P1,1 (S1) is 480,000, q P1,1 (S1) is 500,000, d P1,1 (S1) is 1, and the relevance w P1,1 (S1) is (48-50*0.5) ten thousand / 1 = 230,000. Similarly, the attribute values ​​of the other nodes are w. P1,2 (S1)=240,000, w P2,1 (S1)=15,000,w P2,2 (S1)=-40,000, w P2,3 (S1)=120,000,w P2,4 (S1)=47,500, w P3,1 (S1)=00,000, w P3,2 (S1)=85,000,wP3,3 (S1)=24 million. Thus, node P 2,2 is less likely to be the subject of abnormal resource transfer activities, node P 1,2 and node P 3,3 is more likely to be the subject of abnormal resource transfer activities.

[0105] In addition to step S402, step S403 can also be extended. In some embodiments, based on the attribute values of each node in the resource transfer multi-partite graph, identifying the target cluster from the nodes of the resource transfer multi-partite graph comprises: taking the set of each node in the resource transfer multi-partite graph as a current candidate cluster, and sequentially performing the following steps by iteration to obtain a candidate cluster set: iteration end determination step: in response to the existence of a node-empty subgraph in the resource transfer multi-partite graph corresponding to the current candidate cluster, ending the iteration; cluster feature value calculation step: based on the attribute values of each node in the current candidate cluster, calculating the cluster feature value of the current candidate cluster, and taking the current candidate cluster as a candidate cluster in the candidate cluster set; current candidate cluster updating step: removing the node with the lowest attribute value from the current candidate cluster and updating the attribute values of the adjacent nodes of the removed node to update the current candidate cluster, and going to the iteration end determination step; based on the cluster feature values of each candidate cluster in the candidate cluster set, identifying the target cluster from the candidate cluster set. The cluster feature value represents the likelihood that the corresponding cluster is the target cluster. As an example, formula (2) can be used to determine the cluster feature value of the current candidate cluster:

[0106] Formula (2)

[0107] wherein represents the cluster feature value of the corresponding current candidate cluster of the resource transfer multi-partite graph S, N is the number of nodes in the resource transfer multi-partite graph S, represents the attribute value of node i in the resource transfer multi-partite graph S.

[0108] Formula (2) indicates that the average value of the attribute values of each node in the current candidate cluster is taken as the cluster feature value of the current candidate cluster. In another example, the attribute values of each node can be weighted and summed to calculate the cluster feature value of the current candidate cluster.

[0109] Taking the resource transfer multi-partite graph S1 with Figure 5A as an example, Figures 5B-5HA schematic diagram of the target cluster determined based on the resource transfer multi-directed graph according to this embodiment is shown. In this embodiment, the resource is monetary resource, the attribute value of the node is calculated in the manner described in formula (1), and formula (2) is used to determine the cluster feature value of the current candidate cluster. After step S402, the process of identifying the cluster according to this embodiment is as follows.

[0110] Assuming parameter a is set to 0.5, the set of nodes in the multi-directed resource transfer graph S1 is taken as the current candidate cluster S1. First, in the iteration end determination step, if there is no subgraph with empty nodes in the multi-directed resource transfer graph corresponding to the current candidate cluster S1, subsequent steps can be executed.

[0111] Then, the cluster feature value calculation step is performed to calculate the cluster feature value of the current candidate cluster S1. The attribute value of each node is w. P1,1 (S1)=230,000, w P1,2 (S1)=240,000, w P2,1 (S1)=15,000,w P2,2 (S1)=-40,000, w P2,3 (S1)=120,000,w P2,4 (S1)=47,500, w P3,1 (S1)=00,000, w P3,2 (S1)=85,000,w P3,3 (S1) = 240,000. Therefore, the cluster characteristic value of the current candidate cluster S1 is (23 + 24 + 1.5 - 4 + 12 + 4.75 + 0 + 8.5 + 24) / 9 ≈ 104,200. For ease of explanation, two decimal places are retained here, but this does not constitute a limitation of this application. The current candidate cluster S1 is taken as a candidate cluster in the candidate cluster set.

[0112] Next, the current candidate cluster update step is executed, removing the node P with the lowest attribute value from the current candidate cluster S1. 2,2 Remove node P 2,2 Before, its neighboring node is node P. 1,1 and node P 3,1 Remove node P 2,2 ,like Figure 5B As shown, this is used to update the current candidate cluster S2.

[0113] Proceed to the iteration end determination step. The resource transfer multi-part directed graph corresponding to the current candidate cluster S2 does not have a subgraph with empty nodes, so the iteration can continue.

[0114] In the new iteration, the cluster feature value calculation step is executed first to calculate the cluster feature value of the current candidate cluster S2. Node P 1,1 Resource receipt data remains unchanged, while resource expenditure data decreases from 480,000 to 280,000.P1,1 (S2) = (28 - 50 * 0.5) / 1 = 30,000. Similarly, w P3,1 (S2) = (5 - 8 * 0.5) / 1 = 10,000. The attribute values of other nodes remain unchanged. Therefore, the cluster feature value of the current candidate cluster S2 is (3 + 24 + 1.5 + 12 + 4.75 + 1 + 8.5 + 24) / 8 ≈ 9.84,000. The current candidate cluster S2 is taken as a candidate cluster in the candidate cluster set.

[0115] Then, the current candidate cluster updating step is performed, and the node P 3,1 with the lowest attribute value is removed from the current candidate cluster S2. 3,1 As shown in FIG. 4B, the current candidate cluster S3 is updated. Figure 5C

[0116] Turning to the iteration end determination step, there is no node-empty subgraph in the resource transfer multi-partite graph corresponding to the current candidate cluster S3, and the iteration can continue.

[0117] In a new round of iteration, the cluster feature value calculation step is performed first, and the cluster feature value of the current candidate cluster S3 is calculated. The attribute values of the nodes are w P1,1 (S3) = 30,000, w P1,2 (S3) = 240,000, w P2,1 (S3) = -25,000, w P2,3 (S3) = 120,000, w P2,4 (S3) = 47,500, w P3,2 (S3) = 85,000, w P3,3 (S3) = 240,000. Therefore, the cluster feature value of the current candidate cluster S3 is (3 + 24 - 2.5 + 12 + 4.75 + 8.5 + 24) / 7 ≈ 10.54,000. The current candidate cluster S3 is taken as a candidate cluster in the candidate cluster set.

[0118] Then, the current candidate cluster updating step is performed, and the node P 2,1 with the lowest attribute value is removed from the current candidate cluster S3. 2,1 As shown in FIG. 4C, the current candidate cluster S4 is updated. Figure 5D

[0119] Turning to the iteration end determination step, there is no node-empty subgraph in the resource transfer multi-partite graph corresponding to the current candidate cluster S4, and the iteration can continue.

[0120] In a new round of iteration, the cluster feature value calculation step is performed first, and the cluster feature value of the current candidate cluster S4 is calculated. The attribute values of the nodes are w P1,1 (S4) = -70,000, w P1,2 (S4) = 240,000, w​​P2,3 (S4) = 12 million, w P2,4 (S4) = 4.75 million, w P3,2 (S4) = 8.5 million, w P3,3 (S4) = 24 million. Therefore, the cluster feature value of the current candidate cluster S4 is (-7+24+12+4.75+8.5+24) / 6=11.04 million. The current candidate cluster S4 is taken as a candidate cluster in the candidate cluster set.

[0121] Then, a current candidate cluster updating step is performed, and the node P 1,1 with the lowest attribute value is removed from the current candidate cluster S4. 1,1 As shown in FIG. 4, the current candidate cluster S5 is updated. Figure 5E

[0122] It is turned to an iteration end determining step. There is no subgraph with an empty node in the resource transfer multi-partite graph corresponding to the current candidate cluster S5, and the iteration can continue.

[0123] In a new round of iteration, a cluster feature value calculating step is performed first, and the cluster feature value of the current candidate cluster S5 is calculated. The attribute value of each node is w P1,2 (S5) = 24 million, w P2,3 (S5) = 3 million, w P2,4 (S5) = 4.75 million, w P3,2 (S5) = 8.5 million, w P3,3 (S5) = 24 million. Therefore, the cluster feature value of the current candidate cluster S5 is (24+3+4.75+8.5+24) / 5=12.85 million. The current candidate cluster S5 is taken as a candidate cluster in the candidate cluster set.

[0124] Then, a current candidate cluster updating step is performed, and the node P 2,3 with the lowest attribute value is removed from the current candidate cluster S5. 2,3 As shown in FIG. 4, the current candidate cluster S6 is updated. Figure 5F

[0125] It is turned to an iteration end determining step. There is no subgraph with an empty node in the resource transfer multi-partite graph corresponding to the current candidate cluster S6, and the iteration can continue.

[0126] In a new round of iteration, a cluster feature value calculating step is performed first, and the cluster feature value of the current candidate cluster S6 is calculated. The attribute value of each node is w P1,2 (S6) = -6 million, w P2,4 (S6) = 4.75 million, w P3,2 (S6) = 8.5 million, w P3,3 ​​(S6) = -240,000. Therefore, the cluster characteristic value of the current candidate cluster S6 is (-6 + 4.75 + 8.5 - 24) / 4 ≈ -41,900. The current candidate cluster S6 is taken as a candidate cluster in the candidate cluster set.

[0127] Next, the current candidate cluster update step is executed, removing the node P with the lowest attribute value from the current candidate cluster S6. 3,3 Remove node P 3,3 ,like Figure 5G As shown, this is used to update the current candidate cluster S7.

[0128] Proceed to the iteration end determination step. The resource transfer multi-part directed graph corresponding to the current candidate cluster S7 does not have a subgraph with empty nodes, so the iteration can continue.

[0129] In the new iteration, the cluster feature value calculation step is performed first to calculate the cluster feature value of the current candidate cluster S7. The attribute value of each node is w. P1,2 (S7) = -60,000, w P2,4 (S7) = 47,500, w P3,2 (S7) = 85,000. Therefore, the cluster characteristic value of the current candidate cluster S7 is (-6 + 4.75 + 8.5) / 3 ≈ 24,200. The current candidate cluster S7 is taken as a candidate cluster in the candidate cluster set.

[0130] Next, the current candidate cluster update step is executed, removing the node P with the lowest attribute value from the current candidate cluster S6. 1,2 Remove node P 1,2 ,like Figure 5H As shown, update the current candidate cluster S8.

[0131] Proceed to the iteration end determination step. If the resource transfer multi-part directed graph corresponding to the current candidate cluster S8 has a subgraph with empty nodes, the iteration is determined to be over.

[0132] Comparing candidate clusters S1-S7 in the candidate cluster set, candidate cluster S5 has the largest cluster feature value. Therefore, candidate cluster S5 is identified as the target cluster, which includes node P. 1,2 P 2,3 P 2,4 P 3,2 P 3,3 .

[0133] This application will also provide specific embodiments for determining candidate subjects.

[0134] In some embodiments, the candidate subjects are determined from the plurality of subjects based on the resource transfer data of the plurality of subjects, including: screening a first subject set from the plurality of subjects that has a resource transfer relationship with the second resource system based on the resource receiving data of the plurality of subjects; screening a second subject set from the plurality of subjects that has a resource transfer relationship with the third resource system based on the resource expenditure data of the plurality of subjects; and determining the candidate subjects according to the intersection of the first subject set and the second subject set. Since the candidate subjects are related to both the second resource system and the third resource system, in the specific implementation, the subjects that have a resource transfer relationship with the second resource system and the third resource system can be obtained respectively from the second resource system and the third resource system, and then the intersection of the two sets is obtained to obtain the candidate subjects.

[0135] In particular, in some embodiments, the first subject set is screened from the plurality of subjects that has a resource transfer relationship with the second resource system based on the resource receiving data of the plurality of subjects, including: extracting a direct receiving subject that has a direct resource transfer relationship with the second resource system from the plurality of subjects based on the resource receiving data of the plurality of subjects; extracting an indirect receiving subject that has an indirect resource transfer relationship with the second resource system from the plurality of subjects according to the resource receiving data of the direct receiving subject; and determining the first subject set based on the direct receiving subject and the indirect receiving subject, and / or the second subject set is screened from the plurality of subjects that has a resource transfer relationship with the third resource system based on the resource expenditure data of the plurality of subjects, including: extracting a direct expenditure subject that has a direct resource transfer relationship with the third resource system from the plurality of subjects based on the resource expenditure data of the plurality of subjects; extracting an indirect expenditure subject that has an indirect resource transfer relationship with the third resource system from the plurality of subjects according to the resource expenditure data of the direct expenditure subject; and determining the second subject set based on the direct expenditure subject and the indirect expenditure subject. This embodiment gives more detailed technical details of determining the candidate subjects. In the process of obtaining the first subject set, the direct receiving subject that directly receives resources from the second resource system is obtained first, and then the indirect receiving subject is derived step by step according to the identifier in the resource receiving data, so as to obtain the first subject set related to the second resource system. The same is true for the second subject set. The intersection of the two sets is naturally the candidate subjects, so that the determination of the candidate subjects can be easily and accurately realized, which helps to exclude the interference of irrelevant subjects (such as normal subjects engaged in normal resource transfer activities) or the interference of normal resource transfer process, and improves the efficiency of identifying the target cluster.

[0136] Figure 6An exemplary structural block diagram of a cluster identifying apparatus 600 according to some embodiments of the present application is shown. The apparatus 600 comprises: an obtaining module 601, a first determining module 602, a second determining module 603, and an identifying module 604. The obtaining module 601 is configured to obtain resource transfer data of a plurality of subjects in a first resource system, wherein the resource transfer data comprises resource receiving data and resource spending data of each subject in the plurality of subjects. The first determining module 602 is configured to determine candidate subjects from the plurality of subjects based on the resource transfer data of the plurality of subjects, wherein the resource receiving data of the candidate subjects is related to a second resource system different from the first resource system, and the resource spending data of the candidate subjects is related to a third resource system different from the first resource system, wherein the second resource system is used to spend resources to the candidate subjects in the first resource system, and the third resource system is used to receive resources from the candidate subjects in the first resource system. The second determining module 603 is configured to determine resource transfer levels of each subject in the candidate subjects according to the resource receiving data and the resource spending data of the candidate subjects. The identifying module 604 is configured to identify a target cluster from the candidate subjects according to at least the resource transfer levels of each subject in the candidate subjects, wherein the target cluster comprises at least one candidate subject involved in a resource transfer process from the second resource system via the first resource system to the third resource system.

[0137] It should be noted that the various modules described above can be implemented in software or hardware or a combination of both. Multiple different modules can be implemented in the same software or hardware structure, or one module can be implemented by multiple different software or hardware structures.

[0138] In the cluster identification apparatus according to some embodiments of the present application, from the plurality of subjects in the first resource system, a subject cluster participating in an abnormal resource transfer process (i.e., a resource transfer process in which resources received from another resource system (i.e., a second resource system) flow through the first resource system or are spent to yet another resource system (i.e., a third resource system)) is identified as a target cluster (i.e., an abnormal cluster) by using the resource receiving data and the resource spending data of each subject, so as to consider the complete resource transfer process or resource flow direction of abnormal resources (e.g., certain resources from the second resource system transferred into the third resource system via the first resource system) in the first resource system; moreover, the resource transfer level of each subject is calculated based on the complete resource transfer process, and the resource transfer level of each subject is taken as the identification basis for identifying the target cluster. Such an identification apparatus not only considers the transaction characteristics of each subject itself, but also takes into account the resource transfer relationship and resource flow characteristics between different subjects, thereby overcoming the one-sidedness and limitations of related technologies; meanwhile, since the identification apparatus only involves a subject cluster participating in the complete resource transfer process in the first resource system, it can effectively exclude the interference of irrelevant subjects (e.g., normal subjects engaged in normal resource transfer activities) or the interference of normal resource transfer processes, thereby significantly improving the accuracy of identifying the cluster.

[0139] Figure 7 FIG. 1 illustrates an example system 700 that includes an example computing device 710 that is representative of one or more systems and / or devices that can implement the various methods described herein. The computing device 710 can be, for example, a server of a service provider, a device associated with the server, a system-on-chip, and / or any other suitable computing device or computing system. Figure 6 The cluster identification apparatus 600 described above can take the form of the computing device 710. Alternatively, the cluster identification apparatus 600 can be implemented as a computer program in the form of the application 716.

[0140] The example computing device 710 as illustrated includes a processing system 711, one or more computer readable medium 712, and one or more input / output interface 713 that are communicatively coupled via one or more bus 714. Although not illustrated, the computing device 710 can further include a system bus or other data and command transfer system that couples the various components within the computing device 710. A system bus can include any one or combination of different bus structures, such as a memory bus or memory controller, a peripheral bus, a serial bus, a parallel bus, and / or a local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Advanced Graphics Port (AGP) bus, and / or Video Electronics Standards Association (VESA) local bus.

[0141] The processing system 711 is representative of the functionality performed by a processor as software executes an operation. As such, the processing system 711 is illustrated as including hardware elements 714 that can be configured to perform a processor, functionality, etc. This can include implementation in hardware as an application specific integrated circuit or other logic device formed using one or more semiconductors. The hardware elements 714 are not limited by the materials from which they are formed or the processing mechanisms employed therein. For example, a processor can be comprised of semiconductor(s) and / or transistors (e.g., electronic integrated circuits (ICs)). In such a context, a processor-executable instruction can be an electronic instruction, etc.

[0142] The computer-readable medium 712 is illustrated as including memory / storage 715. The memory / storage 715 represents the memory / storage capacity associated with one or more computer-readable media. The memory / storage 715 can include volatile media (such as random access memory (RAM)) and / or nonvolatile media (such as read only memory (ROM), flash memory, optical disks, magnetic disks, and so forth). The memory / storage 715 can include fixed and removable media, where appropriate. The computer- readable medium 712 can be configured in a variety of other ways as further described below.

[0143] The input and output devices 713 allow a user to enter commands and information to computing device 710, and optionally also allow for the presentation of information from computing device 710 to the user and / or other components or devices using various input / output devices 713. Input devices include but are not limited to keyboard, pointing devices such as a mouse, microphone (e.g., for voice inputs), scanner, touch functionality (e.g., capacitive or other sensors that are configured to detect physical touch), camera (e.g., which can employ visible or non-visible wavelengths such as infrared frequencies to facilitate movement detection that does not involve touch), and so forth. Output devices include but are not limited to display devices, speaker, printer, network card, touch functionality (e.g., capacitive or other sensors that are configured to detect physical touch), and so forth. Thus, the computing device 1610 can be configured in a variety of ways as further described below to support user interaction.

[0144] The computing device 710 also includes applications 716. The applications 716 can be, for example, software instances of the cluster identification apparatus 600 and, in combination with other elements in the computing device 710, implement the techniques described herein.

[0145] The present application provides a computer program product including computer instructions stored in a computer readable storage medium. A processor of a computing device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computing device to perform the cluster identification method provided in various optional implementations described above.

[0146] Various techniques can be described in the general context of software hardware elements or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, and the like that perform particular tasks or implement particular abstract data types. As used herein, the terms "module", "functionality" and "component" generally refer to software, firmware, hardware, or a combination thereof. The features of the techniques described herein are platform-independent, meaning that the techniques can be implemented on a variety of computing platforms having a variety of processors.

[0147] Implementations of the described modules and techniques can be stored or transmitted across some form of computer readable media. Computer readable media can include various media that can be accessed by the computing device 710. By way of example, and not limitation, computer readable media can include "computer readable storage media" and "computer readable signal media".

[0148] In contrast to signal bearing media, "computer readable storage media" refers to media or means that persistently store information and / or have tangible physical form. Thus, computer readable storage media refers to non-signal bearing media. Computer readable storage media includes volatile and non-volatile, removable and non-removable media implemented in a method or technology for storage of information such as computer readable instructions, data structures, program modules, logical elements / circuits, or other data. Examples of computer readable storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, hard disks, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture that are appropriate for storage of desired information and that can be accessed by a computer.

[0149] A "computer-readable signal medium" can refer to a signal-bearing medium that is configured to transmit program code, such that the program code is embodied by the signal medium for

[0150] As previously described, hardware elements 714 and computer-readable media 712 are representative of instructions, modules, programmable device logic and / or fixed device logic implemented in the hardware. Hardware elements can include integrated circuits or

[0151] Combinations of the foregoing can also be employed to implement various techniques and modules described herein. Accordingly, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic embodied on some form of computer-readable storage media and / or by one or more hardware elements 714. The computing device 710 can be configured to implement particular instructions and / or functions corresponding to the software and / or hardware modules. Accordingly, implementation of a module that is executable by the computing device 710 as software can be achieved at least partially in hardware, e.g., through the use of computer-readable storage media and / or hardware elements 714 of the processing system. The

[0152] In various implementations, the computing device 710 can assume a variety of different configurations. For example, the computing device 710 can be implemented as a computer-class device comprising a personal computer, desktop computer, a multi-screen computer, laptop computer, netbook, etc. The computing device 710 can also be implemented as a mobile device-class device comprising a mobile phone, a portable music player, a portable gaming device, a tablet computer, a multi-screen computer, etc. The computing device 710 can also be implemented as a television-class device comprising a device having or connected to a generally larger screen in a casual viewing environment. These devices include televisions, set-top boxes, game consoles, etc.

[0153] The technology described herein can be supported by these various configurations of the computing device 710 and is not limited to the specific examples of the technology described herein. Functionality can also be implemented all or in part within the platform 722, as described below, of the “cloud” 720.

[0154] The cloud 720 includes and / or is representative of the platform 722 for resources 724. The platform 722 abstracts underlying functionality of hardware (e.g., servers) and software resources of the cloud 720. The resources 724 can include applications and / or data that can be utilized while computer processing is executed on servers that are remote from the computing device 710. Resources 724 can also include services provided over the Internet and / or a subscriber network, such as a cellular or Wi-Fi network.

[0155] The platform 722 can abstract resources and functions to connect the computing device 710 with other computing devices. The platform 722 can also serve to abstract scaling of resources to provide a corresponding level of scale to encountered demand for the resources 724 that are implemented via the platform 722. Accordingly, in an interconnected device embodiment, implementation of functionality described herein can be distributed throughout the system 700. For example, the functionality can be implemented in part on the computing device 710 as well as via the platform 722 that abstracts the functionality of the cloud 720.

[0156] It should be understood that the embodiments of the present application described herein are described in reference to different functional units in order to more particularly emphasize the functions of various aspects of the present application. It can be apparent, however, that exemplary embodiments of the present application can be implemented such that each function is not necessarily performed by a separate unit, and that various functions can be performed by the same unit. For example, functions described as being performed by a single unit can be performed by multiple units. Thus, references to specific functional units are only to facilitate understanding of embodiments of the present application, and should not be considered as strictly limiting of the scope of the present application. Accordingly, the present application can be embodied in a single unit or multiple units, and any arrangement of these units can be possible.

[0157] Although the present application has been described in connection with some embodiments, it is not intended to be limited to the particular form set forth herein. Rather, the scope of the present application is limited only by the claims. Additionally, although individual features can be included in different claims, these can possibly be advantageously combined, and the inclusion of different claims does not imply that a combination of features is not feasible and / or advantageous. The order of the features in the claims does not imply any specific order of working of the features. Furthermore, in the claims the word "comprising" does not exclude other elements and the

[0158] It can be understood that in the detailed description of the present application, the resource transfer data (e.g. including resource receiving data and resource spending data) related to the subject is involved. When the embodiments of the present application involving such data are applied to specific products or technologies, the consent or agreement of the user needs to be obtained, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the country and region.

Claims

1. A cluster identification method, comprising: obtaining resource transfer data of a plurality of subjects in a first resource system, wherein the resource transfer data comprises resource receiving data and resource spending data of each subject in the plurality of subjects; determining candidate subjects from the plurality of subjects based on the resource transfer data of the plurality of subjects, wherein the resource receiving data of the candidate subjects is related to a second resource system different from the first resource system, and the resource spending data of the candidate subjects is related to a third resource system different from the first resource system, wherein the second resource system is used to spend resources to the candidate subjects in the first resource system, and the third resource system is used to receive resources from the candidate subjects in the first resource system; determining a resource transfer level of each subject in the candidate subjects according to the resource receiving data and the resource spending data of the candidate subjects; identifying a target cluster from the candidate subjects according to at least the resource transfer level of each subject in the candidate subjects, wherein the target cluster comprises at least one candidate subject involved in a resource transfer process from the second resource system via the first resource system to the third resource system, wherein the identifying the target cluster from the candidate subjects according to at least the resource transfer level of each subject in the candidate subjects comprises: defining directed edges between different nodes based on resource transfer relationships between the nodes, and establishing a resource transfer multi-partite directed graph according to the nodes and the directed edges between the nodes, wherein the resource transfer multi-partite directed graph comprises a plurality of partite graphs, nodes in a same partite graph have a same level, and the level of the nodes indicates a degree of relevance of the nodes to the second resource system or the third resource system; determining an attribute value of each node in the resource transfer multi-partite directed graph according to at least the resource transfer level and the level of each node in the resource transfer multi-partite directed graph, wherein the attribute value of each node indicates a possibility of the node belonging to the target cluster; identifying the target cluster from the nodes of the resource transfer multi-partite directed graph based on the attribute value of each node in the resource transfer multi-partite directed graph. 2.The method of claim 1, wherein the resource transfer level of each node comprises a resource transfer retention value of each node, and the resource transfer retention value of each node represents an absolute value of a difference between a total amount of resource spending and a total amount of resource receiving of the node. 3.The method of claim 2, wherein the resource transfer level of each node further comprises a resource spending-receiving minimum value of each node, and the resource spending-receiving minimum value of each node represents a smaller one of the total amount of resource spending and the total amount of resource receiving of the node. 4.The method of claim 3, wherein the level of the nodes comprises a first type level indicating a degree of relevance of the nodes to the second resource system and a second type level indicating a degree of relevance of the nodes to the third resource system, and the attribute value of each node in the resource transfer multi-partite directed graph satisfies at least one of the following conditions: negatively correlated with the resource transfer retention value of the node. is negatively correlated with the smaller of a first type level and a second type level of the node; and is positively correlated with a resource balance minimum value of the node. 5.The method of claim 1, wherein the levels of the node comprise a first type level indicating a degree of relevance of the node to a second resource system and a second type level indicating a degree of relevance of the node to a third resource system, and determining the attribute value of each node in the resource transfer multi-partite graph based at least on the resource transfer levels and the levels of the nodes in the resource transfer multi-partite graph, comprises: for each node in the resource transfer multi-partite graph, determining the attribute value of the node using the following formula: wherein represents the attribute value of node i in the resource transfer multi-partite graph S, represents the smaller one of the total resource expenditure and the total resource reception of node i, represents the larger one of the total resource expenditure and the total resource reception of node i, represents the smaller one of the first type level and the second type level of node i, and a is a preset parameter greater than or equal to 0 and less than 1. 6.The method of claim 1, wherein identifying the target cluster from the nodes in the resource transfer multi-partite graph based on the attribute values of the nodes in the resource transfer multi-partite graph, comprises: taking a set of nodes in the resource transfer multi-partite graph as a current candidate cluster, sequentially performing the following steps in an iterative manner to obtain a set of candidate clusters: an iteration end determining step: in response to the existence of a node-less subgraph in the resource transfer multi-partite graph corresponding to the current candidate cluster, ending the iteration; a cluster feature value calculating step: based on the attribute values of the nodes in the current candidate cluster, calculating a cluster feature value of the current candidate cluster, and taking the current candidate cluster as a candidate cluster in the set of candidate clusters; a current candidate cluster updating step: removing the node with the lowest attribute value from the current candidate cluster and updating the attribute values of the neighboring nodes of the removed node to update the current candidate cluster, and proceeding to the iteration end determining step, identifying the target cluster from the set of candidate clusters based on the cluster feature values of the candidate clusters in the set of candidate clusters. 7.The method of claim 1, wherein the obtaining the resource transfer data of the plurality of subjects in the first resource system comprises: obtaining the resource transfer data of the plurality of subjects within a preset time period. 8.The method of claim 7, wherein the length of the preset time period is greater than or equal to 3 hours.

9. The method of claim 1, wherein, determining the candidate subject from the plurality of subjects based on the resource transfer data of the plurality of subjects, comprises: based on the resource receiving data of the plurality of subjects, screening a first subject set from the plurality of subjects that has a resource transfer relationship with the second resource system; based on the resource expenditure data of the plurality of subjects, screening a second subject set from the plurality of subjects that has a resource transfer relationship with the third resource system; determining the candidate subject according to the intersection of the first subject set and the second subject set. 10.The method of claim 9, wherein based on the resource receiving data of the plurality of subjects, screening the first subject set from the plurality of subjects that has a resource transfer relationship with the second resource system, comprises: based on the resource receiving data of the plurality of subjects, extracting direct receiving subjects from the plurality of subjects that have a direct resource transfer relationship with the second resource system; extract, from the plurality of subjects, an indirect receiving subject having an indirect resource transfer relationship with the second resource system based on the resource receiving data of the direct receiving subject; determine the first subject set based on the direct receiving subject and the indirect receiving subject, and / or the filtering, from the plurality of subjects, the second subject set having a resource transfer relationship with the third resource system based on the resource transfer data of the plurality of subjects, comprises: extract, from the plurality of subjects, a direct spending subject having a direct resource transfer relationship with the third resource system based on the resource spending data of the plurality of subjects; extract, from the plurality of subjects, an indirect spending subject having an indirect resource transfer relationship with the third resource system based on the resource spending data of the direct spending subject; determine the second subject set based on the direct spending subject and the indirect spending subject.

11. The method of claim 1, wherein the identifying, from the candidate subjects, the target cluster based on at least the resource transfer level of each subject in the candidate subjects, comprises: calculating, based on the resource receiving data and the resource spending data of the candidate subjects, a total amount of resource transfer between pairs of subjects in the candidate subjects having a direct resource transfer relationship; in response to the total amount of resource transfer between the pairs of subjects being less than a resource transfer threshold, deleting the pairs of subjects from the candidate subjects to obtain updated candidate subjects; identifying, from the updated candidate subjects, the target cluster based on at least the resource transfer level of each subject in the updated candidate subjects.

12. A cluster identification apparatus, comprising: an obtaining module configured to obtain resource transfer data of a plurality of subjects in a first resource system, wherein the resource transfer data comprises resource receiving data and resource spending data of each subject in the plurality of subjects; a first determining module configured to determine, from the plurality of subjects, a candidate subject based on the resource transfer data of the plurality of subjects, the resource receiving data of the candidate subject being related to a second resource system different from the first resource system, and the resource spending data of the candidate subject being related to a third resource system different from the first resource system, wherein the second resource system is used to spend resources to the candidate subject in the first resource system, and the third resource system is used to receive resources from the candidate subject in the first resource system; a second determining module configured to determine, based on the resource receiving data and the resource spending data of the candidate subject, a resource transfer level of each subject in the candidate subject; and a third determining module configured to identify, from the candidate subjects, a target cluster based on at least the resource transfer level of each subject in the candidate subjects. The identification module is configured to identify a target cluster from the candidate subjects according to at least a resource transfer level of each subject in the candidate subjects, the target cluster including at least one candidate subject involved in a resource transfer process from the second resource system to the third resource system via the first resource system, and wherein the identification of the target cluster from the candidate subjects according to at least the resource transfer level of each subject in the candidate subjects includes: taking each subject in the candidate subjects as a node, defining a directed edge between different nodes based on a resource transfer relationship between the nodes, and establishing a resource transfer multi-partite directed graph according to the nodes and the directed edges between the nodes, the resource transfer multi-partite directed graph including a plurality of partite graphs, the nodes in a same partite graph having a same level, and the level of the nodes indicating a relevance degree of the nodes to the second resource system or the third resource system; determining an attribute value of each node in the resource transfer multi-partite directed graph according to at least the resource transfer level and the level of each node in the resource transfer multi-partite directed graph, the attribute value of each node indicating a possibility of the node belonging to the target cluster; and identifying the target cluster from the nodes in the resource transfer multi-partite directed graph based on the attribute value of each node in the resource transfer multi-partite directed graph.

13. A computing device comprising: a memory and a processor, wherein the memory has stored therein a computer program that, when executed by the processor, causes the processor to perform the steps of the method of any one of claims 1-11.

14. A computer-readable storage medium having stored thereon computer-readable instructions that, when executed, implement the method of any one of claims 1-11.

15. A computer program product comprising computer instructions that, when executed by a processor, implement the steps of the method of any one of claims 1-11.

Citation Information

Patent Citations

  • Target information determination method, device and equipment

    CN111708897A

  • Method, device and equipment for determining abnormal resource transfer link

    CN113537960A