Abnormal tissue identification method, device, electronic device and medium
By constructing a knowledge graph to identify abnormal transaction links and calculate the average abnormality of the community, the problem of difficulty in identifying highly secretive abnormal organizations in existing technologies is solved, and efficient and accurate abnormal organization identification is achieved.
Patent Information
- Application Number
- CN202210732268.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2042-06-24
AI Technical Summary
Existing technologies are difficult to effectively identify highly secretive abnormal organizations. Traditional methods do not fully cover group behavior, and manual analysis of capital flows is inefficient and prone to errors.
By constructing a knowledge graph, generating communities based on transaction data, identifying abnormal transaction links, and calculating the average abnormality of the community, abnormal organizations are determined.
It achieves comprehensive and efficient identification of highly secretive abnormal tissues, improves identification accuracy and saves computing resources.
Smart Images

Figure CN115062163B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more specifically, to a method, device, electronic device, and medium for identifying abnormal tissue. Background Art
[0002] With the development of network technology, fraudulent transactions disguised through technical means are virtually indistinguishable from normal transactions by ordinary customers, making them extremely secretive and difficult to detect. Existing techniques include analyzing historical fraudulent activity data to develop multiple sets of rules and identifying them accordingly. Another approach involves analyzing fund flow relationships, using special account information provided by financial institutions, and then manually screening transactions on the trading network. Summary of the Invention
[0003] In view of this, the present disclosure provides a comprehensive, accurate, efficient and resource-saving abnormal tissue identification method, device, electronic device and computer-readable storage medium.
[0004] One aspect of the present disclosure provides a method for identifying abnormal organizations, comprising: constructing a knowledge graph based on acquired transaction data of a historical period, the transaction data including account information and transaction information between account information, the nodes of the knowledge graph being constructed based on the account information, and the edges between the nodes being constructed based on the transaction information; dividing the nodes of the knowledge graph based on the degree of association of the nodes to generate m communities, where m is an integer greater than or equal to 1; determining abnormal transaction links from the m communities of the knowledge graph based on abnormal account information, the number of communities in which the nodes in the abnormal transaction links belong being less than or equal to 2, and the abnormal account information being acquired based on preset rules; acquiring data features of each node in the abnormal transaction link, calculating an average abnormality degree of each community based on the data features; and determining a community whose average abnormality degree meets a set threshold as an abnormal organization.
[0005] According to the method for identifying abnormal organizations according to the embodiment of the present disclosure, based on the knowledge graph, abnormal transaction links are determined from m communities, data features of each node in the abnormal transaction links are obtained, the average abnormality of each community is calculated based on the data features, and the community whose average abnormality meets the set threshold is determined as an abnormal organization, so that abnormal transaction links can be easily mined from massive transaction data, and then abnormal organizations can be identified. The identification method of the present disclosure has comprehensive coverage and can efficiently and accurately find highly secretive abnormal organizations. The number of communities in which the nodes in the abnormal transaction links of the present disclosure are located is less than or equal to 2, that is, the method for determining abnormal transaction links does not allow crossing communities, which can prevent the association of invalid groups, further improve the accuracy of abnormal organization identification, and save computing resources at the same time.
[0006] In some embodiments, the nodes of the knowledge graph are divided according to the degree of association of the nodes to generate m communities, including: determining each node in the knowledge graph as a group; traversing each group to determine the intimacy between the group and each group having an edge relationship with it; when the intimacy meets the intimacy threshold, merging the group with the group having an edge relationship with it according to the intimacy; taking the merged group as a new group, repeating the traversal of each group to determine the intimacy between the group and each group having an edge relationship with it; when the intimacy does not meet the intimacy threshold, stopping merging the group with the group having an edge relationship with it; and when the intimacy between each two groups does not meet the intimacy threshold, treating the current m groups as m communities.
[0007] In some embodiments, merging the group and the group having an edge relationship with the group based on the intimacy includes: sorting the intimacy according to numerical values; and merging the two groups ranked first or last in intimacy based on the sorting result.
[0008] In some embodiments, determining an abnormal transaction link from the m communities of the knowledge graph based on abnormal account information includes: determining a first abnormal node in the knowledge graph according to the preset rules; determining a directed connected link in which the first abnormal node in the knowledge graph is located based on the first abnormal node, wherein the directed connected link is a link formed by connecting nodes through directed edges; and taking at least a portion of the link consisting of nodes in two adjacent communities of the directed connected link as an abnormal transaction link, wherein the abnormal transaction link includes the first abnormal node.
[0009] In some embodiments, obtaining the data features of each node in the abnormal transaction link includes: determining node attributes of each node in the abnormal transaction link; and obtaining the data features according to the node attributes.
[0010] In some embodiments, calculating the average abnormality of each community based on the data features includes: constructing a feature vector based on the data features; determining a point abnormality based on the feature vector; and calculating the average abnormality of each community based on the point abnormality.
[0011] In some embodiments, determining the point abnormality based on the feature vector includes: calculating the Euclidean distance based on the feature vector and a preset standard vector, wherein the Euclidean distance is used to measure the similarity between the feature vector and the standard vector; and determining the point abnormality based on the Euclidean distance, wherein the point abnormality is proportional to the similarity.
[0012] In some embodiments, calculating the average abnormality of each community based on the point abnormality includes: calculating the average of the point abnormality of the nodes in the abnormal transaction link included in each community; and using the average as the average abnormality of the community.
[0013] Another aspect of the present disclosure provides an apparatus for identifying abnormal organizations, comprising: a construction module configured to construct a knowledge graph based on acquired transaction data from a historical period, the transaction data including account information and transaction information between the account information, the nodes of the knowledge graph being constructed based on the account information, and the edges between the nodes being constructed based on the transaction information; a generation module configured to partition the nodes of the knowledge graph based on the degree of association of the nodes to generate m communities, where m is an integer greater than or equal to 1; a first determination module configured to determine abnormal transaction links from the m communities of the knowledge graph based on abnormal account information, the number of communities in which the nodes in the abnormal transaction links belong being less than or equal to 2, and the abnormal account information being acquired based on preset rules; a calculation module configured to obtain data features of each node in the abnormal transaction link and calculate an average abnormality degree of each community based on the data features; and a second determination module configured to determine communities whose average abnormality degree meets a set threshold as abnormal organizations.
[0014] Another aspect of the present disclosure provides an electronic device, comprising one or more processors and one or more memories, wherein the memories are used to store executable instructions, and when the executable instructions are executed by the processors, implement the above method.
[0015] Another aspect of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method described above when executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0017] Figure 1 Schematically illustrates an exemplary system architecture to which the method and apparatus according to an embodiment of the present disclosure may be applied;
[0018] Figure 2 A flowchart schematically illustrates a method for identifying abnormal tissue according to an embodiment of the present disclosure;
[0019] Figure 3Schematically shows a flow chart for dividing nodes of a knowledge graph according to the degree of association of the nodes to generate m communities according to an embodiment of the present disclosure;
[0020] Figure 4 A schematic diagram of a knowledge graph according to an embodiment of the present disclosure is schematically shown;
[0021] Figure 5 Schematically shows a flow chart of merging the group and the groups having edge relationships therewith according to intimacy according to an embodiment of the present disclosure;
[0022] Figure 6 A flowchart of determining abnormal transaction links from m communities in a knowledge graph based on abnormal account information according to an embodiment of the present disclosure is schematically shown;
[0023] Figure 7 A flowchart for obtaining data features of each node in an abnormal transaction link according to an embodiment of the present disclosure is schematically shown;
[0024] Figure 8 A flowchart of calculating the average abnormality of each community based on data features according to an embodiment of the present disclosure is schematically shown;
[0025] Figure 9 Schematically shows a flow chart of determining point abnormality based on a feature vector according to an embodiment of the present disclosure;
[0026] Figure 10 Schematically shows a flow chart of calculating the average abnormality of each community based on point abnormality according to an embodiment of the present disclosure;
[0027] Figure 11 Schematically shows a schematic diagram of account fund flow according to an embodiment of the present disclosure;
[0028] Figure 12 A flowchart schematically illustrates a method for identifying abnormal tissue according to an embodiment of the present disclosure;
[0029] Figure 13 A schematic diagram schematically illustrates two accounts and edge attributes according to an embodiment of the present disclosure;
[0030] Figure 14 A schematic diagram of a knowledge graph according to an embodiment of the present disclosure is schematically shown;
[0031] Figure 15 The following schematically shows a structural block diagram of an abnormal tissue identification device according to an embodiment of the present disclosure;
[0032] Figure 16 Schematically shows a structural block diagram of a generation module according to an embodiment of the present disclosure;
[0033] Figure 17 Schematically shows a structural block diagram of a merging unit according to an embodiment of the present disclosure;
[0034] Figure 18 Schematically shows a structural block diagram of a first determination module according to an embodiment of the present disclosure;
[0035] Figure 19 Schematically shows a structural block diagram of a computing module according to an embodiment of the present disclosure;
[0036] Figure 20 Schematically shows a structural block diagram of a computing module according to an embodiment of the present disclosure;
[0037] Figure 21 Schematically shows a structural block diagram of an eighth determining unit according to an embodiment of the present disclosure;
[0038] Figure 22 Schematically shows a structural block diagram of a computing unit according to an embodiment of the present disclosure;
[0039] Figure 23 The block diagram of an electronic device according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0040] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0041] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations, adopt necessary confidentiality measures, and do not violate public order and good morals. In the technical solutions disclosed herein, the acquisition, collection, storage, use, processing, transmission, provision, disclosure, and application of data all comply with the provisions of relevant laws and regulations, adopt necessary confidentiality measures, and do not violate public order and good morals.
[0042] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0043] When using expressions such as "at least one of A, B, or C," they should generally be interpreted in accordance with the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but is not limited to systems having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.). The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly specifying the number of the indicated technical features. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more of the aforementioned features.
[0044] With the development of network technology, fraudulent transactions disguised through technical means are virtually indistinguishable from normal transactions by ordinary customers, making them extremely secretive and difficult to detect. Existing techniques include analyzing historical fraudulent activity data to develop multiple sets of rules and identifying them accordingly. Another approach involves analyzing fund flow relationships, using special account information provided by financial institutions, and then manually screening transactions on the trading network.
[0045] However, when identifying improper behavior based on rules, the rules are difficult to characterize for customers, have incomplete coverage of group behavior, and cannot find group members with strong secretiveness; while manual analysis of fund flows has complex transaction relationships and huge amounts of data, which is prone to errors, omissions, and inefficiency.
[0046] The embodiments of the present disclosure provide a method, device, electronic device, computer-readable storage medium, and computer program product for identifying abnormal organizations. The method for identifying abnormal organizations includes: constructing a knowledge graph based on transaction data acquired during a historical period, the transaction data including account information and transaction information between account information, the nodes of the knowledge graph being constructed based on the account information, and the edges between nodes being constructed based on the transaction information; dividing the nodes of the knowledge graph based on the degree of association of the nodes to generate m communities, where m is an integer greater than or equal to 1; determining abnormal transaction links from the m communities of the knowledge graph based on abnormal account information, the number of communities in which the nodes in the abnormal transaction links are located being less than or equal to 2, and the abnormal account information being obtained based on preset rules; obtaining data features of each node in the abnormal transaction link, calculating the average abnormality of each community based on the data features; and determining the community whose average abnormality meets a set threshold as an abnormal organization.
[0047] It should be noted that the abnormal tissue identification method, device, electronic device, computer-readable storage medium and computer program product disclosed herein can be used in the field of artificial intelligence technology, and can also be used in any field other than the field of artificial intelligence technology, such as the financial field. The field of the present disclosure is not limited here.
[0048] Figure 1 The following schematically illustrates an exemplary system architecture 100 to which the abnormal tissue identification method, apparatus, electronic device, computer-readable storage medium, and computer program product according to an embodiment of the present disclosure can be applied. Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.
[0049] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0050] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0051] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0052] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0053] It should be noted that the abnormal tissue identification method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the abnormal tissue identification device provided in the embodiment of the present disclosure can generally be set in the server 105. The abnormal tissue identification method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the abnormal tissue identification device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0054] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0055] The following will be based on Figure 1 The scene described by Figures 2 to 10 The abnormal tissue identification method according to the embodiment of the present disclosure is described in detail.
[0056] Figure 2 The flowchart of the abnormal tissue identification method according to the embodiment of the present disclosure is schematically shown.
[0057] like Figure 2 As shown, the abnormal tissue identification method of this embodiment includes operations S210 to S250.
[0058] In operation S210, a knowledge graph is constructed based on the acquired transaction data of the historical period. The transaction data includes account information and transaction information between account information. The nodes of the knowledge graph are constructed based on the account information, and the edges between nodes are constructed based on the transaction information.
[0059] It is understood that historical transaction data can be obtained from the business data system. Transaction data includes account information, including account names. Transaction information can also include the transaction initiator's and recipient's account names. Knowledge graph nodes can be constructed based on the account names, and knowledge graph edges can be constructed based on the relationship between the transaction initiator's and recipient's account names in the transaction information.
[0060] In operation S220 , the nodes of the knowledge graph are divided according to the degree of association of the nodes to generate m communities, where m is an integer greater than or equal to 1.
[0061] As an implementable approach, Figure 3As shown, operation S220 divides the nodes of the knowledge graph according to the degree of association of the nodes to generate m communities, including operations S221 to S226.
[0062] In operation S221 , each node in the knowledge graph is determined as a group, thereby initializing the knowledge graph.
[0063] In operation S222, each group is traversed to determine the intimacy between the group and each group with which it has an edge relationship. Figure 4 For example, there are 12 nodes in the knowledge graph, namely nodes a, b, c, d, e, f, g, h, i, j, k, and l. When initializing the knowledge graph, a, b, c, d, e, f, g, h, i, j, k, and l can be treated as a group. The following takes group a as an example to illustrate traversing each group and determining the intimacy between the group and each group with which it has an edge relationship. Groups b, c, and d have edge relationships with group a. Therefore, the intimacy between a and b, the intimacy between a and c, and the intimacy between a and d can be calculated respectively.
[0064] In operation S223 , when the intimacy satisfies an intimacy threshold, the group and the group having an edge relationship therewith are merged according to the intimacy, wherein the intimacy threshold is a standard threshold set as needed.
[0065] In some specific examples, such as Figure 5 As shown, operation S223 merges the group and the group having an edge relationship with it according to the intimacy, including operation S2231 and operation S2232.
[0066] In operation S2231, the intimacy is sorted according to the numerical value.
[0067] In operation S2232, based on the sorting results, the two groups ranked first or last in intimacy are merged. It should be noted that sorting intimacy by numerical value can be done by sorting intimacy in ascending order or in descending order. When sorting intimacy in ascending order, the two groups ranked last in intimacy are merged; when sorting intimacy in descending order, the two groups ranked first in intimacy are merged.
[0068] For example, through operation S222, the intimacy between a and b is calculated as K1, the intimacy between a and c is K2, and the intimacy between a and d is K3. When the order is K1 < K2 < K3, the groups a and d corresponding to K3, which is ranked last, are merged. When the order is K3 > K2 > K1, the groups a and d corresponding to K3, which is ranked first, are merged. Operations S2231 and S2232 facilitate merging the group with the groups that have edge relationships with it based on intimacy.
[0069] In operation S224 , the merged group is treated as a new group, and the process of traversing each group is repeated to determine the intimacy between the group and each group having an edge relationship with the group.
[0070] In operation S225, when the intimacy does not meet the intimacy threshold, the group is stopped from being merged with the group with which it has an edge relationship. For example, the intimacy threshold can be set to a positive number. When the intimacy is negative, the intimacy does not meet the intimacy threshold, and the group is stopped from being merged with the group with which it has an edge relationship.
[0071] In operation S226, when the intimacy between any two groups does not meet the intimacy threshold, the current m groups are treated as m communities. It is understandable that when no two groups in all groups in the knowledge graph can be merged, the community generation is completed, and the current m groups can be treated as m communities.
[0072] Among them, intimacy can be expressed as Q, and the value of intimacy can be obtained by formula (1).
[0073]
[0074] Where m represents the number of edges between the group being traversed and other groups, ki,in represents the sum of the weights of the edges from the group being traversed to the target group, ∑ tot represents the total weight of the edges incident to the target group, and ki represents the total weight of the edges of the group being traversed.
[0075] Therefore, through operations S221 to S226, it is easy to divide the nodes of the knowledge graph according to the degree of association of the nodes and generate m communities.
[0076] In operation S230, an abnormal transaction link is determined from m communities of the knowledge graph according to the abnormal account information. The number of communities where the nodes in the abnormal transaction link belong is less than or equal to 2. The abnormal account information is obtained based on preset rules.
[0077] It is understandable that since the members who engage in improper financial activities are financially connected and transfer money to each other frequently, their capital flows are mainly composed of three parts, namely upstream, midstream and downstream. Among them, the upstream is mainly used to absorb funds, with multiple accounts transferring in and out in a dispersed manner or in large amounts within a short period of time, basically leaving no balance, and having obvious characteristics of fund collection; the midstream is mainly used to transfer upstream funds. The intermediary characteristics of this transaction are relatively strong, with the characteristics of fast in and out, multiple in and out, and relatively concentrated time; the downstream is mainly used to split the midstream funds into various small funds.
[0078] Therefore, the preset rules for judging whether account information is abnormal account information include absorbing funds, transferring them into multiple accounts in a short period of time, transferring them out in a dispersed manner or in large amounts, leaving basically no balance, and having obvious characteristics of fund collection, and the accounts that meet the abnormal account information are regarded as upstream nodes on the abnormal transaction chain; the preset rules for judging whether account information is abnormal account information include transferring upstream funds with strong intermediary characteristics, fast in and out, multiple in and out, and relatively concentrated time, and the accounts that meet the abnormal account information are regarded as midstream nodes on the abnormal transaction chain; the preset rules for judging whether account information is abnormal account information include splitting the midstream funds into various small amounts of funds, and the accounts that meet the abnormal account information are regarded as downstream nodes on the abnormal transaction chain.
[0079] As a possible way to achieve this, Figure 6 As shown, operation S230 determines abnormal transaction links from m communities in the knowledge graph according to abnormal account information, including operations S231 to S233.
[0080] In operation S231, a first abnormal node in the knowledge graph is determined according to a preset rule. It is understood that, for example, abnormal account information can be determined based on a preset rule that includes funds absorbed, dispersed inflows and outflows from multiple accounts within a short period of time, dispersed or large-amount transfers, essentially no balance, and significant fund collection characteristics, and accounts meeting this abnormal account information are designated as the first abnormal node in the knowledge graph. Abnormal account information can be determined based on a preset rule that includes funds transferred upstream, exhibiting strong intermediary characteristics, with rapid inflows and outflows, multiple inflows and multiple outflows, and relatively concentrated time periods, and accounts meeting this abnormal account information are designated as the first abnormal node in the knowledge graph. Abnormal account information can be determined based on a preset rule that includes funds transferred through midstream channels into various small amounts, and accounts meeting this abnormal account information are designated as the first abnormal node in the knowledge graph.
[0081] In operation S232, based on the first abnormal node, a directed connected link in the knowledge graph where the first abnormal node is located is determined, wherein the directed connected link is a link formed by connecting nodes through directed edges. Figure 4Assume that node a is determined as the first abnormal node through operation S231, and there are 6 directed connected links connected by directed edges to node a, namely: adfe; adfg-1-ki; adfglkj; adfgli; ab; acihj.
[0082] In operation S233, at least a portion of the directed connected link consisting of nodes in two adjacent communities is identified as an abnormal transaction link. The abnormal transaction link includes the first abnormal node. Specifically, in the directed connected link adfe, nodes a and d exist in community A, and nodes f and e exist in community B. Since communities A and B have an edge relationship, they are adjacent communities. Therefore, link adfe can be identified as an abnormal transaction link.
[0083] In the directed connected link adfglki, nodes a and d exist in community A, nodes f and g exist in community B, and nodes l, k, and i exist in community C. The first abnormal node a exists in community A. There is an edge relationship between communities A and B, so communities A and B are adjacent and cannot cross into community C. Therefore, link adfg can be identified as an abnormal transaction link. The method for identifying abnormal transaction links in the directed connected links adfglkj, adfgli, ab, and acihj is similar and will not be repeated here.
[0084] The above method of not allowing abnormal transaction links to be determined across communities can prevent the association of invalid groups, improve the accuracy of identifying abnormal organizations, and save computing resources.
[0085] Operations S231 to S233 can facilitate the determination of abnormal transaction links from m communities in the knowledge graph based on abnormal account information.
[0086] In operation S240 , data features of each node in the abnormal transaction link are obtained, and the average abnormality degree of each community is calculated based on the data features.
[0087] As a possible way to achieve this, Figure 7 As shown, operation S240 obtains data features of each node in the abnormal transaction link, including operation S241 and operation S242.
[0088] In operation S241, the node attributes of each node in the abnormal transaction link are determined. It is understood that transaction data includes account information. As a possible implementation, account information may also include at least one of the following: card number, account holder, account opening agent, account opening time, account opening branch, whether the account is opened in a different location, whether online banking is enabled, and account opening amount. Transaction information may also include at least one of the following: transaction initiator's card number, transaction amount, transaction time, transaction recipient's card number, transaction method, and transaction address.
[0089] Specifically, at least one of the following from the account information: card number, account holder, account opening agent, account opening time, account opening location, remote account opening, online banking availability, and account opening amount can be used as a node attribute; and at least one of the following from the transaction information: transaction initiator card number, transaction amount, transaction time, recipient card number, transaction method, and transaction address can be used as an edge attribute. After determining the abnormal transaction link, determining the node attributes of each node in the abnormal transaction link eliminates the need to obtain the account information of all nodes in the knowledge graph; only the account information of the nodes in the abnormal transaction link can be obtained. This saves computing resources and speeds up data processing.
[0090] In operation S242, data features are obtained based on the node attributes. For example, based on the attributes of the node and the attributes of the edge, it is possible to analyze whether the funds are transferred in in a dispersed manner, whether the funds are transferred out in a dispersed manner, whether the funds are quickly transferred in and out, whether the account enters a dormant period after opening, whether the account frequently crosses regions, whether it is cross-bank transactions, whether the time, location, and outlets for opening multiple bank cards are relatively concentrated, and whether the accounts are opened in different places. Of course, the data features are not limited to this. This is only an example and cannot be understood as a limitation of the present disclosure. Through operations S241 and S242, it is easy to obtain the data features of each node in the abnormal transaction link.
[0091] As a possible way to achieve this, Figure 8 As shown, operation S240 calculates the average abnormality of each community according to the data features, including operations S243 to S245.
[0092] In operation S243, a feature vector is constructed based on the data features. Examples of these features include whether funds are transferred in dispersedly, whether funds are transferred out dispersedly, whether funds are transferred in and out quickly, whether the account enters a dormant period after opening, whether the account frequently crosses regions, whether transactions occur across banks, whether multiple bank card accounts are opened at a relatively concentrated time, location, and branch, and whether the accounts are opened in different locations. These data features are represented in a structured form, with a 1 representing a condition if satisfied and a 0 representing an error if not.
[0093] Assume that the data characteristics of node a are dispersed transfer-in, dispersed transfer-out, fast inflow and outflow of funds, no dormant period after account opening, no frequent cross-regional account, no cross-bank transactions, non-concentrated time, location, and outlets for opening multiple bank accounts, and no out-of-town account opening. Dispersed transfer-in, dispersed transfer-out, and fast inflow and outflow of funds meet the conditions and are represented by 1. No dormant period after account opening, no frequent cross-regional account, no cross-bank transactions, non-concentrated time, location, and outlets for opening multiple bank accounts, and no out-of-town account opening do not meet the conditions and are represented by 0. Therefore, the feature vector θ(1, 1, 1, 0, 0, 0, 0, 0, 0) of node a is constructed.
[0094] In operation S244 , the point abnormality degree is determined based on the feature vector.
[0095] As a feasible way, Figure 9 As shown, operation S244 determines the point abnormality degree according to the feature vector, including operation S2441 and operation S2442.
[0096] In operation S2441, the Euclidean distance is calculated based on the feature vector and the preset standard vector, where the Euclidean distance is used to measure the similarity between the feature vector and the standard vector. It is understandable that the vector constructed when all the above data features meet the conditions can be used as the standard vector to obtain the standard vector β (1, 1, 1, 1, 1, 1, 1, 1, 1). The Euclidean distance is represented by D and can be calculated using formula (2).
[0097]
[0098] Wherein, i represents the number of the node in the abnormal transaction link.
[0099] In operation S2442, the point outlier is determined based on the Euclidean distance, where the point outlier is proportional to the similarity. Assuming that the proportionality coefficient is set to c, the point outlier can be cD. Operations S2441 and S2442 can facilitate the determination of the point outlier based on the feature vector.
[0100] In operation S245 , the average abnormality degree of each community is calculated based on the point abnormality degrees.
[0101] As a feasible way, Figure 10 As shown, operation S245 calculates the average abnormality of each community based on the point abnormality, including operation S2451 and operation S2452.
[0102] In operation S2451 , the average value of the point abnormality of the nodes in the abnormal transaction link included in each community is calculated.
[0103] In operation S2452 , the average value is used as the average abnormality of the community.
[0104] It is understandable that each community may contain multiple nodes on the abnormal transaction chain, refer to Figure 4 , combined with the abnormal transaction links determined in operation S233: adfe; adfg; ab and acihj, we can obtain the nodes a, b, c and d in the abnormal transaction link in community A, so the average abnormality degree of community A is the average value of the point abnormality degrees of nodes a, b, c and d; we can obtain the nodes f, e and g in the abnormal transaction link in community B, so the average abnormality degree of community B is the average value of the point abnormality degrees of nodes f, e and g; we can obtain the nodes i, h and j in the abnormal transaction link in community C, so the average abnormality degree of community C is the average value of the point abnormality degrees of nodes i, h and j.
[0105] Operations S2451 and S2452 can be used to calculate the average abnormality of each community based on the point abnormality. Operations S243 to S245 can be used to calculate the average abnormality of each community based on data features.
[0106] In operation S250 , a community whose average abnormality degree meets a set threshold is determined to be an abnormal organization.
[0107] According to the method for identifying abnormal organizations according to the embodiment of the present disclosure, based on the knowledge graph, abnormal transaction links are determined from m communities, data features of each node in the abnormal transaction links are obtained, the average abnormality of each community is calculated based on the data features, and the community whose average abnormality meets the set threshold is determined as an abnormal organization, so that abnormal transaction links can be easily mined from massive transaction data, and then abnormal organizations can be identified. The identification method of the present disclosure has comprehensive coverage and can efficiently and accurately find highly secretive abnormal organizations. The number of communities in which the nodes in the abnormal transaction links of the present disclosure are located is less than or equal to 2, that is, the method for determining abnormal transaction links does not allow crossing communities, which can prevent the association of invalid groups, further improve the accuracy of abnormal organization identification, and save computing resources at the same time.
[0108] Refer to the following Figure 11-14 The abnormal tissue identification method according to the embodiment of the present disclosure is described in detail. It is worth noting that the following description is only for illustrative purposes and is not intended to limit the present disclosure.
[0109] This paper proposes a method for identifying abnormal organizations based on knowledge graphs, which is suitable for discovering abnormal bank transaction fund transfers. Since abnormal accounts are related to each other, they are also related in terms of funds. The composition relationship of account groups is relatively complex, and transfers between them are frequent. Using these known fund flows can accurately identify abnormal organizations. The fund flow mainly consists of three parts: upstream, midstream, and downstream. Figure 11 As shown in the figure, upstream funds are primarily used to absorb funds, with dispersed transfers into and out of multiple accounts within a short period of time, or large transfers, leaving little balance and exhibiting a distinct characteristic of fund concentration. Midstream funds primarily transfer upstream funds, and this type of transaction has a strong intermediary nature, characterized by rapid inflows and outflows, multiple inflows and multiple outflows, and a relatively concentrated timeframe. Downstream funds are primarily split into various small amounts through midstream funds.
[0110] In view of the above ideas, the present invention finds individual abnormal accounts and then abnormal organizations through the flow of capital chains: first, the transaction flow data within a time window is preprocessed to construct a knowledge graph of capital flow; then the account group is divided into closely related groups through the Louvian algorithm, and then abnormal accounts are found according to historical rules and marked as important suspicious nodes. The transaction group discovery algorithm is used to find the upstream and downstream abnormal account sets closely related to them, and the abnormal account sets are judged as abnormal based on the degree of suspicion.
[0111] The knowledge graph-based abnormal tissue identification method disclosed herein includes:
[0112] (1) Constructing a knowledge graph of fund flows: After processing bank transaction data, a knowledge graph is constructed based on the transaction relationship between accounts and fund flows to analyze fund flows.
[0113] (2) Constructing transaction groups: Use the Louvian community discovery algorithm to divide account groups and find closely related account groups based on capital flows. The Louvian community discovery algorithm is based on modularity and can discover hierarchical community structures, making nodes within a community closely connected and nodes between communities as few as possible. Modularity is a metric for evaluating the quality of a community network.
[0114] (3) Abnormal transaction link identification: find abnormal nodes through historical rules, find the account nodes associated with the upstream, midstream and downstream in the knowledge graph, and form a complete transaction link. In order to avoid associating all transaction groups, at most one layer of groups is allowed to be crossed when constructing the transaction link.
[0115] (4) Abnormal account identification: Analyze the abnormal account set, calculate the suspiciousness of each account in the group, and calculate the average suspiciousness among the groups. Then sort the suspiciousness from high to low and select the group with the highest suspiciousness.
[0116] The flowchart of the abnormal tissue identification method is as follows Figure 12 shown.
[0117] 1. Data collection: Obtain bank transaction data within a certain time window and perform structured processing on the data, including account name, card number, transaction amount, transaction time, counterparty account name, counterparty card number, transaction method, IP address, MAC address, and transaction outlet, which are recorded as raw data, as shown in Table 1.
[0118] Table 1
[0119]
[0120] 2. Data cleaning: Clean dirty data and filter out invalid, incomplete, and failed transaction data.
[0121] 3. Build a knowledge graph of capital flows:
[0122] 1) Import transaction data into a graph database, with the transaction account and the counterparty's account as nodes, the transaction relationship as an edge, and the capital flow as the direction of the edge.
[0123] 2) Graph edge attribute construction: There are multiple transaction records between two accounts, so the feature array consisting of "transaction number, transaction amount, time, method, outlet, IP address, MAC address" is constructed as the edge attribute, such as Figure 13 Shown are two accounts d and f and edge attributes.
[0124] 4. Transaction group construction: After the knowledge graph is constructed, Figure 14 As shown in the figure, the Louvian algorithm is used to divide the nodes in the graph into different groups according to the degree of association. This is very helpful for identifying abnormal tissues. The steps of group division are as follows:
[0125] 1) Initialize, Figure 14 Each node in the network is considered as an independent group. The number of groups is the same as the number of nodes, and the weights of all edges are considered the same.
[0126] 2) Start transferring nodes between groups. For each node i, try to assign node i to the group where each of its neighboring nodes is located. Calculate the change in modularity before and after the assignment. The calculation formula is as follows:
[0127]
[0128] Where m represents the number of edges in the network, ki,in represents the sum of the weights of the incident groups C from node i, ∑ tot represents the total weight of the incident group C, and ki represents the total weight of the incident node i.
[0129] 3) Repeat 2) and continue to evaluate the node transfer between groups until the groups to which all nodes belong no longer change, that is, the node transfer between groups is completed.
[0130] 4) Reconstruct the graph, reconstruct all nodes in the same group into a new group, update the edge weights between nodes in the group to the ring weights of the new nodes, and update the edge weights between groups to the edge weights between the new nodes.
[0131] 5) Repeat 2) until the modularity of the entire graph no longer changes. The modularity calculation formula is as follows:
[0132]
[0133] Among them, ∑in represents the sum of the weights of the edges in group C, and ∑tot represents the sum of the weights of all edges connected to the nodes in group C. The constructed group is as follows Figure 14 As shown, each circle constitutes a group, and there may be connections between groups.
[0134] 5. Identification of abnormal transaction links:
[0135] 1) Find abnormal account information in the original data based on historical rules.
[0136] 2) Find the corresponding nodes in the knowledge graph through abnormal transaction records and abnormal transaction accounts, and mark them as important abnormal nodes.
[0137] 3) Based on the abnormal nodes, find the associated account nodes in the knowledge graph for the upstream, midstream, and downstream, and build a transaction chain node. The specific steps of the transaction chain node discovery algorithm are as follows:
[0138] a. Initialize all nodes and treat each node as a separate transaction chain.
[0139] b. Based on the abnormal node, find its inflow and outflow nodes. Based on the characteristics of the upstream funds, use the rules to determine whether it is an upstream node and mark it as an upstream node. If a risk-free account is encountered, remove it and add its inflow and outflow nodes to the transaction chain.
[0140] c. Repeat step 3) b for the new transaction chain node. If it is an upstream node, stop tracking its incoming nodes and only add its outgoing nodes until no incoming or outgoing nodes can be found for the new node.
[0141] d. If the upstream and downstream nodes directly associated with the abnormal node are not in the same group, the algorithm only allows one level of cross-grouping. For example, if abnormal node a is in group A, b and c are in group B, and d is in group D, and there is a transaction relationship of abcd, then a's transaction link is abc and is not allowed to cross to group D to prevent invalid group association.
[0142] 6.Account data acquisition:
[0143] 1) For the nodes in the abnormal transaction link in step 5, obtain the account opening data, including: account holder, account opening agent, account opening time, account opening branch, whether the account is opened in a different location, whether online banking is enabled, and account opening amount.
[0144] 2) Use the account opening data as the label of each account node and attach it to the node of the knowledge graph to facilitate data analysis.
[0145] 7. Fraudulent account identification:
[0146] 1) Based on the transaction link nodes, node attributes, edges, and edge attributes found on the knowledge graph, relevant data features are analyzed, including:
[0147] a. Account transaction behavior characteristics: whether funds are transferred in and out in a dispersed manner, whether funds are transferred in and out quickly.
[0148] b. Account behavior characteristics: whether the account enters a dormant period after opening, and whether the account frequently conducts cross-regional and cross-bank transactions.
[0149] c. Characteristics of inter-account linkage: whether the time, location, and branches of multiple bank card account openings are relatively concentrated, and whether the accounts are opened in different locations.
[0150] 2) Calculation of the abnormality level of fraudulent accounts:
[0151] a. Select the nine features of the account above to form a feature vector and express these features in a structured form. If the conditions are met, it is represented as 1, otherwise it is represented as 0. Set a standard feature vector β(1, 1, 1, 1, 1, 1, 1, 1, 1). For example, there is an abnormal account a. The behavioral data over a certain period is that it has the following abnormal behaviors: scattered transfers in and out, and fast in and fast out. Its feature vector is θ(1, 1, 1, 0, 0, 0, 0, 0, 0). The abnormality of account a is the Euclidean distance between vector β and θ, calculated as follows:
[0152]
[0153] Wherein, i represents the number of the node in the abnormal transaction link.
[0154] b. Use the Euclidean distance formula to calculate the similarity between each account's feature vector and the standard feature vector. The higher the similarity, the higher the abnormality of the node. Then calculate the average abnormality of each group node.
[0155] c. Sort the top N groups of nodes with the highest degree of abnormality from high to low, and finally visualize the abnormal group nodes through the knowledge graph.
[0156] The disclosed identification method can identify unusual team behavior from hidden actions by analyzing the flow of funds. Compared to the misconduct of a single individual, uncovering unusual organizations can more comprehensively uncover the entire abnormal chain, effectively enhancing the identification of unknown risks in financial activities and improving the risk management capabilities of financial institutions. Furthermore, the disclosed method visualizes complex transaction relationships through a knowledge graph, making risk analysis more intuitive for staff and improving work efficiency.
[0157] Based on the above abnormal tissue identification method, the present disclosure also provides an abnormal tissue identification device 10. Figure 15-Figure 22 The abnormal tissue identification device 10 is described in detail.
[0158] Figure 15 The structure block diagram of the abnormal tissue identification device 10 according to an embodiment of the present disclosure is schematically shown.
[0159] The abnormal tissue identification device 10 includes a construction module 1 , a generation module 2 , a first determination module 3 , a calculation module 4 and a second determination module 5 .
[0160] Construction module 1, construction module 1 is used to perform operation S210: construct a knowledge graph based on the acquired transaction data of the historical period, the transaction data includes account information and transaction information between account information, the nodes of the knowledge graph are constructed based on the account information, and the edges between nodes are constructed based on the transaction information.
[0161] Generation module 2, generation module 2 is used to perform operation S220: divide the nodes of the knowledge graph according to the degree of association of the nodes, and generate m communities, where m is an integer greater than or equal to 1.
[0162] The first determination module 3 is used to perform operation S230: determine an abnormal transaction link from m communities in the knowledge graph based on the abnormal account information, the number of communities where the nodes in the abnormal transaction link are located is less than or equal to 2, and the abnormal account information is obtained based on preset rules.
[0163] The calculation module 4 is used to perform operation S240: obtaining data features of each node in the abnormal transaction link, and calculating the average abnormality degree of each community based on the data features.
[0164] The second determining module 5 is used to perform operation S250: determining a community whose average abnormality degree meets a set threshold as an abnormal organization.
[0165] Figure 16The schematic diagram shows a structural block diagram of a generation module 2 according to an embodiment of the present disclosure. The generation module 2 includes a first determination unit 21, a second determination unit 22, a merging unit 23, a repeat execution unit 24, a termination unit 25 and a third determination unit 26.
[0166] The first determining unit 21 is used to determine each node in the knowledge graph as a group.
[0167] The second determining unit 22 is configured to traverse each group and determine the intimacy between the group and each group having an edge relationship with the group.
[0168] The merging unit 23 is configured to merge the group and the group having an edge relationship with the group according to the intimacy when the intimacy meets the intimacy threshold.
[0169] The repetitive execution unit 24 is used to treat the merged group as a new group, repeatedly traverse each group, and determine the intimacy between the group and each group having an edge relationship with the group.
[0170] The termination unit 25 is configured to stop merging the group with the group having an edge relationship therewith when the intimacy does not meet the intimacy threshold.
[0171] The third determining unit 26 is configured to determine the current m groups as m communities when the intimacy between any two groups does not satisfy the intimacy threshold.
[0172] Figure 17 The structure block diagram of the merging unit 23 according to an embodiment of the present disclosure is schematically shown. The merging unit 23 includes a sorting component 231 and a merging component 232.
[0173] The sorting component 231 is used to sort the intimacy according to the numerical value.
[0174] The merging component 232 is used to merge the two groups ranked first or last in terms of intimacy according to the sorting result.
[0175] Figure 18 The structure block diagram of the first determination module 3 according to an embodiment of the present disclosure is schematically shown. The first determination module 3 includes a fourth determination unit 31 , a fifth determination unit 32 and a sixth determination unit 33 .
[0176] The fourth determining unit 31 is used to determine the first abnormal node in the knowledge graph according to a preset rule.
[0177] The fifth determination unit 32 is used to determine the directed connected link where the first abnormal node in the knowledge graph is located based on the first abnormal node, wherein the directed connected link is a link formed by connecting nodes through directed edges.
[0178] The sixth determining unit 33 is configured to determine whether at least a portion of the directed connected link, which is formed by nodes in two adjacent communities, is an abnormal transaction link, wherein the abnormal transaction link includes the first abnormal node.
[0179] Figure 19 The structural block diagram of the calculation module 4 according to the embodiment of the present disclosure is schematically shown. The calculation module 4 includes a seventh determination unit 41 and an acquisition unit 42.
[0180] The seventh determining unit 41 is used to determine the node attributes of each node in the abnormal transaction link.
[0181] The acquisition unit 42 is used to acquire data features according to node attributes.
[0182] Figure 20 The structural block diagram of the calculation module 4 according to the embodiment of the present disclosure is schematically shown. The calculation module 4 includes a construction unit 43, an eighth determination unit 44 and a calculation unit 45.
[0183] The construction unit 43 is used to construct a feature vector according to data features.
[0184] The eighth determining unit 44 is configured to determine the point abnormality degree according to the feature vector.
[0185] The calculation unit 45 is used to calculate the average abnormality of each community based on the point abnormality.
[0186] Figure 21 The eighth determining unit 44 according to an embodiment of the present disclosure is schematically shown in FIG. The eighth determining unit 44 includes a first calculating element 441 and a first determining element 442 .
[0187] The first calculation component 441 is used to calculate the Euclidean distance according to the feature vector and a preset standard vector, wherein the Euclidean distance is used to measure the similarity between the feature vector and the standard vector.
[0188] The first determining component 442 is used to determine the point abnormality according to the Euclidean distance, wherein the point abnormality is proportional to the similarity.
[0189] Figure 22The block diagram of the calculation unit 45 according to the embodiment of the present disclosure is schematically shown. The calculation unit 45 includes a second calculation element 451 and a second determination element 452.
[0190] The second calculation component 451 is used to calculate the average value of the point abnormality of the nodes in the abnormal transaction links included in each community.
[0191] The second determining component 452 is used to use the average value as the average abnormality of the community.
[0192] According to the abnormal organization identification device 10 of the embodiment of the present disclosure, based on the knowledge graph, abnormal transaction links are determined from m communities, data features of each node in the abnormal transaction link are obtained, the average abnormality of each community is calculated according to the data features, and the community whose average abnormality meets the set threshold is determined as an abnormal organization, so that abnormal transaction links can be easily mined from massive transaction data, and then abnormal organizations can be identified. The identification method of the present disclosure has comprehensive coverage and can efficiently and accurately find highly secretive abnormal organizations. The number of communities in which the nodes in the abnormal transaction link of the present disclosure are located is less than or equal to 2, that is, the determination method of the abnormal transaction link does not allow crossing communities, which can prevent the association of invalid groups, further improve the accuracy of abnormal organization identification, and save computing resources at the same time.
[0193] In addition, according to an embodiment of the present disclosure, any multiple modules among the construction module 1, the generation module 2, the first determination module 3, the calculation module 4, and the second determination module 5 can be combined into a single module for implementation, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module.
[0194] According to an embodiment of the present disclosure, at least one of the construction module 1, the generation module 2, the first determination module 3, the calculation module 4, and the second determination module 5 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in an appropriate combination of any of them.
[0195] Alternatively, at least one of the construction module 1, the generation module 2, the first determination module 3, the calculation module 4 and the second determination module 5 can be at least partially implemented as a computer program module, which can perform corresponding functions when executed.
[0196] Figure 23 A block diagram of an electronic device suitable for implementing the above method according to an embodiment of the present disclosure is schematically shown.
[0197] like Figure 23 As shown, the electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 into a random access memory (RAM) 903. The processor 901 may, for example, include a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include an onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0198] Various programs and data required for the operation of the electronic device 900 are stored in the RAM 903. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0199] According to an embodiment of the present disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the I / O interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage portion 908 including a hard disk; and a communication portion 909 including a network interface card such as a LAN card or a modem. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in the drive 910 as needed, so that a computer program read therefrom can be installed into the storage portion 908 as needed.
[0200] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0201] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 902 and / or RAM 903 described above and / or one or more memories other than ROM 902 and RAM 903.
[0202] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to cause the computer system to implement the method of the embodiments of the present disclosure.
[0203] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the processor 901 executes the computer program. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0204] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0205] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from a removable medium 911. When the computer program is executed by the processor 901, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0206] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0207] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0208] Those skilled in the art will appreciate that various combinations and / or combinations of features described in the various embodiments and / or claims of this disclosure may be made, even if such combinations or combinations are not explicitly described in this disclosure. In particular, various combinations and / or combinations of features described in the various embodiments and / or claims of this disclosure may be made, without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0209] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A method for identifying abnormal tissue, characterized in that: include: Constructing a knowledge graph based on the acquired transaction data for a historical period, wherein the transaction data includes account information and transaction information between account information, nodes of the knowledge graph are constructed based on the account information, and edges between nodes are constructed based on the transaction information; Divide the nodes of the knowledge graph according to the degree of association of the nodes to generate m communities, where m is an integer greater than or equal to 1; Determine an abnormal transaction link from the m communities of the knowledge graph based on the abnormal account information, where the number of communities in which the nodes in the abnormal transaction link belong is less than or equal to 2, and the abnormal account information is obtained based on preset rules; Obtaining data features of each node in the abnormal transaction link, and calculating the average abnormality degree of each community based on the data features; as well as Determine the community whose average abnormality degree meets the set threshold as an abnormal organization; The nodes of the knowledge graph are divided according to the degree of association of the nodes to generate m communities, including: Identify each node in the knowledge graph as a group; Traverse each group and determine the closeness between the group and each group with which it has an edge relationship; When the intimacy meets the intimacy threshold, merging the group with the group having an edge relationship with the group according to the intimacy; Treat the merged group as a new group, repeatedly perform the traversal of each group, and determine the intimacy between the group and each group having an edge relationship with it; When the intimacy does not meet the intimacy threshold, stop merging the group with the group having an edge relationship with it; and When the intimacy between any two groups does not meet the intimacy threshold, the current m groups are regarded as m communities; The step of determining abnormal transaction links from the m communities in the knowledge graph based on abnormal account information includes: Determine a first abnormal node in the knowledge graph according to the preset rule; Determining, based on the first abnormal node, a directed connected link in the knowledge graph where the first abnormal node is located, wherein the directed connected link is a link formed by connecting nodes through directed edges; and At least a portion of the directed connected link consisting of nodes in two adjacent communities is used as an abnormal transaction link, wherein the abnormal transaction link includes the first abnormal node; Calculating the average abnormality of each community based on the data features includes: Constructing a feature vector according to the data features; Determining the point abnormality based on the feature vector; and Based on the point abnormality, the average abnormality of each community is calculated; Wherein, determining the point abnormality degree according to the feature vector includes: Calculating a Euclidean distance based on the feature vector and a preset standard vector, wherein the Euclidean distance is used to measure the similarity between the feature vector and the standard vector; and A point outlier degree is determined according to the Euclidean distance, wherein the point outlier degree is proportional to the similarity.
2. The method according to claim 1, characterized in that The merging of the group and the group having an edge relationship with the group according to the intimacy includes: Sort the intimacy according to numerical values; and According to the sorting results, the two groups ranked first or last in terms of intimacy are merged.
3. The method according to claim 1, characterized in that The obtaining of data features of each node in the abnormal transaction link includes: Determining the node attributes of each node of the abnormal transaction link; and Data features are obtained according to the node attributes.
4. The method according to claim 1, wherein Calculating the average abnormality of each community based on the point abnormality includes: Calculating the average value of the point abnormality of the nodes in the abnormal transaction link contained in each community; and The average value is taken as the average abnormality of the community.
5. A device for identifying abnormal tissue, characterized in that: include: a construction module, the construction module being configured to construct a knowledge graph based on the acquired transaction data for a historical period, the transaction data including account information and transaction information between the account information, the nodes of the knowledge graph being constructed based on the account information, and the edges between the nodes being constructed based on the transaction information; A generation module, the generation module is used to divide the nodes of the knowledge graph according to the degree of association of the nodes to generate m communities, where m is an integer greater than or equal to 1; A first determination module, configured to determine an abnormal transaction link from m communities in the knowledge graph based on abnormal account information, wherein the number of communities in which nodes in the abnormal transaction link reside is less than or equal to 2, and the abnormal account information is obtained based on preset rules; a calculation module, configured to obtain data features of each node in the abnormal transaction link and calculate an average abnormality degree of each community based on the data features; as well as a second determination module, configured to determine a community whose average abnormality degree meets a set threshold as an abnormal organization; The nodes of the knowledge graph are divided according to the degree of association of the nodes to generate m communities, including: Identify each node in the knowledge graph as a group; Traverse each group and determine the closeness between the group and each group with which it has an edge relationship; When the intimacy meets the intimacy threshold, merging the group with the group having an edge relationship with the group according to the intimacy; Treat the merged group as a new group, repeatedly perform the traversal of each group, and determine the intimacy between the group and each group having an edge relationship with it; When the intimacy does not meet the intimacy threshold, stop merging the group with the group having an edge relationship with it; and When the intimacy between any two groups does not meet the intimacy threshold, the current m groups are regarded as m communities; The step of determining abnormal transaction links from the m communities in the knowledge graph based on abnormal account information includes: Determine a first abnormal node in the knowledge graph according to the preset rule; Determining, based on the first abnormal node, a directed connected link in the knowledge graph where the first abnormal node is located, wherein the directed connected link is a link formed by connecting nodes through directed edges; and At least a portion of the directed connected link consisting of nodes in two adjacent communities is used as an abnormal transaction link, wherein the abnormal transaction link includes the first abnormal node; Calculating the average abnormality of each community based on the data features includes: Constructing a feature vector according to the data features; Determining the point abnormality based on the feature vector; and Based on the point abnormality, the average abnormality of each community is calculated; Wherein, determining the point abnormality degree according to the feature vector includes: Calculating a Euclidean distance based on the feature vector and a preset standard vector, wherein the Euclidean distance is used to measure the similarity between the feature vector and the standard vector; and A point outlier degree is determined according to the Euclidean distance, wherein the point outlier degree is proportional to the similarity.
6. An electronic device, characterized in that: include: one or more processors; One or more memories, for storing executable instructions, wherein when the executable instructions are executed by the processor, the method according to any one of claims 1 to 4 is implemented.
7. A computer-readable storage medium, characterized in that The storage medium stores executable instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Multilayer fund abnormal flow direction monitoring method based on knowledge graph
CN111126828A
Anti-money laundering crime recognition method based on community division and graph convolution
CN112463983A