User account association method and device, electronic equipment and storage medium

By constructing a multidimensional fingerprint association graph and dividing it into groups, the problem of insufficient user account association in traditional risk identification methods is solved, and more efficient risk identification and user behavior analysis are achieved.

CN120804562APending Publication Date: 2025-10-17BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510820981.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional risk identification methods are based on a single-dimensional relationship, which makes it difficult to effectively identify potential connections between user accounts, resulting in insufficient accuracy and coverage of risk identification.

Method used

By constructing a multidimensional fingerprint association graph of user accounts, behavioral fingerprints and auxiliary fingerprints are used to generate an association graph between user accounts, and a graph account grouping method is adopted to identify potential related accounts.

Benefits of technology

It improves the accuracy and coverage of risk identification, has strong robustness and scalability, enhances the accuracy of account behavior and correlation analysis, and protects user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804562A_ABST
    Figure CN120804562A_ABST
Patent Text Reader

Abstract

The invention provides a user account association method and device, electronic equipment and a storage medium, and the method comprises the steps: determining a fingerprint set of a user account, the fingerprint set comprising behavior fingerprints of the user account and at least one type of auxiliary fingerprints; using the user accounts as nodes in the association graph, and determining a connection relationship between the nodes according to the fingerprint set so as to generate an association graph between the user accounts; and according to the association graph, carrying out account group division on the user account to obtain a group division result. According to the account graph mining algorithm based on the multilateral features, the user association graph is constructed by fusing multidimensional fingerprints, a graph account group division method is adopted, potential association accounts are effectively identified, the accuracy and coverage range of risk identification are improved, high robustness, expansibility and practicability are achieved, meanwhile, by establishing behavior fingerprints of the user accounts, the risk identification accuracy is improved, and the risk identification efficiency is improved. The accuracy of account behavior and account relevance analysis can be further improved through the behavior characteristics between the accounts and the multi-fingerprint crossing relation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data processing, in particular to the field of artificial intelligence and big data, and in particular to a user account association method and device, an electronic device and a storage medium. BACKGROUND

[0002] In the context of increasingly complex Internet applications and escalating attack methods, user behavior analysis and risk identification technology is widely used in financial risk control, social platform governance, e-commerce anti-fraud and other fields. Traditional risk identification methods are usually based on single-dimensional association relationships, such as focusing only on the one-way mapping relationship between user ID (UID) and IP address, i.e., by counting the number of active accounts under the same IP, account behavior patterns and other information to determine whether there is abnormal behavior. SUMMARY

[0003] The present disclosure provides a user account association method, device, electronic device and storage medium.

[0004] According to a first aspect of the present disclosure, a user account association method is provided, comprising: determining a fingerprint set of a user account, the fingerprint set comprising a behavior fingerprint of the user account and at least one type of auxiliary fingerprint; taking the user account as a node in an association graph, and determining the connection relationship between nodes according to the fingerprint set to generate an association graph between user accounts; and performing account group division on the user account according to the association graph to obtain a group division result.

[0005] According to a second aspect of the present disclosure, a user account association device is provided, comprising: a determination module configured to determine a fingerprint set of a user account, the fingerprint set comprising a behavior fingerprint of the user account and at least one type of auxiliary fingerprint; a generation module configured to take the user account as a node in an association graph, and determine the connection relationship between nodes according to the fingerprint set to generate an association graph between user accounts; and a division module configured to perform account group division on the user account according to the association graph to obtain a group division result.

[0006] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the user account association method of the above-mentioned first aspect embodiment.

[0007] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause the computer to perform the association method of the user account according to the embodiment of the above aspect.

[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, and the computer program product comprises computer program / instructions, and the computer program / instructions are executed by a processor to implement the association method of the user account according to the embodiment of the above aspect.

[0009] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description.

[0010] Beneficial effects: the account graph mining algorithm based on the multi-edge feature constructs a user association graph by fusing multi-dimensional fingerprints, and adopts a graph account group division method to effectively identify potential associated accounts, improve the accuracy and coverage of risk identification, and has strong robustness, expansibility and practicality. At the same time, by establishing the behavior fingerprint of the user account, the accuracy of the analysis of the account behavior and the association between accounts can be further increased through the relationship between the behavior characteristics and the multi-fingerprint intersection between accounts. BRIEF DESCRIPTION OF DRAWINGS

[0011] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them:

[0012] Figure 1 A flowchart of a user account association method provided by an embodiment of the present disclosure is shown;

[0013] Figure 2 A flowchart of another user account association method provided by an embodiment of the present disclosure is shown;

[0014] Figure 3 A flowchart of another user account association method provided by an embodiment of the present disclosure is shown;

[0015] Figure 4 A user account association graph provided by an embodiment of the present disclosure is shown;

[0016] Figure 5 A flowchart of another user account association method provided by an embodiment of the present disclosure is shown;

[0017] Figure 6 A flowchart of another user account association method provided by an embodiment of the present disclosure is shown;

[0018] Figure 7A structural schematic diagram of an association device of a user account provided by an embodiment of the present disclosure is provided.

[0019] Figure 8 A block diagram of an electronic device for a user account association method according to an embodiment of the present disclosure is provided. DETAILED DESCRIPTION

[0020] Exemplary embodiments of the present disclosure are described herein below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, and should be considered as merely exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the present disclosure. Also, in the following description, descriptions of well-known functions and constructions are omitted for clarity and conciseness.

[0021] A user account association method, device and electronic device of an embodiment of the present disclosure are described below with reference to the accompanying drawings.

[0022] Data processing is a technical process of analyzing and processing data (including numerical and non-numerical). It includes processing and handling of various raw data such as analysis, arrangement, calculation, editing, etc. It is broader than data analysis. With the increasing popularity of computers, in the field of computer application, numerical calculation accounts for a very small proportion, and information management through computer data processing has become the main application. For example, mapping management, warehouse management, financial management, transportation management, technical information management, office automation, etc. In the field of geographic data, there are a large amount of natural environment data (land, water, climate, biological resources data, etc.), and a large amount of social and economic data (population, transportation, industry and agriculture, etc.), which often require comprehensive data processing. Therefore, it is necessary to establish a geographic database to systematically arrange and store geographic data, reduce redundancy, develop data processing software, and make full use of database technology for data management and processing.

[0023] Artificial intelligence (AI) is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) of human life, including both hardware and software technologies. Artificial intelligence hardware technologies generally include computer vision technology, speech recognition technology, natural language processing technology, and learning / deep learning, big data processing technology, knowledge graph technology, etc.

[0024] The field of big data refers to the technology and application scope of data processing and analysis involving massive, high-speed, diverse and low-value-density data. Its core includes five directions of data storage and computing, data management, data circulation, data application and data security, and is widely used in business, finance, medical care, government and other fields, and continues to develop in the direction of intelligence, real-time and standardization.

[0025] Figure 1 A flowchart of a user account association method provided by an embodiment of the present disclosure is shown.

[0026] As shown in Figure 1 The user account association method can include the following steps.

[0027] S101, determine a fingerprint set of a user account, the fingerprint set including a behavior fingerprint of the user account and at least one type of auxiliary fingerprint.

[0028] In an embodiment of the present disclosure, the behavior fingerprint of the user account refers to identifiable and comparable behavior characteristic data generated by collecting and modeling a series of operation behaviors of a user in a system. It does not depend on traditional identity credentials (such as username, password, IP address), but establishes a behavior analysis by analyzing the operation mode, access path, interaction habits and other dynamic behaviors of the user.

[0029] It should be noted that the behavior fingerprint can include various data, which is not limited here. For example, it can include the data shown in the following table:

[0030]

[0031] It should be noted that the auxiliary fingerprint can include various, for example, one or more of Ja3 fingerprint, canvas fingerprint, fp fingerprint, etc.

[0032] Wherein, JA3 is a fingerprint generated by capturing the "ClientHello" information sent by the client and the server when performing the Transport Layer Security (TLS) / Secure Sockets Layer (SSL) handshake, which can be used to identify the TLS configuration characteristics of different clients.

[0033] Canvas fingerprint is a fingerprint generated by HTML5 <canvas>The element draws an image and reads the pixel value generated by the fingerprint. Due to the slight difference in image rendering of different browsers / devices, it can be used to identify the browser environment.

[0034] The FP fingerprint is a set of features generated based on the combination of functions supported by the browser (such as Web Real-Time Communication (WebRTC), Web Graphics Library (WebGL), Canvas, Audio Context, plug-ins, fonts, etc.), and is usually generated by using a hash or string concatenation method to generate a unique identifier.

[0035] In the embodiments of the present disclosure, each user account can be taken as a node in the association graph, and the connection relationship between the nodes is determined according to the fingerprint set, so as to generate the association graph between the user accounts.

[0036] In the embodiments of the present disclosure, each user account can be taken as a node in the association graph, and the connection relationship between the nodes is determined according to the fingerprint set, so as to generate the association graph between the user accounts.

[0037] In a possible implementation manner, if the fingerprint set between two nodes has a similar or common part, a connection line or a connection edge can be established based on the similar or common part.

[0038] In another possible implementation manner, since the fingerprint set includes the behavior fingerprint and at least one type of auxiliary fingerprint, when establishing the edge connection line or the connection edge, different connection lines or connection edges can be established according to different fingerprints. For example, when the behavior fingerprint between two nodes has a similar or common part, a first edge can be established, when the JA3 fingerprint between two nodes has a similar or common part, a second edge can be established, and when the FP fingerprint between two nodes has a similar or common part, a third edge can be established.

[0039] Preferably, different edges can also be assigned weights, and the weight of the different edge is used to represent the similarity or association strength of the corresponding fingerprints of the two nodes.

[0040] S103, according to the association graph, performing account group division on the user accounts to obtain a group division result.

[0041] In the embodiments of the present disclosure, the account group division divides the nodes in the graph into several groups (i.e., groups), so that the nodes in the same group are closely connected, while the connection between different groups is relatively sparse. By performing account group division on the association graph of user accounts, a user group with similar behavior patterns or potential associations can be found. It should be noted that a group can also be regarded as a community.

[0042] In the embodiments of the present disclosure, the method for performing account group division on user accounts according to the association graph can be various, which is not limited here.

[0043] Optionally, the association graph can be calculated based on a preset account group division algorithm to determine the group division result. The community planning algorithm can be various, for example, it can be a Girvan-Newman algorithm, a label propagation algorithm, etc.

[0044] Optionally, the association score between different nodes can also be calculated according to the connection relationship between different nodes in the association graph, for example, the score between different nodes can be calculated according to a preset weight and association relationship, and then whether two nodes belong to the same community is determined according to the score. In this way, the group division result is determined by analyzing all nodes in the association graph.

[0045] In the embodiments of the present disclosure, first, a fingerprint set of a user account is determined, the fingerprint set including a behavior fingerprint and at least one type of auxiliary fingerprint of the user account, then the user account is taken as a node in the association graph, and the connection relationship between the nodes is determined according to the fingerprint set to generate an association graph between user accounts, and finally, the user accounts are divided into account groups according to the association graph to obtain a group division result. Therefore, the account graph mining algorithm based on multi-edge features fuses multi-dimensional fingerprints to construct a user association graph, and adopts a graph account group division method to effectively identify potential associated accounts, improve the accuracy and coverage of risk identification, and has strong robustness, scalability and practicality. At the same time, by establishing the behavior fingerprint of the user account, the accuracy of the analysis of the behavior of the account and the association between the accounts can be further improved through the relationship between the behavior characteristics and the multi-fingerprint intersection of the accounts.

[0046] In the above embodiments, the process of determining the behavior fingerprint of the user account can also be performed by Figure 2 Further explanation, the method comprises:

[0047] S201, collecting behavior data of a user account, and performing feature extraction from the behavior data.

[0048] The user behavior data refers to various operation records generated by the user in the process of using a product or service. For example, the user behavior data can include the behaviors shown in the following table:

[0049]

[0050]

[0051] It should be noted that after obtaining certain user behavior data, the user behavior data needs to be encrypted and desensitized to prevent user information leakage and protect the data security of the user.

[0052] In the embodiments of the present disclosure, the method of collecting behavior data of a user account can be various, and no limitation is made herein.

[0053] Optionally, the behavior data of the user account can be collected by embedding a data monitoring script in a webpage or an application.

[0054] Optionally, the behavior data corresponding to the user account can also be determined by processing the interface request, login record, operation log and the like of the backend service record.

[0055] S202, performing feature engineering on the extracted behavior features to obtain target feature information.

[0056] Feature engineering refers to the process of extracting, transforming and constructing more meaningful and representative features from raw data, and the purpose is to improve the performance of the model, enhance the expression ability of the graph structure, and improve the accuracy of risk identification.

[0057] The feature engineering in the embodiments of the present disclosure can include various manners, and in one possible implementation manner, the extracted feature engineering can be processed by a binning processing method to obtain target feature information. Binning is a common data preprocessing technique. It converts continuous numerical data into discrete data by dividing it into several intervals (i.e. "barrels"), thereby simplifying the data complexity and handling outliers and missing values. In the fields of user behavior analysis, risk control system, device fingerprint identification, etc., the binning technology also has important application value.

[0058] In another possible implementation manner, the extracted behavior features can also be normalized to obtain target feature information. For example, multiple dimensions of behavior features (such as click frequency, page dwell time, operation path similarity, etc.) can be normalized to uniform values with comparability, and then classified according to the generated values, each class as a target feature information.

[0059] S203, performing hash calculation on the target feature information to obtain the behavior fingerprint of the user account.

[0060] In the embodiments of the present disclosure, after the target feature information is acquired, in order to further improve the security of the behavior fingerprint, the target feature information can be subjected to hash calculation, and the finally generated behavior fingerprint can hide the original behavior details and avoid sensitive information leakage. The multi-dimensional features are compressed into a fixed-length string or number, which is convenient for storage and comparison.

[0061] Meanwhile, the accounts of similar behaviors can generate the same or similar hash values, which can be used to quickly determine whether they belong to the same user or use the same tool.

[0062] In the embodiments of the present disclosure, first, the behavior data of the user account is collected, and the features are extracted from the behavior data, then the extracted behavior features are subjected to feature engineering to obtain the target feature information, and finally the target feature information is subjected to hash calculation to obtain the behavior fingerprint of the user account. Therefore, through the collection of behavior data, the extraction of features, the engineering processing and the generation of behavior fingerprints, the flow can effectively improve the user recognition accuracy, enhance the graph modeling capability, and strengthen the robustness of the risk control system, and bring significant beneficial effects in privacy protection, cross-platform identification, and rapid matching.

[0063] In the embodiments of the present disclosure, the process of determining the auxiliary fingerprint of the user account, determining the IP address of the user account, and determining the IP fingerprint of the user account according to the IP address, and / or determining the network features of the user account, and determining the Transport Layer Security (TLS) fingerprint of the user account based on the network features, and / or determining the browser associated with the user account, and generating the canvas fingerprint of the user account according to the canvas information of the browser, and / or generating the device fingerprint of the user account according to the hardware device information of the browser. Therefore, by collecting and analyzing the IP fingerprint, TLS fingerprint, canvas fingerprint, and device fingerprint of the user account, the user recognition accuracy can be significantly improved, the risk control system capability can be enhanced, and the graph modeling quality can be optimized, while protecting the privacy of the user and realizing consistent identity recognition across platforms.

[0064] In the above embodiments, according to the fingerprint set, the connection relationship between the nodes is determined to generate the association graph between the user accounts, and the association graph can also be generated by Figure 3 Further explanation, the method comprises:

[0065] S301, for a user account U i and a user account U j , the fingerprints in the fingerprint set are traversed and compared.

[0066] In the embodiments of the present disclosure, by comparing the fingerprint information (such as behavior fingerprint, IP fingerprint, TLS fingerprint, and device fingerprint) of two user accounts in multiple dimensions, it is determined whether they have similarity or association. In this way, user association analysis, device sharing identification, and gang discovery can be realized.

[0067] It should be noted that the traversal comparison of the fingerprints in the fingerprint set refers to comparing corresponding same type fingerprints of two user accounts U i and U j one by one, calculating the similarity or matching degree, and determining whether the two accounts are possibly used by the same person, share a device, or have an abnormal association. For example, the fingerprint features of two user accounts U i and U j are shown in the following table:

[0068] Fingerprint type User U i ]] User U j ]] Behavioural fingerprint abc123 abc125 IP fingerprint 192.168.1.1 192.168.1.1 TLS fingerprint ja3_abc ja3_abc Device fingerprint dev_xyz dev_xyz Canvas fingerprint canvas_123 canvas_125

[0069] As can be seen from the above table, the IP fingerprints of user accounts U i and U j are the same, the TLS fingerprints are the same, the device fingerprints are the same, the behavior fingerprints are highly similar, and the canvas fingerprints are different.

[0070] S302, in response to the fact that the current traversal fingerprint comparison result satisfies the association condition, a connection edge corresponding to the current traversal fingerprint is formed between the node N i of the user account U i and the node N j of the user account U j until the traversal of the fingerprints in the fingerprint set is completed, and the edge set between the node N i and the node N j is obtained.

[0071] It should be noted that based on the table in S301, the IP fingerprints of user accounts U i and U j are the same, the TLS fingerprints are the same, the device fingerprints are the same, the behavior fingerprints are highly similar, and the canvas fingerprints are different. The edge corresponding to the behavior fingerprint can be generated between the node N i and the node N j , and the edge corresponding to the IP fingerprint can be generated, and the edge corresponding to the TLS fingerprint can be generated, and the edge corresponding to the device fingerprint can be generated.

[0072] In one possible implementation manner, according to the fingerprints corresponding to the connection edges in the edge set, the edge combination type between the node N i and the node N j is determined, and according to the edge combination type, the edge weight between the node N i and the node N j is determined. Since the behavior fingerprints of user accounts U i and U j are highly similar but not completely the same, in order to distinguish from the completely same edge, the edges with different similarities can be valued or different weights are set, to provide a data basis for subsequent calculation and analysis.

[0073] S303, in response to the determination of the connection edges of all nodes being completed, generating the association graph between the user accounts.

[0074] In one possible implementation manner, taking the embodiments shown in S301 and S302 as an example, the final generated association graph between the user accounts is as shown in Figure 4 The edges associated with each of the user account i and the user account j correspond to the connection, and the edges are valued. It should be noted that the above embodiment only takes two user accounts as an example, and in the implementation, there can be multiple user accounts, and the edge set between the multiple user accounts can be determined through the above method, and then the association graph between the multiple user accounts is generated.

[0075] In another possible implementation manner, taking the embodiments shown in S301 and S302 as an example, the association graph between the user accounts can also be generated based on the comparison results of the behavior fingerprint, the IP fingerprint, the TLS fingerprint, the device fingerprint, and the canvas fingerprint of the two user accounts, and the value is assigned to the connection edge between the two user accounts, so as to generate the association graph between the user accounts. The value is used to describe the association relationship between the two accounts in the five fingerprint dimensions.

[0076] In the embodiments of the present disclosure, first, the user accounts U i and the user accounts U j are traversed and compared for the fingerprints in the fingerprint set, and then in response to the fact that the comparison result of the currently traversed fingerprint satisfies the association condition, a connection edge corresponding to the currently traversed fingerprint is formed between the node N i of the user account U i and the node N j of the user account U j , until the traversal of the fingerprints in the fingerprint set is completed, the edge set between the node N i and the node N j is obtained, and finally in response to the fact that the determination of the connection edges of all nodes is completed, the association graph between the user accounts is generated. Therefore, by converting the multi-dimensional fingerprint information of the user accounts into the nodes and edges in the graph structure, the association graph between the user accounts is finally generated, and this process can significantly improve the user identification accuracy, enhance the risk control system capability, and bring significant practical value in gang identification, account group division, and anomaly detection.

[0077] In the above embodiment, according to the association graph, the user accounts are divided into account groups to obtain a group division result, and the method can also be further explained as follows. Figure 5

[0078] S501, merging nodes according to the local optimization modularity of the association graph to obtain the group structure in the association graph.

[0079] ​Modularity (Q) is a measure of the quality of structure in a graph partitioned into communities. The core idea is that if a graph has more edges within communities than would be expected in a random graph, then the community structure is reasonable.

[0080] In the embodiments of the present disclosure, the local optimization modularity to merge nodes can identify the sub-graphs (i.e. communities) in the graph that are tightly connected by maximizing the modularity, thereby revealing the potential group structure in the graph.

[0081] In one possible implementation, the local optimization can be performed first, and each node is traversed to try to move it to a neighboring community, and if the movement can improve the modularity, the movement is accepted, until the modularity can no longer be improved. Then a super-node graph is constructed, i.e. each community is merged into a "super-node", and then a new graph is constructed, with edge weights being the connection strength between communities in the original graph, and the process of local optimization is continued on the new graph until the modularity can no longer be improved, and the group structure in the association graph is output.

[0082] S502, based on the group structure, updating the association graph to generate a candidate graph, wherein the nodes in the candidate graph are used to represent the group structure.

[0083] The candidate graph is a graph structure reconstructed on the basis of the original graph by a certain rule (such as account group division). The nodes are no longer original individuals, but communities or groups composed of multiple original nodes.

[0084] After obtaining the group structure, each community can be merged into a new node, and in one possible implementation, the edge weights between communities can also be calculated. The candidate graph is established based on the edge weights and the new nodes.

[0085] S503, continuing to perform local optimization modularity to merge nodes on the candidate graph to optimize the group structure, to obtain the group division result.

[0086] It should be noted that the optimization of the group structure to obtain the group division result is to further refine and optimize these communities on the basis of the existing account group division, so as to achieve a higher modularity.

[0087] It should be noted that the specific optimization process can refer to the process of local optimization in S401, which will not be described here.

[0088] In the embodiments of the present disclosure, firstly, the local optimization modularity of the association graph is performed to merge nodes, and the group structure in the association graph is obtained, then the association graph is updated based on the group structure to generate a candidate graph, wherein the nodes in the candidate graph are used to represent the group structure, and finally the local optimization modularity of the candidate graph is continued to be performed to merge nodes, so as to optimize the group structure and obtain the group division result. Thus, by performing twice optimization on the association graph, the accuracy of the finally generated group division result can be improved.

[0089] In the above embodiments, after the user account is divided into groups according to the association graph to obtain the group division result, the association graph can be further mined in different dimensions to obtain mining information of the association graph, and then target description information of the association graph is generated according to the group division result and the mining information.

[0090] It should be noted that the dimensions include but are not limited to user behavior patterns, interest preferences, social network structures, etc.

[0091] By mining the association graph in different temperatures, target description information about each community can be generated. These descriptions can cover the common characteristics of community members, their interaction patterns, and their relationships with other communities, etc. In this way, not only can the user group be better understood, but also more accurate and effective decisions can be made, resource allocation can be optimized, and competitiveness can be improved.

[0092] It should be noted that mining the association graph in different dimensions to obtain mining information of the association graph at least includes one of the following steps:

[0093] (1) The association graph is mined for connected components to obtain connected subgraphs of the association graph.

[0094] The connected subgraph of an undirected graph refers to the graph composed of the largest group of nodes in the graph, wherein there is a path between every two nodes. For a directed graph, the connected subgraph is the largest group of nodes in the graph, and there is a path from u to v and a path from v to u between any two nodes u and v in the group.

[0095] (2) The association graph is mined for centrality nodes to obtain centrality nodes of the association graph.

[0096] Centrality is a measure of the importance of a node in a network, which can be calculated by different methods, each of which emphasizes different aspects of the importance of the node.

[0097] The centrality node is the node with the largest centrality, and the types of centrality can be various, which are not limited here.

[0098] In the embodiments of the present disclosure, the method for determining the centrality of a node can be various, and the centrality calculation method corresponding to different types of nodes is different. For example, the degree centrality is simply based on the degree of a node, that is, the number of other nodes directly connected to the node. The closeness centrality is defined based on the sum of the reciprocals of the shortest path lengths from a node to all other nodes. The closer a node is to all other nodes, the higher its closeness centrality is. The betweenness centrality is defined based on the proportion of the shortest paths passing through a certain node.

[0099] (3) determining the node data in the association graph.

[0100] It should be noted that the node data in the association graph involves extracting relevant information from the original data. The node data can include various types, for example, it can include user identity information: a unique identifier for each user, interaction information: for example, transaction records, communication records, jointly participated activities, etc., attribute information: such as age, gender, geographic location, etc.

[0101] It should be noted that the target description information includes at least one of the number of nodes, the maximum group and the community size corresponding to the maximum group, the head centrality node and the group division result. Wherein, the number of nodes is the total number of nodes in the graph, the maximum group and the community size corresponding to the maximum group can be determined by the maximum group and its size through the community detection algorithm (such as Louvain method, Girvan-Newman algorithm, etc.), the head centrality node can be identified as the most important node according to a certain centrality measure, such as degree centrality, closeness centrality, betweenness centrality, etc., and the group division result is used to identify which communities the entire graph is divided into.

[0102] In the above embodiments, after the account group division of the user account is performed and the group division result is obtained, the following steps are further included, as shown in the following table: Figure 6

[0103] S601, determining the group features of different groups according to the group division result.

[0104] In the embodiments of the present disclosure, after the group division result is obtained, the community can be analyzed to determine the group features.

[0105] It should be noted that the group features can include various features. In one possible implementation, the group features can be structural features, for example, community size: calculating the number of nodes in each community. Density: evaluating the density of internal connections in the community, that is, the ratio of the actual existing edges in the community to the theoretically maximum possible edges. Centrality: identifying the central nodes in the community.

[0106] ​In another possible implementation manner, the group feature can also be a member attribute feature, for example, attribute information (such as age, gender, geographical location, and the like), and distribution of the attributes in the community can be analyzed. Average age, gender ratio, and other demographic characteristics. Common interests or behavior patterns: for example, the distribution of interest tags of corresponding community members can be analyzed.

[0107] In another possible implementation manner, the group feature can also be a behavior feature, for example, activity frequency: analyzing the frequency of community members participating in specific activities. Interaction mode: understanding the interaction mode between members in the community, such as whether a core group is formed or a uniform interaction mode.

[0108] S602, identifying a safe group and a risk group from different groups according to the group feature.

[0109] In the embodiment of the present disclosure, by analyzing the structural features, member attributes, and behavior patterns within the community, the risk level of each community can be effectively evaluated, and classification can be performed accordingly.

[0110] In one possible implementation manner, a judgment rule for judging whether a community is a safe group or a risk group can be established in advance. The judgment rule can be limited according to actual operation needs or judgment accuracy.

[0111] For example, the features of the safe group can include:

[0112] High-density connection: frequent and close interaction between members.

[0113] Low volatility rate: relatively fixed community members, with fewer new members joining.

[0114] Positive behavior pattern: such as positive interaction, positive evaluation, and the like.

[0115] Compliant behavior: complying with platform rules, no violation records.

[0116] The features of the risk group can include:

[0117] Low-density connection: loose connection between members.

[0118] High volatility rate: high member mobility, with frequent new members joining or leaving.

[0119] Abnormal behavior pattern: for example, frequent account creation and deletion, abnormal fund flow, and the like.

[0120] Violating behavior: including but not limited to fraud, money laundering, malicious attack, and the like.

[0121] According to the group division result, the group characteristics are determined, and the safe group and the risk group are identified, which can effectively improve the risk management and safety protection ability, and bring significant beneficial effects in privacy protection, cross-platform identification, rapid matching and the like.

[0122] After identifying the safe group and the risk group from different groups, the safe group can be taken as a positive sample, the risk group can be taken as a negative sample, and the recall model is optimized and trained according to the group characteristics of the positive sample and the negative sample. Thus, through community-level modeling, abnormal patterns can be more effectively captured, which helps to discover potential risk groups and improves the perception ability of the system to hidden risks.

[0123] Corresponding to the association method of the user account provided in the above several embodiments, an embodiment of the present disclosure also provides an association device of a user account. Since the association device of the user account provided in the embodiment of the present disclosure corresponds to the association method of the user account provided in the above several embodiments, the implementation of the association method of the user account is also applicable to the association device of the user account provided in the embodiment of the present disclosure, which will not be described in detail in the following embodiments.

[0124] Figure 7 A structural schematic diagram of an association device of a user account provided in the embodiment of the present disclosure. The association device 700 of the user account includes a determination module 710, a generation module 720 and a division module 730.

[0125] The determination module 710 is configured to determine a fingerprint set of the user account, the fingerprint set including a behavior fingerprint of the user account and at least one type of auxiliary fingerprint.

[0126] The generation module 720 is configured to take the user account as a node in an association graph, and determine a connection relationship between the nodes according to the fingerprint set, to generate an association graph between the user accounts.

[0127] The division module 730 is configured to divide the user accounts into account groups according to the association graph, to obtain a group division result.

[0128] In the embodiment of the present disclosure, the process of determining the behavior fingerprint of the user account includes: collecting behavior data of the user account, and performing feature extraction from the behavior data; performing feature engineering on the extracted behavior features to obtain target feature information; and performing hash calculation on the target feature information to obtain the behavior fingerprint of the user account.

[0129] In the embodiments of the present disclosure, the process of determining the auxiliary fingerprint of the user account includes at least one of the following operations: determining the IP address of the user account, and determining the IP fingerprint of the user account according to the IP address; determining the network feature of the user account, and determining the Transport Layer Security (TLS) fingerprint of the user account based on the network feature; determining the browser associated with the user account, and generating the canvas fingerprint of the user account according to the canvas information of the browser; and generating the device fingerprint of the user account according to the hardware device information of the browser.

[0130] In the embodiments of the present disclosure, the connection relationship between the nodes is determined according to the fingerprint set, so as to generate the association graph between the user accounts, including: for the user account U i and the user account U j , the fingerprints in the fingerprint set are iterated and compared; in response to the fact that the comparison result of the currently iterated fingerprint meets the association condition, a connection edge corresponding to the currently iterated fingerprint is formed between the node N i of the user account U i and the node N j of the user account U j , until the iteration of the fingerprints in the fingerprint set is completed, and the edge set between the node N i and the node N j is obtained; in response to the fact that the determination of the connection edges of all nodes is completed, the association graph between the user accounts is generated.

[0131] In the embodiments of the present disclosure, after the association graph between the user accounts is generated, the following operations are further included: determining the edge combination type between the node N i and the node N j according to the fingerprint corresponding to the connection edge in the edge set; and determining the edge weight between the node N i and the node N j according to the edge combination type.

[0132] In the embodiments of the present disclosure, the user accounts are divided into account groups to obtain a group division result, including: merging nodes by local optimization modularity of the association graph to obtain a group structure in the association graph; updating the association graph based on the group structure to generate a candidate graph, wherein the nodes in the candidate graph are used to represent the group structure; and continuing to merge nodes by local optimization modularity of the candidate graph to optimize the group structure, so as to obtain the group division result.

[0133] In the embodiments of the present disclosure, after the user accounts are divided into account groups according to the association graph to obtain a group division result, the following operations are further included: mining the association graph in different dimensions to obtain mining information of the association graph; and generating target description information of the association graph according to the group division result and the mining information.

[0134] In the embodiments of the present disclosure, the association graph is mined in different dimensions to obtain mining information of the association graph, including at least one of the following operations: performing connected component mining on the association graph to obtain a connected subgraph of the association graph; performing centrality node mining on the association graph to obtain a centrality node of the association graph; and determining node data in the association graph.

[0135] In the embodiments of the present disclosure, the target description information includes at least one of the following information: a number of nodes; a maximum group and a community size corresponding to the maximum group; a head centrality node; and a group division result.

[0136] In the embodiments of the present disclosure, after the user account is divided into groups according to the association graph to obtain the group division result, the method further includes: determining group features of different groups according to the group division result; and identifying a safe group and a risk group from the different groups according to the group features.

[0137] In the embodiments of the present disclosure, after the safe group and the risk group are identified from the different groups according to the group features, the method further includes: taking the safe group as a positive sample, taking the risk group as a negative sample, and optimizing and training a recall model based on group features of the positive sample and the negative sample.

[0138] Therefore, the account graph mining algorithm based on multi-edge features can effectively identify potential associated accounts by fusing multi-dimensional fingerprints to construct a user association graph and using a graph account group division method, thereby improving the accuracy and coverage of risk identification, and having strong robustness, expansibility and practicality. Meanwhile, by establishing a behavior fingerprint of a user account, the accuracy of analysis on account behavior and association between accounts can be further improved through the relationship between behavior features of the accounts and the multi-fingerprint intersection.

[0139] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0140] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0141] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0142] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to computer programs / instructions stored in a read-only memory (ROM) 802 or computer programs / instructions loaded from a storage unit 806 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0143] Various components in device 800 are connected to I / O interface 805, including: an input unit 806 such as a keyboard, mouse, etc.; an output unit 807 such as various types of displays, speakers, etc.; a storage unit 808 such as a magnetic disk, optical disk, etc.; and a communication unit 809 such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0144] The computing unit 801 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs various methods and processes described above, such as the association method of user accounts. For example, in some embodiments, the association method of user accounts can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 806. In some embodiments, portions or all of the computer program / instructions can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program / instructions are loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the association method of user accounts described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the association method of user accounts by any other suitable means, such as by means of firmware.

[0145] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0146] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0147] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0148] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0149] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0150] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0151] It should be understood that the various forms of flow shown above can be re-ordered, added to, or have steps deleted, using the steps disclosed in the present disclosure. For example, the steps disclosed in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure are achieved, and the present disclosure is not limited herein.

[0152] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.< / canvas>

Claims

1. A method for associating user accounts, wherein: The method comprises the following steps: Determining a fingerprint set of a user account, the fingerprint set including a behavioral fingerprint of the user account and at least one type of auxiliary fingerprint; Taking the user accounts as nodes in the association graph and determining the connection relationships between the nodes based on the fingerprint set to generate an association graph between the user accounts; The user accounts are divided into account groups according to the association graph to obtain group division results.

2. The method according to claim 1, wherein The process of determining the behavioral fingerprint of the user account includes: Collecting behavioral data of the user account and extracting features from the behavioral data; Perform feature engineering on the extracted behavioral features to obtain target feature information; Perform hash calculation on the target feature information to obtain the behavioral fingerprint of the user account.

3. The method according to claim 1, wherein The process of determining the auxiliary fingerprint of the user account includes at least one of the following operations: Determining the IP address of the user account, and determining the IP fingerprint of the user account based on the IP address; Determining a network characteristic of the user account, and determining a Transport Layer Security (TLS) fingerprint of the user account based on the network characteristic; Determine the browser associated with the user account, and generate a canvas fingerprint of the user account based on the canvas information of the browser; A device fingerprint of the user account is generated based on the hardware device information of the browser.

4. The method according to any one of claims 1 to 3, wherein Determining the connection relationship between nodes based on the fingerprint set to generate a relationship graph between user accounts includes: For user account U i and user account U j , traverse and compare the fingerprints in the fingerprint set; In response to the fingerprint comparison result of the current traversal meeting the association condition, the user account U i Node N i and user account U j Node N j A connection edge corresponding to the currently traversed fingerprint is formed between them until the fingerprint traversal in the fingerprint set is completed and the node N is obtained. i and node N j The set of edges between In response to the connection edges of all nodes being determined to be complete, a relationship graph between the user accounts is generated.

5. The method according to claim 4, wherein After generating the association graph between the user accounts, the method further includes: Determine the node N according to the fingerprint corresponding to the connecting edge in the edge set i and node N j The type of edge combination between them; According to the edge combination type, determine the N i and node N j The edge weights between .

6. The method according to claim 1, wherein The dividing the user accounts into account groups according to the association graph to obtain group division results includes: Performing local optimization on the modularity of the association graph to merge nodes and obtain a group structure in the association graph; Based on the group structure, the association graph is updated to generate a candidate graph, wherein the nodes in the candidate graph are used to represent the group structure; The modularity of the candidate graph is further locally optimized to merge nodes, thereby optimizing the group structure and obtaining the group division result.

7. The method according to claim 1, wherein After dividing the user accounts into account groups according to the association graph and obtaining the group division results, the method further includes: Mining the association graph in different dimensions to obtain mining information of the association graph; Target description information of the association graph is generated according to the group division result and the mining information.

8. The method according to claim 7, wherein: The mining of the association graph in different dimensions to obtain mining information of the association graph includes at least one of the following operations: Performing connected component mining on the association graph to obtain a connected subgraph of the association graph; Performing central node mining on the association graph to obtain the central nodes of the association graph; Node data in the association graph is determined.

9. The method according to claim 8, wherein The target description information includes at least one of the following information: Number of nodes; The largest group and the group size corresponding to the largest group; Head centrality node; The group division result.

10. The method according to claim 1, wherein After dividing the user accounts into account groups according to the association graph and obtaining the group division results, the method further includes: determining group characteristics of different groups according to the group division results; A safe group and a risky group are identified from the different groups according to the group characteristics.

11. The method according to claim 10, wherein: After identifying the safe group and the risky group from the different groups according to the group characteristics, the method further includes: The safe group is used as a positive sample, the risk group is used as a negative sample, and based on the group characteristics of the positive and negative samples, the precision-recall rate of the recall model is optimized and trained.

12. A user account association device, wherein: The device comprises: a determination module, configured to determine a fingerprint set of a user account, the fingerprint set comprising a behavioral fingerprint of the user account and at least one type of auxiliary fingerprint; a generating module, configured to use the user accounts as nodes in an association graph and, based on the fingerprint set, determine the connection relationships between the nodes to generate an association graph between the user accounts; The division module is used to divide the user accounts into account groups according to the association graph to obtain group division results.

13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the user account association method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the method for associating user accounts according to any one of claims 1 to 11 is implemented.