User identifier determination method and device and electronic equipment
By obtaining and classifying user usage data, and using community discovery algorithms to generate unified user IDs, the problem of insufficient UID uniformity in the existing technology is solved, and more efficient user management and service provision are achieved.
Patent Information
- Application Number
- CN202311510598.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
The existing user ID unified mapping method has poor UID uniformity and cannot meet higher user management needs.
By obtaining multiple user usage data, classifying and merging based on user information, a community discovery algorithm is used to generate a unified user identity, and the accurate mapping of user ID is achieved.
It improves the unity of UID management, can meet higher user management needs, and achieves better provision of user services.
Smart Images

Figure CN120011818A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a method, device and electronic equipment for determining a user identity. Background Art
[0002] With the development of the digital age, we should consider building our own customer relationship management system (CRMS) and customer data platform (CDP) to connect various user-related data platforms. Considering that the user ID (UID) should be unique, it represents the unique identity that the user can identify. For example, CRMS, CDP, etc. can find out the user's personal information, preferences and other information through the management of UID, and provide better services to users. Usually, for corporate official websites, membership systems, APPs, etc., CRMS, CDP, etc. can store various types of information of end customers. Considering that these data may be inconsistent in different systems, and data from different sources often correspond to different UIDs, it will be inconvenient for CRMS, CDP, etc. to conduct unified user management to provide better services.
[0003] In order to achieve unified management of users, it is necessary to use the method of UID unified mapping to map the ID information extracted from the data collected from multiple sources to form data connection. Among them, ID unified mapping refers to mapping user IDs in different systems. The same user identity of different channels and systems can be uniformly identified to achieve user association between systems. ID unified mapping includes two parts. The first is the matching between IDs, that is, extracting the source ID information in the collected data. At this time, multiple user IDs will be formed, and multiple user IDs need to be mapped to ID uniformly; then, the data strings such as behaviors and attributes under multiple user IDs are marked on the unified user ID. This step is data mapping; that is, ID unified mapping and data mapping are the connection process. However, the current method of UID unified mapping is usually based on the fact that the collected data contains the same mobile phone number, the same ID number, the same email address, or the same device ID, etc., and the data from different sources are mapped to UID uniformly. This method has poor UID uniformity and cannot meet higher user management requirements. Summary of the invention
[0004] The purpose of the present invention is to provide a user identification determination method, device and electronic device to solve the problem that the current UID unified mapping method has poor UID uniformity and cannot meet higher user management requirements.
[0005] To achieve the above object, an embodiment of the present invention provides a method for determining a user identity, comprising:
[0006] Acquire multiple user usage data; wherein the user usage data includes one or more user information related to the user identification;
[0007] According to the user information, the plurality of user usage data are classified to determine a plurality of initial data groups; wherein each of the initial data groups corresponds to at least one of the user usage data, and different user usage data in each of the initial data groups contain the same user information and / or the user information similarity between different user usage data meets a preset condition;
[0008] According to a community discovery algorithm, the multiple initial data groups are merged to obtain a target data group;
[0009] For each target data group, a unified user identifier corresponding to the target data group is generated.
[0010] Optionally, the classifying the plurality of user usage data according to the user information to determine a plurality of initial data groups includes:
[0011] According to the user information, a plurality of initial data groups are determined based on a first rule by classifying the plurality of user usage data; wherein the first rule is: classifying the user usage data containing the same user information into the same initial data group;
[0012] and / or,
[0013] According to the user information, multiple initial data groups obtained by classifying the multiple user usage data are determined based on a second rule; wherein the second rule is: for user usage data containing different user information, user usage data whose user information similarity meets preset conditions are classified into the same initial data group.
[0014] Optionally, the determining, based on the user information and based on a second rule, a plurality of initial data groups obtained by classifying the plurality of user usage data includes:
[0015] Classify the user information in the target user usage data according to the type of the user information, and determine at least one type group; wherein the target user usage data is user usage data containing different user information, and each type group corresponds to at least one user information;
[0016] For each type group, respectively determine the similarity between different user information in the type group;
[0017] For different user usage data, determining the correlation between different user usage data according to the similarity between different user information corresponding to at least one type group;
[0018] Different user usage data with correlation greater than a first threshold are classified into the same initial data group, and a plurality of initial data groups obtained by classifying the plurality of user usage data are determined.
[0019] Optionally, for each type group, determining the similarity between different user information in the type group includes:
[0020] For each type group, each user information in the type group is structurally divided to obtain multiple fields;
[0021] For the multiple fields obtained by dividing each user information, each field corresponding to different user information is respectively compared for similarity, and the similarity between the different user information is determined based on the comparison results of the multiple fields.
[0022] Optionally, the determining, for different user usage data, the relevance between different user usage data according to the similarity between different user information corresponding to at least one type group, includes:
[0023] For different user usage data, the weighted sum of the similarities between different user information corresponding to each type group is determined as the correlation between different user usage data.
[0024] Optionally, merging the multiple initial data groups according to a community discovery algorithm to obtain a target data group includes:
[0025] Each initial data group is used as a community node, and different community nodes are connected based on the correlation between user usage data in different initial data groups to build a community network;
[0026] Based on the community network, a community discovery algorithm is used to merge multiple community nodes into communities;
[0027] When any two community nodes do not satisfy the merging condition, it is determined that the target data group is obtained.
[0028] Optionally, based on the community network, using a community discovery algorithm to merge multiple community nodes includes:
[0029] Based on the community network, a community discovery algorithm is used to perform K-stage community merging on multiple community nodes;
[0030] The community merger at each stage is performed according to the following steps:
[0031] Select multiple community nodes from the community network in the i-th stage as seed nodes;
[0032] A community discovery algorithm is used to merge neighbor nodes that meet the merging conditions with seed nodes corresponding to the neighbor nodes; wherein the neighbor nodes are community nodes other than the seed nodes in the community nodes after the community is merged in the i-th stage, K and i are positive integers, and i≤K.
[0033] Optionally, the selecting a plurality of community nodes as seed nodes from the community network in the i-th stage includes:
[0034] For each community node in the community network of the i-th stage, the influence value of each community node is calculated according to the degree of the community node, the degree of the neighboring nodes of the community node, and the correlation between the usage data of different users corresponding to the community node and each neighboring node;
[0035] According to the influence value of each community node, multiple community nodes are used as seed nodes; wherein the influence value of the seed node is greater than the influence value of the neighboring node of the seed node.
[0036] Optionally, the adopting of a community discovery algorithm to merge neighboring nodes that meet a merging condition with seed nodes corresponding to the neighboring nodes includes:
[0037] The community discovery algorithm is used to perform multiple rounds of merging. Each round of merging is performed according to the following steps:
[0038] For each seed node, respectively calculate the modularity gain of each neighboring node of the seed node merged into the seed node;
[0039] If the modularity gain corresponding to the target neighboring node is greater than a second threshold, the target neighboring node is merged with the seed node; otherwise, the target neighboring node is not merged; wherein the target neighboring node is any neighboring node of the seed node.
[0040] Optionally, for each seed node, respectively calculating the modularity gain of each neighboring node of the seed node merged into the seed node includes:
[0041] For each neighboring node of the seed node, the modularity gain of merging the neighboring node into the seed node is calculated according to the weight sum of the first edge, the weight sum of the second edge, the weight sum of the third edge, and the total community degree of the seed node;
[0042] Among them, the sum of the weights of the first edges is: the sum of the weights of the edges connecting the neighboring node to the seed node; the sum of the weights of the second edges is: the sum of the weights of all edges connected to the neighboring node; the sum of the weights of the third edges is: the sum of the weights of all edges in the community network of the i-th stage.
[0043] Optionally, the calculation formula of the modularity gain is:
[0044]
[0045] Where ΔQ is the modularity gain, k i,in is the weight sum of the first edge, k i is the weight sum of the second edge, m is the weight sum of the third edge, Σ tot is the total number of degrees.
[0046] Optionally, the any two community nodes do not satisfy the merging condition if: a modularity gain of merging one of the any two community nodes into the other community node is less than or equal to a second threshold.
[0047] Optionally, the method further comprises:
[0048] Acquire new user usage data; wherein the new user usage data includes one or more user information related to the user identifier;
[0049] If the new user usage data and the user usage data in the first target data group contain the same user information, or the user information similarity between the new user usage data and the user usage data in the first target data group meets a preset condition, then the unified user identifier corresponding to the first target data group is determined as the unified user identifier of the new user usage data; otherwise, the new user usage data is used as a new community node;
[0050] If, based on the community discovery algorithm, the new community node and the community node corresponding to the second target data group meet the merging condition, the unified user identifier corresponding to the second target data group is determined as the unified user identifier of the new user usage data; otherwise, a unified user identifier corresponding to the new user usage data is generated.
[0051] To achieve the above object, an embodiment of the present invention provides a user identification determination device, comprising:
[0052] An acquisition module, used to acquire a plurality of user usage data; wherein the user usage data includes one or more user information related to the user identification;
[0053] a classification module, configured to classify the plurality of user usage data according to the user information, and determine a plurality of initial data groups; wherein each of the initial data groups corresponds to at least one of the user usage data, and different user usage data in each of the initial data groups contain the same user information and / or the user information similarity between different user usage data meets a preset condition;
[0054] A merging module, used to merge the multiple initial data groups according to a community discovery algorithm to obtain a target data group;
[0055] The generating module is used to generate a unified user identification corresponding to each target data group.
[0056] To achieve the above-mentioned purpose, an embodiment of the present invention provides an electronic device, comprising: a transceiver, a processor, a memory, and a program or instruction stored in the memory and executable on the processor; when the processor executes the program or instruction, the steps of the user identification determination method as described above are implemented.
[0057] To achieve the above objective, an embodiment of the present invention provides a readable storage medium on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the user identification determination method described above are implemented.
[0058] The beneficial effects of the above technical solution of the present invention are as follows:
[0059] In the embodiment of the present invention, the multiple user usage data are classified according to the user information in the multiple user usage data, multiple initial data groups are determined, and the multiple initial data groups are merged according to the community discovery algorithm to obtain the target data group. In this way, a unified user identification is generated based on the target data group, that is, accurate user ID mapping for multiple user usage data is achieved, and user services are provided based on the unified user identification generated by the accurate user ID mapping, that is, UID management is guaranteed to have high uniformity and can meet higher user management requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 A flow chart showing a method for determining a user identity according to an embodiment of the present invention;
[0061] Figure 2 A schematic diagram showing a classification method of an embodiment of the present invention;
[0062] Figure 3 A block diagram showing a user identification determination device according to an embodiment of the present invention;
[0063] Figure 4 A block diagram showing an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0065] It should be understood that the references to "one embodiment" or "an embodiment" throughout the specification mean that the specific features, structures, or characteristics associated with the embodiment are included in at least one embodiment of the present invention. Therefore, the references to "in one embodiment" or "in an embodiment" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0066] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the following processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0067] Additionally, the terms "system" and "network" are often used interchangeably herein.
[0068] In the embodiments provided by the present invention, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0069] like Figure 1 As shown, an embodiment of the present invention provides a method for determining a user identity, comprising the following steps:
[0070] Step 11: Acquire multiple user usage data; wherein the user usage data includes one or more user information related to the user identification.
[0071] Optionally, each user usage data corresponds to a user identifier of a data source, or each user usage data carries a user identifier of a data source. The user identifier of the data source can be generated by a terminal, server, etc. of the data source. For example, the relevant identification information (i.e., user usage data) of the user's login usage data can be obtained through multiple channels such as applications (Application, APP), mobile web pages (HTML5, H5), mini programs, public accounts, etc.
[0072] Optionally, the user information includes but is not limited to at least one of the following: login information, transaction information, communication information, such as mobile phone number, email address, device unique identifier (device_id), International Mobile Equipment Identity (IMEI), base station information, Internet Protocol (IP) address, Media Access Control (Mac) address, etc. This information can be written into a graph database, such as a distributed data graph library (NebulaGraph).
[0073] Step 12: Classify the multiple user usage data according to the user information to determine multiple initial data groups; wherein each of the initial data groups corresponds to at least one of the user usage data, and different user usage data in each of the initial data groups contain the same user information and / or the user information similarity between different user usage data meets a preset condition.
[0074] Optionally, the user usage data may be classified using existing technologies, or may be classified using determination rules, and / or the user usage data may be classified according to implicit association rules, and multiple initial data groups obtained by classifying the multiple user usage data are determined (i.e., preliminary user ID mapping relationships are obtained). The determination rule is a rule for classification based on completely identical user information (i.e., different user usage data in each of the initial data groups contain the same user information), and the implicit association rule is a rule for classification based on the correlation of different user information (i.e., the user information similarity between different user usage data meets a preset condition).
[0075] Step 13: According to the community discovery algorithm, the multiple initial data groups are merged to obtain a target data group.
[0076] Optionally, based on the initial data grouping obtained by the above classification (that is, the ID mapping relationship initially obtained), a directed network graph is constructed, and the multiple initial data groups are merged through a community discovery algorithm to obtain a target data grouping (that is, accurate mapping of user IDs).
[0077] Step 14: For each target data group, a unified user identifier corresponding to the target data group is generated, that is, different unified user identifiers are generated for different target data groups.
[0078] In the embodiment of the present invention, the multiple user usage data are classified according to the user information in the multiple user usage data, multiple initial data groups are determined, and the multiple initial data groups are merged according to the community discovery algorithm to obtain the target data group. In this way, a unified user identification is generated based on the target data group, that is, accurate user ID mapping for multiple user usage data is achieved, and user services are provided based on the unified user identification generated by the accurate user ID mapping, that is, UID management is guaranteed to have high uniformity and can meet higher user management requirements.
[0079] Optionally, the classifying the plurality of user usage data according to the user information to determine a plurality of initial data groups includes:
[0080] According to the user information, a plurality of initial data groups obtained by classifying the plurality of user usage data based on a first rule are determined; wherein the first rule is: classifying the user usage data containing the same user information into the same initial data group; for example: classifying the first user usage data into the same initial data group, and determining the plurality of initial data groups obtained by classifying the plurality of user usage data; wherein the first user usage data is a plurality of user usage data containing the same user information;
[0081] and / or,
[0082] According to the user information, multiple initial data groups obtained by classifying the multiple user usage data are determined based on a second rule; wherein the second rule is: for user usage data containing different user information, user usage data whose similarity of user information meets a preset condition is classified into the same initial data group; for example: for the second user usage data, the similarity between user information is classified, and user usage data whose similarity of user information meets a preset condition is classified into the same initial data group, and multiple initial data groups obtained by classifying the multiple user usage data are determined; wherein the second user usage data is a plurality of user usage data containing different user information.
[0083] For example, multiple user usage data may be classified by only classifying multiple user usage data containing the same user information into the same initial data group, and multiple user usage data containing different user information into different initial data groups (or obtaining ID mapping relationships according to a determined rule). Alternatively, multiple user usage data may be classified by only classifying according to the similarity between user information (or obtaining ID mapping relationships according to implicit association rules). Alternatively, multiple user usage data may be first classified by classifying multiple user usage data containing the same user information into the same initial data group, and multiple user usage data containing different user information into different initial data groups, and then classified by classifying according to the similarity between user information, etc. The embodiments of the present invention are not limited thereto.
[0084] For example, obtaining the ID mapping relationship according to a certain rule means: giving priority to the simplest and most direct method, that is, obtaining the ID mapping relationship through a certain rule. For example, email addresses, IP addresses, Internet access devices and other strictly structured data, when they become points in the graph, a simple equality relationship is sufficient to find such a correspondence, for example: having the same email address, mobile phone number, IMEI, etc. is mapped to the same user ID; for example: IP addresses and device information may be shared by Internet cafes, and this information can be used as a non-deterministic rule in the embodiments of the present invention. Figure 2 As shown, user_7 and user_22 are associated because they share the same email account and can be mapped to the same user ID.
[0085] For example, obtaining ID mapping relationships according to implicit association rules means: processing uncertain conditions, mining implicit equal and similar association relationships, that is, combining characteristics such as attribute similarity between user information, and determining multiple initial data groups obtained by classifying the usage data of the multiple users.
[0086] Optionally, the determining, based on the user information and based on a second rule, a plurality of initial data groups obtained by classifying the plurality of user usage data includes:
[0087] Classify the user information in the target user usage data according to the type of the user information, and determine at least one type group; wherein the target user usage data is user usage data containing different user information, and each type group corresponds to at least one user information;
[0088] For each type group, respectively determine the similarity between different user information in the type group;
[0089] For different user usage data, determining the correlation between different user usage data according to the similarity between different user information corresponding to at least one type group;
[0090] Different user usage data with correlation greater than a first threshold are classified into the same initial data group, and a plurality of initial data groups obtained by classifying the plurality of user usage data are determined.
[0091] It should be noted that the correlation between different user usage data is determined based on the similarity of user information between the different user usage data. Here, different user usage data with a correlation greater than the first threshold is user usage data whose user information similarity meets preset conditions.
[0092] Optionally, the user information may be a mobile phone number, an email address, a device_id, an IMEI, base station information, an IP address, a Mac address, etc. The user information is classified according to the type, such as an email address belongs to the email type, a mobile phone number belongs to the mobile phone number type, etc. For example, the structure of an email address is such as Hinton99@gmail.com and Hinton99+001@gmail.com, etc. The structure of a mobile phone number is such as +86-XXXX-XXXX-XXX and (+86)-XXXX-XXXX-XXX, etc.
[0093] For the above user information that can be subjected to structured analysis, the similarity between two values can be directly calculated, or the similarity between different user information can be calculated by determining substrings and operating the Jaccard similarity algorithm, etc., but the embodiments of the present invention are not limited thereto.
[0094] Optionally, the similarity is calculated in a structured analysis manner, that is, for each type group, the similarity between different user information in the type group is determined respectively, including:
[0095] For each type group, each user information in the type group is structurally divided to obtain multiple fields;
[0096] For the multiple fields obtained by dividing each user information, each field corresponding to different user information is respectively compared for similarity, and the similarity between the different user information is determined based on the comparison results of the multiple fields.
[0097] Specifically, according to the attributes and writing methods of user information, the user information is split into multiple fields with finer granularity, such as splitting the email address: Hinton99@gmail.com into three fields, such as the email_handle field is: Hinton99, the email_alias field is: 99, and the email_domain field is: gmail.com. Based on this, detailed deterministic rules can be designed to compare each field to obtain the similarity between different user information. For example, the email_handle field is equal, and other non-deterministic rules can even be applied on this basis, such as the email_domain field is preset to be equivalent, such as gmail.com and googlemail.com are equivalent, to see whether there is similarity between different user information, that is, to split the user information, and determine whether there is similarity between different user information by comparing multiple fields, so as to determine whether the user usage data belongs to the same user ID.
[0098] For example, string-based telephone numbers can be uniformly converted into structured data such as <country code>+<area code>+<local number>+<extension number>, and then further compared with each other in the structured data format to determine the corresponding similarity. That is, user information is split, and multiple fields are compared to determine whether different user information is similar, so as to determine whether the user usage data belongs to the same user ID.
[0099] For another example, if the user information contains pictures, such as avatars, Uniform Resource Locator (URL) information, etc., equivalent analysis can be performed by comparing the similarities of the pictures to determine the corresponding similarities, but the embodiments of the present invention are not limited thereto.
[0100] Optionally, the determining, for different user usage data, the relevance between different user usage data according to the similarity between different user information corresponding to at least one type group, includes:
[0101] For different user usage data, the weighted sum of the similarities between different user information corresponding to each type group is determined as the correlation between different user usage data.
[0102] Optionally, the type of user information may be an email address, a mobile phone number, or a login account, etc., but the embodiment of the present invention is not limited thereto.
[0103] Specifically, considering the implicit association rules or non-deterministic rules between user information, comprehensive consideration and judgment can be made through multiple rules, that is, through the mechanism of multi-factor scoring, weighting is performed according to the importance of multiple conditions to give the correlation between different user usage data. For example: correlation = email similarity * weight 1 + phone number similarity * weight 2 + base station location similarity * weight 3 + ... The specific conditions or user information types used in the present invention are not limited to this.
[0104] Optionally, merging the multiple initial data groups according to a community discovery algorithm to obtain a target data group includes:
[0105] Each initial data group is used as a community node, and different community nodes are connected based on the correlation between user usage data in different initial data groups to build a community network;
[0106] Based on the community network, a community discovery algorithm is used to merge multiple community nodes into communities;
[0107] When any two community nodes do not satisfy the merging condition, it is determined that the target data group is obtained.
[0108] Optionally, when connecting different first nodes based on the correlation between user usage data in different initial data groups, the correlation includes but is not limited to, for example: frequency of contact with the same mobile phone number, email address, etc., geographic location, base station information, etc., the embodiments of the present invention are not limited to this.
[0109] In this embodiment, a community discovery algorithm is used to merge multiple community nodes in a community network. Multiple merges may be performed until any two community nodes do not meet the merge condition. It is considered that the community merge has achieved the ultimate goal, so that user usage data belonging to the same user can be merged as much as possible to achieve accurate and unified mapping of user IDs. Any two community nodes do not meet the merge condition here means that after multiple merges, the community nodes in the community network can no longer be merged, and each community node after the final merge is a target data group.
[0110] Optionally, based on the community network, using a community discovery algorithm to merge multiple community nodes includes:
[0111] Based on the community network, a community discovery algorithm is used to perform K-stage community merging on multiple community nodes;
[0112] The community merger at each stage is performed according to the following steps:
[0113] Select multiple community nodes from the community network in the i-th stage as seed nodes;
[0114] A community discovery algorithm is used to merge neighbor nodes that meet the merging conditions with seed nodes corresponding to the neighbor nodes; wherein the neighbor nodes are community nodes other than the seed nodes in the community nodes after the community is merged in the i-th stage, K and i are positive integers, and i≤K.
[0115] It should be noted that in the community discovery process based on the community network, the community merging process in each stage is similar. The difference is that the community discovery process in the i+1th stage is a community discovery process further performed on the community network after the community merging in the i-th stage.
[0116] In the embodiment of the present invention, there is no specific limit on the number of stages of community merging, and the specific value of K depends on whether any two community nodes in the community network meet the merging conditions. For example, when the amount of user usage data is small or the source of user usage data is small, it may be possible to use a community discovery algorithm in one stage to obtain the final ID unified mapping relationship, or to use a community discovery algorithm in two stages to obtain the final ID unified mapping relationship; when the amount of user usage data is large or the source of user usage data is large, it may also be possible to use more times of community discovery algorithms to obtain the final ID unified mapping relationship, etc. The embodiment of the present invention is not limited to this.
[0117] Optionally, the selecting a plurality of community nodes as seed nodes from the community network in the i-th stage includes:
[0118] For each community node in the community network of the i-th stage, the influence value of each community node is calculated according to the degree of the community node, the degree of the neighboring nodes of the community node, and the correlation between the usage data of different users corresponding to the community node and each neighboring node;
[0119] According to the influence value of each community node, multiple community nodes are used as seed nodes; wherein the influence value of the seed node is greater than the influence value of the neighboring node of the seed node.
[0120] In this embodiment, by selecting seed nodes and taking them as the core for overlapping community discovery, the computational complexity on large-scale data sets is solved; and when the community nodes in the community network change dynamically due to new users joining the network, new services being launched, etc., an incremental dynamic community discovery model can be used to achieve that when adding or deleting nodes in the user network, only the newly added nodes or related nodes are recalculated, thereby reducing the amount of calculation and ensuring the accuracy of the community discovery algorithm.
[0121] Specifically, for the community network of the i-th stage, which includes N community nodes (N is a positive integer), Node2Vec can be used to calculate the node vector of the graph network, that is, to calculate the influence value of each community node. For example, a custom influence function can be used to calculate the influence value of each community node, so as to select multiple seed nodes from the N community nodes according to the influence value.
[0122] Among them, the influence function F(u,v) of node u on node v is expressed as:
[0123]
[0124] Wherein, D(u) is the degree of node u, D(v) is the degree of node v, sum(u,v) is the correlation between node u and node v (which can be determined by the above-mentioned method of calculating the correlation between usage data of different users, which will not be repeated here), 1-sum(u,v) represents the distance between node u and node v. The smaller the distance, the smaller the influence of node u on node v.
[0125] Based on the above influence function F(u,v), for each community node in the community network, the influence value of each community node can be calculated according to the degree of the community node, the degree of the neighboring nodes of the community node, and the correlation between the community node and each neighboring node. The calculation formula of the influence value of the community node is as follows:
[0126]
[0127] Among them, F(u) is the influence value of node u, and ν∈N(μ) represents the total number of neighboring nodes v of node u.
[0128] In the process of selecting seed nodes, if the influence value of a node is greater than the influence values of all its neighboring nodes, then the node will be used as the seed node; in other words, for a node, if there is no node with a greater influence value than this node among all its neighboring nodes, then the node will be used as the seed node.
[0129] Optionally, the adopting of a community discovery algorithm to merge neighboring nodes that meet a merging condition with seed nodes corresponding to the neighboring nodes includes:
[0130] The community discovery algorithm is used to perform multiple rounds of merging. Each round of merging is performed according to the following steps:
[0131] For each seed node, respectively calculate the modularity gain of each neighboring node of the seed node merged into the seed node;
[0132] If the modularity gain corresponding to the target neighboring node is greater than the second threshold, the target neighboring node is merged with the seed node; otherwise, the target neighboring node is not merged; wherein the target neighboring node is any neighboring node of the seed node.
[0133] The following is an explanation of the community discovery algorithm based on the community network in Phase 1 (for example, including N community nodes):
[0134] The community discovery process based on the community network in the first phase is that for a community network with N community nodes, each community node belongs to an independent community containing only one node at the time of initialization, that is, there are N communities. For example: taking each initial data group as an independent node, connecting different nodes based on the correlation between user usage data in different initial data groups, that is, forming the community network in the first phase, and each community corresponds to an initial data group.
[0135] Furthermore, M seed nodes are selected from N community nodes, and the remaining community nodes are used as neighbor nodes of these M seed nodes. Each neighbor node i tries to consider whether to join the community of its corresponding seed node j. If joining the community of a certain seed node j obtains the maximum modularity gain, and it is a positive number (that is, the modularity gain is greater than the second threshold), then the neighbor node i can join the community of the seed node j, that is, it is confirmed that the neighbor node i and the seed node j can be merged, that is, unified as the same unified user ID; if the modularity gain is negative regardless of which seed node community the neighbor node i joins, then the neighbor node i will not join any community this round, that is, it is confirmed that the neighbor node i and any seed node cannot be unified as the same unified user ID. In this way, it is determined whether each neighbor node can join the community of the seed node until any neighbor node does not meet the merging conditions, and the community discovery process of this stage is determined to be over.
[0136] The following is an introduction to the calculation process of modularity gain:
[0137] Optionally, for each seed node, respectively calculating the modularity gain of each neighboring node of the seed node merged into the seed node includes:
[0138] For each neighboring node of the seed node, the modularity gain of merging the neighboring node into the seed node is calculated according to the weight sum of the first edge, the weight sum of the second edge, the weight sum of the third edge, and the total community degree of the seed node;
[0139] Among them, the sum of the weights of the first edges is: the sum of the weights of the edges connecting the neighboring node to the seed node; the sum of the weights of the second edges is: the sum of the weights of all edges connected to the neighboring node; the sum of the weights of the third edges is: the sum of the weights of all edges in the community network of the i-th stage.
[0140] Optionally, the calculation formula of the modularity gain is:
[0141]
[0142] Where ΔQ is the modularity gain, k i,in is the weight sum of the first edge, k i is the weight sum of the second edge, m is the weight sum of the third edge, ∑ tot is the total number of degrees.
[0143] For example, taking the modularity gain of a neighbor node i joining the community C of the seed node as an example, the calculation process of the modularity gain is explained:
[0144] ΔQ = the modularity of community C after neighbor node i joins community C - (the modularity of community C before neighbor node i joins community C + the change in modularity of the community after neighbor node i leaves the community).
[0145] The calculation formula of modularity is:
[0146]
[0147] Among them, Q is modularity;
[0148] c v =c w indicates that node v and node w are in the same community, c v ! =c w Indicates that node v and node w are in different communities;
[0149] A vw represents the number of actual edges between node v and node w, k v and k w are the degrees of node v and node w respectively;
[0150] Indicates that both node v and node w belong to community C. Sum.
[0151] Furthermore, vw k v k w =∑ v k v ∑ w kw , the calculation formula of the above modularity can be transformed into:
[0152]
[0153] Among them, ∑ in It represents twice the total number of edges in community C (i.e., the sum of the weights of the edges of nodes in the community before neighbor node i joins community C. Here, each edge is considered to connect two nodes, that is, an edge will be counted twice, so it is "2 times"). The edges in the community refer to the nodes connected at both ends of the edge belonging to the same community;
[0154] ∑ tot represents the total degree of all current nodes in community C (or, twice the total number of internal edges in community C and external edges connecting other communities);
[0155] ∑ v k v Indicates that node v belongs to community C, for k v sum;
[0156] ∑ w k w Indicates that node w belongs to community C, for k w sum;
[0157] Indicates that neighbor node i belongs to community C. sum;
[0158] After neighbor node i joins community C, community C
[0159] k i,in represents the sum of the weights of the edges connecting neighboring node i to the nodes in community C;
[0160] k i represents the sum of the weights of all edges connected to neighboring node i;
[0161] represents the sum of the weights of the edges within the community where neighbor node i is located, and m represents the sum of the weights of the edges in the entire graph (that is, the sum of the weights of all edges in the current community network);
[0162] Before neighbor node i joins community C,
[0163] After neighbor node i leaves this community, the community
[0164] Therefore, the calculation formula for the modularity gain of community C after a neighbor node i joins the community C can be obtained as follows:
[0165]
[0166] After the community discovery process in the first phase is completed, the community network after the community is merged in the first phase is used as the initial community network of the second round of community discovery process to repeat the above steps, and so on, until any two communities no longer meet the merging conditions, and the final community node is determined (that is, each community node is a target data group).
[0167] Specifically, the community discovery process based on the community network of the i+1th stage is to treat the community network after merging the communities of the i-th stage as a new network. The independent nodes in the network are each community in the i-th stage, and the edge weights between the nodes are the total number of edges connecting the two communities. That is, all the nodes in the community divided by the community discovery process based on the i-th stage community network are merged into a new node in the community discovery process based on the i+1th stage community network, and the connection between the new nodes is the connection between the communities of the i-th stage. In this way, the community network based on the i+1th stage continues to repeat the same steps of the community discovery process of the i-th stage until the community where the node is located no longer changes (that is, any two communities no longer meet the merging conditions), which also means that the modularity reaches the maximum value, and then the iteration stops.
[0168] Optionally, the method further comprises:
[0169] Acquire new user usage data; wherein the new user usage data includes one or more user information related to the user identifier;
[0170] If the new user usage data and the user usage data in the first target data group contain the same user information, or the user information similarity between the new user usage data and the user usage data in the first target data group meets a preset condition, then the unified user identifier corresponding to the first target data group is determined as the unified user identifier of the new user usage data; otherwise, the new user usage data is used as a new community node;
[0171] If, based on the community discovery algorithm, the new community node and the community node corresponding to the second target data group meet the merging condition, the unified user identifier corresponding to the second target data group is determined as the unified user identifier of the new user usage data; otherwise, a unified user identifier corresponding to the new user usage data is generated.
[0172] In this embodiment, for the user usage data of the new user ID brought by the subsequent new users or other new services, the user ID association matching is first performed with the user usage data of the generated user unified ID through the determination rules and implicit rules (that is, judging whether the new user usage data and the user usage data in the first target data group contain the same user information, or whether the user information similarity between the new user usage data and the user usage data in the first target data group meets the preset conditions). If the match is successful, the generated user unified ID is used as the user unified ID of the user usage data of the new user ID; if the match is not successful, the user usage data of the new user ID is used as a new community node, and the user ID association matching is performed based on the community discovery algorithm. If the user usage data of the new user ID is successfully matched with the user usage data of the generated user unified ID by the community discovery algorithm, the generated user unified ID is used as the user unified ID of the user usage data of the new user ID; if the match fails, a corresponding unified user ID is generated for the user usage data of the new user ID, and finally the user ID unified mapping of the user usage data of the new user ID is realized.
[0173] In an embodiment of the present invention, a message passing mechanism in a distributed graph processing framework (GraphX) is used to aggregate community information around each node, and the modularity gain is used to calculate whether the node needs to join a new community. In order to achieve large-scale processing, a parallel computing platform such as a big data parallel framework (Spark) can also be used to execute the algorithm; when performing full-graph operations, local sensitive hashing (Min Hash) can also be used to reduce the dimensionality of pairwise comparisons. According to the final graph network, the nodes in each community correspond to the same user ID, forming a final unified user identification. Optionally, if there is new anonymous user information (i.e., user usage information that does not conform to the above-mentioned unified user identification, etc.) in the future, the above steps can also be repeated for the anonymous user information to determine the corresponding same user identification, etc. The embodiment of the present invention is not limited to this.
[0174] The embodiment of the present invention provides a unified user identity mapping model based on a community discovery algorithm. By selecting seed nodes and taking them as the core to perform overlapping community discovery, the problem of complex calculations on large-scale data sets is solved. The model is suitable for changes in user dynamic networks and realizes dynamic community discovery when adding or deleting nodes in the user network through an incremental dynamic community discovery model.
[0175] In the embodiment of the present application, the user usage data is gradually matched in a way from easy to difficult for a large amount of acquired user usage data. On the basis of identifying the user usage data belonging to the same user ID based on the definite rules and implicit rules, the community discovery algorithm is then performed to improve the accuracy of the community discovery algorithm and solve the problem of difficulty in identifying non-definite relationships in the community discovery algorithm. At the same time, when selecting seed nodes, the correlation between different user usage data is calculated when using implicit rules to identify the same user ID, which is used for direct calculation of the influence function in the selected seed nodes to reduce the calculation cost.
[0176] The actual application effect of the embodiment of the present invention: it can solve the ID matching problem between multi-channel users and existing users, which is conducive to the efficient identification of information of each channel; the unified UID mapping implemented based on the above scheme can ensure the fine-grained restoration of user portraits, facilitate the unified authentication management of multi-channel users, improve the accuracy of the recommendation algorithm, and enhance the user experience through intelligent recommendation.
[0177] like Figure 3 As shown, a user identification determination device 300 according to an embodiment of the present invention includes:
[0178] A first acquisition module 310 is used to acquire a plurality of user usage data; wherein the user usage data includes one or more user information related to the user identification;
[0179] The classification module 320 is used to classify the plurality of user usage data according to the user information to determine a plurality of initial data groups; wherein each of the initial data groups corresponds to at least one of the user usage data, and different user usage data in each of the initial data groups contain the same user information and / or the user information similarity between different user usage data meets a preset condition;
[0180] A merging module 330, configured to merge the multiple initial data groups according to a community discovery algorithm to obtain a target data group;
[0181] The generating module 340 is used to generate a unified user identifier corresponding to each target data group.
[0182] Optionally, the classification module 320 includes:
[0183] A first classification unit is used to determine, according to the user information, a plurality of initial data groups obtained by classifying the plurality of user usage data based on a first rule; wherein the first rule is: classifying the user usage data containing the same user information into the same initial data group;
[0184] and / or,
[0185] A second classification unit is used to determine, according to the user information, a plurality of initial data groups obtained by classifying the plurality of user usage data based on a second rule; wherein the second rule is: for user usage data containing different user information, user usage data whose similarity of user information meets a preset condition is classified into the same initial data group.
[0186] Optionally, the second classification unit is further used for:
[0187] Classify the user information in the target user usage data according to the type of the user information, and determine at least one type group; wherein the target user usage data is user usage data containing different user information, and each type group corresponds to at least one user information;
[0188] For each type group, respectively determine the similarity between different user information in the type group;
[0189] For different user usage data, determining the correlation between different user usage data according to the similarity between different user information corresponding to at least one type group;
[0190] Different user usage data with correlation greater than a first threshold are classified into the same initial data group, and a plurality of initial data groups obtained by classifying the plurality of user usage data are determined.
[0191] Optionally, the second classification unit is further used for:
[0192] For each type group, each user information in the type group is structurally divided to obtain multiple fields;
[0193] For the multiple fields obtained by dividing each user information, each field corresponding to different user information is respectively compared for similarity, and the similarity between the different user information is determined based on the comparison results of the multiple fields.
[0194] Optionally, the second classification unit is further used for:
[0195] For different user usage data, the weighted sum of the similarities between different user information corresponding to each type group is determined as the correlation between different user usage data.
[0196] Optionally, the merging module 330 includes:
[0197] A construction unit, used to use each initial data group as a community node, connect different community nodes based on the correlation between user usage data in different initial data groups, and construct a community network;
[0198] A merging unit, configured to merge multiple community nodes using a community discovery algorithm based on the community network;
[0199] The determination unit is used to determine the target data group when any two community nodes do not meet the merging condition.
[0200] Optionally, the merging unit is used for:
[0201] Based on the community network, a community discovery algorithm is used to perform K-stage community merging on multiple community nodes;
[0202] The community merger at each stage is performed according to the following steps:
[0203] Select multiple community nodes from the community network in the i-th stage as seed nodes;
[0204] A community discovery algorithm is used to merge neighbor nodes that meet the merging conditions with seed nodes corresponding to the neighbor nodes; wherein the neighbor nodes are community nodes other than the seed nodes in the community nodes after the community is merged in the i-th stage, K and i are positive integers, and i≤K.
[0205] Optionally, the merging unit is further used for:
[0206] For each community node in the community network of the i-th stage, the influence value of each community node is calculated according to the degree of the community node, the degree of the neighboring nodes of the community node, and the correlation between the usage data of different users corresponding to the community node and each neighboring node;
[0207] According to the influence value of each community node, multiple community nodes are used as seed nodes; wherein the influence value of the seed node is greater than the influence value of the neighboring node of the seed node.
[0208] Optionally, the merging unit is further used for:
[0209] The community discovery algorithm is used to perform multiple rounds of merging. Each round of merging is performed according to the following steps:
[0210] For each seed node, respectively calculate the modularity gain of each neighboring node of the seed node merged into the seed node;
[0211] If the modularity gain corresponding to the target neighboring node is greater than the second threshold, the target neighboring node is merged with the seed node; otherwise, the target neighboring node is not merged; wherein the target neighboring node is any neighboring node of the seed node.
[0212] Optionally, the merging unit is further used for:
[0213] For each neighboring node of the seed node, the modularity gain of merging the neighboring node into the seed node is calculated according to the weight sum of the first edge, the weight sum of the second edge, the weight sum of the third edge, and the total community degree of the seed node;
[0214] Among them, the sum of the weights of the first edges is: the sum of the weights of the edges connecting the neighboring node to the seed node; the sum of the weights of the second edges is: the sum of the weights of all edges connected to the neighboring node; the sum of the weights of the third edges is: the sum of the weights of all edges in the community network of the i-th stage.
[0215] Optionally, the calculation formula of the modularity gain is:
[0216]
[0217] Where ΔQ is the modularity gain, k i,in is the weight sum of the first edge, k i is the weight sum of the second edge, m is the weight sum of the third edge, Σ tot is the total number of degrees.
[0218] Optionally, the any two community nodes do not satisfy the merging condition if: a modularity gain of merging one of the any two community nodes into the other community node is less than or equal to a second threshold.
[0219] Optionally, the method further comprises:
[0220] A second acquisition module, used to acquire new user usage data; wherein the new user usage data includes one or more user information related to the user identification;
[0221] a first processing module, configured to determine the unified user identifier corresponding to the first target data group as the unified user identifier of the new user usage data if the new user usage data and the user usage data in the first target data group contain the same user information, or the user information similarity between the new user usage data and the user usage data in the first target data group meets a preset condition; otherwise, use the new user usage data as a new community node;
[0222] The second processing module is configured to determine the unified user identifier corresponding to the second target data group as the unified user identifier of the new user usage data if, based on the community discovery algorithm, the new community node and the community node corresponding to the second target data group meet the merging condition; otherwise, generate a unified user identifier corresponding to the new user usage data.
[0223] The above-mentioned device in the embodiment of the present invention can implement each step of the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here.
[0224] The embodiment of the present invention further provides an electronic device, such as Figure 4 As shown, it includes a transceiver 410, a processor 400, a memory 420, and a program or instruction stored in the memory 420 and executable on the processor 400; when the processor 400 executes the program or instruction, the steps of the above-mentioned user identification determination method are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0225] The transceiver 410 is used to receive and send data under the control of the processor 400 .
[0226] Among them, Figure 4 In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically linking together various circuits of one or more processors represented by processor 400 and memory represented by memory 420. The bus architecture may also link together various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface. The transceiver 410 may be a plurality of components, i.e., including a transmitter and a receiver, providing a unit for communicating with various other devices on a transmission medium. The processor 400 is responsible for managing the bus architecture and general processing, and the memory 420 may store data used by the processor 400 when performing operations.
[0227] A readable storage medium according to an embodiment of the present invention stores a program or instruction thereon. When the program or instruction is executed by a processor, the steps in the user identification determination method described above are implemented and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0228] The processor is a processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0229] It should be further explained that the terminals described in this specification include but are not limited to smart phones, tablet computers, etc., and many of the functional components described are called modules in order to more particularly emphasize the independence of their implementation methods.
[0230] In the embodiment of the present invention, module can be implemented with software so that it can be executed by various types of processors. For example, an executable code module of an identification can include one or more physical or logical blocks of computer instructions, for example, it can be constructed as an object, process or function. Nevertheless, the executable code of the identified module does not need to be physically located together, but can include different instructions stored in different positions, and when these instructions are logically combined together, it constitutes a module and realizes the specified purpose of the module.
[0231] In fact, executable code module can be a single instruction or many instructions, and can even be distributed on a plurality of different code segments, distributed among different programs, and distributed across a plurality of memory devices. Similarly, operating data can be identified in the module, and can be implemented and organized in the data structure of any appropriate type according to any appropriate form. The operating data can be collected as a single data set, or can be distributed in different locations (including on different storage devices), and can only be present on a system or network as an electronic signal at least in part.
[0232] When a module can be implemented by software, considering the level of existing hardware technology, a person skilled in the art can build a corresponding hardware circuit to implement the corresponding function of the module that can be implemented by software without considering the cost. The hardware circuit includes a conventional very large scale integration (VLSI) circuit or gate array and existing semiconductors such as logic chips, transistors, or other discrete components. The module can also be implemented by a programmable hardware device, such as a field programmable gate array, a programmable array logic, a programmable logic device, etc.
[0233] The above exemplary embodiments are described with reference to the accompanying drawings, and many different forms and embodiments are feasible without departing from the spirit and teachings of the present invention. Therefore, the present invention should not be constructed as a limitation of the exemplary embodiments proposed herein. More specifically, these exemplary embodiments are provided so that the present invention will be perfect and complete, and the scope of the present invention will be conveyed to those who are familiar with the technology. In these figures, the component sizes and relative sizes may be exaggerated for clarity. The terms used here are only based on the purpose of describing specific exemplary embodiments and are not intended to be limiting. As used herein, unless the text clearly indicates otherwise, the singular forms "one", "an" and "the" are intended to include these multiple forms. It will be further understood that the terms "including" and / or "comprising" when used in this specification indicate the presence of the features, integers, steps, operations, components and / or components, but do not exclude the presence or increase of one or more other features, integers, steps, operations, components, components and / or their groups. Unless otherwise indicated, when stated, a range of values includes the upper and lower limits of that range and any subranges therebetween.
[0234] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for determining a user identity, characterized in that: include: Acquire multiple user usage data; wherein the user usage data includes one or more user information related to the user identification; According to the user information, the plurality of user usage data are classified to determine a plurality of initial data groups; wherein each of the initial data groups corresponds to at least one of the user usage data, and different user usage data in each of the initial data groups contain the same user information and / or the user information similarity between different user usage data meets a preset condition; According to a community discovery algorithm, the multiple initial data groups are merged to obtain a target data group; For each target data group, a unified user identifier corresponding to the target data group is generated.
2. The method according to claim 1, characterized in that The step of classifying the plurality of user usage data according to the user information to determine a plurality of initial data groups includes: According to the user information, a plurality of initial data groups are determined based on a first rule by classifying the plurality of user usage data; wherein the first rule is: classifying the user usage data containing the same user information into the same initial data group; and / or, According to the user information, multiple initial data groups obtained by classifying the multiple user usage data are determined based on a second rule; wherein the second rule is: for user usage data containing different user information, user usage data whose user information similarity meets preset conditions are classified into the same initial data group.
3. The method according to claim 2, characterized in that The determining, based on the user information and based on a second rule, a plurality of initial data groups obtained by classifying the plurality of user usage data comprises: Classify the user information in the target user usage data according to the type of the user information, and determine at least one type group; wherein the target user usage data is user usage data containing different user information, and each type group corresponds to at least one user information; For each type group, respectively determine the similarity between different user information in the type group; For different user usage data, determining the correlation between different user usage data according to the similarity between different user information corresponding to at least one type group; Different user usage data with correlation greater than a first threshold are classified into the same initial data group, and a plurality of initial data groups obtained by classifying the plurality of user usage data are determined.
4. The method according to claim 3, characterized in that The step of determining, for each type group, the similarity between different user information in the type group comprises: For each type group, each user information in the type group is structurally divided to obtain multiple fields; For the multiple fields obtained by dividing each user information, each field corresponding to different user information is respectively compared for similarity, and the similarity between the different user information is determined based on the comparison results of the multiple fields.
5. The method according to claim 3, characterized in that: The determining, for different user usage data, the relevance between different user usage data according to the similarity between different user information corresponding to at least one type group, includes: For different user usage data, the weighted sum of the similarities between different user information corresponding to each type group is determined as the correlation between different user usage data.
6. The method according to any one of claims 1 to 5, characterized in that The step of merging the multiple initial data groups according to the community discovery algorithm to obtain a target data group includes: Each initial data group is used as a community node, and different community nodes are connected based on the correlation between user usage data in different initial data groups to build a community network; Based on the community network, a community discovery algorithm is used to merge multiple community nodes into communities; When any two community nodes do not satisfy the merging condition, it is determined that the target data group is obtained.
7. The method according to claim 6, characterized in that Based on the community network, a community discovery algorithm is used to merge multiple community nodes, including: Based on the community network, a community discovery algorithm is used to perform K-stage community merging on multiple community nodes; The community merger at each stage is performed according to the following steps: Select multiple community nodes from the community network in the i-th stage as seed nodes; A community discovery algorithm is used to merge neighbor nodes that meet the merging conditions with seed nodes corresponding to the neighbor nodes; wherein the neighbor nodes are community nodes other than the seed nodes in the community nodes after the community is merged in the i-th stage, K and i are positive integers, and i≤K.
8. The method according to claim 7, characterized in that The step of selecting a plurality of community nodes as seed nodes from the community network in the i-th stage includes: For each community node in the community network of the i-th stage, the influence value of each community node is calculated according to the degree of the community node, the degree of the neighboring nodes of the community node, and the correlation between the usage data of different users corresponding to the community node and each neighboring node; According to the influence value of each community node, multiple community nodes are used as seed nodes; wherein the influence value of the seed node is greater than the influence value of the neighboring node of the seed node.
9. The method according to claim 7, characterized in that: The adopting of a community discovery algorithm to merge neighboring nodes that meet a merging condition with seed nodes corresponding to the neighboring nodes includes: The community discovery algorithm is used to perform multiple rounds of merging. Each round of merging is performed according to the following steps: For each seed node, respectively calculate the modularity gain of each neighboring node of the seed node merged into the seed node; If the modularity gain corresponding to the target neighboring node is greater than a second threshold, the target neighboring node is merged with the seed node; otherwise, the target neighboring node is not merged; wherein the target neighboring node is any neighboring node of the seed node.
10. The method according to claim 9, characterized in that The step of respectively calculating, for each seed node, the modularity gain of each neighboring node of the seed node merged into the seed node comprises: For each neighboring node of the seed node, the modularity gain of merging the neighboring node into the seed node is calculated according to the weight sum of the first edge, the weight sum of the second edge, the weight sum of the third edge, and the total community degree of the seed node; Among them, the sum of the weights of the first edges is: the sum of the weights of the edges connecting the neighboring node to the seed node; the sum of the weights of the second edges is: the sum of the weights of all edges connected to the neighboring node; the sum of the weights of the third edges is: the sum of the weights of all edges in the community network of the i-th stage.
11. The method according to claim 9, characterized in that The arbitrary two community nodes do not satisfy the merging condition when: the modularity gain of merging one of the arbitrary two community nodes into the other community node is less than or equal to the second threshold.
12. The method according to claim 1, characterized in that Also includes: Acquire new user usage data; wherein the new user usage data includes one or more user information related to the user identifier; If the new user usage data and the user usage data in the first target data group contain the same user information, or the user information similarity between the new user usage data and the user usage data in the first target data group meets a preset condition, then the unified user identifier corresponding to the first target data group is determined as the unified user identifier of the new user usage data; otherwise, the new user usage data is used as a new community node; If, based on the community discovery algorithm, the new community node and the community node corresponding to the second target data group meet the merging condition, the unified user identifier corresponding to the second target data group is determined as the unified user identifier of the new user usage data; otherwise, a unified user identifier corresponding to the new user usage data is generated.
13. A user identification determination device, characterized in that: include: An acquisition module, used to acquire a plurality of user usage data; wherein the user usage data includes one or more user information related to the user identification; a classification module, configured to classify the plurality of user usage data according to the user information, and determine a plurality of initial data groups; wherein each of the initial data groups corresponds to at least one of the user usage data, and different user usage data in each of the initial data groups contain the same user information and / or the user information similarity between different user usage data meets a preset condition; A merging module, used to merge the multiple initial data groups according to a community discovery algorithm to obtain a target data group; The generating module is used to generate a unified user identification corresponding to each target data group.
14. An electronic device comprising: A transceiver, a processor, a memory, and a program or instruction stored in the memory and executable on the processor; characterized in that when the processor executes the program or instruction, the steps of the user identification determination method as described in any one of claims 1 to 12 are implemented.
15. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or instruction is executed by a processor, the steps of the user identification determination method according to any one of claims 1 to 12 are implemented.