Social network identity correlation method and system based on big data

CN119357691BActive Publication Date: 2026-09-15GUANGDONG ZHIAN INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411354535.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2026-09-15
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

[0004]上述方法虽然能够实现社交网络身份关联,但是在面对复杂的数据和多变的用户行为时存在一定的局限性,无法全面的分析用户数据且无法精准地进行身份关联,因此,当前的社交网络关联方法存在关联精准度低的问题

Benefits of technology

[0088]To address the problems described in the background section, this invention achieves accurate association of social network identities by combining collected user identity big data from two network domains, constructing multiple user feature models, and comprehensively processing the user identity big data. First, by collecting user identity big data from two network domains, an active time period model, a browsing preference model, a friend relationship model, and a follower relationship model are constructed, enabling a comprehensive capture of user behavior and characteristics. This accurately describes user behavior patterns in different network domains, improving the comprehensiveness and representativeness of the data. Then, a first activity feature map, a first browsing feature map, a second activity feature map, and a second browsing feature map are plotted and homogeneity processing is performed, clearly demonstrating the similarity of two user behaviors in different network domains. This is achieved by calculating the activity... The homogeneity of user features and browsing features, along with determining whether they meet a preset homogeneity threshold, can effectively improve the accuracy of identity recognition. Subsequently, if the homogeneity of activity features and browsing features does not fully meet the preset threshold, the second user feature set is iteratively updated. This iterative update continuously optimizes user features and improves the efficiency of social network identity association. Furthermore, by calculating the homogeneity of friend features and following features, and calculating mixed homogeneity based on preset feature weights, multiple feature dimensions of the user are comprehensively considered, providing more comprehensive and accurate data for final identity confirmation. Finally, by constructing a feature function and inputting the mixed homogeneity into the feature function, the matching status of user identities can be automatically determined, resulting in high accuracy in social network identity association. Therefore, this invention can improve the accuracy of social network identity association.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357691B_ABST
    Figure CN119357691B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of social network analysis, and relates to a social network identity correlation method based on big data, which comprises the following steps: constructing an active time period model, a browsing preference model, a friend relationship model and an attention relationship model; calculating active feature homogeneity and browsing feature homogeneity; judging whether all of the homogeneities are greater than a preset homogeneity threshold value; if not, extracting second iteration user features and recalculating the active feature homogeneity and the browsing feature homogeneity; if all of the homogeneities are greater than the homogeneity threshold value, calculating friend feature homogeneity and attention feature homogeneity; calculating mixed homogeneity based on feature weight; constructing a feature function; inputting the mixed homogeneity into the feature function to obtain a function output value; judging whether the function output value is 1; if the function output value is 1, confirming social network identity correlation. The application can improve the correlation accuracy of social network identity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of social network analysis technology, and in particular to a method, system, electronic device, and computer-readable storage medium for social network identity association based on big data. Background Technology

[0002] Social network identity association refers to the process of analyzing relevant information of users on different online platforms to determine whether two accounts on different online platforms belong to the same individual.

[0003] Traditional social network identity association methods rely on simple feature matching, such as usernames, email addresses, or phone numbers, to determine social network identity by matching a single feature.

[0004] While the methods described above can achieve identity association on social networks, they have certain limitations when faced with complex data and ever-changing user behavior. They cannot comprehensively analyze user data or accurately associate identities. Therefore, current social network association methods suffer from low accuracy. Summary of the Invention

[0005] This invention provides a method for associating social network identities based on big data and a computer-readable storage medium, the main purpose of which is to improve the accuracy of social network identity association.

[0006] To achieve the above objectives, this invention provides a social network identity association method based on big data, comprising:

[0007] Collect user identity big data from two network domains to construct active time period models, browsing preference models, friend relationship models, and follow relationship models;

[0008] Based on the active time period model, browsing preference model, friend relationship model and follow relationship model, data processing is performed on the user identity big data to obtain a first user feature set and a second user feature set.

[0009] First user features are extracted based on the first user feature set, wherein the first user features include: first activity features, first browsing features, first friend features, and first follow features;

[0010] The second user features are extracted sequentially from the second user feature set, wherein the second user features include: second activity features, second browsing features, second friend features, and second follow features;

[0011] Using the first active feature, the first browsing feature, the second active feature, and the second browsing feature, respectively draw the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map;

[0012] Based on the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map, homogeneity processing is performed to obtain the homogeneity of active features and the homogeneity of browsing features;

[0013] Determine whether the homogeneity of the active features and the homogeneity of the browsing features are all greater than a preset homogeneity threshold;

[0014] If the homogeneity of the active feature and the homogeneity of the browsing feature are not all greater than the homogeneity threshold, then return to the above steps of sequentially extracting the second user features from the second user feature set.

[0015] If the homogeneity of the active feature and the homogeneity of the browsing feature are all greater than the homogeneity threshold, then the homogeneity of the friend feature and the homogeneity of the follow feature are calculated based on the first friend feature, the first follow feature, the second friend feature and the second follow feature.

[0016] The mixed homogeneity is calculated based on the homogeneity of the active feature, the browsing feature, the friend feature, the following feature, and the preset feature weights.

[0017] Construct a feature function, input the homogeneity of the mixture into the feature function, and obtain the function output value;

[0018] Determine whether the function output value is 1;

[0019] If the function output value is not 1, then return to the steps of sequentially extracting the second user features from the second user feature set.

[0020] If the function outputs a value of 1, it confirms that the first user feature and the second user feature have completed the social network identity association.

[0021] Optionally, the collection of user identity big data from two network domains includes:

[0022] The API authentication mechanisms of two network domains are obtained, and a preset API request is sent to the two network domains based on the API authentication mechanisms of the two network domains and the pre-obtained API key. The request result is obtained, and the API request is used to determine whether the API request successfully connects to the two network domains.

[0023] If the API request fails to connect to the two network domains, return to the steps described above for obtaining the API authentication mechanism for the two network domains.

[0024] If the API request successfully connects the two network domains, a first API request is constructed using preset initial pagination parameters, and first request data is obtained based on the initial pagination parameters and the first API request.

[0025] Determine if the preset total pagination parameter is equal to 1;

[0026] If the total pagination parameter is equal to 1, then the first request data will be used as the user identity big data.

[0027] If the total pagination parameter is not equal to 1, then determine whether the initial pagination parameter is equal to the total pagination parameter;

[0028] If the initial pagination parameter is not equal to the total pagination parameter, then the initial pagination parameter is incremented by one to obtain the updated pagination parameter. The updated pagination parameter is used to update the initial pagination parameter. Then, the initial pagination parameter is used to construct an update API request. Based on the initial pagination parameter and the update API request, update request data is obtained, and the steps of determining whether the initial pagination parameter is equal to the total pagination parameter are returned.

[0029] If the initial pagination parameter is equal to the total pagination parameter, then the update request data and the first request data are merged in the form of the first request data first and the update request data last to obtain the user identity big data.

[0030] Optionally, the construction process of the active time period model is as follows:

[0031] Obtain the active state weight, work state weight, state activity level, and work activity level for each active time period. The active time period can be: [0,6)h, [6,16)h, [16,20)h, [20,24)h.

[0032] The active time period model is constructed using the active state weight, work state weight, state activity, and work activity:

[0033] T e =ε×X e +∈×Y e

[0034] Among them, T e This represents the first or second active feature value in the e-th active time period, where e represents the sequence number of the active time period, ε represents the active state weight, and X... e Y represents the state activity level during the e-th active time period, ∈ represents the working state weight, and Y e This represents the work activity level during the e-th active time period.

[0035] Optionally, the construction process of the browsing preference model, friend relationship model, and follow relationship model is as follows:

[0036] The system acquires the browsing time period, content type, browsing duration range, browsing frequency, browsing device, browsing location, and content nature. The browsing time period includes [0,6)h, [6,16)h, [16,20)h, and [20,24)h. The content type includes news content, entertainment content, technology content, literature content, and educational content. The browsing duration range includes [10,20)min, [20,40)min, [40,60)min, and [60,+∞)min. The browsing frequency includes: greater than 3 times / day and less than or equal to 5 times / day, greater than 5 times / day and less than or equal to 7 times / day, greater than 7 times / day and less than or equal to 9 times / day, and greater than 9 times / day. The browsing device includes mobile devices and PCs. The browsing location includes urban and suburban areas. The content nature includes the latest news and past content.

[0037] The browsing dimensions are obtained by arranging and combining the browsing time period, content type, browsing duration range, browsing frequency, browsing device, browsing location, and content nature.

[0038] Obtain the browsing preference degree, the dimension ordinal number, and the dimension weight of the browsing dimension, and construct a browsing preference model based on the browsing dimension, the browsing preference degree, the dimension ordinal number, and the dimension weight of the browsing dimension.

[0039] The browsing preference model is as follows:

[0040]

[0041] Among them, B τ (j,k,l,v,w,f) represents the first or second browsing feature value in the τ-th browsing time period when the browsing dimension is (j,k,l,v,w,f), where τ represents the index of the browsing time period, j represents the content type, k represents the browsing duration interval, l represents the browsing frequency, v represents the browsing device, w represents the browsing location, f represents the content nature, n represents the ordinal number of the browsing dimension, and α represents the dimensional ordinal number. i β represents the dimensional weight of the i-th browsing dimension. i This represents the browsing preference level for the i-th browsing dimension;

[0042] The system acquires the interaction frequency, mutual friend relationships, friend interaction degree of interaction frequency, interaction frequency weight, interaction weight of interaction frequency, friend relationship degree of mutual friend relationships, mutual friend weight, and relationship weight of mutual friend relationships. The interaction frequency includes: greater than 1 time / day and less than or equal to 3 times / day, greater than 3 times / day and less than or equal to 5 times / day, greater than 5 times / day and less than or equal to 10 times / day, and greater than 10 times / day. The mutual friend relationships include: greater than 1 mutual friend and less than or equal to 3 mutual friends, greater than 3 mutual friends and less than or equal to 5 mutual friends, greater than 5 mutual friends and less than or equal to 10 mutual friends, and greater than 10 mutual friends.

[0043] A friend relationship model is constructed based on the interaction frequency, mutual friend relationships, friend interaction degree of interaction frequency, interaction frequency weight, interaction frequency interaction weight, friend relationship degree of mutual friend relationships, mutual friend weight, and relationship weight of mutual friend relationships.

[0044] The friend relationship model is as follows:

[0045]

[0046] Where A represents the feature value of the first friend or the feature value of the second friend, θ represents the interaction frequency weight, and B b B represents the friend interaction frequency at the b-th interaction frequency. b1 Let μ represent the interaction weight for the b-th interaction frequency, and O represent the weight of mutual friends. m O represents the friendship degree of the m-th mutual friend. m1 This represents the relationship weight of the m-th mutual friend relationship;

[0047] The system obtains the following parameters: common interest relationships, content similarity intervals, common interest weights, common interest degree of common interest relationships, attention weight of common interest relationships, weight of content similarity intervals, content similarity degree of content similarity intervals, and similarity interval weight of content similarity intervals. The common interest relationships include: more than 1 common interest relationship and less than or equal to 3 common interest relationships, more than 3 common interest relationships and less than or equal to 5 common interest relationships, more than 5 common interest relationships and less than or equal to 10 common interest relationships, and more than 10 common interest relationships. The content similarity intervals include: [0, 10%), [10%, 30%), [30%, 60%), and [60%, 100%).

[0048] The attention relationship model is constructed using the common attention relationship, content similarity interval, common attention weight, common attention degree of common attention relationship, attention weight of common attention relationship, weight of content similarity interval, content similarity of content similarity interval, and similarity interval weight of content similarity interval;

[0049] The attention relationship model is shown below:

[0050]

[0051] Where Y represents the first or second feature value of interest, ρ represents the common interest weight, and M m M represents the degree of common interest in the m-th common interest relationship. m1 This represents the attention weight of the m-th common attention relationship. U represents the weight of content similarity intervals. u U represents the content similarity of the u-th content similarity interval. u1 This represents the weight of the similarity interval for the u-th content similarity interval.

[0052] Optionally, the step of drawing a first activity feature map, a first browsing feature map, a second activity feature map, and a second browsing feature map using the first activity feature, the first browsing feature, the second activity feature, and the second browsing feature, respectively, includes:

[0053] A first active time data matrix is ​​constructed based on the active time period model and the first active feature value;

[0054] A first active feature map is drawn based on the first active time data matrix, wherein the x-axis of the first active feature map represents the active time period and the y-axis represents the first active feature value;

[0055] A first browsing preference data matrix is ​​constructed based on the browsing preference model and the first browsing feature value;

[0056] A first browsing feature map is drawn based on the first browsing preference data matrix, wherein the x-axis of the first browsing feature map represents the browsing time period and the y-axis represents the first browsing feature value;

[0057] A second active time data matrix is ​​constructed based on the active time period model and the second active feature value;

[0058] A second active feature map is plotted based on the second active time data matrix, wherein the x-axis of the second active feature map represents the active time period and the y-axis represents the second active feature value;

[0059] A second browsing preference data matrix is ​​constructed based on the browsing preference model and the second browsing feature value;

[0060] A second browsing feature map is plotted based on the second browsing preference data matrix, wherein the x-axis of the second browsing feature map represents the browsing time period and the y-axis represents the second browsing feature value.

[0061] Optionally, constructing the first active time data matrix based on the active time period model and the first active feature value includes:

[0062] Acquire active status and working status, wherein the active status includes: high active status, medium active status, low active status and inactive status, and the working status includes: working and resting;

[0063] Construct a first active time data matrix using the active state and working state;

[0064] The first active time data matrix is ​​shown below:

[0065]

[0066] Where R refers to the first active time data matrix, T 11 The first active feature value, T, refers to the active state being highly active and the working state being working during the e-th active time period. 12 T refers to the first active characteristic value during the e-th active time period when the active state is high and the working state is resting. 41 The first active feature value, T, refers to the value of a character who is inactive during the e-th active time period and is active during the working period. 42 This refers to the first active feature value during the e-th active time period when the active state is inactive and the working state is resting.

[0067] Optionally, the step of performing homogeneity processing based on the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map to obtain the homogeneity of active features and browsing features includes:

[0068] Obtain the time weight of the active time period, the second active feature value of the active time period, and the maximum possible distance of the feature values. Calculate the homogeneity of the active features based on the time weight of the active time period, the second active feature value, and the maximum possible distance of the feature values.

[0069]

[0070] Where H1 represents the homogeneity of active features, Q e T represents the time weight of the e-th active time period. e T represents the first active feature value in the e-th active time period. e ' represents the second active feature value of the e-th active time period, D max This represents the maximum possible distance of the feature value.

[0071] Optionally, the step of calculating the mixed homogeneity based on the homogeneity of the activity feature, browsing feature, friend feature, following feature, and preset feature weights includes:

[0072] Using the aforementioned feature weights, the homogeneity of the mixture is calculated according to the following formula, wherein the feature weights include: activity feature weight, browsing feature weight, friend feature weight, and following feature weight:

[0073] H m =ω d ×H1+ω z ×H2+ω f ×H3+ω c ×H4

[0074] Among them, H m Indicates the homogeneity of the mixture, ω d ω represents the active feature weight. z ω represents the browsing feature weights. f ω represents the weight of the friend feature. c H1 represents the weight of the features of interest, H2 represents the homogeneity of browsing features, H3 represents the homogeneity of friend features, and H4 represents the homogeneity of the features of interest.

[0075] Optionally, the construction of the characteristic function includes:

[0076] Construct a feature function based on a preset threshold for homogeneity of mixtures:

[0077]

[0078] Where, f(H) m H represents the characteristic function. y This represents the threshold for homogeneity in the mixture.

[0079] To achieve the above objectives, the present invention also provides a social network identity association system based on big data, comprising:

[0080] The first user feature acquisition module is used to collect user identity big data from two network domains, construct an active time period model, a browsing preference model, a friend relationship model, and a follow relationship model, perform data processing on the user identity big data based on the active time period model, browsing preference model, friend relationship model, and follow relationship model to obtain a first user feature set and a second user feature set, and extract a first user feature based on the first user feature set, wherein the first user feature includes: a first activity feature, a first browsing feature, a first friend feature, and a first follow feature;

[0081] The homogeneity calculation module is used to sequentially extract second user features from the second user feature set. The second user features include: second activity features, second browsing features, second friend features, and second following features. The module draws a first activity feature map, a first browsing feature map, a second activity feature map, and a second browsing feature map using the first activity feature map, the first browsing feature map, the second activity feature map, and the second browsing feature map, respectively. Homogeneity processing is performed based on the first activity feature map, the first browsing feature map, the second activity feature map, and the second browsing feature map to obtain the homogeneity of the activity features and the homogeneity of the browsing features.

[0082] The homogeneity threshold judgment module is used to determine whether the homogeneity of the active feature and the homogeneity of the browsing feature are all greater than a preset homogeneity threshold. If the homogeneity of the active feature and the homogeneity of the browsing feature are not all greater than the homogeneity threshold, the process returns to the above steps of sequentially extracting the second user feature in the second user feature set. If the homogeneity of the active feature and the homogeneity of the browsing feature are all greater than the homogeneity threshold, the homogeneity of the friend feature and the homogeneity of the follow feature are calculated based on the first friend feature, the first follow feature, the second friend feature, and the second follow feature.

[0083] The social network identity association module is used to calculate a mixed homogeneity based on the homogeneity of the active features, browsing features, friend features, and following features, as well as preset feature weights. It constructs a feature function, inputs the mixed homogeneity into the feature function, obtains the function output value, and determines whether the function output value is 1. If the function output value is not 1, it returns to the above steps of sequentially extracting the second user features from the second user feature set. If the function output value is 1, it confirms that the first user feature and the second user feature have completed the social network identity association.

[0084] To address the above problems, the present invention also provides an electronic device, the electronic device comprising:

[0085] Memory, storing at least one instruction;

[0086] The processor executes the instructions stored in the memory to implement the aforementioned big data-based social network identity association method.

[0087] To address the aforementioned problems, the present invention also provides a computer-readable storage medium storing at least one instruction, which is executed by a processor in an electronic device to implement the aforementioned big data-based social network identity association method.

[0088] To address the problems described in the background section, this invention achieves accurate association of social network identities by combining collected user identity big data from two network domains, constructing multiple user feature models, and comprehensively processing the user identity big data. First, by collecting user identity big data from two network domains, an active time period model, a browsing preference model, a friend relationship model, and a follower relationship model are constructed, enabling a comprehensive capture of user behavior and characteristics. This accurately describes user behavior patterns in different network domains, improving the comprehensiveness and representativeness of the data. Then, a first activity feature map, a first browsing feature map, a second activity feature map, and a second browsing feature map are plotted and homogeneity processing is performed, clearly demonstrating the similarity of two user behaviors in different network domains. This is achieved by calculating the activity... The homogeneity of user features and browsing features, along with determining whether they meet a preset homogeneity threshold, can effectively improve the accuracy of identity recognition. Subsequently, if the homogeneity of activity features and browsing features does not fully meet the preset threshold, the second user feature set is iteratively updated. This iterative update continuously optimizes user features and improves the efficiency of social network identity association. Furthermore, by calculating the homogeneity of friend features and following features, and calculating mixed homogeneity based on preset feature weights, multiple feature dimensions of the user are comprehensively considered, providing more comprehensive and accurate data for final identity confirmation. Finally, by constructing a feature function and inputting the mixed homogeneity into the feature function, the matching status of user identities can be automatically determined, resulting in high accuracy in social network identity association. Therefore, this invention can improve the accuracy of social network identity association. Attached Figure Description

[0089] Figure 1 This is a flowchart illustrating a social network identity association method based on big data, provided in an embodiment of the present invention.

[0090] Figure 2 A functional block diagram of a social network identity association system based on big data provided in an embodiment of the present invention;

[0091] Figure 3 This is a schematic diagram of the structure of an electronic device that implements the big data-based social network identity association method according to an embodiment of the present invention.

[0092] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0093] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0094] This application provides a method for social network identity association based on big data. The executing entity of the method includes, but is not limited to, at least one of the following: a server, a terminal, or other electronic device that can be configured to execute the method provided in this application. In other words, the method can be executed by software or hardware installed on a terminal device or a server device, and the software may be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.

[0095] Reference Figure 1 The diagram shown is a flowchart illustrating a big data-based social network identity association method according to an embodiment of the present invention. In this embodiment, the big data-based social network identity association method includes:

[0096] S1. Collect user identity big data from two network domains to construct active time period models, browsing preference models, friend relationship models, and follower relationship models.

[0097] Explainable terms include: network domain (referring to online social networking platforms such as Weibo, Twitter, and Baidu Tieba); user identity big data (referring to publicly available data when users register on online social networking platforms and data generated during their use of these platforms, such as content viewed, browsing duration, and browsing frequency); active time period model (referring to models used to analyze user activity characteristics at different times of the day, where activity characteristics indicate the degree of user activity at different times of the day); browsing preference model (referring to models used to analyze user browsing preferences at different times of the day, where browsing preferences indicate the content users prefer to browse); friend relationship model (referring to models used to analyze user friend relationships, where friend relationships refer to the relationships between friends added by users online on online social networking platforms); and follow relationship model (referring to models used to analyze user follow relationships, where follow relationships refer to the relationships between other users followed by users on online social networking platforms).

[0098] In detail, the collection of user identity big data from the two network domains includes:

[0099] The API authentication mechanisms of two network domains are obtained, and a preset API request is sent to the two network domains based on the API authentication mechanisms of the two network domains and the pre-obtained API key. The request result is obtained, and the API request is used to determine whether the API request successfully connects to the two network domains.

[0100] If the API request fails to connect to the two network domains, return to the steps described above for obtaining the API authentication mechanism for the two network domains.

[0101] If the API request successfully connects the two network domains, a first API request is constructed using preset initial pagination parameters, and first request data is obtained based on the initial pagination parameters and the first API request.

[0102] Determine if the preset total pagination parameter is equal to 1;

[0103] If the total pagination parameter is equal to 1, then the first request data will be used as the user identity big data.

[0104] If the total pagination parameter is not equal to 1, then determine whether the initial pagination parameter is equal to the total pagination parameter;

[0105] If the initial pagination parameter is not equal to the total pagination parameter, then the initial pagination parameter is incremented by one to obtain the updated pagination parameter. The updated pagination parameter is used to update the initial pagination parameter. Then, the initial pagination parameter is used to construct an update API request. Based on the initial pagination parameter and the update API request, update request data is obtained, and the steps of determining whether the initial pagination parameter is equal to the total pagination parameter are returned.

[0106] If the initial pagination parameter is equal to the total pagination parameter, then the update request data and the first request data are merged in the form of the first request data first and the update request data last to obtain the user identity big data.

[0107] Explainable terms: API authentication mechanism refers to the mechanism for verifying identity in an application programming interface (API) to ensure that API requests come from legitimate users or systems. It is typically used to protect sensitive data and prevent unauthorized access. API key refers to a unique string or code generated by the API provider to verify the identity of the API caller and authorize their access. API request refers to a request sent to a network domain through the application programming interface to access data within that network domain. Request result refers to the content returned by the network domain based on the API request. Initial paging parameter refers to the initial page number used to control the amount of identity data retrieved when acquiring user identity data. The initial paging parameter starts from 1 and is used to handle large amounts of data, avoiding excessive server or network pressure caused by retrieving all data at once. First API request refers to the actual request used to retrieve data. The process of retrieving user identity data begins by constructing this request. First requested data refers to the first page of user identity data retrieved according to the initial paging parameter. Total paging parameter refers to the maximum number of pages specified. For example, if retrieving user identity data starts from the first page and the maximum number of pages is specified as 100, the retrieval operation ends after retrieving the 100th page of user identity data. 100 is the total paging parameter.

[0108] Understandably, user identity big data refers to data obtained based on the first request data or data obtained based on the first request data, update request data, and merging rules. Merging rules refer to the rule that the first request data comes first and the update request data comes later. Update pagination parameters refer to the pagination parameters obtained by adding one to the initial pagination parameters when the number of pages is less than or equal to the maximum number of pages specified by the total pagination parameters. For example, if the maximum pagination parameter is 10 pages and the initial pagination parameter is 1 page, then the update pagination parameter is 2 pages. Update API request refers to an API request reconstructed based on the update pagination parameters. The update API request uses the updated page number to obtain a new page of user identity data. Update request data refers to the new page of user identity data obtained using the updated page number.

[0109] For example, if the initial pagination parameter is 1 and the total pagination parameter is 2, the first request data is obtained based on the initial pagination parameter and the first API request. The initial pagination parameter is incremented by one to obtain the updated pagination parameter of 2. The initial pagination parameter is updated using the updated pagination parameter. Then, the updated request data is obtained using the initial pagination parameter and the updated API request. Since the initial pagination parameter is equal to the total pagination parameter, the first request data is placed before the updated request data to obtain the user identity big data.

[0110] S2. Based on the active time period model, browsing preference model, friend relationship model and attention relationship model, perform data processing on the user identity big data to obtain the first user feature set and the second user feature set.

[0111] Understandably, performing data processing on the aforementioned user identity big data refers to the action of inputting the user identity big data into the active time period model, browsing preference model, friend relationship model, and follow relationship model. The first user feature set refers to the set of data obtained after processing user identity big data collected from one online social platform through the active time period model, browsing preference model, friend relationship model, and follow relationship model. The second user feature set refers to the set of data obtained after processing user identity big data collected from another online social platform through the active time period model, browsing preference model, friend relationship model, and follow relationship model.

[0112] The detailed process of constructing the active time period model is as follows:

[0113] Obtain the active state weight, work state weight, state activity level, and work activity level for each active time period. The active time period can be: [0,6)h, [6,16)h, [16,20)h, [20,24)h.

[0114] The active time period model is constructed using the active state weight, work state weight, state activity, and work activity:

[0115] T e =ε×X e +∈×Y e

[0116] Among them, T e This represents the first or second active feature value in the e-th active time period, where e represents the sequence number of the active time period, ε represents the active state weight, and X... e Y represents the state activity level during the e-th active time period, ∈ represents the working state weight, and Y e This represents the work activity level during the e-th active time period.

[0117] Understandably, an active time period refers to the time interval during which a user is active within a day. For example, [16, 20)h and [20, 24)h are active time periods. The active state weight refers to the degree of influence of the active state on the first and second active feature values. The work state weight refers to the degree of influence of the work state on the first and second active feature values. For example, if the influence of the active state on the first and second active feature values ​​is 40%, and the influence of the work state on the first and second active feature values ​​is 60%, then the active state weight is 0.4, and the work state weight is 0.6. Active states include: high activity, medium activity, low activity, and inactive states. High activity, medium activity, low activity, and inactive states all refer to the user's activity level within the current active time period. The higher the activity level, the more active the user is within that active time period. High activity, medium activity, low activity, and inactive states each represent a range. When a user's activity level reaches... When the user's activity level is categorized into high-activity, medium-activity, low-activity, and inactive states, the activity level of the user during the current active time period is confirmed. For example, the high-activity state represents an activity level of (80%, 100%). Work status includes both work and rest, referring to the user's state during the current active time period. Work means the user is working during the current active time period, and rest means the user is resting during the current active time period. Status activity refers to the numerical value of the user's activity status during the current active time period, and work activity refers to the numerical value of the user's work status during the current active time period. The first activity characteristic value refers to the comprehensive activity level of a user in one network domain under different activity and work status conditions. The second activity characteristic value refers to the comprehensive activity level of a user in another network domain under different activity and work status conditions. The comprehensive activity level combines the activity levels of both activity and work status. The activity level is used to determine whether users in two network domains belong to the same user.

[0118] The detailed construction process of the browsing preference model, friend relationship model, and follow relationship model is as follows:

[0119] The system acquires the browsing time period, content type, browsing duration range, browsing frequency, browsing device, browsing location, and content nature. The browsing time period includes [0,6)h, [6,16)h, [16,20)h, and [20,24)h. The content type includes news content, entertainment content, technology content, literature content, and educational content. The browsing duration range includes [10,20)min, [20,40)min, [40,60)min, and [60,+∞)min. The browsing frequency includes: greater than 3 times / day and less than or equal to 5 times / day, greater than 5 times / day and less than or equal to 7 times / day, greater than 7 times / day and less than or equal to 9 times / day, and greater than 9 times / day. The browsing device includes mobile devices and PCs. The browsing location includes urban and suburban areas. The content nature includes the latest news and past content.

[0120] The browsing dimensions are obtained by arranging and combining the browsing time period, content type, browsing duration range, browsing frequency, browsing device, browsing location, and content nature.

[0121] Obtain the browsing preference degree, the dimension ordinal number, and the dimension weight of the browsing dimension, and construct a browsing preference model based on the browsing dimension, the browsing preference degree, the dimension ordinal number, and the dimension weight of the browsing dimension.

[0122] The browsing preference model is as follows:

[0123]

[0124] Among them, B τ (j,k,l,v,w,f) represents the first or second browsing feature value in the τ-th browsing time period when the browsing dimension is (j,k,l,v,w,f), where τ represents the index of the browsing time period, j represents the content type, k represents the browsing duration interval, l represents the browsing frequency, v represents the browsing device, w represents the browsing location, f represents the content nature, n represents the ordinal number of the browsing dimension, and α represents the dimensional ordinal number. i β represents the dimensional weight of the i-th browsing dimension. i This represents the browsing preference level for the i-th browsing dimension;

[0125] The system acquires the interaction frequency, mutual friend relationships, friend interaction degree of interaction frequency, interaction frequency weight, interaction weight of interaction frequency, friend relationship degree of mutual friend relationships, mutual friend weight, and relationship weight of mutual friend relationships. The interaction frequency includes: greater than 1 time / day and less than or equal to 3 times / day, greater than 3 times / day and less than or equal to 5 times / day, greater than 5 times / day and less than or equal to 10 times / day, and greater than 10 times / day. The mutual friend relationships include: greater than 1 mutual friend and less than or equal to 3 mutual friends, greater than 3 mutual friends and less than or equal to 5 mutual friends, greater than 5 mutual friends and less than or equal to 10 mutual friends, and greater than 10 mutual friends.

[0126] A friend relationship model is constructed based on the interaction frequency, mutual friend relationships, friend interaction degree of interaction frequency, interaction frequency weight, interaction frequency interaction weight, friend relationship degree of mutual friend relationships, mutual friend weight, and relationship weight of mutual friend relationships.

[0127] The friend relationship model is as follows:

[0128]

[0129] Where A represents the feature value of the first friend or the feature value of the second friend, θ represents the interaction frequency weight, and B b B represents the friend interaction frequency at the b-th interaction frequency. b1 Let μ represent the interaction weight for the b-th interaction frequency, and O represent the weight of mutual friends. m O represents the friendship degree of the m-th mutual friend. m1 This represents the relationship weight of the m-th mutual friend relationship;

[0130] The system obtains the following parameters: common interest relationships, content similarity intervals, common interest weights, common interest degree of common interest relationships, attention weight of common interest relationships, weight of content similarity intervals, content similarity degree of content similarity intervals, and similarity interval weight of content similarity intervals. The common interest relationships include: more than 1 common interest relationship and less than or equal to 3 common interest relationships, more than 3 common interest relationships and less than or equal to 5 common interest relationships, more than 5 common interest relationships and less than or equal to 10 common interest relationships, and more than 10 common interest relationships. The content similarity intervals include: [0, 10%), [10%, 30%), [30%, 60%), and [60%, 100%).

[0131] The attention relationship model is constructed using the common attention relationship, content similarity interval, common attention weight, common attention degree of common attention relationship, attention weight of common attention relationship, weight of content similarity interval, content similarity of content similarity interval, and similarity interval weight of content similarity interval;

[0132] The attention relationship model is shown below:

[0133]

[0134] Where Y represents the first or second feature value of interest, ρ represents the common interest weight, and M m M represents the degree of common interest in the m-th common interest relationship. m1 This represents the attention weight of the m-th common attention relationship. U represents the weight of content similarity intervals. u U represents the content similarity of the u-th content similarity interval. u1 This represents the weight of the similarity interval for the u-th content similarity interval.

[0135] Understandably, browsing time period refers to the time interval during which a user browses content within a day; content type refers to the type of content the user browses, such as news or sports; browsing duration interval refers to the time a user spends browsing a particular content type; browsing frequency refers to the number of times a user browses content per day; browsing device refers to the device used to browse content; browsing location refers to the geographical location when browsing content; content nature refers to the freshness or age of the content the user browses; browsing dimension refers to the combination of content type, browsing duration interval, browsing frequency, browsing device, browsing location, and content nature. For example, if the content type is news, the browsing duration interval is [10, 20) min, the browsing frequency is more than 3 times / day, the browsing device is a mobile device, the browsing location is a city, and the content nature is latest news, this constitutes one browsing dimension. Browsing preference for all browsing dimensions refers to the degree of user preference for each browsing dimension. Preference degree refers to the degree to which the user likes it. The first browsing characteristic value refers to the comprehensive preference degree of a user under different content types, browsing frequencies, browsing devices, browsing locations, and content nature conditions in one network domain. The second browsing characteristic value refers to the comprehensive preference degree of a user under different content types, browsing frequencies, browsing devices, browsing locations, and content nature conditions in another network domain.

[0136] For example, suppose there are three browsing dimensions, τ = 1, corresponding to a browsing time period of [0, 6)h, the dimension weight of the first browsing dimension α1 = 0.3, the browsing preference of the first browsing dimension β1 = 0.5, the dimension weight of the second browsing dimension α2 = 0.3, the browsing preference of the second browsing dimension β2 = 0.3, the dimension weight of the third browsing dimension α3 = 0.4, the browsing preference of the third browsing dimension β3 = 0.2, B1(j,k,l,v,w,f) = 0.3×0.5 + 0.3×0.3 + 0.4×0.2 = 0.44, and the first browsing feature value or the second browsing feature value under the first browsing time period is 0.44.

[0137] Interpretable, interaction frequency refers to the number of times a user interacts daily; mutual friend relationships refer to the number of mutual friends between two users in one network domain and another; friend interaction degree of interaction frequency refers to the user's activity level at each interaction frequency; interaction frequency weight refers to the degree of influence of interaction frequency on the first and second friend feature values; corresponding interaction weight refers to the proportion of each interaction frequency to all interaction frequencies. For example, the corresponding interaction weight is 0.1 for more than 1 interaction / day, 0.2 for more than 3 interactions / day, 0.3 for more than 5 interactions / day, and 0.4 for more than 10 interactions / day; mutual friend weight refers to the influence of mutual friend relationships on the first and second friend feature values. The influence of the second friend characteristic value is as follows: the friend relationship degree of mutual friend relationships refers to the user's activity level under each mutual friend relationship; the relationship weight of mutual friend relationships refers to the proportion of each mutual friend relationship to all mutual friend relationships. For example, the relationship weight of more than 1 mutual friend relationship is 0.1, the relationship weight of more than 3 mutual friend relationships is 0.2, the relationship weight of more than 5 mutual friend relationships is 0.3, and the relationship weight of more than 10 mutual friend relationships is 0.4. The first friend characteristic value refers to the user's comprehensive activity level under different interaction frequencies and different mutual friend relationship conditions in one network domain; the second friend characteristic value refers to the user's comprehensive activity level under different interaction frequencies and different mutual friend relationship conditions in another network domain.

[0138] Interpretable, common interest relationship refers to the number of interests shared by two users in one network domain and another network domain; content similarity interval refers to the range of similarity in the content viewed by two users in one network domain and another network domain; common interest weight refers to the degree of influence of common interest relationship on the first and second interest feature values; common interest degree of common interest relationship refers to the user's activity level under each common interest relationship; attention weight of common interest relationship refers to the proportion of each common interest relationship to all common interest relationships; content similarity interval weight refers to the degree of influence of content similarity interval on the first and second interest feature values; content similarity of content similarity interval refers to the similarity in the content viewed by two users in one network domain and another network domain; corresponding similarity interval weight of content similarity interval refers to the proportion of each content similarity interval to all content similarity intervals.

[0139] S3. Extract first user features based on the first user feature set, wherein the first user features include: first activity features, first browsing features, first friend features, and first follow features.

[0140] Understandably, the first user feature refers to the activity features, browsing features, friend features, and following features of each user in the first user feature set; the second user feature refers to the activity features, browsing features, friend features, and following features of each user in the second user feature set; the first activity feature refers to the activity features of each user in the first user feature set; the first browsing feature refers to the browsing features of each user in the first user feature set; the first friend feature refers to the friend features of each user in the first user feature set; and the first following feature refers to the following features of each user in the first user feature set.

[0141] S4. Extract the second user features sequentially from the second user feature set, wherein the second user features include: second activity features, second browsing features, second friend features, and second follow features.

[0142] Understandably, the second activity feature refers to the activity features of each user in the second user feature set, the second browsing feature refers to the browsing features of each user in the second user feature set, the second friend feature refers to the friend features of each user in the second user feature set, and the second following feature refers to the following features of each user in the second user feature set.

[0143] S5. Using the first active feature, the first browsing feature, the second active feature, and the second browsing feature, respectively draw the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map.

[0144] Understandably, the first active feature graph refers to a line graph that divides a day into four time periods, using the four time periods as the x-axis and the first active feature value as the y-axis; the first browsing feature graph refers to a line graph that divides a day into four time periods, using the four time periods as the x-axis and the first browsing feature value as the y-axis; the second active feature graph refers to a line graph that divides a day into four time periods, using the four time periods as the x-axis and the second active feature value as the y-axis; and the second browsing feature graph refers to a line graph that divides a day into four time periods, using the four time periods as the x-axis and the second browsing feature value as the y-axis.

[0145] In detail, the step of drawing a first activity feature map, a first browsing feature map, a second activity feature map, and a second browsing feature map using the first activity feature, the first browsing feature, the second activity feature, and the second browsing feature, respectively, includes:

[0146] A first active time data matrix is ​​constructed based on the active time period model and the first active feature value;

[0147] A first active feature map is drawn based on the first active time data matrix, wherein the x-axis of the first active feature map represents the active time period and the y-axis represents the first active feature value;

[0148] A first browsing preference data matrix is ​​constructed based on the browsing preference model and the first browsing feature value;

[0149] A first browsing feature map is drawn based on the first browsing preference data matrix, wherein the x-axis of the first browsing feature map represents the browsing time period and the y-axis represents the first browsing feature value;

[0150] A second active time data matrix is ​​constructed based on the active time period model and the second active feature value;

[0151] A second active feature map is plotted based on the second active time data matrix, wherein the x-axis of the second active feature map represents the active time period and the y-axis represents the second active feature value;

[0152] A second browsing preference data matrix is ​​constructed based on the browsing preference model and the second browsing feature value;

[0153] A second browsing feature map is plotted based on the second browsing preference data matrix, wherein the x-axis of the second browsing feature map represents the browsing time period and the y-axis represents the second browsing feature value.

[0154] Understandably, the first active time data matrix refers to the matrix composed of the first active feature values, the first browsing preference data matrix refers to the matrix composed of the first browsing feature values, the second active time data matrix refers to the matrix composed of the second active feature values, and the second browsing preference data matrix refers to the matrix composed of the second browsing feature values.

[0155] Specifically, the construction of the first active time data matrix based on the active time period model and the first active feature value includes:

[0156] Acquire active status and working status, wherein the active status includes: high active status, medium active status, low active status and inactive status, and the working status includes: working and resting;

[0157] Construct a first active time data matrix using the active state and working state;

[0158] The first active time data matrix is ​​shown below:

[0159]

[0160] Where R refers to the first active time data matrix, T 11 The first active feature value, T, refers to the active state being highly active and the working state being working during the e-th active time period. 12 T refers to the first active characteristic value during the e-th active time period when the active state is high and the working state is resting. 41The first active feature value, T, refers to the value of a character who is inactive during the e-th active time period and is active during the working period. 42 This refers to the first active feature value during the e-th active time period when the active state is inactive and the working state is resting.

[0161] S6. Perform homogeneity processing based on the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map to obtain the homogeneity of active features and the homogeneity of browsing features.

[0162] Understandably, performing homogeneity processing refers to the action of calculating the homogeneity of active features and browsing features based on the first active feature map and the second active feature map, and based on the first browsing feature map and the second browsing feature map. The homogeneity of active features and browsing features refers to the degree of similarity between the first active feature map and the second active feature map, and between the first browsing feature map and the second browsing feature map. The higher the homogeneity, the more likely the two users in the two network domains are to be the same natural person.

[0163] In detail, the homogeneity processing based on the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map to obtain the homogeneity of active features and browsing features includes:

[0164] Obtain the time weight of the active time period, the second active feature value of the active time period, and the maximum possible distance of the feature values. Calculate the homogeneity of the active features based on the time weight of the active time period, the second active feature value, and the maximum possible distance of the feature values.

[0165]

[0166] Where H1 represents the homogeneity of active features, Q e T represents the time weight of the e-th active time period. e T represents the first active feature value in the e-th active time period. e ' represents the second active feature value of the e-th active time period, D max This represents the maximum possible distance of the feature value.

[0167] Understandably, the time weight of all active time periods refers to the weight of the four active time periods. For example, the weight of the first active time period is 0.2, the weight of the first active time period is 0.4, the weight of the first active time period is 0.3, and the weight of the first active time period is 0.1. The second active feature value of all active time periods refers to the second active feature value of all active time periods under another network domain. The maximum possible distance of the feature value refers to the maximum difference between the first active feature value and the second active feature value under all active event segments.

[0168] S7. Determine whether the homogeneity of the active feature and the homogeneity of the browsing feature are both greater than the preset homogeneity threshold.

[0169] Understandably, the homogeneity threshold refers to a preset critical value used to measure the similarity between active features and browsing features.

[0170] If the homogeneity of the active feature and the homogeneity of the browsing feature are not all greater than the homogeneity threshold, then return to the steps described above of sequentially extracting the second user features from the second user feature set.

[0171] If the homogeneity of the active feature and the homogeneity of the browsing feature are both greater than the homogeneity threshold, then execute S8 to calculate the homogeneity of the friend feature and the homogeneity of the follow feature based on the first friend feature, the first follow feature, the second friend feature and the second follow feature.

[0172] Explainable, the second-iteration user features refer to the activity features, browsing features, friend features, and following features of the second user in the second user feature set. Friend feature homogeneity refers to the degree of similarity between the first friend features and the second friend features, and following feature homogeneity refers to the degree of similarity between the first following features and the second following features.

[0173] S9. Calculate the mixed homogeneity based on the homogeneity of the active feature, the homogeneity of the browsing feature, the homogeneity of the friend feature, the homogeneity of the following feature, and the preset feature weights.

[0174] Understandably, feature weight refers to the importance of homogeneity of friend features, homogeneity of following features, homogeneity of browsing features, and homogeneity of activity features. Mixed homogeneity refers to an overall homogeneity that is synthesized by assigning different weights to friend features, homogeneity of following features, homogeneity of browsing features, and homogeneity of activity features.

[0175] In detail, the calculation of mixed homogeneity based on the homogeneity of activity features, browsing features, friend features, following features, and preset feature weights includes:

[0176] Using the aforementioned feature weights, the homogeneity of the mixture is calculated according to the following formula, wherein the feature weights include: activity feature weight, browsing feature weight, friend feature weight, and following feature weight:

[0177] H m =ω d ×H1+ω z ×H2+ω f ×H3+ω c ×H4

[0178] Among them, H m Indicates the homogeneity of the mixture, ω d ω represents the active feature weight.z ω represents the browsing feature weights. f ω represents the weight of the friend feature. c H1 represents the weight of the features of interest, H2 represents the homogeneity of browsing features, H3 represents the homogeneity of friend features, and H4 represents the homogeneity of the features of interest.

[0179] Understandably, the active feature weight refers to the degree of influence of the homogeneity of the active feature on the mixed homogeneity, the browsing feature weight refers to the degree of influence of the homogeneity of the browsing feature on the mixed homogeneity, the friend feature weight refers to the degree of influence of the homogeneity of the friend feature on the mixed homogeneity, and the attention feature weight refers to the degree of influence of the homogeneity of the attention feature on the mixed homogeneity.

[0180] S10. Construct a feature function, input the homogeneity of the mixture into the feature function, and obtain the function output value.

[0181] Explainable, the feature function refers to the function that determines whether two users in two network domains are the same natural person based on the homogeneity of the mixture and the homogeneity threshold. The homogeneity threshold refers to a preset threshold value used to measure the degree of similarity. The function output value refers to the output used to determine whether two users in two network domains are the same natural person.

[0182] In detail, the construction of the characteristic function includes:

[0183] Construct a feature function based on a preset threshold for homogeneity of mixtures:

[0184]

[0185] Where, f(H) m H represents the characteristic function. y This represents the threshold for homogeneity in the mixture.

[0186] For example, assuming the homogeneity threshold is 0.9 and the homogeneity is 0.92, the feature function outputs 1.

[0187] S11. Determine whether the output value of the function is 1.

[0188] If the function output value is not 1, then return to the steps described above of sequentially extracting the second user features from the second user feature set.

[0189] If the function output value is 1, then execute S12 to confirm that the first user feature and the second user feature have completed the social network identity association.

[0190] To address the problems described in the background section, this invention achieves accurate association of social network identities by combining collected user identity big data from two network domains, constructing multiple user feature models, and comprehensively processing the user identity big data. First, by collecting user identity big data from two network domains, an active time period model, a browsing preference model, a friend relationship model, and a follower relationship model are constructed, enabling a comprehensive capture of user behavior and characteristics. This accurately describes user behavior patterns in different network domains, improving the comprehensiveness and representativeness of the data. Then, a first activity feature map, a first browsing feature map, a second activity feature map, and a second browsing feature map are plotted and homogeneity processing is performed, clearly demonstrating the similarity of two user behaviors in different network domains. This is achieved by calculating the activity... The homogeneity of user features and browsing features, along with determining whether they meet a preset homogeneity threshold, can effectively improve the accuracy of identity recognition. Subsequently, if the homogeneity of activity features and browsing features does not fully meet the preset threshold, the second user feature set is iteratively updated. This iterative update continuously optimizes user features and improves the efficiency of social network identity association. Furthermore, by calculating the homogeneity of friend features and following features, and calculating mixed homogeneity based on preset feature weights, multiple feature dimensions of the user are comprehensively considered, providing more comprehensive and accurate data for final identity confirmation. Finally, by constructing a feature function and inputting the mixed homogeneity into the feature function, the matching status of user identities can be automatically determined, resulting in high accuracy in social network identity association. Therefore, this invention can improve the accuracy of social network identity association.

[0191] like Figure 2 The diagram shown is a functional block diagram of a social network identity association system based on big data provided in an embodiment of the present invention.

[0192] The big data-based social network identity association system 100 described in this invention can be installed in an electronic device. Depending on the functions implemented, the big data-based social network identity association system 100 may include a first user feature acquisition module 101, a homogeneity calculation module 102, a homogeneity threshold judgment module 103, and a social network identity association module 104. The module described in this invention can also be called a unit, referring to a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, stored in the memory of the electronic device.

[0193] The first user feature acquisition module 101 is used to collect user identity big data from two network domains, construct an active time period model, a browsing preference model, a friend relationship model, and a follow relationship model, perform data processing on the user identity big data based on the active time period model, browsing preference model, friend relationship model, and follow relationship model to obtain a first user feature set and a second user feature set, and extract a first user feature based on the first user feature set, wherein the first user feature includes: a first activity feature, a first browsing feature, a first friend feature, and a first follow feature;

[0194] The homogeneity calculation module 102 is used to sequentially extract second user features from the second user feature set, wherein the second user features include: second activity features, second browsing features, second friend features, and second following features. The module draws a first activity feature map, a first browsing feature map, a second activity feature map, and a second browsing feature map using the first activity feature map, the first browsing feature map, the second activity feature map, and the second browsing feature map, respectively. Homogeneity processing is performed based on the first activity feature map, the first browsing feature map, the second activity feature map, and the second browsing feature map to obtain the homogeneity of the activity features and the homogeneity of the browsing features.

[0195] The homogeneity threshold judgment module 103 is used to determine whether the homogeneity of the active feature and the homogeneity of the browsing feature are all greater than a preset homogeneity threshold. If the homogeneity of the active feature and the homogeneity of the browsing feature are not all greater than the homogeneity threshold, the step of sequentially extracting the second user feature in the second user feature set is returned. If the homogeneity of the active feature and the homogeneity of the browsing feature are all greater than the homogeneity threshold, the homogeneity of the friend feature and the homogeneity of the follow feature are calculated based on the first friend feature, the first follow feature, the second friend feature and the second follow feature.

[0196] The social network identity association module 104 is used to calculate a mixed homogeneity based on the homogeneity of the active features, browsing features, friend features, following features, and preset feature weights, construct a feature function, input the mixed homogeneity into the feature function, obtain the function output value, determine whether the function output value is 1, if the function output value is not 1, then return to the above steps of sequentially extracting the second user features in the second user feature set, if the function output value is 1, then confirm that the first user feature and the second user feature have completed the social network identity association.

[0197] In detail, the modules in the big data-based social network identity association system 100 described in this embodiment of the invention employ the same methods as described above. Figure 1 The method uses the same technical means as the big data-based social network identity association method described in the article and can produce the same technical effect, so it will not be repeated here.

[0198] like Figure 3 The diagram shown is a structural schematic of an electronic device that implements a big data-based social network identity association method according to an embodiment of the present invention.

[0199] The electronic device 1 may include a processor 10, a memory 11 and a bus 12, and may also include a computer program stored in the memory 11 and capable of running on the processor 10, such as a social network identity association method program based on big data.

[0200] The memory 11 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as a portable hard drive. In other embodiments, the memory 11 can be an external storage device of the electronic device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device 1. Furthermore, the memory 11 includes both internal storage units and external storage devices of the electronic device 1. The memory 11 can be used not only to store application software and various types of data installed on the electronic device 1, such as the code of a social network identity association method program based on big data, but also to temporarily store data that has been output or will be output.

[0201] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 11 (e.g., a social network identity association method program based on big data) and calls data stored in the memory 11 to perform various functions of the electronic device 1 and process data.

[0202] The bus 12 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is configured to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0203] Figure 3 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 3 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0204] For example, although not shown, the electronic device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0205] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the electronic device 1 and other electronic devices.

[0206] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), or a standard wired or wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device 1 and to display a visual user interface.

[0207] The social network identity association method program based on big data, stored in the memory 11 of the electronic device 1, is a combination of multiple instructions. When run in the processor 10, it can achieve the following:

[0208] Collect user identity big data from two network domains to construct active time period models, browsing preference models, friend relationship models, and follow relationship models;

[0209] Based on the active time period model, browsing preference model, friend relationship model and follow relationship model, data processing is performed on the user identity big data to obtain a first user feature set and a second user feature set.

[0210] First user features are extracted based on the first user feature set, wherein the first user features include: first activity features, first browsing features, first friend features, and first follow features;

[0211] The second user features are extracted sequentially from the second user feature set, wherein the second user features include: second activity features, second browsing features, second friend features, and second follow features;

[0212] Using the first active feature, the first browsing feature, the second active feature, and the second browsing feature, respectively draw the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map;

[0213] Based on the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map, homogeneity processing is performed to obtain the homogeneity of active features and the homogeneity of browsing features;

[0214] Determine whether the homogeneity of the active features and the homogeneity of the browsing features are all greater than a preset homogeneity threshold;

[0215] If the homogeneity of the active feature and the homogeneity of the browsing feature are not all greater than the homogeneity threshold, then return to the above steps of sequentially extracting the second user features from the second user feature set.

[0216] If the homogeneity of the active feature and the homogeneity of the browsing feature are all greater than the homogeneity threshold, then the homogeneity of the friend feature and the homogeneity of the follow feature are calculated based on the first friend feature, the first follow feature, the second friend feature and the second follow feature.

[0217] The mixed homogeneity is calculated based on the homogeneity of the active feature, the browsing feature, the friend feature, the following feature, and the preset feature weights.

[0218] Construct a feature function, input the homogeneity of the mixture into the feature function, and obtain the function output value;

[0219] Determine whether the function output value is 1;

[0220] If the function output value is not 1, then return to the steps of sequentially extracting the second user features from the second user feature set.

[0221] If the function outputs a value of 1, it confirms that the first user feature and the second user feature have completed the social network identity association.

[0222] Specifically, the processor 10's implementation method for the above instructions can be found in [reference needed]. Figures 1 to 3 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.

[0223] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).

[0224] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following:

[0225] Collect user identity big data from two network domains to construct active time period models, browsing preference models, friend relationship models, and follow relationship models;

[0226] Based on the active time period model, browsing preference model, friend relationship model and follow relationship model, data processing is performed on the user identity big data to obtain a first user feature set and a second user feature set.

[0227] First user features are extracted based on the first user feature set, wherein the first user features include: first activity features, first browsing features, first friend features, and first follow features;

[0228] The second user features are extracted sequentially from the second user feature set, wherein the second user features include: second activity features, second browsing features, second friend features, and second follow features;

[0229] Using the first active feature, the first browsing feature, the second active feature, and the second browsing feature, respectively draw the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map;

[0230] Based on the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map, homogeneity processing is performed to obtain the homogeneity of active features and the homogeneity of browsing features;

[0231] Determine whether the homogeneity of the active features and the homogeneity of the browsing features are all greater than a preset homogeneity threshold;

[0232] If the homogeneity of the active feature and the homogeneity of the browsing feature are not all greater than the homogeneity threshold, then return to the above steps of sequentially extracting the second user features from the second user feature set.

[0233] If the homogeneity of the active feature and the homogeneity of the browsing feature are all greater than the homogeneity threshold, then the homogeneity of the friend feature and the homogeneity of the follow feature are calculated based on the first friend feature, the first follow feature, the second friend feature and the second follow feature.

[0234] The mixed homogeneity is calculated based on the homogeneity of the active feature, the browsing feature, the friend feature, the following feature, and the preset feature weights.

[0235] Construct a feature function, input the homogeneity of the mixture into the feature function, and obtain the function output value;

[0236] Determine whether the function output value is 1;

[0237] If the function output value is not 1, then return to the steps of sequentially extracting the second user features from the second user feature set.

[0238] If the function outputs a value of 1, it confirms that the first user feature and the second user feature have completed the social network identity association.

[0239] In the embodiments provided by this invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and actual implementations may have other classification methods.

[0240] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0241] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0242] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0243] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for social network identity association based on big data, characterized in that, The method includes: Collect user identity big data from two network domains to construct active time period models, browsing preference models, friend relationship models, and follow relationship models; Among them, the active time period model refers to the model used to analyze the active feature values ​​of users in different time periods of the day. The active feature value refers to the degree of user activity in different time periods of the day. The browsing preference model refers to the model used to analyze the browsing preferences of users in different time periods of the day. The browsing preference refers to the content that users prefer to browse. The friend relationship model refers to the model used to analyze the friend relationships of users. The friend relationship refers to the relationship between friends that users have added online on the social networking platform. The follow relationship model refers to the model used to analyze the follow relationship of users. The follow relationship refers to the relationship between users who follow other users on the social networking platform. Based on the active time period model, browsing preference model, friend relationship model and follow relationship model, data processing is performed on the user identity big data to obtain a first user feature set and a second user feature set. First user features are extracted based on the first user feature set, wherein the first user features include: first activity features, first browsing features, first friend features, and first follow features; The second user features are extracted sequentially from the second user feature set, wherein the second user features include: second activity features, second browsing features, second friend features, and second follow features; Using the first active feature, the first browsing feature, the second active feature, and the second browsing feature, respectively draw the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map; Based on the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map, homogeneity processing is performed to obtain the homogeneity of active features and the homogeneity of browsing features; Determine whether the homogeneity of the active features and the homogeneity of the browsing features are all greater than a preset homogeneity threshold; If the homogeneity of the active feature and the homogeneity of the browsing feature are not all greater than the homogeneity threshold, then return to the above steps of sequentially extracting the second user features from the second user feature set. If the homogeneity of the active feature and the homogeneity of the browsing feature are all greater than the homogeneity threshold, then the homogeneity of the friend feature and the homogeneity of the follow feature are calculated based on the first friend feature, the first follow feature, the second friend feature and the second follow feature. The mixed homogeneity is calculated based on the homogeneity of the active feature, the browsing feature, the friend feature, the following feature, and the preset feature weights. Construct a feature function, input the homogeneity of the mixture into the feature function, and obtain the function output value; The construction of the characteristic function includes: Construct a feature function based on a preset threshold for homogeneity of mixtures: in, Represents the characteristic function. Indicates the threshold of homogeneity in the mixture. Indicates the homogeneity of the mixture; Determine whether the function output value is 1; If the function output value is not 1, then return to the steps of sequentially extracting the second user features from the second user feature set. If the function outputs a value of 1, it confirms that the first user feature and the second user feature have completed the social network identity association.

2. The social network identity association method based on big data as described in claim 1, characterized in that, The collected user identity big data from the two network domains includes: The API authentication mechanisms of two network domains are obtained, and a preset API request is sent to the two network domains based on the API authentication mechanisms of the two network domains and the pre-obtained API key. The request result is obtained, and the API request is used to determine whether the API request successfully connects to the two network domains. If the API request fails to connect to the two network domains, return to the steps described above for obtaining the API authentication mechanism for the two network domains. If the API request successfully connects the two network domains, a first API request is constructed using preset initial pagination parameters, and first request data is obtained based on the initial pagination parameters and the first API request. Determine if the preset total pagination parameter is equal to 1; If the total pagination parameter is equal to 1, then the first request data will be used as the user identity big data. If the total pagination parameter is not equal to 1, then determine whether the initial pagination parameter is equal to the total pagination parameter; If the initial pagination parameter is not equal to the total pagination parameter, then the initial pagination parameter is incremented by one to obtain the updated pagination parameter. The updated pagination parameter is used to update the initial pagination parameter. Then, the initial pagination parameter is used to construct an update API request. Based on the initial pagination parameter and the update API request, update request data is obtained, and the steps of determining whether the initial pagination parameter is equal to the total pagination parameter are returned. If the initial pagination parameter is equal to the total pagination parameter, then the update request data and the first request data are merged in the form of the first request data first and the update request data last to obtain the user identity big data.

3. The social network identity association method based on big data as described in claim 2, characterized in that, The construction process of the active time period model is as follows: Obtain the activity status weight, work status weight, status activity level, and work activity level for each active time period. The active time period can be: h, h, h, h; The active time period model is constructed using the active state weight, work state weight, state activity, and work activity: in, This represents the first or second active feature value within the e-th active time period, where e represents the sequence number of the active time period. Indicates the weight of the active state. Indicates the first Status activity level during a specific active time period Indicates the weight of the working status. Indicates the first Work activity level within an active time period. Status activity level refers to the user's active status within the current active time period, while work activity level refers to the user's work status within the current active time period. Active statuses include: high activity, medium activity, low activity, and inactive. Each of these states indicates the user's level of activity within the current active time period; higher activity levels indicate greater user activity during that period. Each of these states represents a range. When a user's activity level reaches the range represented by a high activity, medium activity, low activity, or inactive state, their active status within the current active time period is confirmed.

4. The social network identity association method based on big data as described in claim 3, characterized in that, The construction process of the browsing preference model, friend relationship model, and follow relationship model is as follows: The system acquires the browsing time period, content type, browsing duration range, browsing frequency, browsing device, browsing location, and content nature, wherein the browsing time period includes: h, h, h, h, the content types include: news content, entertainment content, technology content, literary content, and educational content, and the browsing time range includes: min, min, min, The browsing frequency includes: greater than 3 times / day and less than or equal to 5 times / day, greater than 5 times / day and less than or equal to 7 times / day, greater than 7 times / day and less than or equal to 9 times / day, and greater than 9 times / day. The browsing devices include: mobile devices and PCs. The browsing locations include: cities and suburbs. The content nature includes: latest news and past content. The browsing dimensions are obtained by arranging and combining the browsing time period, content type, browsing duration range, browsing frequency, browsing device, browsing location, and content nature. Obtain the browsing preference degree, the dimension ordinal number, and the dimension weight of the browsing dimension, and construct a browsing preference model based on the browsing dimension, the browsing preference degree, the dimension ordinal number, and the dimension weight of the browsing dimension. The browsing preference model is as follows: in, Indicates the first Within a browsing time period and with browsing dimensions as ( The first or second browsing feature value at the time of browsing. Indicates the sequence number of the browsing time period. Indicates the content type. Indicates the browsing time range. Indicates browsing frequency. Indicates the browsing device. Indicates the browsing location. Indicates the nature of the content. The ordinal number of the browsing dimension. This represents the dimensional weight of the i-th browsing dimension. This represents the browsing preference level for the i-th browsing dimension; The system acquires interaction frequency, mutual friend relationships, friend interaction degree of interaction frequency, interaction frequency weight, interaction weight of interaction frequency, friend relationship degree of mutual friend relationships, mutual friend weight, and relationship weight of mutual friend relationships. The interaction frequency includes: greater than 1 time / day and less than or equal to 3 times / day, greater than 3 times / day and less than or equal to 5 times / day, greater than 5 times / day and less than or equal to 10 times / day, and greater than 10 times / day. The mutual friend relationships include: greater than 1 mutual friend and less than or equal to 3 mutual friends, greater than 3 mutual friends and less than or equal to 5 mutual friends, greater than 5 mutual friends and less than or equal to 10 mutual friends, and greater than 10 mutual friends. The friend interaction degree of interaction frequency refers to the user's activity level at each interaction frequency. A friend relationship model is constructed based on the interaction frequency, mutual friend relationships, friend interaction degree of interaction frequency, interaction frequency weight, interaction frequency interaction weight, friend relationship degree of mutual friend relationships, mutual friend weight, and relationship weight of mutual friend relationships. The friend relationship degree of mutual friend relationships refers to the user's activity level under each mutual friend relationship. The friend relationship model is as follows: in, This represents the feature value of the first friend or the feature value of the second friend. Indicates the interaction frequency weight. This represents the friend interaction frequency at the b-th interaction frequency. This represents the interaction weight for the b-th interaction frequency. Indicates the weight of mutual friends. This represents the friendship degree of the m-th mutual friend. This represents the relationship weight of the m-th mutual friend relationship; This document retrieves the following parameters: mutual following relationships, content similarity intervals, mutual following weights, mutual following degree of mutual following relationships, mutual following weight of mutual following relationships, weight of content similarity intervals, content similarity of content similarity intervals, and similarity interval weight of content similarity intervals. The mutual following relationships include: greater than 1 and less than or equal to 3 mutual following relationships, greater than 3 and less than or equal to 5 mutual following relationships, greater than 5 and less than or equal to 10 mutual following relationships, and greater than 10 mutual following relationships. The content similarity intervals include: , , and ; The attention relationship model is constructed using the common attention relationship, content similarity interval, common attention weight, common attention degree of common attention relationship, attention weight of common attention relationship, weight of content similarity interval, content similarity of content similarity interval, and similarity interval weight of content similarity interval; The attention relationship model is shown below: in, This indicates the first or second feature value of interest. This indicates a shared focus on weights. This represents the degree of common interest in the m-th common interest relationship. This represents the attention weight of the m-th common attention relationship. Indicates the weight of content similarity intervals. This represents the content similarity of the u-th content similarity interval. This represents the weight of the similarity interval for the u-th content similarity interval.

5. The social network identity association method based on big data as described in claim 4, characterized in that, The step of drawing a first activity feature map, a first browsing feature map, a second activity feature map, and a second browsing feature map using the first activity feature, the first browsing feature, the second activity feature, and the second browsing feature, respectively, includes: A first active time data matrix is ​​constructed based on the active time period model and the first active feature value; A first active feature map is drawn based on the first active time data matrix, wherein the x-axis of the first active feature map represents the active time period and the y-axis represents the first active feature value; A first browsing preference data matrix is ​​constructed based on the browsing preference model and the first browsing feature value; A first browsing feature map is drawn based on the first browsing preference data matrix, wherein the x-axis of the first browsing feature map represents the browsing time period and the y-axis represents the first browsing feature value; A second active time data matrix is ​​constructed based on the active time period model and the second active feature value; A second active feature map is plotted based on the second active time data matrix, wherein the x-axis of the second active feature map represents the active time period and the y-axis represents the second active feature value; A second browsing preference data matrix is ​​constructed based on the browsing preference model and the second browsing feature value; A second browsing feature map is plotted based on the second browsing preference data matrix, wherein the x-axis of the second browsing feature map represents the browsing time period and the y-axis represents the second browsing feature value.

6. The social network identity association method based on big data as described in claim 5, characterized in that, The construction of the first active time data matrix based on the active time period model and the first active feature value includes: Acquire active status and working status, wherein the active status includes: high active status, medium active status, low active status and inactive status, and the working status includes: working and resting; Construct a first active time data matrix using the active state and working state; The first active time data matrix is ​​shown below: in, Refers to the first active time data matrix. This refers to the first active feature value during the e-th active time period when the active state is high-activity and the working state is working. This refers to the first active feature value during the e-th active time period when the active state is high-activity and the working state is resting. This refers to the first active feature value during the e-th active time period when the active state is inactive and the working state is active. This refers to the first active feature value during the e-th active time period when the active state is inactive and the working state is resting.

7. The social network identity association method based on big data as described in claim 6, characterized in that, The process of performing homogeneity processing based on the first active feature map, the first browsing feature map, the second active feature map, and the second browsing feature map to obtain the homogeneity of active features and browsing features includes: Obtain the time weight of the active time period, the second active feature value of the active time period, and the maximum possible distance of the feature values. Calculate the homogeneity of the active features based on the time weight of the active time period, the second active feature value, and the maximum possible distance of the feature values. in, Indicates the homogeneity of active features. This represents the time weight of the e-th active time period. This represents the first active feature value during the e-th active time period. This represents the second active feature value during the e-th active time period. This represents the maximum possible distance of the feature value.

8. The social network identity association method based on big data as described in claim 7, characterized in that, The calculation of mixed homogeneity based on the homogeneity of activity features, browsing features, friend features, following features, and preset feature weights includes: Using the aforementioned feature weights, the homogeneity of the mixture is calculated according to the following formula, wherein the feature weights include: activity feature weight, browsing feature weight, friend feature weight, and following feature weight: in, Indicates the degree of homogeneity of the mixture. Represents the active feature weights. Indicates the browsing feature weights. Indicates the weight of friend features. This indicates that the focus is on the feature weights. Indicates the homogeneity of browsing features. Indicates the homogeneity of friends' characteristics. This indicates a focus on the homogeneity of features.

9. A social network identity association system based on big data, characterized in that, The system includes: The first user feature acquisition module is used to collect user identity big data from two network domains, construct an active time period model, a browsing preference model, a friend relationship model, and an attention relationship model, and perform data processing on the user identity big data based on the active time period model, browsing preference model, friend relationship model, and attention relationship model to obtain a first user feature set and a second user feature set. Based on the first user feature set, a first user feature is extracted, wherein the first user feature includes: a first activity feature, a first browsing feature, a first friend feature, and a first attention feature. The active time period model refers to a model used to analyze the user's activity feature values ​​at different times of the day, where the activity feature value refers to the degree of user activity at different times of the day; the browsing preference model refers to a model used to analyze the user's browsing preferences at different times of the day, where browsing preferences refer to the content that the user prefers to browse; the friend relationship model refers to a model used to analyze the user's friend relationships, where friend relationships refer to the relationships between friends added by the user online on the network social platform; and the attention relationship model refers to a model used to analyze the user's attention relationships, where attention relationships refer to the relationships between other users that the user follows on the network social platform. The homogeneity calculation module is used to sequentially extract second user features from the second user feature set. The second user features include: second activity features, second browsing features, second friend features, and second following features. The module draws a first activity feature map, a first browsing feature map, a second activity feature map, and a second browsing feature map using the first activity feature map, the first browsing feature map, the second activity feature map, and the second browsing feature map, respectively. Homogeneity processing is performed based on the first activity feature map, the first browsing feature map, the second activity feature map, and the second browsing feature map to obtain the homogeneity of the activity features and the homogeneity of the browsing features. The homogeneity threshold judgment module is used to determine whether the homogeneity of the active feature and the homogeneity of the browsing feature are all greater than a preset homogeneity threshold. If the homogeneity of the active feature and the homogeneity of the browsing feature are not all greater than the homogeneity threshold, the process returns to the above steps of sequentially extracting the second user feature in the second user feature set. If the homogeneity of the active feature and the homogeneity of the browsing feature are all greater than the homogeneity threshold, the homogeneity of the friend feature and the homogeneity of the follow feature are calculated based on the first friend feature, the first follow feature, the second friend feature, and the second follow feature. The social network identity association module is used to calculate a mixed homogeneity based on the homogeneity of the activity feature, browsing feature, friend feature, and following feature, as well as preset feature weights; construct a feature function; input the mixed homogeneity into the feature function to obtain the function output value; determine whether the function output value is 1; if the function output value is not 1, return to the steps of sequentially extracting the second user features from the second user feature set; if the function output value is 1, confirm that the first user feature and the second user feature have completed the social network identity association. The construction of the feature function includes: Construct a feature function based on a preset threshold for homogeneity of mixtures: in, Represents the characteristic function. Indicates the threshold of homogeneity in the mixture. Indicates the degree of homogeneity of the mixture.

Citation Information

Patent Citations

  • User identity association method and device

    CN110046293A

  • Cross-domain identity association method and system based on multivariate relation representation

    CN112288007A