A Social Network Community Discovery Method Based on Density Peak Clustering with Neighborhood Fuzzy Kernel

By obtaining the impact index of the neighborhood fuzzy kernel in real time in social networks, and updating the neighborhood fuzzy kernel for clustering, the problem of low timeliness of social network community discovery is solved, and faster and more accurate community hot spot discovery is achieved.

CN119939045BActive Publication Date: 2025-07-11HUNAN UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510436189.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The prior art social network community is found to be in low timeliness in dynamically changing social networks, and key parameters need to be set in advance, making it difficult to quickly update community divisions.

Method used

By collecting social network data within the preset community discovery cycle, obtaining the impact index of the neighborhood fuzzy kernel, determining whether to update the neighborhood fuzzy kernel, performing clustering processing and calculating similarity deviations, and updating the community collection in real time.

Benefits of technology

It improves the timeliness and accuracy of hot community discovery in social networks, can adapt to changes in network structure more quickly, and achieve more accurate community division.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939045B_ABST
    Figure CN119939045B_ABST
Patent Text Reader

Abstract

The present invention discloses a social network community discovery method based on density peak clustering with neighborhood fuzzy kernel, which relates to the technical field of data clustering analysis. The social network community discovery method based on density peak clustering with neighborhood fuzzy kernel includes the following steps: obtaining the neighborhood fuzzy kernel, obtaining the community set, and updating the community set. The present invention obtains the corresponding neighborhood fuzzy kernel through all user data collected in the social network, and obtains the applicable influence index of the neighborhood fuzzy kernel by real-time obtaining community change data, thereby judging whether to update the neighborhood fuzzy kernel, performing clustering processing on the neighborhood fuzzy kernel to obtain the community set, and obtaining the community division set based on the community set and the community set in the previous preset time period, and at the same time obtaining the similarity deviation to update the community set, achieving the effect of improving the timeliness of hot community discovery in the social network and solving the problem of low timeliness of social network community discovery in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data clustering analysis, and in particular to a social network community discovery method based on density peak clustering of neighborhood fuzzy kernels. Background Art

[0002] As urbanization accelerates, understanding and managing urban communities becomes increasingly complex. Community hotspots refer to areas where user activities are frequent and people gather in large numbers. These areas usually have high commercial value and social influence, and are crucial to urban management and service optimization. Identifying community hotspots can help city managers allocate resources more effectively, improve the quality of public services, and promote economic development. However, traditional clustering algorithms such as K-means or hierarchical clustering often perform poorly when processing social network data, mainly because they have difficulty adapting to the high dimensionality and complex structure of the data.

[0003] The existing density peak clustering method based on neighborhood fuzzy kernel is mainly used for community discovery in social networks. This method calculates the similarity between nodes through fuzzy kernel function and combines the idea of ​​density peak clustering to effectively identify high-density areas in social networks and divide communities. By adjusting the parameters of the fuzzy kernel (such as radius and kernel width), this method can handle complex network structures, solve the limitations of traditional methods on noisy data and edge nodes, and achieve better community division results.

[0004] For example, the invention patent announcement with announcement number: CN108090197B discloses a community discovery method for a multi-dimensional social network, including: obtaining the total correlation between users by integrating the friend relationship network, comment relationship network, recommendation and forwarding relationship network, and interest similarity network in the social network at multiple levels, and then treating each user as a node, using the total correlation between users as the transmission probability, and using the label propagation algorithm to divide the community, thereby completing social discovery.

[0005] For example, the invention patent announcement with announcement number: CN103729467B discloses a method for discovering community structure in a social network, including: step 1: converting the social network into an adjacency matrix form, if there is an edge between two nodes, then the corresponding element is 1, otherwise it is 0; step 2: using random walk theory to process the adjacency matrix to obtain a new node degree P-degree and edge weight P-weight; step 3: obtaining the leader node in the social network according to the new node degree P-degree; step 4: generating a subcommunity based on the leader node, and performing community discovery through a series of operations on the subcommunity.

[0006] However, in the process of implementing the inventive technical solution in the embodiments of the present application, it is found that the above technologies have at least the following technical problems:

[0007] In the prior art, when applied to a dynamically changing social network, some challenges still exist. For example, many methods require presetting key parameters in advance, which limits their degree of automation. At the same time, when the social network structure changes, how to quickly update the community division is also an urgent problem to be solved, and there is a problem of low timeliness in social network community discovery. Summary of the Invention

[0008] The embodiments of the present application provide a method for discovering social network communities based on density peak clustering of neighborhood fuzzy kernels, which solves the problem of low timeliness in social network community discovery in the prior art and realizes the improvement of the timeliness of discovering hot communities in social networks.

[0009] The embodiments of the present application provide a method for discovering social network communities based on density peak clustering of neighborhood fuzzy kernels, including the following steps: S1, obtaining the corresponding neighborhood fuzzy kernel through all user data in the social network collected within a preset community discovery period, and obtaining the neighborhood fuzzy kernel applicability influence index by real-time acquiring community change data, and judging whether to update the neighborhood fuzzy kernel according to the neighborhood fuzzy kernel applicability influence index; S2, performing clustering processing on the neighborhood fuzzy kernel obtained in S1 to obtain a community set, and performing an intersection process on the community set and the community set in the previous preset time period to obtain a community division set; S3, calculating the similarity deviation between the community set and the community set in the previous preset time period, and updating the community set according to the similarity deviation for the community division set.

[0010] Further, the process of obtaining the corresponding neighborhood fuzzy kernel through all user data in the social network collected within a preset community discovery period is as follows: collecting all user data in the social network within a preset community discovery period, and initializing the community set; performing dimensionality reduction processing on the user data and performing geodesic distance calculation to obtain the corresponding geodesic distance, where the dimensionality reduction processing means mapping the complex structure of the user data in the high-dimensional space to the low-dimensional space through a dimensionality reduction tool; constructing a geodesic distance matrix based on the geodesic distances of all users, and obtaining the neighborhood fuzzy membership degree between users according to the geodesic distance matrix and the neighborhood set; constructing a corresponding neighborhood fuzzy membership degree matrix based on the neighborhood fuzzy membership degree, and obtaining the corresponding neighborhood fuzzy kernel based on the neighborhood fuzzy membership degree matrix and the kernel parameters, where the kernel parameters include a kernel width and a kernel radius; calculating the local density of each user according to the neighborhood fuzzy membership degree matrix.

[0011] Further, the specific steps of the dimensionality reduction process are as follows: Step 1, set a user set in the high-dimensional space and calculate the Euclidean distance between each user and the remaining users; Step 2, sort each user in ascending order according to the Euclidean distance, and select the users with the nearest preset neighbor value to construct a neighbor set; Step 3, when a user is within the neighbor set of the remaining users, it means that the two users are neighbors to each other and an edge is established between them; Step 4, construct a mapping set of the weight of the edge and the Euclidean distance between two nodes to obtain the weight of the edge corresponding to the two nodes, and construct a neighbor graph based on the nodes corresponding to the users, the established edges, and the weights of the edges; Step 5, obtain the shortest path distance between any two nodes on the neighbor graph through a distance tool; Step 6, construct a geodesic distance matrix between users based on all the obtained shortest path distances, and transform the geodesic distance matrix into a centered matrix; Step 7, perform eigenvalue decomposition on the centered matrix and extract its preset number of largest eigenvalues and their corresponding eigenvectors to obtain the coordinates of the users in the low-dimensional space.

[0012] Further, the specific method for obtaining the neighborhood fuzzy membership degree is as follows: Obtain a given user set based on user data and number the users in the given user set; Obtain the corresponding Euclidean distance according to the dimensionality-reduced social data set, and obtain the expected value of the Euclidean distance between all users based on the Euclidean distance to obtain a sparsity factor; Perform classification analysis operations based on the sparsity factor, the Euclidean distance, and the corresponding neighbor set to obtain the neighborhood fuzzy membership degree.

[0013] Further, the real-time acquisition of community change data to obtain the applicable influence index of the neighborhood fuzzy kernel is as follows: Real-time acquisition of community change data within a preset community discovery period, where the community change data includes the user data volume, the relationship data volume, and the user attribute change rate; Obtain the community change weights from a preset database, where the community change weights include user weights, relationship weights, and user attribute weights; Obtain the user attribute influence factor through data change operations on the user data volume, where the data change operations are used to quantify the change of the user data volume and perform mapping; Perform user change fluctuation operations on the user data volume up to the current preset community discovery period to obtain the user data fluctuation value, where the user change fluctuation operations represent a way to quantify the change fluctuation of the user data volume up to the current preset community discovery period; Perform relationship change fluctuation operations on the relationship data volume up to the current preset community discovery period to obtain the relationship data fluctuation value, where the relationship change fluctuation operations represent a way to quantify the change fluctuation of the relationship data volume up to the current preset community discovery period; Combine the user data fluctuation value, the relationship data fluctuation value, the user attribute influence factor, and the corresponding community change weights to perform community change influence degree allocation to obtain the applicable influence index of the neighborhood fuzzy kernel, where the community change influence degree allocation is used to comprehensively quantify the timeliness of the community change data on the neighborhood fuzzy kernel in the current preset community discovery period, and the applicable influence index of the neighborhood fuzzy kernel represents the data that quantifies the influence degree of the community user data change on the timeliness of the neighborhood fuzzy kernel.

[0014] Further, the judgment of whether to update the neighborhood fuzzy kernel according to the applicable influence index of the neighborhood fuzzy kernel is as follows: Compare the applicable influence index of the neighborhood fuzzy kernel with the influence critical value in the preset database: If the applicable influence index of the neighborhood fuzzy kernel is less than the influence critical value, then do not update the neighborhood fuzzy kernel; If the applicable influence index of the neighborhood fuzzy kernel is not less than the influence critical value, then update the neighborhood fuzzy kernel; The specific process of updating the neighborhood fuzzy kernel is as follows: Construct a mapping set between the applicable influence index of the neighborhood fuzzy kernel and the kernel parameter adjustment factor, input the real-time applicable influence index of the neighborhood fuzzy kernel into the mapping set to obtain the corresponding kernel parameter adjustment factor, obtain the updated kernel parameters according to the kernel parameter adjustment factor, and update the neighborhood fuzzy kernel with the updated kernel parameters.

[0015] Further, the process of clustering the neighborhood fuzzy kernel to obtain the community set is as follows: Obtain the local density based on the neighborhood fuzzy membership degree, a preset neighbor value, and the total number of users; Obtain the relative distance by getting the distance between each user and the users whose local density is higher than that of the current user and whose Euclidean distance is lower than that of the current user; Select a preset number of hot users with local density higher than the preset density and relative distance higher than the preset distance as social centers for marking, and initialize the community set as empty; For unmarked users, obtain the corresponding shared neighbor similarity by performing a similarity operation between the unmarked users and the remaining marked users; Obtain the user proximity by performing a proximity operation through the neighbor sets of two users and the corresponding Euclidean distance; Perform a similarity influence operation based on the shared neighbor similarity and user proximity between users to obtain the weighted shared neighbor similarity; Sort the weighted shared neighbor similarity in descending order, and sequentially assign the unmarked users to the community corresponding to the marked user with the highest weighted shared neighbor similarity and mark them as marked users; If there are still unmarked users, assign them to the community where the marked user has a larger local density than the unmarked user and the smallest Euclidean distance from the unmarked user and mark them as marked users; Output the corresponding community set based on the above assignment process.

[0016] Further, the specific process of obtaining the community division set is as follows: Perform an intersection operation on the community set and the community set of the previous preset community discovery period to obtain the community division set. The community division set includes a stable community set, a set to be optimized, and a dynamically updated set. The stable community set represents the intersection part of the community set in the current community discovery period and the community set of the previous preset community discovery period. The set to be optimized represents the remaining set part after removing the intersection part from the community set of the previous preset community discovery period. The dynamically updated set represents the remaining set after removing the intersection part from the community set; Calculate the similarity deviation of the similarity operation between the community set and the community set of the previous preset community discovery period. The similarity operation is used to measure the similarity degree of the communities in the community set and the community set of the previous preset community discovery period.

[0017] Further, the process of updating the community set according to the similarity deviation is as follows: Construct a community set update mapping set of the similarity deviation and a preset division update ratio combination, and input the real-time similarity deviation into the community set update mapping set to obtain the corresponding division update ratio combination. The division update ratio combination includes an optimization set ratio and a dynamic set ratio; Sort the local densities corresponding to the users in the set to be optimized in descending order, and update the community set of the set to be optimized according to the optimization set ratio; Sort the local densities corresponding to the users in the dynamically updated set in descending order, and update the community set of the dynamically updated set according to the dynamic set ratio.

[0018] Further, the specific content of the community set update is as follows: The community where the user is located corresponding to the to-be-optimized set obtained by sorting the local density in descending order is retained according to the optimization set ratio in sequence, and the remaining communities in the to-be-optimized set are denoted as the to-be-updated set; The community where the user is located corresponding to the dynamically updated set obtained by sorting the local density in descending order is retained according to the dynamic set ratio in sequence and marked as the updated set; The to-be-updated set is updated by the updated set and merged with the stable community set to obtain the updated community set.

[0019] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0020] 1. By obtaining the corresponding neighborhood fuzzy core from all user data in the social network collected within the preset community discovery period, and obtaining the neighborhood fuzzy core applicability influence index by real-time acquiring community change data, thereby determining whether to update the neighborhood fuzzy core, clustering the neighborhood fuzzy core to obtain a community set, performing an intersection process on the community set and the community set of the previous preset time period to obtain a community division set, and calculating the similarity deviation at the same time to update the community set of the community division set, so as to update the community hotspots more real-time, and further improving the timeliness of discovering hot communities in the social network, effectively solving the problem of low timeliness of social network community discovery in the prior art;

[0021] 2. By obtaining a given user set from user data and numbering the users in the given user set, then obtaining the corresponding Euclidean distance according to the dimensionality-reduced social data set, and obtaining the expected value of the Euclidean distance between all users based on the Euclidean distance to obtain a sparsity factor, and then performing a classification analysis operation based on the sparsity factor, Euclidean distance and the corresponding neighbor set to obtain the neighborhood fuzzy membership degree, so as to obtain a more accurate neighborhood fuzzy core, and further realizing the discovery of more accurate community hotspots;

[0022] 3. By real-time acquiring community change data and obtaining the community change weight within the preset community discovery period, then obtaining the user attribute influence factor through the user data volume, obtaining the user data fluctuation value based on the user data volume up to the current preset community discovery period, and obtaining the relationship data fluctuation value according to the relationship data volume up to the current preset community discovery period, and finally obtaining the neighborhood fuzzy core applicability influence index by combining the above data, so as to more accurately evaluate the applicability of the current neighborhood fuzzy core, and further realizing the obtaining of more accurate community hotspots. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a flowchart of a social network community discovery method based on density peak clustering of neighborhood fuzzy core provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The embodiment of the present application provides a social network community discovery method based on density peak clustering of neighborhood fuzzy kernels, which solves the problem of low timeliness in social network community discovery in the prior art. By collecting all user data in the social network within a preset community discovery period, a given user set is obtained and the users in the given user set are numbered. Then, the corresponding Euclidean distance is obtained according to the dimensionality-reduced social data set, and the expected value of the Euclidean distance between all users is obtained based on the Euclidean distance to obtain a sparsity factor. Then, classification analysis operations are performed based on the sparsity factor, Euclidean distance, and corresponding neighbor set to obtain a neighborhood fuzzy membership degree to obtain a corresponding neighborhood fuzzy kernel. By obtaining community change data in real time within the preset community discovery period and obtaining a community change weight, then obtaining a user attribute influence factor based on the user data volume, obtaining a user data fluctuation value based on the user data volume up to the current preset community discovery period, obtaining a relationship data fluctuation value based on the relationship data volume up to the current preset community discovery period, and then combining the above data to obtain a neighborhood fuzzy kernel applicability influence index, thereby judging whether to update the neighborhood fuzzy kernel, performing clustering processing on the neighborhood fuzzy kernel to obtain a community set, and performing an intersection process on the community set and the community set of the previous preset time period to obtain a community division set, and at the same time calculating a similarity deviation to update the community set for the community division set, which realizes the improvement of the timeliness of hot community discovery in the social network.

[0025] The technical solution in the embodiment of the present application for solving the problem of low timeliness in social network community discovery is generally as follows:

[0026] By obtaining a corresponding neighborhood fuzzy kernel from all user data in the social network collected within a preset community discovery period, obtaining a neighborhood fuzzy kernel applicability influence index by obtaining community change data in real time, thereby judging whether to update the neighborhood fuzzy kernel, performing clustering processing on the neighborhood fuzzy kernel to obtain a community set, and performing an intersection process on the community set and the community set of the previous preset time period to obtain a community division set, and at the same time calculating a similarity deviation to update the community set for the community division set, the effect of improving the timeliness of hot community discovery in the social network is achieved.

[0027] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0028] As Figure 1As shown in the figure, it is a flowchart of a social network community discovery method based on density peak clustering with a neighborhood fuzzy kernel provided by an embodiment of the present application. The method includes the following steps: S1, obtaining a corresponding neighborhood fuzzy kernel from all user data in the social network collected within a preset community discovery period, and obtaining a neighborhood fuzzy kernel applicability influence index by real-time acquiring community change data, and determining whether to update the neighborhood fuzzy kernel according to the neighborhood fuzzy kernel applicability influence index; S2, performing clustering processing on the neighborhood fuzzy kernel obtained in S1 to obtain a community set, and performing an intersection process on the community set and the community set of the previous preset time period to obtain a community division set; S3, calculating the similarity deviation between the community set and the community set of the previous preset time period, and updating the community set according to the similarity deviation for the community division set.

[0029] In this embodiment, the behavior data of users is obtained from social network platforms (such as Facebook, Instagram, Twitter, WeChat, and TikTok, etc.), including check-in records, interaction behaviors, and geographical location information. The above data usually contains a large amount of noise and redundant information, so it is necessary to perform preprocessing first. In the data cleaning stage, invalid or duplicate data entries are removed to ensure the accuracy and consistency of the data. And the activity trajectories of users are segmented, and the continuous movement paths of users are divided into meaningful segments, for example, the frequently visited locations are determined according to the residence time and movement patterns of users; for missing values, linear interpolation or other smoothing algorithms are used for filling and smoothing processing to improve the data quality; through the process of the above method, the complex structure of social data in high-dimensional space is mapped into low-dimensional space, which can improve the data visualization effect and the timeliness of discovering hot communities in the social network while retaining the global structure and non-linear relationship of the data.

[0030] It should be added that the Density Peaks (DP) clustering algorithm provides a novel and effective clustering method by selecting points with higher local density and larger distance as clustering centers. However, when facing high-dimensional data sets, the traditional DP algorithm may have problems such as low computational efficiency and difficult parameter setting. By introducing the concept of a neighborhood fuzzy kernel, the neighborhood fuzzy membership degree is obtained, which enables better capturing of data distribution characteristics and improving the robustness to noise and outliers.

[0031] Furthermore, the corresponding neighborhood fuzzy kernel is obtained from all user data in the social network collected within the preset community discovery period. The specific process is as follows: Collect all user data in the social network within the preset community discovery period and initialize the community set; Perform dimensionality reduction processing on the user data and perform geodesic distance calculation to obtain the corresponding geodesic distance. Dimensionality reduction processing means mapping the complex structure of the user data in the high-dimensional space to the low-dimensional space through a dimensionality reduction tool; Construct a geodesic distance matrix based on the geodesic distances of all users, and obtain the neighborhood fuzzy membership degree between users according to the geodesic distance matrix and the neighborhood set; Construct the corresponding neighborhood fuzzy membership degree matrix based on the neighborhood fuzzy membership degree, and obtain the corresponding neighborhood fuzzy kernel based on the neighborhood fuzzy membership degree matrix and the kernel parameters. The kernel parameters include the kernel width and the kernel radius; Calculate the local density of each user according to the neighborhood fuzzy membership degree matrix.

[0032] In this embodiment, the neighborhood fuzzy membership degree measures the membership degree or similarity of each data point to other points within its neighborhood. The neighborhood fuzzy membership degree matrix contains the neighborhood fuzzy membership degree values between all nodes, representing the neighborhood similarity between each node and other nodes. And the neighborhood fuzzy membership degree matrix is an n*n matrix, where n represents the total number of nodes in the community set; while the neighborhood fuzzy kernel is a method of mapping the neighborhood fuzzy membership degree to the high-dimensional space, usually mapping the neighborhood fuzzy membership degree matrix to a new space through a certain kernel function (such as Gaussian kernel). For example: Assume we define a kernel function , through the neighborhood fuzzy membership degree matrix to calculate:

[0033] ;

[0034] where, is the kernel parameter used to control the range of similarity, represents the i-th user, represents the j-th user; By using the neighborhood fuzzy membership degree matrix , calculate a fuzzy kernel value for each pair of nodes, representing the similarity between data nodes, to obtain the neighborhood fuzzy kernel matrix , and its elements are defined as:

[0035] ;

[0036] where, represents the neighborhood fuzzy kernel matrix between the i-th user and the j-th user, It represents the neighborhood fuzzy membership matrix between the i-th user and the j-th user; through the above process, the finally obtained neighborhood fuzzy kernel matrix is used for subsequent machine learning or social network analysis tasks, such as clustering, classification, dimensionality reduction, etc.; the neighborhood fuzzy kernel matrix is calculated based on the neighborhood fuzzy membership matrix, and at the same time combines the kernel function to obtain the similarity measure between each pair of nodes; through the acquisition of the neighborhood fuzzy kernel, the processing efficiency of subsequent clustering analysis is improved.

[0037] Specifically, the ISOMAP (Isometric Feature Mapping) algorithm is used to perform dimensionality reduction on the preprocessed high-dimensional social data. First, an adjacency graph is constructed based on the geodesic distance between users to ensure that each node is connected to its nearest node. Then, the shortest path distance between each pair of nodes is calculated by the Dijkstra algorithm. Finally, the classical multidimensional scaling technique is used to convert the distance matrix in the high-dimensional space into a low-dimensional embedding representation, thereby obtaining the manifold structure. Through the above process, it helps to retain the global structure and non-linear relationship of the user social relationship, making the dimensionality-reduced social data easier to analyze and visualize.

[0038] Furthermore, the specific steps of the dimensionality reduction process are as follows: Step 1, set the user set in the high-dimensional space and calculate the Euclidean distance between each user and the remaining users; Step 2, sort each user in ascending order according to the Euclidean distance from the remaining users, and select the users with the preset nearest neighbor value to construct a nearest neighbor set; Step 3, when a user is within the nearest neighbor set of the remaining users, it means that the two users are neighbors to each other and an edge is established between them; Step 4, construct a mapping set of the edge weight and the Euclidean distance between two nodes to obtain the weight of the edge corresponding to the two nodes, and construct a nearest neighbor graph according to the nodes corresponding to the users, the established edges and the edge weights; Step 5, obtain the shortest path distance between any two nodes on the nearest neighbor graph through the distance tool, and the calculation formula of the shortest path distance is as follows:

[0039] ;

[0040] ;

[0041] In the formula, n represents the user number, , represents the total number of users, represents the given user set, represents the i-th user, represents the j-th user, represents the nearest neighbor graph, represents the preset nearest neighbor value, represents on the nearest neighbor graph the Euclidean distance between the i-th user and the j-th user, Denote the Euclidean distance between the $i$-th user and the $k$-th user on the adjacent graph , and denote the Euclidean distance between the $k$-th user and the $j$-th user on the adjacent graph . Denote the shortest path distance between the $i$-th user and the $j$-th user on the adjacent graph ;

[0042] Step 6: Construct the geodesic distance matrix between users based on all the obtained shortest path distances, and transform the geodesic distance matrix into a centered matrix. The specific method for obtaining the centered matrix is as follows:

[0043] ;

[0044] ;

[0045] where denotes the total number of users, denotes an identity matrix, denotes an $N$-dimensional vector all of whose elements are 1, denotes the transformation matrix, denotes the geodesic distance matrix, denotes the centered matrix;

[0046] Step 7: Perform eigenvalue decomposition on the centered matrix and extract its largest preset number of eigenvalues and their corresponding eigenvectors to obtain the coordinates of the users in the low-dimensional space;

[0047] The coordinates of the users in the low-dimensional space are calculated according to the following formula:

[0048] ;

[0049] where denotes the coordinates of the users in the low-dimensional space, denotes the preset number, denotes the first eigenvector, denotes the second eigenvector, denotes the $d$-th eigenvector, denotes the largest eigenvalue.

[0050] In this embodiment, the preset neighbor value is specifically set by a preset professional according to the social network situation; the preset number is specifically set by a preset professional according to the social network situation; the distance tool includes the Dijkstra algorithm or the Floyd algorithm; by introducing the neighborhood fuzzy membership degree, the similarity of users in dense and sparse areas is balanced, the dependence on the cut-off distance is reduced, and the similarity of user data in dense and sparse areas can be better balanced, thereby improving the accuracy and adaptability of local density measurement, improving the accuracy of the neighborhood fuzzy kernel, and further improving the accuracy of community hot spot discovery.

[0051] Further, the specific method for obtaining the neighborhood fuzzy membership degree is as follows: a given user set is obtained based on user data, and the users in the given user set are numbered; the corresponding Euclidean distance is obtained according to the dimensionality-reduced social data set, and the expected Euclidean distance between all users is obtained based on the Euclidean distance to obtain the sparsity factor; classification analysis operations are performed based on the sparsity factor, the Euclidean distance, and the corresponding neighbor set to obtain the neighborhood fuzzy membership degree. The classification analysis operations are used to obtain the neighborhood fuzzy membership degree in different situations according to the relationship between the user and the neighbor set. The specific constraint expression of the neighborhood fuzzy membership degree is as follows:

[0052] ;

[0053] ;

[0054] ;

[0055] ;

[0056] In the formula, n represents the number of the user, , represents the total number of users, represents the Euclidean distance between the i-th user and the j-th user, represents the sparsity factor, represents the expected Euclidean distance, represents the j-th user, represents the neighbor set of the i-th user, represents the given user set, represents the neighborhood fuzzy membership degree between the i-th user and the j-th user.

[0057] In this embodiment, for the neighbor points and non-neighbor points of the user, the calculation methods of their neighborhood fuzzy membership degrees are different; in the case of neighbor points (i.e., when), the neighborhood fuzzy membership degree is positively correlated with the distance between the user and its neighbor points; while in the case of non-neighbor points (i.e., ), the calculation of neighborhood fuzzy membership not only depends on the distance between users, but also is affected by the sparsity factor, which is used to measure the distribution dispersion degree between users; the nearest neighbor fuzzy kernel can more accurately describe the neighborhood fuzzy membership relationship between user data. The closer its value is to 1, the stronger the neighborhood fuzzy membership of the sample. Through the above analysis, a more accurate neighborhood fuzzy membership is obtained, thereby improving the accuracy of the neighborhood fuzzy kernel, and further improving the accuracy of community hot spot discovery.

[0058] Furthermore, real-time acquisition of community change data to obtain the applicable influence index of the neighborhood fuzzy kernel, the specific process is as follows: real-time acquisition of community change data within the preset community discovery period, and the community change data includes the user data volume, relationship data volume, and user attribute change rate; obtain the community change weight from the preset database, and the community change weight includes user weight, relationship weight, and user attribute weight; obtain the user attribute influence factor through data change operation on the user data volume, and the data change operation is used to quantify the change of the user data volume and perform mapping; perform user change fluctuation operation on the user data volume up to the current preset community discovery period to obtain the user data fluctuation value, and the user change fluctuation operation represents a way to quantify the change fluctuation of the user data volume up to the current preset community discovery period; perform relationship change fluctuation operation on the relationship data volume up to the current preset community discovery period to obtain the relationship data fluctuation value, and the relationship change fluctuation operation represents a way to quantify the change fluctuation of the relationship data volume up to the current preset community discovery period; combine the user data fluctuation value, relationship data fluctuation value, user attribute influence factor, and the corresponding community change weight to perform community change influence degree allocation to obtain the applicable influence index of the neighborhood fuzzy kernel. The community change influence degree allocation is used to comprehensively quantify the timeliness of the community change data on the neighborhood fuzzy kernel in the current preset community discovery period. The applicable influence index of the neighborhood fuzzy kernel represents the data that quantifies the influence degree of the community user data change on the timeliness of the neighborhood fuzzy kernel;

[0059] The specific limit expression of the applicable influence index of the neighborhood fuzzy kernel is as follows:

[0060] ;

[0061] ;

[0062] In the formula, t represents the number of the preset community discovery period, , represents the total number of preset community discovery periods, represents the user data volume of the t-th preset community discovery period, represents the relationship data volume of the t-th preset community discovery period, represents the user attribute change rate of the t-th preset community discovery period, represents the amount of user data for the th preset community discovery period, represents the amount of relationship data for the th preset community discovery period, represents the influence factor of user attributes for the t-th preset community discovery period, represents the user weight, represents the relationship weight, represents the user attribute weight, represents the applicable influence index of the neighborhood fuzzy kernel for the t-th preset community discovery period.

[0063] In this embodiment, the algorithm combines community change data and community change weights for comprehensive analysis to obtain the applicable influence index of the neighborhood fuzzy kernel. In the formula, the applicable influence index of the neighborhood fuzzy kernel increases with the change deviation of the user data volume (i.e., ), the change deviation of the relationship data volume (i.e., ), and the user attribute change rate (i.e., ). The larger the applicable influence index of the neighborhood fuzzy kernel, the less applicable the neighborhood fuzzy kernel of the current preset community discovery period is to the current community network. Moreover, the user attribute change rate is affected by the change of the user data volume. As the user data volume increases or decreases, the corresponding user attribute change rate also increases or decreases. Therefore, a user attribute influence factor for the user attribute change rate is obtained through the change deviation of the user data volume. Through the above analysis, it can be seen that according to the applicable influence index of the neighborhood fuzzy kernel, it is helpful to more accurately evaluate the applicability of the neighborhood fuzzy kernel, thereby improving the timeliness of the neighborhood fuzzy kernel, and further analyzing and discovering more accurate community hotspots, improving the efficiency of community hotspot discovery.

[0064] It should be added that the corresponding amount of user data is obtained by counting the number of newly added and deleted users through event logs. In a social network, the relationships between users (such as following, friends, likes, comments, etc.) are usually represented by edges. Monitoring the addition or deletion of edges obtains the corresponding amount of relationship data. The user behavior logs are used to record the changes in user attributes, and the ratio of the amount of user data with changed user attributes to the total amount of user data is statistically compared to obtain the corresponding user attribute change rate.

[0065] Specifically, the community change weight is obtained from a preset database, and the community change weight represents the influence degree of community change data on the influence index of the neighborhood fuzzy kernel applicability. There is a unique mapping relationship between each community change data and the community change weight, and the value range is between 0 and 1. For example, a mapping set of community change data and preset community change weights is constructed, and the real-time user data volume, relationship data volume, and user attribute change rate are input into the mapping set to obtain the user weight, relationship weight, and user attribute weight respectively, which represent the influence degrees of the user data volume, relationship data volume, and user attribute change rate on the influence index of the neighborhood fuzzy kernel applicability, and the sum of the three is 1.

[0066] Further, it is judged whether to update the neighborhood fuzzy kernel according to the neighborhood fuzzy kernel applicability influence index. The specific steps are as follows: compare the neighborhood fuzzy kernel applicability influence index with the influence critical value in the preset database. If the neighborhood fuzzy kernel applicability influence index is less than the influence critical value, the neighborhood fuzzy kernel is not updated. If the neighborhood fuzzy kernel applicability influence index is not less than the influence critical value, the neighborhood fuzzy kernel is updated. The specific process of updating the neighborhood fuzzy kernel is as follows: construct a mapping set of the neighborhood fuzzy kernel applicability influence index and the kernel parameter adjustment factor, input the real-time neighborhood fuzzy kernel applicability influence index into the mapping set to obtain the corresponding kernel parameter adjustment factor, obtain the updated kernel parameter according to the kernel parameter adjustment factor, and update the neighborhood fuzzy kernel through the updated kernel parameter.

[0067] In this embodiment, the updated kernel parameter is obtained by performing a product operation on the kernel parameter adjustment factor and the kernel parameter respectively, and the kernel parameter of the neighborhood fuzzy kernel is adjusted to the updated kernel parameter to obtain the updated neighborhood fuzzy kernel. By adjusting the neighborhood fuzzy kernel according to the actual community change data, the neighborhood fuzzy kernel can better adapt to the changes in the complex dynamic environment of the social network. For example, with the change of user behavior, the update of the neighborhood fuzzy kernel can reflect more accurate user similarity.

[0068] Specifically, the influence critical value is obtained from a preset database. In a specific embodiment, the community change data that obtains the accurate community hot spot corresponding situation in the historical data is substituted into the specific limit expression of the neighborhood fuzzy kernel applicability influence index to obtain the corresponding data set, and the average value operation is performed on the data set to obtain the corresponding influence critical value.

[0069] Further, clustering processing is performed on the neighborhood fuzzy kernel to obtain a community set. The specific process is as follows: based on the neighborhood fuzzy membership degree, the preset neighbor value, and the total number of users, the local density is obtained. The specific way to obtain the local density is as follows:

[0070] ;

[0071] Among them, represents the neighborhood fuzzy membership degree between the i-th user and the j-th user. represents the neighborhood fuzzy membership degree between the \(i\)-th user and the \(m\)-th user, represents a preset near-neighbor value, represents the total number of users, represents the local density of the \(i\)-th user;

[0072] Obtain the relative distance by getting the distance between each user and the users whose local density is higher than that of the current user and the Euclidean distance is lower than that of the current user. The way to obtain the relative distance is as follows:

[0073] ;

[0074] Among them, represents the Euclidean distance between the \(i\)-th user and the \(j\)-th user, represents the set of local densities of all users, represents the local density of the \(i\)-th user, represents the local density of the \(j\)-th user, represents the geodesic distance matrix, represents the relative distance of the \(i\)-th user;

[0075] Select the users with the preset number of hotspots whose local density is higher than the preset density and relative distance is higher than the preset distance as the social centers for marking, and initialize the community set to be empty; for the unmarked users, obtain the corresponding shared neighbor similarity by performing a similarity operation on the unmarked users and the remaining marked users. The specific calculation method is as follows:

[0076] ;

[0077] Among them, represents the neighbor set of the \(i\)-th user, represents the neighbor set of the \(j\)-th user, represents a given set of users, represents the shared neighbor similarity between the \(i\)-th user and the \(j\)-th user;

[0078] Obtain the user proximity by performing a proximity operation through the neighbor sets of two users and the corresponding Euclidean distances. The specific calculation method is as follows:

[0079] ;

[0080] Among them, represents the neighbor set of the \(i\)-th user, represents the neighbor set of the \(j\)-th user, represents the Euclidean distance between the \(j\)-th user and the \(u\)-th user, represents the Euclidean distance between the \(i\)-th user and the \(v\)-th user, represents the user proximity between the i-th user and the j-th user;

[0081] Based on the shared neighbor similarity and user proximity between users, similarity influence calculation is performed to obtain weighted shared neighbor similarity. The specific calculation method of weighted shared neighbor similarity is as follows:

[0082] ;

[0083] in, represents the user proximity between the i-th user and the j-th user, represents the shared neighbor similarity between the i-th user and the j-th user, represents the weighted shared neighbor similarity between the i-th user and the j-th user;

[0084] Arrange the weighted shared neighbor similarities in descending order, and assign the unlabeled users to the communities corresponding to the labeled users with the highest weighted shared neighbor similarity and mark them as labeled users; if there are still unlabeled users, assign them to the communities where the labeled users with larger local density than the unlabeled users and the smallest Euclidean distance to the unlabeled users are located and mark them as labeled users; output the corresponding community set based on the above assignment process.

[0085] In this embodiment, the community set is finally output, in which each community is a community hotspot in the social network. These hotspot areas reflect specific areas where user activities are frequent and where a large number of people gather. For example, in the analysis of a certain city, the results obtained by the above method show that the city center business district, the surrounding areas of large shopping malls and the vicinity of major transportation hubs are the community hotspot areas with the most users. These areas are not only the core areas of urban economic vitality, but also an indispensable part of residents' daily life.

[0086] The above method helps to accurately identify community hotspots in social networks, helping city managers to better understand citizens' behavior patterns and develop corresponding optimization strategies. At the same time, it also provides merchants with accurate market analysis tools to help business development. For example, city managers can reasonably plan public transportation routes, increase bus frequencies or set up more public bicycle stations based on the distribution of community hotspots; merchants can choose appropriate store locations or develop targeted promotional activities based on the flow of people in hotspot areas, thereby increasing sales and market share.

[0087] In addition, identifying community hot spots can also help social media platforms better understand and optimize user experience. For example, the platform can push relevant localized content and services based on the user's active area to enhance user stickiness and satisfaction. At the same time, it can also provide accurate target audience positioning for advertising, improving advertising effectiveness and return on investment.

[0088] Further, the specific process of obtaining the community division set is as follows: The community set is intersected with the community set of the previous preset community discovery period to obtain the community division set. The community division set includes a stable community set, an optimization-to-be set, and a dynamic update set. The stable community set represents the intersection part of the community set and the community set of the previous preset community discovery period. The optimization-to-be set represents the remaining set part after removing the intersection part from the community set of the previous preset community discovery period. The dynamic update set represents the remaining set after removing the intersection part from the community set; the similarity deviation of the similarity operation between the community set and the community set of the previous preset community discovery period is calculated. The similarity operation is used to measure the similarity degree of the communities in the community set and the community set of the previous preset community discovery period.

[0089] In this embodiment, the similarity operation includes, but is not limited to, the cosine similarity algorithm, the Jaccard similarity algorithm, and the Pearson correlation coefficient algorithm; the result of the similarity operation between the community set and the community set of the previous preset community discovery period is subjected to a deviation operation to obtain the similarity deviation. The deviation operation represents the ratio operation of the difference between the result of the similarity operation between the community set of the current community discovery period and the community set of the previous preset community discovery period and the result of the similarity operation of the community set of the previous preset community discovery period; based on the division methods of stable communities, communities to be optimized, and dynamically updated communities, it is beneficial to improve the accuracy, efficiency, and flexibility of community division, and further improve the efficiency of community set update.

[0090] Further, the community set is updated according to the similarity deviation for the community division set, and the specific process is as follows: A community set update mapping set of the similarity deviation and the preset division update ratio combination is constructed. The real-time similarity deviation is input into the community set update mapping set to obtain the corresponding division update ratio combination. The division update ratio combination includes an optimization set ratio and a dynamic set ratio; the local densities corresponding to the users in the optimization-to-be set are sorted in descending order, and the community set is updated for the optimization-to-be set according to the optimization set ratio; the local densities corresponding to the users in the dynamic update set are sorted in descending order, and the community set is updated for the dynamic update set according to the dynamic set ratio.

[0091] In this embodiment, through the calculation of the similarity deviation, the use of the division update ratio combination, and the priority sorting based on local density, the accuracy, efficiency, and adaptability of the community update process can be ensured; and by combining the adjustment of the optimization set ratio and the dynamic set ratio, the community set update is made more flexible and accurate, and at the same time, the frequency of community set update is effectively controlled.

[0092] Further, the specific content of the community set update is as follows: the community where the corresponding user is located in the set to be optimized obtained by sorting the local density in descending order is retained according to the optimization set ratio in sequence, and the remaining communities in the set to be optimized are recorded as the set to be updated; the community where the corresponding user is located in the dynamic update set obtained by sorting the local density in descending order is retained according to the dynamic set ratio in sequence and marked as the update set; the set to be updated is updated by the update set and merged with the stable community set to obtain the updated community set.

[0093] In this embodiment, through local density sorting, the priority processing of the part that needs to be updated in the community set is ensured, and the combination of the optimization set ratio and the dynamic set ratio makes the update of the community set more targeted, improving the accuracy of the community set update; and according to the adjustment of the optimization set ratio and the dynamic set ratio, the intensity of optimization and update is more flexibly controlled, enabling the division of the community set to adaptively adjust to network changes; at the same time, the merged community set can not only maintain the original stability but also absorb real-time updated changes, ensuring the accuracy of community division and the timeliness of the community set.

[0094] In summary, in the embodiment of the present application, the corresponding neighborhood fuzzy kernel is obtained from all user data in the social network collected within the preset community discovery period, and the neighborhood fuzzy kernel applicable influence index is obtained by real-time acquiring community change data, and based on this, it is judged whether to update the neighborhood fuzzy kernel. The neighborhood fuzzy kernel is clustered to obtain a community set, and the community set is intersected with the community set in the previous preset time period to obtain a community division set. At the same time, similarity calculation is performed to obtain a similarity deviation to update the community set of the community division set, thereby updating community hotspots more real-time, and further improving the timeliness of hotspot community discovery in the social network, effectively solving the problem of low timeliness of social network community discovery in the prior art.

[0095] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0096] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0097] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0099] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0100] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A method for discovering social network communities based on density peak clustering with neighborhood fuzzy kernel, characterized in that, Including the following steps: S1. Obtain a corresponding neighborhood fuzzy kernel from all user data in the social network collected within a preset community discovery period, and obtain a neighborhood fuzzy kernel applicability impact index in real time from community change data. Determine whether to update the neighborhood fuzzy kernel according to the neighborhood fuzzy kernel applicability impact index. The user data includes check-in records, interaction behaviors, and geographical location information; S2. Perform clustering processing on the neighborhood fuzzy kernel obtained in S1 to obtain a community set, and perform an intersection operation on the community set and the community set in the previous preset time period to obtain a community division set; S3. Calculate the similarity deviation between the community set and the community set in the previous preset time period, update the community set according to the similarity deviation for the community division set, and finally output the community set, where each community is a community hotspot in the social network, reflecting a specific area with frequent user activities and a large gathering of people; The process of obtaining the neighborhood fuzzy kernel applicability impact index by obtaining community change data in real time is as follows: Obtain community change data in real time within a preset community discovery period. The community change data includes the amount of user data, the amount of relationship data, and the user attribute change rate; Obtain community change weights from a preset database. The community change weights include user weights, relationship weights, and user attribute weights; Obtain a user attribute influence factor through data change operations on the amount of user data. The data change operations are used to quantify the change of the amount of user data and perform mapping; Perform user change fluctuation operations on the amount of user data up to the current preset community discovery period to obtain a user data fluctuation value. The user change fluctuation operations represent a way to quantify the change fluctuation of the amount of user data up to the current preset community discovery period; Perform relationship change fluctuation operations on the amount of relationship data up to the current preset community discovery period to obtain a relationship data fluctuation value. The relationship change fluctuation operations represent a way to quantify the change fluctuation of the amount of relationship data up to the current preset community discovery period; Allocate the community change impact degree by combining the user data fluctuation value, the relationship data fluctuation value, the user attribute influence factor, and the corresponding community change weights to obtain the neighborhood fuzzy kernel applicability impact index. The community change impact degree allocation is used to comprehensively quantify the timeliness of community change data on the neighborhood fuzzy kernel in the current preset community discovery period. The neighborhood fuzzy kernel applicability impact index represents data that quantifies the impact degree of community user data change on the timeliness of the neighborhood fuzzy kernel; 2. The social network community discovery method based on density peak clustering of neighborhood fuzzy kernel according to claim 1, wherein: The process of obtaining the corresponding neighborhood fuzzy kernel from all user data in the social network collected within a preset community discovery period is as follows: Collect all user data in the social network within a preset community discovery period, and initialize the community set; Perform dimensionality reduction processing on the user data and perform geodesic distance operations to obtain the corresponding geodesic distance. The dimensionality reduction processing means mapping the complex structure of the user data in the high-dimensional space to the low-dimensional space through a dimensionality reduction tool; Construct a geodesic distance matrix based on the geodesic distances of all users, and obtain the neighborhood fuzzy membership degree between users according to the geodesic distance matrix and the neighborhood set; Construct based on neighborhood fuzzy membership degree Build The corresponding neighborhood fuzzy membership degree matrix, and obtain the corresponding neighborhood fuzzy kernel based on the neighborhood fuzzy membership degree matrix and kernel parameters, where the kernel parameters include kernel width and kernel radius; Calculate the local density of each user according to the neighborhood fuzzy membership matrix.

3. The method for discovering social network communities based on density peak clustering with neighborhood fuzzy kernel as claimed in claim 2, wherein: The specific steps of the dimensionality reduction process are as follows: Step 1, set a user set in the high-dimensional space and calculate the Euclidean distance between each user and the rest of the users; Step 2, sort each user in ascending order according to the Euclidean distance, and select the users with the preset nearest neighbor value to construct a nearest neighbor set; Step 3, when a user is within the nearest neighbor set of the rest of the users, it means that the two users are mutual nearest neighbors and an edge is established between them; Step 4, construct a mapping set of the weight of the edge and the Euclidean distance between two nodes to obtain the weight of the edge corresponding to the two nodes, and construct a nearest neighbor graph according to the nodes corresponding to the users, the established edges and the weights of the edges; Step 5, obtain the shortest path distance between any two nodes on the nearest neighbor graph through a distance tool; Step 6, construct a geodesic distance matrix between users according to all the obtained shortest path distances, and convert the geodesic distance matrix into a centralized matrix; Step 7, perform eigenvalue decomposition on the centralized matrix and extract its preset number of largest eigenvalues and their corresponding eigenvectors to obtain the coordinates of the users in the low-dimensional space.

4. The method for discovering social network communities based on density peak clustering with neighborhood fuzzy kernel according to claim 2, wherein: The specific way to obtain the neighborhood fuzzy membership is as follows: Based on user data, obtain a given user set and number the users in the given user set; Obtain the corresponding Euclidean distance according to the dimensionality-reduced social data set, and obtain the expected value of the Euclidean distance between all users based on the Euclidean distance to obtain a sparsity factor; Perform classification analysis operations based on the sparsity factor, Euclidean distance and the corresponding nearest neighbor set to obtain the neighborhood fuzzy membership.

5. The social network community discovery method based on density peak clustering with neighborhood fuzzy kernel as claimed in claim 1, wherein: The specific steps to determine whether to update the neighborhood fuzzy kernel according to the neighborhood fuzzy kernel applicability influence index are as follows: Compare the neighborhood fuzzy kernel applicability influence index with the influence critical value in the preset database: If the neighborhood fuzzy kernel applicability influence index is less than the influence critical value, do not update the neighborhood fuzzy kernel; If the neighborhood fuzzy kernel applicability influence index is not less than the influence critical value, update the neighborhood fuzzy kernel; The specific process of updating the neighborhood fuzzy kernel is as follows: Construct a mapping set of the neighborhood fuzzy kernel applicability influence index and the kernel parameter adjustment factor, input the real-time neighborhood fuzzy kernel applicability influence index into the mapping set to obtain the corresponding kernel parameter adjustment factor, obtain the updated kernel parameter according to the kernel parameter adjustment factor, and update the neighborhood fuzzy kernel through the updated kernel parameter.

6. The method for discovering social network communities based on density peak clustering with neighborhood fuzzy kernel according to claim 1, wherein: The specific process of clustering the neighborhood fuzzy kernel to obtain a community set is as follows: Obtain the local density based on the neighborhood fuzzy membership, preset nearest neighbor value and the total number of users; Obtain the relative distance by getting the distance between each user and the users whose local density is higher than the local density of the current user and the Euclidean distance is lower than the Euclidean distance of the current user; Select the preset number of hot spot users with local density higher than the preset density and relative distance higher than the preset distance as social centers for marking, and initialize the community set to be empty; For unmarked users, obtain the corresponding shared nearest neighbor similarity by performing a similarity operation on the unmarked users and the rest of the marked users; Obtain the user proximity by performing a proximity operation on the nearest neighbor sets of two users and the corresponding Euclidean distance; Perform a similarity influence operation based on the shared neighbor similarity and user proximity between users to obtain the weighted shared neighbor similarity; Sort the weighted shared neighbor similarity in descending order, and sequentially assign the unlabeled users to the communities corresponding to the labeled users with the highest weighted shared neighbor similarity and label them as labeled users; If there are still unlabeled users, assign them to the community where the labeled user has a larger local density than the unlabeled user and the smallest Euclidean distance from the unlabeled user and label them as labeled users; Output the corresponding community set based on the above assignment process.

7. The method for discovering social network communities based on density peak clustering with neighborhood fuzzy kernel according to claim 1, characterized in that: The specific process of obtaining the community partition set is as follows: Perform an intersection operation on the community set and the community set of the previous preset community discovery period to obtain the community partition set. The community partition set includes a stable community set, an optimization required set, and a dynamic update set. The stable community set represents the intersection part of the community set of the current community discovery period and the community set of the previous preset community discovery period. The optimization required set represents the remaining set part of the community set of the previous preset community discovery period after removing the intersection part. The dynamic update set represents the remaining set of the community set after removing the intersection part; Calculate the similarity deviation of the similarity operation between the community set and the community set of the previous preset community discovery period. The similarity operation is used to measure the similarity degree of the communities in the community set and the community set of the previous preset community discovery period.

8. The method for discovering social network communities based on density peak clustering with neighborhood fuzzy kernel according to claim 7, characterized in that: The specific process of updating the community set according to the similarity deviation is as follows: Construct a community set update mapping set of the similarity deviation and the preset partition update ratio combination, and input the real-time similarity deviation into the community set update mapping set to obtain the corresponding partition update ratio combination. The partition update ratio combination includes an optimization set ratio and a dynamic set ratio; Sort the local densities corresponding to the users in the optimization required set in descending order, and update the community set of the optimization required set according to the optimization set ratio; Sort the local densities corresponding to the users in the dynamic update set in descending order, and update the community set of the dynamic update set according to the dynamic set ratio.

9. The method for discovering social network communities based on density peak clustering with neighborhood fuzzy kernel according to claim 1, characterized in that: The specific content of the community set update is as follows: Retain the communities where the corresponding users of the optimization required set obtained by sorting the local density in descending order are located according to the optimization set ratio in sequence, and record the remaining communities in the optimization required set as the set to be updated; Retain the communities where the corresponding users of the dynamic update set obtained by sorting the local density in descending order are located according to the dynamic set ratio in sequence and label them as the update set; Update the set to be updated through the update set and merge it with the stable community set to obtain the updated community set.

Citation Information

Patent Citations

  • A Community Structure Discovery Method in Social Networks

    CN103729467B

  • A Community Discovery Method for Multidimensional Social Networks

    CN108090197B

  • Improved density peak clustering-based social network community discovery method

    CN108647739A

  • An overlapping community detection method based on density peak and community belonging degree

    CN108959652A