Social network community discovery method based on density peak clustering of neighborhood fuzzy kernel

By calculating and updating the density peak clustering method based on neighborhood fuzzy kernels in social networks, the problem of low timeliness of social network community discovery is solved, and more efficient community discovery and update are achieved.

CN119939045AActive Publication Date: 2025-05-06HUNAN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510436189.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The prior art social network community found that the timeliness is low in dynamically changing social networks and requires preset key parameters, limiting its automation level.

Method used

By collecting social network data during the preset community discovery cycle, calculating neighborhood fuzzy kernels and applicable impact indexes, determining whether to update the fuzzy kernels, and performing clustering and community division updates, to improve the timeliness of community discovery.

Benefits of technology

The timeliness of discovery of hot communities in social networks have been improved, the problem of low timeliness of discovery of community is solved, and the degree of automation has been improved by updating the fuzzy core in real time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939045A_ABST
    Figure CN119939045A_ABST
Patent Text Reader

Abstract

The invention discloses a social network community discovery method based on density peak clustering of a neighborhood fuzzy kernel, and relates to the technical field of data clustering analysis. The social network community discovery method based on the density peak clustering of the neighborhood fuzzy kernel comprises the following steps of obtaining the neighborhood fuzzy kernel, obtaining a community set and updating the community set. According to the method, the corresponding neighborhood fuzzy kernel is obtained through all collected user data in the social network, the community change data is obtained in real time to obtain the neighborhood fuzzy kernel applicable influence index, whether the neighborhood fuzzy kernel is updated or not is judged accordingly, the neighborhood fuzzy kernel is clustered to obtain the community set, and the community set is updated. A community division set is obtained based on the community set and the community set in the previous preset time period, and meanwhile, similarity deviation is obtained to perform community set updating on the community division set, so that the effect of improving the timeliness of hotspot community discovery in the social network is achieved; the problem of low timeliness of social network community discovery in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data clustering analysis, and in particular to a social network community discovery method based on density peak clustering of neighborhood fuzzy kernels. Background Art

[0002] As urbanization accelerates, understanding and managing urban communities becomes increasingly complex. Community hotspots refer to areas where user activities are frequent and people gather in large numbers. These areas usually have high commercial value and social influence, and are crucial to urban management and service optimization. Identifying community hotspots can help city managers allocate resources more effectively, improve the quality of public services, and promote economic development. However, traditional clustering algorithms such as K-means or hierarchical clustering often perform poorly when processing social network data, mainly because they have difficulty adapting to the high dimensionality and complex structure of the data.

[0003] The existing density peak clustering method based on neighborhood fuzzy kernel is mainly used for community discovery in social networks. This method calculates the similarity between nodes through fuzzy kernel function and combines the idea of ​​density peak clustering to effectively identify high-density areas in social networks and divide communities. By adjusting the parameters of the fuzzy kernel (such as radius and kernel width), this method can handle complex network structures, solve the limitations of traditional methods on noisy data and edge nodes, and achieve better community division results.

[0004] For example, the invention patent announcement with announcement number: CN108090197B discloses a community discovery method for a multi-dimensional social network, including: obtaining the total correlation between users by integrating the friend relationship network, comment relationship network, recommendation and forwarding relationship network, and interest similarity network in the social network at multiple levels, and then treating each user as a node, using the total correlation between users as the transmission probability, and using the label propagation algorithm to divide the community, thereby completing social discovery.

[0005] For example, the invention patent announcement with announcement number: CN103729467B discloses a method for discovering community structure in a social network, including: step 1: converting the social network into an adjacency matrix form, if there is an edge between two nodes, then the corresponding element is 1, otherwise it is 0; step 2: using random walk theory to process the adjacency matrix to obtain a new node degree P-degree and edge weight P-weight; step 3: obtaining the leader node in the social network according to the new node degree P-degree; step 4: generating a subcommunity based on the leader node, and performing community discovery through a series of operations on the subcommunity.

[0006] However, in the process of implementing the technical solution of the invention in the embodiments of the present application, the present application found that the above technology has at least the following technical problems: Existing technologies still face some challenges when applied to dynamically changing social networks. For example, many methods require key parameters to be set in advance, which limits their degree of automation. At the same time, when the social network structure changes, how to quickly update the community division is also an urgent problem to be solved, and there is a problem of low timeliness in social network community discovery. Summary of the invention

[0007] The embodiment of the present application solves the problem of low timeliness of social network community discovery in the prior art by providing a social network community discovery method based on density peak clustering of neighborhood fuzzy kernels, and improves the timeliness of hot community discovery in social networks.

[0008] The embodiment of the present application provides a community discovery method for a social network based on density peak clustering of a neighborhood fuzzy kernel, comprising the following steps: S1, obtaining a corresponding neighborhood fuzzy kernel by using all user data in a social network collected within a preset community discovery period, and obtaining community change data in real time to obtain an applicable impact index of the neighborhood fuzzy kernel, and judging whether to update the neighborhood fuzzy kernel according to the applicable impact index of the neighborhood fuzzy kernel; S2, clustering the neighborhood fuzzy kernel obtained in S1 to obtain a community set, and performing intersection processing on the community set and the community set of the previous preset time period to obtain a community partition set; S3, performing similarity calculation on the community set and the community set of the previous preset time period to obtain a similarity deviation, and performing community set update on the community partition set according to the similarity deviation.

[0009] Furthermore, the corresponding neighborhood fuzzy kernel is obtained by collecting all user data in the social network within a preset community discovery cycle, and the specific process is as follows: all user data in the social network is collected within a preset community discovery cycle, and the community set is initialized; the user data is subjected to dimensionality reduction processing and a geodesic distance operation is performed to obtain the corresponding geodesic distance, and the dimensionality reduction processing means mapping the complex structure of the user data in a high-dimensional space to a low-dimensional space through a dimensionality reduction tool; a geodesic distance matrix is ​​constructed based on the geodesic distances of all users, and the neighborhood fuzzy membership between users is obtained according to the geodesic distance matrix and the neighborhood set; a corresponding neighborhood fuzzy membership matrix is ​​constructed based on the neighborhood fuzzy membership, and a corresponding neighborhood fuzzy kernel is obtained based on the neighborhood fuzzy membership matrix and kernel parameters, and the kernel parameters include kernel width and kernel radius; and the local density of each user is calculated according to the neighborhood fuzzy membership matrix.

[0010] Furthermore, the specific steps of the dimensionality reduction processing are as follows: Step 1, set a user set in a high-dimensional space, and calculate the Euclidean distance between each user and the remaining users; Step 2, sort each user in ascending order according to the Euclidean distance, and select the user with the closest preset neighbor value to construct a neighbor set; Step 3, when the user is within the neighbor set of the remaining users, it means that the two users are neighbors of each other and an edge is established between them; Step 4, construct a mapping set of edge weights and Euclidean distances between two nodes to obtain the weights of the edges corresponding to the two nodes, and construct a neighbor graph according to the nodes corresponding to the users, the established edges and the edge weights; Step 5, obtain the shortest path distance between any two nodes on the neighbor graph through a distance tool; Step 6, construct a geodesic distance matrix between users based on all the shortest path distances obtained, and convert the geodesic distance matrix into a centralized matrix; Step 7, perform eigenvalue decomposition on the centralized matrix and extract its preset number of maximum eigenvalues ​​and their corresponding eigenvectors to obtain the coordinates of the user in the low-dimensional space.

[0011] Furthermore, the specific method for obtaining the neighborhood fuzzy membership is as follows: a given user set is obtained based on user data, and the users in the given user set are numbered; the corresponding Euclidean distance is obtained according to the social data set after dimensionality reduction, and the expected value of the Euclidean distance between all users is obtained based on the Euclidean distance to obtain the sparsity factor; a classification analysis operation is performed based on the sparsity factor, the Euclidean distance and the corresponding neighbor set to obtain the neighborhood fuzzy membership.

[0012] Furthermore, the real-time acquisition of community change data obtains the neighborhood fuzzy kernel applicable impact index, and the specific process is as follows: real-time acquisition of community change data within a preset community discovery cycle, the community change data including user data volume, relationship data volume and user attribute change rate; obtaining community change weights from a preset database, the community change weights including user weights, relationship weights and user attribute weights; obtaining user attribute impact factors by performing data change operations on user data volume, the data change operations are used to quantify changes in user data volume and map them; performing user change fluctuation operations on user data volume up to the current preset community discovery cycle to obtain user data fluctuation values, the user change fluctuation operations represent a method of quantifying the impact of user attributes up to the current preset community discovery cycle. The method for quantifying the fluctuation of the amount of user data in a period; performing a relationship change fluctuation operation on the amount of relationship data up to the current preset community discovery period to obtain a relationship data fluctuation value, wherein the relationship change fluctuation operation represents a method for quantifying the fluctuation of the amount of relationship data up to the current preset community discovery period; combining the user data fluctuation value, the relationship data fluctuation value, the user attribute influencing factor and the corresponding community change weight to distribute the degree of community change impact to obtain the neighborhood fuzzy kernel applicable impact index, wherein the community change impact distribution is used to comprehensively quantify the timeliness of the community change data on the neighborhood fuzzy kernel in the current preset community discovery period, and the neighborhood fuzzy kernel applicable impact index represents data that quantifies the degree of impact of community user data changes on the timeliness of the neighborhood fuzzy kernel.

[0013] Furthermore, the specific steps of judging whether to update the neighborhood fuzzy kernel based on the neighborhood fuzzy kernel applicability influence index are as follows: comparing the neighborhood fuzzy kernel applicability influence index with the influence critical value from a preset database: if the neighborhood fuzzy kernel applicability influence index is less than the influence critical value, the neighborhood fuzzy kernel is not updated; if the neighborhood fuzzy kernel applicability influence index is not less than the influence critical value, the neighborhood fuzzy kernel is updated; the specific process of updating the neighborhood fuzzy kernel is as follows: constructing a mapping set of the neighborhood fuzzy kernel applicability influence index and the kernel parameter adjustment factor, inputting the real-time neighborhood fuzzy kernel applicability influence index into the mapping set to obtain the corresponding kernel parameter adjustment factor, obtaining the updated kernel parameters according to the kernel parameter adjustment factor, and updating the neighborhood fuzzy kernel through the updated kernel parameters.

[0014] Furthermore, the neighborhood fuzzy kernel is clustered to obtain a community set, and the specific process is as follows: local density is obtained based on neighborhood fuzzy membership, preset neighbor value and the total number of users; the distance between each user and the user whose local density is higher than the local density of the current user and whose Euclidean distance is lower than the Euclidean distance of the current user is obtained to obtain the relative distance; users whose local density is higher than the preset density and whose relative distance is higher than the preset distance and the preset number of hot spots are selected as social centers for marking, and the community set is initialized to be empty; for unmarked users, the corresponding shared neighbor relationship is obtained by performing similarity operations on the unmarked users and the remaining marked users. similarity; perform proximity operation through the neighbor sets of two users and the corresponding Euclidean distance to obtain user proximity; perform similarity influence operation based on the shared neighbor similarity and user proximity between users to obtain weighted shared neighbor similarity; sort the weighted shared neighbor similarities in descending order, and assign the unlabeled users to the communities corresponding to the labeled users with the highest weighted shared neighbor similarity and mark them as labeled users; if there are still unlabeled users, assign them to the community where the labeled users with a larger local density than the unlabeled users and the smallest Euclidean distance to the unlabeled users are located and mark them as labeled users; output the corresponding community set based on the above assignment process.

[0015] Furthermore, the specific process of obtaining the community partition set is as follows: performing intersection processing on the community set and the community set of the previous preset community discovery cycle to obtain the community partition set, the community partition set includes a stable community set, a set to be optimized and a dynamically updated set, the stable community set represents the intersection part of the community set of the current community discovery cycle and the community set of the previous preset community discovery cycle, the set to be optimized represents the remaining set part of the community set of the previous preset community discovery cycle minus the intersection part, and the dynamically updated set represents the remaining set of the community set minus the intersection part; performing a similarity operation on the community set and the community set of the previous preset community discovery cycle to obtain a similarity deviation, the similarity operation is used to measure the similarity between the community set and the communities in the community set of the previous preset community discovery cycle.

[0016] Furthermore, the community set is updated according to the similarity deviation. The specific process is as follows: constructing a community set update mapping set of similarity deviation and preset division update ratio combination, inputting the real-time similarity deviation into the community set update mapping set to obtain a corresponding division update ratio combination, wherein the division update ratio combination includes an optimization set ratio and a dynamic set ratio; arranging the local densities corresponding to users in the optimization set in descending order, and updating the community set for the optimization set according to the optimization set ratio; arranging the local densities corresponding to users in the dynamic update set in descending order, and updating the community set for the dynamic update set according to the dynamic set ratio.

[0017] Furthermore, the specific content of the community set update is as follows: the set to be optimized obtained after arranging the local density in descending order retains the community where the corresponding user is located in order according to the optimization set ratio, and the remaining communities in the set to be optimized are recorded as the set to be updated; the dynamic update set obtained after arranging the local density in descending order retains the community where the corresponding user is located in order according to the dynamic set ratio and marks it as the update set; the set to be updated is updated through the update set, and merged with the stable community set to obtain the updated community set.

[0018] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. The corresponding neighborhood fuzzy kernel is obtained by collecting all user data in the social network within a preset community discovery cycle, and the community change data is obtained in real time to obtain the neighborhood fuzzy kernel applicable impact index, based on which it is judged whether to update the neighborhood fuzzy kernel, clustering the neighborhood fuzzy kernel to obtain a community set, and performing intersection processing on the community set and the community set of the previous preset time period to obtain a community partition set, and at the same time performing similarity calculation to obtain a similarity deviation to update the community partition set, so as to update the community hotspots in more real time, thereby improving the timeliness of hot community discovery in the social network, and effectively solving the problem of low timeliness of community discovery in the social network in the prior art; 2. Obtain a given user set through user data and number the users in the given user set. Then, obtain the corresponding Euclidean distance according to the social data set after dimensionality reduction, and obtain the expected value of the Euclidean distance between all users based on the Euclidean distance to obtain the sparse factor. Then, perform classification analysis based on the sparse factor, the Euclidean distance and the corresponding neighbor set to obtain the neighborhood fuzzy membership, thereby obtaining a more accurate neighborhood fuzzy kernel, thereby realizing the discovery of more accurate community hotspots. 3. By obtaining community change data and community change weights in real time within the preset community discovery cycle, and then obtaining user attribute influence factors through the amount of user data, and obtaining user data fluctuation values ​​based on the amount of user data up to the current preset community discovery cycle, and then obtaining relationship data fluctuation values ​​based on the amount of relationship data up to the current preset community discovery cycle, and finally combining the above data to obtain the neighborhood fuzzy kernel applicability impact index, it is possible to more accurately evaluate the applicability of the current neighborhood fuzzy kernel, thereby achieving more accurate community hotspots. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A flowchart of a social network community discovery method based on density peak clustering using neighborhood fuzzy kernels provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] The embodiment of the present application solves the problem of low timeliness of social network community discovery in the prior art by providing a social network community discovery method based on density peak clustering of neighborhood fuzzy kernels. All user data in the social network is collected within a preset community discovery period, and a given user set is obtained based on the obtained data and the users in the given user set are numbered. Then, the corresponding Euclidean distance is obtained according to the social data set after dimensionality reduction, and the expected value of the Euclidean distance between all users is obtained based on the Euclidean distance to obtain a sparse factor. Then, a classification analysis operation is performed based on the sparse factor, the Euclidean distance and the corresponding nearest neighbor set to obtain a neighborhood fuzzy membership to obtain a corresponding neighborhood fuzzy kernel, and the community variation is obtained in real time within the preset community discovery period. The data is quantified and the community change weight is obtained. Then, the user attribute influence factor is obtained through the user data volume, and the user data fluctuation value is obtained based on the user data volume up to the current preset community discovery cycle. Then, the relationship data fluctuation value is obtained according to the relationship data volume up to the current preset community discovery cycle. Then, the neighborhood fuzzy kernel applicable influence index is obtained by combining the above data, and it is judged whether to update the neighborhood fuzzy kernel based on this. The neighborhood fuzzy kernel is clustered to obtain a community set, and the community set is intersected with the community set of the previous preset time period to obtain a community partition set. At the same time, similarity calculation is performed to obtain a similarity deviation to update the community partition set, thereby improving the timeliness of hot community discovery in social networks.

[0021] The technical solution in the embodiment of the present application is to solve the problem of low timeliness of discovery of social network communities. The overall idea is as follows: The corresponding neighborhood fuzzy kernel is obtained by collecting all user data in the social network within the preset community discovery cycle, and the community change data is obtained in real time to obtain the neighborhood fuzzy kernel applicable impact index. Based on this, it is judged whether to update the neighborhood fuzzy kernel. The neighborhood fuzzy kernel is clustered to obtain a community set, and the community set is intersected with the community set of the previous preset time period to obtain a community partition set. At the same time, similarity calculation is performed to obtain the similarity deviation to update the community partition set, thereby achieving the effect of improving the timeliness of hot community discovery in social networks.

[0022] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0023] like Figure 1As shown, it is a flow chart of a method for discovering a community in a social network based on density peak clustering of a neighborhood fuzzy kernel provided by an embodiment of the present application, and the method comprises the following steps: S1, obtaining a corresponding neighborhood fuzzy kernel by using all user data in a social network collected within a preset community discovery period, and obtaining community change data in real time to obtain an applicable impact index of the neighborhood fuzzy kernel, and judging whether to update the neighborhood fuzzy kernel according to the applicable impact index of the neighborhood fuzzy kernel; S2, clustering the neighborhood fuzzy kernel obtained in S1 to obtain a community set, and performing intersection processing on the community set with the community set of the previous preset time period to obtain a community partition set; S3, performing similarity calculation on the community set and the community set of the previous preset time period to obtain a similarity deviation, and updating the community partition set according to the similarity deviation.

[0024] In this embodiment, the user's behavior data, including check-in records, interactive behaviors and geographic location information, is obtained from social networking platforms (such as Facebook, Instagram, Twitter, WeChat and TikTok, etc.). The above data usually contains a lot of noise and redundant information, so it needs to be preprocessed first. In the data cleaning stage, invalid or duplicate data entries are removed to ensure the accuracy and consistency of the data. The user's activity trajectory is segmented, and the user's continuous movement path is divided into meaningful segments, such as determining the frequented places according to the user's stay time and movement mode; for missing values, linear interpolation or other smoothing algorithms are used for filling and smoothing to improve data quality; through the process of the above method, the complex structure of social data in high-dimensional space is mapped to low-dimensional space, which can preserve the global structure and nonlinear relationship of the data while mapping it to low-dimensional space, thereby improving the data visualization effect and the timeliness of hot community discovery in social networks.

[0025] It should be added that the Density Peaks (DP) clustering algorithm provides a novel and effective clustering method by selecting points with higher local density and larger distance as cluster centers. However, when facing high-dimensional data sets, the traditional DP algorithm may have problems of low computational efficiency and difficult parameter setting. By introducing the concept of neighborhood fuzzy kernel, the neighborhood fuzzy membership is obtained, which can better capture the data distribution characteristics and improve the robustness to noise and outliers.

[0026] Furthermore, the corresponding neighborhood fuzzy kernel is obtained by collecting all the user data in the social network within the preset community discovery cycle. The specific process is as follows: all the user data in the social network is collected within the preset community discovery cycle, and the community set is initialized; the user data is subjected to dimensionality reduction processing and a geodesic distance operation is performed to obtain the corresponding geodesic distance, and the dimensionality reduction processing means mapping the complex structure of the user data in the high-dimensional space to the low-dimensional space through the dimensionality reduction tool; a geodesic distance matrix is ​​constructed based on the geodesic distances of all users, and the neighborhood fuzzy membership between users is obtained according to the geodesic distance matrix and the neighborhood set; a corresponding neighborhood fuzzy membership matrix is ​​constructed based on the neighborhood fuzzy membership, and the corresponding neighborhood fuzzy kernel is obtained based on the neighborhood fuzzy membership matrix and kernel parameters, and the kernel parameters include kernel width and kernel radius; the local density of each user is calculated according to the neighborhood fuzzy membership matrix.

[0027] In this embodiment, the neighborhood fuzzy membership measures the membership or similarity of each data point to other points in its neighborhood. The neighborhood fuzzy membership matrix contains the neighborhood fuzzy membership values ​​between all nodes, indicating the neighborhood similarity of each node to other nodes, and the neighborhood fuzzy membership matrix is ​​an n*n matrix, where n represents the total number of nodes in the community set. The neighborhood fuzzy kernel is a method of mapping the neighborhood fuzzy membership to a high-dimensional space. Usually, the neighborhood fuzzy membership matrix is ​​mapped to a new space through a kernel function (such as a Gaussian kernel). For example, suppose we define a kernel function , through the neighborhood fuzzy membership matrix To calculate: ; in, The kernel parameter is used to control the range of similarity. represents the i-th user, represents the jth user; by using the neighborhood fuzzy membership matrix , calculate a fuzzy kernel value for each pair of nodes, indicating the similarity between data nodes, and obtain the neighborhood fuzzy kernel matrix , whose elements are defined as: ; in, represents the neighborhood fuzzy kernel matrix between the i-th user and the j-th user, Represents the neighborhood fuzzy membership matrix of the i-th user and the j-th user; through the above process, the neighborhood fuzzy kernel matrix finally obtained is used for subsequent machine learning or social network analysis tasks, such as clustering, classification, dimensionality reduction, etc.; the neighborhood fuzzy kernel matrix is ​​calculated based on the neighborhood fuzzy membership matrix, and the kernel function is combined to obtain the similarity measure between each pair of nodes; by obtaining the neighborhood fuzzy kernel, the processing efficiency of subsequent clustering analysis is improved.

[0028] Specifically, the ISOMAP (Isometric Feature Mapping) algorithm is used to reduce the dimensionality of the pre-processed high-dimensional social data. First, an adjacency graph is constructed based on the geodesic distance between users to ensure that each node is connected to its nearest node. Then, the shortest path distance between each pair of nodes is calculated using the Dijkstra algorithm. Finally, the distance matrix in the high-dimensional space is converted into a low-dimensional embedding representation using the classic multidimensional scaling technique to obtain the manifold structure. The above process helps to preserve the global structure and nonlinear relationships of user social relationships, making the reduced-dimensional social data easier to analyze and visualize.

[0029] Furthermore, the specific steps of dimensionality reduction processing are as follows: Step 1, set a user set in a high-dimensional space, and calculate the Euclidean distance between each user and the remaining users; Step 2, sort each user in ascending order according to the Euclidean distance to the remaining users, and select the user with the closest preset neighbor value to build a neighbor set; Step 3, when a user is within the neighbor set of the remaining users, it means that the two users are neighbors and an edge is established between them; Step 4, construct a mapping set of edge weights and Euclidean distances between two nodes to obtain the weights of the edges corresponding to the two nodes, and construct a neighbor graph according to the nodes corresponding to the users, the established edges and the weights of the edges; Step 5, obtain the shortest path distance between any two nodes on the neighbor graph through the distance tool, and the calculation formula of the shortest path distance is as follows: ; ; In the formula, n represents the user number, , Represents the total number of users, Represents a given set of users, represents the i-th user, represents the jth user, represents the neighbor graph, represents the preset neighbor value, Represented in the neighborhood graph The Euclidean distance between the i-th user and the j-th user, Represented in the neighborhood graph The Euclidean distance between the i-th user and the k-th user, Represented in the neighborhood graph The Euclidean distance between the kth user and the jth user, Represented in the neighborhood graph The shortest path distance between the i-th user and the j-th user; Step 6: construct a geodesic distance matrix between users based on all the shortest path distances obtained, and convert the geodesic distance matrix into a centralized matrix. The centralized matrix is ​​obtained as follows: ; ; in, Represents the total number of users, Indicates a The identity matrix of represents an N-dimensional vector with all 1s, represents the transformation matrix, represents the geodesic distance matrix, represents the centralization matrix; Step 7: Perform eigenvalue decomposition on the centralized matrix and extract the preset number of maximum eigenvalues ​​and their corresponding eigenvectors to obtain the coordinates of the user in the low-dimensional space; The user's coordinates in low-dimensional space are calculated according to the following formula: ; in, represents the coordinates of the user in the low-dimensional space, Indicates the preset number. represents the first eigenvector, represents the second eigenvector, represents the dth eigenvector, represents the maximum eigenvalue.

[0030] In this embodiment, the preset nearest neighbor value is specifically set by a preset professional according to the social network situation; the preset number is specifically set by a preset professional according to the social network situation; the distance tool includes the Dijkstra algorithm or the Floyd algorithm; by introducing the neighborhood fuzzy membership to balance the similarity of users in dense and sparse areas, the dependence on the cutoff distance is reduced, and the similarity of user data in dense and sparse areas can be better balanced, thereby improving the accuracy and adaptability of the local density measurement, improving the accuracy of the neighborhood fuzzy kernel, and then improving the accuracy of community hotspot discovery.

[0031] Furthermore, the specific method of obtaining the neighborhood fuzzy membership is as follows: a given user set is obtained based on user data, and the users in the given user set are numbered; the corresponding Euclidean distance is obtained according to the social data set after dimensionality reduction, and the expected value of the Euclidean distance between all users is obtained based on the Euclidean distance to obtain the sparsity factor; a classification analysis operation is performed based on the sparsity factor, the Euclidean distance and the corresponding neighbor set to obtain the neighborhood fuzzy membership, and the classification analysis operation is used to obtain the neighborhood fuzzy membership of different situations according to the relationship between the user and the neighbor set. The specific restriction expression of the neighborhood fuzzy membership is as follows: ; ; ; ; In the formula, n represents the user number, , Represents the total number of users, represents the Euclidean distance between the i-th user and the j-th user, represents the sparse factor, represents the expected value of the Euclidean distance, represents the jth user, represents the neighbor set of the i-th user, Represents a given set of users, Represents the neighborhood fuzzy membership between the i-th user and the j-th user.

[0032] In this embodiment, there are differences in the calculation methods of the neighborhood fuzzy membership for the user's neighboring points and non-neighboring points; in the case of neighboring points (i.e. ), the neighborhood fuzzy membership is positively correlated with the distance between the user and its neighboring points; while in the case of non-neighboring points (i.e. ), the calculation of neighborhood fuzzy membership depends not only on the distance between users, but also on the sparse factor, which is used to measure the degree of distribution dispersion between users; the nearest neighbor fuzzy kernel can more accurately describe the neighborhood fuzzy membership relationship between user data. The closer its value is to 1, the stronger the neighborhood fuzzy membership of the sample is; through the above analysis, a more accurate neighborhood fuzzy membership is obtained, thereby improving the accuracy of the neighborhood fuzzy kernel, and then improving the accuracy of community hotspot discovery.

[0033] Furthermore, the community change data is obtained in real time to obtain the neighborhood fuzzy kernel applicable impact index, and the specific process is as follows: the community change data is obtained in real time within the preset community discovery cycle, and the community change data includes the user data volume, the relationship data volume and the user attribute change rate; the community change weight is obtained from the preset database, and the community change weight includes the user weight, the relationship weight and the user attribute weight; the user attribute impact factor is obtained by performing a data change operation on the user data volume, and the data change operation is used to quantify the change of the user data volume and map it; the user change fluctuation operation is performed on the user data volume as of the current preset community discovery cycle to obtain the user data fluctuation value, and the user change fluctuation operation represents a method of quantifying the user attribute impact factor as of the current preset community discovery cycle. The relationship change fluctuation method is used to calculate the fluctuation of the amount of user data; the relationship change fluctuation calculation is performed on the relationship data volume up to the current preset community discovery cycle to obtain the relationship data fluctuation value, and the relationship change fluctuation calculation represents a method for quantifying the fluctuation of the relationship data volume up to the current preset community discovery cycle; the community change impact degree is allocated by combining the user data fluctuation value, the relationship data fluctuation value, the user attribute influencing factor and the corresponding community change weight to obtain the neighborhood fuzzy kernel applicable impact index, and the community change impact degree allocation is used to comprehensively quantify the timeliness of the community change data on the neighborhood fuzzy kernel in the current preset community discovery cycle. The neighborhood fuzzy kernel applicable impact index represents data that quantifies the degree of influence of community user data changes on the timeliness of the neighborhood fuzzy kernel; The specific restriction expression of the applicable influence index of the neighborhood fuzzy kernel is as follows: ; ; Where t represents the number of the preset community discovery cycle, , Indicates the total number of preset community discovery cycles, represents the amount of user data in the tth preset community discovery cycle, represents the amount of relationship data in the t-th preset community discovery cycle, represents the user attribute change rate in the tth preset community discovery cycle, Indicates The amount of user data for a preset community discovery cycle, Indicates The amount of relationship data for a preset community discovery cycle, represents the user attribute influencing factor of the t-th preset community discovery cycle, Represents the user weight, represents the relationship weight, represents the user attribute weight, Represents the neighborhood fuzzy kernel applicability impact index of the t-th preset community discovery cycle.

[0034] In this embodiment, the algorithm combines the community change data and the community change weight for comprehensive analysis to obtain the neighborhood fuzzy kernel applicability impact index, where the neighborhood fuzzy kernel applicability impact index varies with the change deviation of the user data volume (i.e. ), the variation deviation of the relationship data volume (i.e. ) and user attribute change rate (i.e. ) increases. The larger the neighborhood fuzzy kernel applicability impact index is, the less suitable the domain fuzzy kernel of the current preset community discovery cycle is for the current community network. In addition, the user attribute change rate is affected by the change in the amount of user data. As the amount of user data increases or decreases, the corresponding user attribute change rate also increases or decreases. Therefore, a user attribute influence factor on the user attribute change rate is obtained through the change deviation of the user data amount. Through the above analysis, it can be seen that the neighborhood fuzzy kernel applicability impact index is helpful to more accurately evaluate the applicability of the domain fuzzy kernel, thereby improving the timeliness of the domain fuzzy kernel, and then analyzing and discovering more accurate community hotspots, thereby improving the efficiency of community hotspot discovery.

[0035] It should be added that the corresponding user data volume is obtained by counting the number of new and deleted users through event logs. In social networks, the relationships between users (such as follow, friends, likes, comments, etc.) are usually represented by edges. The addition or deletion of edges is monitored to obtain the corresponding relationship data volume. The user's attribute changes are recorded through user behavior logs, and the amount of user data whose user attributes have changed is compared with the total user data volume to obtain the corresponding user attribute change rate.

[0036] Specifically, the community change weight is obtained from a preset database, and the community change weight indicates the degree of influence of the community change data on the neighborhood fuzzy kernel applicability impact index. Each community change data has a unique mapping relationship with the community change weight, and the value range is between 0 and 1; for example, a mapping set of community change data and preset community change weights is constructed, and the real-time user data volume, relationship data volume, and user attribute change rate are input into the mapping set to obtain user weights, relationship weights, and user attribute weights, respectively, indicating the degree of influence of the user data volume, relationship data volume, and user attribute change rate on the neighborhood fuzzy kernel applicability impact index, and the sum of the three is 1.

[0037] Furthermore, whether to update the neighborhood fuzzy kernel is determined according to the neighborhood fuzzy kernel applicability influence index. The specific steps are as follows: compare the neighborhood fuzzy kernel applicability influence index with the influence critical value from the preset database: if the neighborhood fuzzy kernel applicability influence index is less than the influence critical value, the neighborhood fuzzy kernel is not updated; if the neighborhood fuzzy kernel applicability influence index is not less than the influence critical value, the neighborhood fuzzy kernel is updated; the specific process of updating the neighborhood fuzzy kernel is as follows: construct a mapping set of the neighborhood fuzzy kernel applicability influence index and the kernel parameter adjustment factor, input the real-time neighborhood fuzzy kernel applicability influence index into the mapping set to obtain the corresponding kernel parameter adjustment factor, obtain the updated kernel parameter according to the kernel parameter adjustment factor, and update the neighborhood fuzzy kernel through the updated kernel parameter.

[0038] In this embodiment, the kernel parameter adjustment factor and the kernel parameter are respectively multiplied to obtain the updated kernel parameter, and the kernel parameter of the neighborhood fuzzy kernel is adjusted to the updated kernel parameter to obtain the updated neighborhood fuzzy kernel; by adjusting the neighborhood fuzzy kernel according to the actual community change data, the neighborhood fuzzy kernel is made more adaptable to the changes in the complex dynamic environment of the social network; for example, as the user behavior changes, the update of the neighborhood fuzzy kernel can reflect more accurate user similarity.

[0039] Specifically, the impact critical value is obtained from a preset database. In a specific embodiment, the community change data of the accurate community hotspot corresponding situation obtained from the historical data is substituted into the specific restriction expression of the neighborhood fuzzy kernel applicable impact index to obtain the corresponding data set, and the data set is averaged to obtain the corresponding impact critical value.

[0040] Furthermore, the neighborhood fuzzy kernel is clustered to obtain a community set. The specific process is as follows: Based on the neighborhood fuzzy membership, the preset neighbor value and the total number of users, the local density is obtained. The specific method of obtaining the local density is as follows: ; in, represents the neighborhood fuzzy membership between the i-th user and the j-th user, represents the neighborhood fuzzy membership between the i-th user and the m-th user, represents the preset neighbor value, Represents the total number of users, represents the local density of the i-th user; The relative distance is obtained by obtaining the distance between each user and the user whose local density is higher than the local density of the current user and whose Euclidean distance is lower than the Euclidean distance of the current user. The relative distance is obtained as follows: ; in, represents the Euclidean distance between the i-th user and the j-th user, represents the set of local densities of all users, represents the local density of the i-th user, represents the local density of the jth user, represents the geodesic distance matrix, represents the relative distance of the i-th user; Select the users with a preset number of hot spots whose local density is higher than the preset density and whose relative distance is higher than the preset distance as social centers and mark them, and initialize the community set to be empty; for unmarked users, obtain the corresponding shared neighbor similarity by performing similarity operations on the unmarked users and the remaining marked users. The specific calculation method is as follows: ; in, represents the neighbor set of the i-th user, represents the neighbor set of the jth user, Represents a given set of users, represents the shared neighbor similarity between the i-th user and the j-th user; The user proximity is obtained by performing a proximity operation on the neighbor sets of two users and the corresponding Euclidean distance. The specific calculation method is as follows: ; in, represents the neighbor set of the i-th user, represents the neighbor set of the jth user, represents the Euclidean distance between the jth user and the uth user, represents the Euclidean distance between the i-th user and the v-th user, represents the user proximity between the i-th user and the j-th user; Based on the shared neighbor similarity and user proximity between users, similarity influence calculation is performed to obtain weighted shared neighbor similarity. The specific calculation method of weighted shared neighbor similarity is as follows: ; in, represents the user proximity between the i-th user and the j-th user, represents the shared neighbor similarity between the i-th user and the j-th user, represents the weighted shared neighbor similarity between the i-th user and the j-th user; Arrange the weighted shared neighbor similarities in descending order, and assign the unlabeled users to the communities corresponding to the labeled users with the highest weighted shared neighbor similarity and mark them as labeled users; if there are still unlabeled users, assign them to the communities where the labeled users with larger local density than the unlabeled users and the smallest Euclidean distance to the unlabeled users are located and mark them as labeled users; output the corresponding community set based on the above assignment process.

[0041] In this embodiment, the community set is finally output, in which each community is a community hotspot in the social network. These hotspot areas reflect specific areas where user activities are frequent and where a large number of people gather. For example, in the analysis of a certain city, the results obtained by the above method show that the city center business district, the surrounding areas of large shopping malls and the vicinity of major transportation hubs are the community hotspot areas with the most users. These areas are not only the core areas of urban economic vitality, but also an indispensable part of residents' daily life.

[0042] The above method helps to accurately identify community hotspots in social networks, helping city managers to better understand citizens' behavior patterns and develop corresponding optimization strategies. At the same time, it also provides merchants with accurate market analysis tools to help business development. For example, city managers can reasonably plan public transportation routes, increase bus frequencies or set up more public bicycle stations based on the distribution of community hotspots; merchants can choose appropriate store locations or develop targeted promotional activities based on the flow of people in hotspot areas, thereby increasing sales and market share.

[0043] In addition, identifying community hot spots can also help social media platforms better understand and optimize user experience. For example, the platform can push relevant localized content and services based on the user's active area to enhance user stickiness and satisfaction. At the same time, it can also provide accurate target audience positioning for advertising, improving advertising effectiveness and return on investment.

[0044] Furthermore, the specific process of obtaining the community partition set is as follows: performing intersection processing on the community set and the community set of the previous preset community discovery cycle to obtain the community partition set, the community partition set includes a stable community set, a set to be optimized and a dynamically updated set, the stable community set represents the intersection part of the community set and the community set of the previous preset community discovery cycle, the set to be optimized represents the remaining set part of the community set of the previous preset community discovery cycle minus the intersection part, and the dynamically updated set represents the remaining set of the community set minus the intersection part; performing a similarity operation on the community set and the community set of the previous preset community discovery cycle to obtain a similarity deviation, and the similarity operation is used to measure the similarity between the community set and the communities in the community set of the previous preset community discovery cycle.

[0045] In this embodiment, the similarity operation includes but is not limited to a cosine similarity algorithm, a Jaccard similarity algorithm, and a Pearson correlation coefficient algorithm; a deviation operation is performed on the result of the similarity operation between the community set and the community set of the previous preset community discovery cycle to obtain a similarity deviation, and the deviation operation represents a ratio operation between the difference between the result of the similarity operation between the community set of the current community discovery cycle and the community set of the previous preset community discovery cycle and the result of the similarity operation between the community set of the previous preset community discovery cycle; the division method based on stable communities, communities to be optimized, and dynamically updated communities is conducive to improving the accuracy, efficiency, and flexibility of community division, thereby improving the efficiency of community set update.

[0046] Furthermore, the community set is updated according to the similarity deviation. The specific process is as follows: construct a community set update mapping set of similarity deviation and preset partition update ratio combination, input the real-time similarity deviation into the community set update mapping set to obtain the corresponding partition update ratio combination, and the partition update ratio combination includes the optimization set ratio and the dynamic set ratio; arrange the local densities corresponding to the users in the optimization set in descending order, and update the community set of the optimization set according to the optimization set ratio; arrange the local densities corresponding to the users in the dynamic update set in descending order, and update the community set of the dynamic update set according to the dynamic set ratio.

[0047] In this embodiment, the accuracy, efficiency and adaptability of the community update process can be ensured by calculating the similarity deviation, using the combination of divided update ratios, and prioritizing based on local density; and by combining the adjustment of the optimized set ratio and the dynamic set ratio, the community set update is made more flexible and accurate, while the frequency of the community set update is effectively controlled.

[0048] Furthermore, the specific content of the community set update is as follows: the set to be optimized obtained after arranging the local density in descending order retains the community where the corresponding user is located in order according to the optimization set ratio, and the remaining communities in the set to be optimized are recorded as the set to be updated; the dynamic update set obtained after arranging the local density in descending order retains the community where the corresponding user is located in order according to the dynamic set ratio and marks it as the update set; the set to be updated is updated through the update set, and merged with the stable community set to obtain the updated community set.

[0049] In this embodiment, by sorting by local density, priority is ensured for the part of the community set that needs to be updated, and the combination of the optimized set ratio and the dynamic set ratio makes the update of the community set more targeted, thereby improving the accuracy of the community set update; and according to the adjustment of the optimized set ratio and the dynamic set ratio, the intensity of optimization and update is more flexibly controlled, so that the division of the community set can adaptively adjust the network changes; at the same time, the merged community set can not only maintain the original stability, but also absorb the changes of real-time updates, thereby ensuring the accuracy of community division and the timeliness of the community set.

[0050] In summary, the embodiment of the present application obtains the corresponding neighborhood fuzzy kernel by collecting all user data in the social network within a preset community discovery period, and obtains the community change data in real time to obtain the neighborhood fuzzy kernel applicable impact index, and judges whether to update the neighborhood fuzzy kernel based on this, clusters the neighborhood fuzzy kernel to obtain a community set, and performs intersection processing on the community set and the community set of the previous preset time period to obtain a community partition set, and simultaneously performs similarity calculation to obtain a similarity deviation to update the community partition set for the community set, thereby updating the community hotspots in more real time, thereby improving the timeliness of hot community discovery in the social network, and effectively solving the problem of low timeliness of community discovery in the social network in the prior art.

[0051] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0052] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0053] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0054] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0055] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0056] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A social network community discovery method based on density peak clustering of neighborhood fuzzy kernel, characterized in that: The following steps are involved: S1, obtain the corresponding neighborhood fuzzy kernel by collecting all user data in the social network within a preset community discovery cycle, and obtain the neighborhood fuzzy kernel applicable impact index in real time by obtaining the community change data, and determine whether to update the neighborhood fuzzy kernel according to the neighborhood fuzzy kernel applicable impact index; S2, clustering the neighborhood fuzzy kernel obtained in S1 to obtain a community set, and performing intersection processing on the community set and the community set of the previous preset time period to obtain a community partition set; S3, calculating the similarity between the community set and the community set in the previous preset time period to obtain a similarity deviation, and updating the community set for the community division set according to the similarity deviation.

2. The method for discovering social network communities based on density peak clustering of neighborhood fuzzy kernels as claimed in claim 1, characterized in that: The corresponding neighborhood fuzzy kernel is obtained by collecting all user data in the social network within the preset community discovery cycle. The specific process is as follows: Collect all user data in the social network within the preset community discovery cycle and initialize the community set; Performing dimensionality reduction processing on the user data and performing geodesic distance calculation to obtain corresponding geodesic distances, wherein the dimensionality reduction processing means mapping the complex structure of the user data in a high-dimensional space to a low-dimensional space by using a dimensionality reduction tool; A geodesic distance matrix is ​​constructed based on the geodesic distances of all users, and the neighborhood fuzzy membership between users is obtained according to the geodesic distance matrix and the neighborhood set; Based on the neighborhood fuzzy membership, a corresponding neighborhood fuzzy membership matrix is ​​constructed, and based on the neighborhood fuzzy membership matrix and kernel parameters, a corresponding neighborhood fuzzy kernel is obtained, wherein the kernel parameters include a kernel width and a kernel radius; The local density of each user is calculated based on the neighborhood fuzzy membership matrix.

3. The social network community discovery method based on density peak clustering of neighborhood fuzzy kernel as claimed in claim 2, characterized in that: The specific steps of the dimensionality reduction process are as follows: Step 1: Set a user set in a high-dimensional space and calculate the Euclidean distance between each user and the other users; Step 2: sort each user in ascending order according to the Euclidean distance, and select the user with the closest preset neighbor value to build a neighbor set; Step 3: When a user is in the neighbor set of other users, it means that the two users are neighbors and an edge is established between them; Step 4: construct a mapping set of edge weights and Euclidean distances between two nodes to obtain the edge weights corresponding to the two nodes, and construct a neighbor graph based on the user's corresponding nodes, the established edges, and the edge weights; Step 5: Obtain the shortest path distance between any two nodes using the distance tool on the neighbor graph; Step 6: construct a geodesic distance matrix between users based on all the shortest path distances obtained, and convert the geodesic distance matrix into a centralized matrix; Step seven, perform eigenvalue decomposition on the centralized matrix and extract a preset number of maximum eigenvalues ​​and their corresponding eigenvectors to obtain the coordinates of the user in the low-dimensional space.

4. The social network community discovery method based on density peak clustering of neighborhood fuzzy kernel as claimed in claim 2, characterized in that: The specific method of obtaining the neighborhood fuzzy membership is as follows: A given user set is obtained based on the user data, and users in the given user set are numbered; Obtain the corresponding Euclidean distance according to the social data set after dimensionality reduction, and obtain the expected value of the Euclidean distance between all users based on the Euclidean distance to obtain the sparse factor; The neighborhood fuzzy membership is obtained by performing classification analysis based on the sparse factor, Euclidean distance and the corresponding neighbor set.

5. The method for discovering social network communities based on density peak clustering of neighborhood fuzzy kernels as claimed in claim 1, characterized in that: The real-time acquisition of community change data to obtain the neighborhood fuzzy kernel applicable impact index is as follows: Acquire community change data in real time within a preset community discovery cycle, wherein the community change data includes user data volume, relationship data volume, and user attribute change rate; Obtaining a community change weight from a preset database, wherein the community change weight includes a user weight, a relationship weight, and a user attribute weight; The user attribute influencing factor is obtained by performing a data change operation on the user data volume, wherein the data change operation is used to quantify the change of the user data volume and perform mapping; Performing a user change fluctuation calculation on the amount of user data up to the current preset community discovery period to obtain a user data fluctuation value, wherein the user change fluctuation calculation represents a method of quantifying the change fluctuation of the amount of user data up to the current preset community discovery period; Performing a relationship change fluctuation operation on the relationship data volume up to the current preset community discovery period to obtain a relationship data fluctuation value, wherein the relationship change fluctuation operation represents a method of quantifying the relationship data volume change fluctuation up to the current preset community discovery period; The community change impact degree is distributed by combining the user data fluctuation value, the relationship data fluctuation value, the user attribute influencing factor and the corresponding community change weight to obtain the neighborhood fuzzy kernel applicable impact index. The community change impact degree distribution is used to comprehensively quantify the timeliness of the community change data on the neighborhood fuzzy kernel in the current preset community discovery cycle. The neighborhood fuzzy kernel applicable impact index represents data that quantifies the impact of community user data changes on the timeliness of the neighborhood fuzzy kernel.

6. The method for discovering social network communities based on density peak clustering of neighborhood fuzzy kernels as claimed in claim 5, characterized in that: The specific steps of judging whether to update the neighborhood fuzzy kernel according to the neighborhood fuzzy kernel applicable influence index are as follows: Compare the neighborhood fuzzy kernel applicability impact index with the impact threshold from the preset database: If the applicable influence index of the neighborhood fuzzy kernel is less than the critical influence value, the neighborhood fuzzy kernel will not be updated; If the applicable influence index of the neighborhood fuzzy kernel is not less than the critical influence value, the neighborhood fuzzy kernel is updated; The specific process of updating the neighborhood fuzzy kernel is as follows: A mapping set of neighborhood fuzzy kernel applicability impact index and kernel parameter adjustment factor is constructed, and the real-time neighborhood fuzzy kernel applicability impact index is input into the mapping set to obtain the corresponding kernel parameter adjustment factor. The updated kernel parameters are obtained according to the kernel parameter adjustment factor, and the neighborhood fuzzy kernel is updated by the updated kernel parameters.

7. The method for discovering social network communities based on density peak clustering of neighborhood fuzzy kernels as claimed in claim 1, characterized in that: The clustering process of the neighborhood fuzzy kernel to obtain the community set is as follows: The local density is obtained based on the neighborhood fuzzy membership, the preset neighbor value and the total number of users; Obtain the distance between each user and the user whose local density is higher than the local density of the current user and whose Euclidean distance is lower than the Euclidean distance of the current user to obtain the relative distance; Select the users with a preset number of hotspots whose local density is higher than the preset density and whose relative distance is higher than the preset distance as social centers, and initialize the community set to be empty; For unlabeled users, the corresponding shared neighbor similarity is obtained by performing similarity operations between the unlabeled users and the remaining labeled users; The user proximity is obtained by performing a proximity operation through the neighbor sets of two users and the corresponding Euclidean distance; Based on the shared neighbor similarity and user proximity between users, similarity influence calculation is performed to obtain weighted shared neighbor similarity; Arrange the weighted shared neighbor similarities in descending order, assign the unmarked users to the communities corresponding to the marked users with the highest weighted shared neighbor similarities, and mark them as marked users; If there are still unmarked users, they will be assigned to the community where the marked users with a larger local density than the unmarked users and the smallest Euclidean distance to the unmarked users are located and marked as marked users; Based on the above allocation process, the corresponding community set is output.

8. The method for discovering social network communities based on density peak clustering based on neighborhood fuzzy kernel as claimed in claim 1, characterized in that: The specific process of obtaining the community division set is as follows: Performing intersection processing on the community set and the community set of the last preset community discovery cycle to obtain a community partition set, wherein the community partition set includes a stable community set, a set to be optimized, and a dynamically updated set, wherein the stable community set represents the intersection of the community set of the current community discovery cycle and the community set of the last preset community discovery cycle, the set to be optimized represents the remaining set portion after the community set of the last preset community discovery cycle is removed from the intersection portion, and the dynamically updated set represents the remaining set after the community set is removed from the intersection portion; A similarity deviation of a similarity operation performed on the community set and the community set of the last preset community discovery cycle, wherein the similarity operation is used to measure the similarity between the community set and the communities in the community set of the last preset community discovery cycle.

9. The social network community discovery method based on density peak clustering of neighborhood fuzzy kernel as claimed in claim 8, characterized in that: The specific process of updating the community set according to the similarity deviation is as follows: Constructing a community set update mapping set of similarity deviation and preset partition update ratio combination, inputting the real-time similarity deviation into the community set update mapping set to obtain a corresponding partition update ratio combination, wherein the partition update ratio combination includes an optimized set ratio and a dynamic set ratio; The local densities corresponding to the users in the set to be optimized are sorted in descending order, and the community set of the set to be optimized is updated according to the proportion of the optimized set; The local densities corresponding to the users in the dynamic update set are sorted in descending order, and the community set is updated for the dynamic update set according to the dynamic set ratio.

10. The method for discovering social network communities based on density peak clustering of neighborhood fuzzy kernels as claimed in claim 1, characterized in that: The specific content of the community collection update is as follows: The set to be optimized obtained after sorting the local density in descending order retains the communities where the corresponding users are located in order according to the optimization set ratio, and the remaining communities in the set to be optimized are recorded as the set to be updated; The dynamic update set obtained by sorting the local density in descending order is marked as the update set by retaining the community where the corresponding user is located in order according to the dynamic set ratio; The set to be updated is updated through the update set, and is merged with the stable community set to obtain the updated community set.

Citation Information

Patent Citations

  • A Community Structure Discovery Method in Social Networks

    CN103729467B

  • A Community Discovery Method for Multidimensional Social Networks

    CN108090197B

  • Improved density peak clustering-based social network community discovery method

    CN108647739A

  • An overlapping community detection method based on density peak and community belonging degree

    CN108959652A

  • Dynamic weighted hybrid clustering algorithm based circuit breaker fault diagnosis method

    CN109444728A

Cited By

  • Network community discovery system and method based on matrix analysis

    CN121213273A