Target community discovery method and device fusing content structure rule and time rule

By identifying the user's content structure rules and time rules, calculating similarity and using spectral clustering methods, the existing technology is not efficient and difficult to consider user preferences in large-scale network environments, and the effect of accurately positioning the target community is achieved.

CN119991328AActive Publication Date: 2025-05-13NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510151780.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-19
Filing Date
2025-02-12
Publication Date
2025-05-13
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

Existing community detection technology is not efficient in large-scale network environments and is difficult to consider users' personalized preferences, making it difficult to accurately locate the target community based on user preferences.

Method used

By obtaining the user's post information and the number of posts, identifying the user's content structure rules and time rules, calculating the similarity between content structure rules and time rules between users, establishing an undirected weighted graph, and using the spectral clustering method for community discovery.

Benefits of technology

This method can reveal potential connections and document writing habits between users, help discover hidden communities and influence networks in social platforms, and accurately locate target communities based on user preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991328A_ABST
    Figure CN119991328A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of complex network analysis, in particular to a content structure rule and time rule fused target community discovery method and device, and the method comprises the steps: obtaining the message sending information and message sending times of a user; user content structure rules are identified from the text sending information, and the similarity of the content structure rules among the users is calculated through a Jaccard similarity coefficient; constructing a user document sending time rule matrix based on the document sending times, and calculating time rule similarity among users through a Pearson's correlation coefficient; establishing a network undirected weighted graph based on the inter-user content structure rule similarity and the inter-user time rule similarity; and performing community discovery on the network undirected weighted graph by using a spectral clustering method to obtain a community division result. According to the technical scheme, the hidden communities and the influence network in the social platform can be found, and the target community based on the user preference can be accurately positioned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of complex network analysis, and in particular to a method and device for discovering a target community by integrating content structure rules and time regularities. Background Art

[0002] Community detection has become a very important research direction in the field of complex network analysis. Existing community detection methods can be roughly divided into two categories: global community detection methods for the analysis of the entire complex network and local community detection methods designed for specific goals. Global community detection, as a traditional community detection method, is mostly based on global information. As the scale of the network continues to increase, it is very difficult to obtain the global information of the network. At the same time, the amount of calculation in the acquisition process is also very huge. Therefore, traditional community detection methods are not efficient in large-scale network environments. Community detection methods based on local community detection often use seeds as the initial community and expand the community by running the greedy optimization process of the quality function. Therefore, the quality function and the expansion method determine the effectiveness of the local community detection method.

[0003] Although current community detection technology has alleviated the efficiency problems faced by community detection to a certain extent, it often focuses on the topological structure of the network and pays less attention to the personalized preferences of users. It is difficult to accurately locate the target community based on user preferences for specific applications. Summary of the invention

[0004] In order to solve the problems in the related art, the embodiments of the present disclosure provide a method and device for discovering a target community by integrating content structure rules and time regularities.

[0005] In a first aspect, the present disclosure provides a method for discovering a target community by integrating content structure rules and time regularities, including:

[0006] Get the user's posting information and posting times;

[0007] Identifying user content structure rules from the posting information, and calculating the similarity of content structure rules between users by using the Jaccard similarity coefficient;

[0008] Based on the number of posts, a user posting time regularity matrix is ​​constructed, and the time regularity similarity between users is calculated by the Pearson correlation coefficient;

[0009] Establishing a network undirected weighted graph based on the similarity of content structure rules between users and the similarity of time rules between users;

[0010] The spectral clustering method is used to perform community discovery on the undirected weighted graph of the network to obtain a community division result.

[0011] In one implementation of the present disclosure, the identifying the user content structure rule from the posting information includes:

[0012] Constructing a pattern set for content structure rule recognition; the pattern set includes preset pattern characters;

[0013] Recognizing preset pattern characters in the posted information;

[0014] All recognized preset pattern characters are used as user content structure rules.

[0015] In an implementation of the present disclosure, constructing a user posting time regularity matrix based on the posting times includes:

[0016] Count the number of posts in a preset time period;

[0017] Based on the preset time period, the number of posts in the preset time period is used to construct a user posting time regularity matrix.

[0018] In an implementation of the present disclosure, constructing a user posting time regularity matrix based on the posting times further includes:

[0019] If there is interaction between users, a bias is added to the corresponding element in the user posting time regularity matrix.

[0020] In an implementation of the present disclosure, before calculating the temporal regularity similarity between users by using the Pearson correlation coefficient, the method further includes:

[0021] The user posting time regularity matrix is ​​normalized.

[0022] In an implementation of the present disclosure, the establishing of a network undirected weighted graph based on the content structure rule similarity between users and the time regularity similarity between users includes:

[0023] The social network is represented as an undirected unweighted graph; the undirected unweighted graph includes user nodes and edges connecting the user nodes;

[0024] Calculate the weight WJR of the edge in the undirected unweighted graph ij , the formula is as follows:

[0025] Wl ij =αJ(v i ,v j )+(1-α)R(v i ,v j );

[0026] Based on the weights WJR ij and undirected unweighted graphs to build network undirected weighted graphs;

[0027] Among them, α is an adjustable parameter, 0<α<1, J(v i ,v j ) represents user v i With user v j The similarity of content structure rules between users, R(v i ,v j ) represents user v i With user v j The temporal regularity similarity between users.

[0028] In one implementation of the present disclosure, it further includes:

[0029] The modularity of the community is calculated based on the community division result.

[0030] In a second aspect, the present disclosure provides a target community discovery device integrating content structure rules and time regularities, including:

[0031] An acquisition module is configured to acquire the user's posting information and posting times;

[0032] A first calculation module is configured to identify user content structure rules from the posting information and calculate the similarity of content structure rules between users through the Jaccard similarity coefficient;

[0033] A second calculation module is configured to construct a user posting time regularity matrix based on the posting times, and calculate the time regularity similarity between users through the Pearson correlation coefficient;

[0034] A construction module configured to establish a network undirected weighted graph based on the similarity of content structure rules between users and the similarity of time rules between users;

[0035] The community discovery module is configured to use a spectral clustering method to perform community discovery on the undirected weighted graph of the network to obtain a community division result.

[0036] In an implementation of the present disclosure, the part of the first calculation module that identifies the user content structure rule from the posting information is configured as follows:

[0037] Constructing a pattern set for content structure rule recognition; the pattern set includes preset pattern characters;

[0038] Recognizing preset pattern characters in the posted information;

[0039] All recognized preset pattern characters are used as user content structure rules.

[0040] In an implementation of the present disclosure, the part of the second calculation module that constructs the user posting time regularity matrix based on the posting times is configured as follows:

[0041] Count the number of posts in a preset time period;

[0042] Based on the preset time period, the number of posts in the preset time period is used to construct a user posting time regularity matrix.

[0043] In an implementation of the present disclosure, the part of the second calculation module that constructs the user posting time regularity matrix based on the posting times is further configured as follows:

[0044] If there is interaction between users, a bias is added to the corresponding element in the user posting time regularity matrix.

[0045] In one implementation of the present disclosure, the device further includes:

[0046] The processing module is configured to normalize the user posting time regularity matrix.

[0047] In one implementation of the present disclosure, the building blocks include:

[0048] A construction unit is configured to represent a social network as an undirected unweighted graph; the undirected unweighted graph includes user nodes and edges connecting the user nodes;

[0049] A computing unit configured to calculate the weight WJR of the edge in the undirected unweighted graph ij , the formula is as follows:

[0050] Wl ij =αJ(v i ,v j )+(1-α)R(v i ,v j );

[0051] Establishing unit, configured to be based on the weight WJR ij and undirected unweighted graphs to build network undirected weighted graphs;

[0052] Among them, α is an adjustable parameter, 0<α<1, J(v i ,v j ) represents user v i With user v j The similarity of content structure rules between users, R(v i ,v j ) represents user v i With user v j The temporal regularity similarity between users.

[0053] In one implementation of the present disclosure, it further includes:

[0054] The third calculation module is configured to calculate the modularity of the community based on the community division result.

[0055] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising a memory and a processor, wherein the memory is used to store one or more computer instructions, and wherein the one or more computer instructions are executed by the processor to implement a method as described in any one of the first aspects.

[0056] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the method described in any one of the first aspects is implemented.

[0057] The technical effects provided by the embodiments of the present disclosure may include the following beneficial effects:

[0058] According to the technical solution provided by the embodiment of the present disclosure, the target community discovery method integrating content structure rules and time regularities includes: obtaining the user's posting information and posting times; identifying the user's content structure rules from the posting information, and calculating the content structure rule similarity between users through the Jaccard similarity coefficient; constructing a user posting time regularity matrix based on the posting times, and calculating the time regularity similarity between users through the Pearson correlation coefficient; establishing a network undirected weighted graph based on the content structure rule similarity between users and the time regularity similarity between users; using the spectral clustering method to perform community discovery on the network undirected weighted graph to obtain community division results. In the above technical solution, by calculating the content structure rule similarity between users and the time regularity similarity between users, the potential connections and posting writing habits between users can be revealed, which helps to discover hidden communities and influence networks in social platforms, accurately locate target communities based on user preferences, provide valuable insights for researchers, and provide strong support for research and application in the field of community discovery.

[0059] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Other features, objectives and advantages of the present disclosure will become more apparent through the following detailed description of non-limiting embodiments in conjunction with the accompanying drawings.

[0061] Figure 1 A flow chart of a target community discovery method integrating content structure rules and time regularities according to an embodiment of the present disclosure is shown.

[0062] Figure 2 A flowchart of a target community discovery method integrating content structure rules and time regularities according to a specific embodiment of the present disclosure is shown.

[0063] Figure 3 A schematic diagram showing a user posting time regularity matrix according to an embodiment of the present disclosure.

[0064] Figure 4 A structural block diagram of a target community discovery device integrating content structure rules and time regularities according to an embodiment of the present disclosure is shown.

[0065] Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0066] Figure 6 A schematic diagram showing the structure of a computer system suitable for implementing the method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0067] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for the sake of clarity, parts not related to the description of the exemplary embodiments are omitted in the accompanying drawings.

[0068] In the present disclosure, it should be understood that terms such as "include" or "have" are intended to indicate the presence of features, numbers, steps, behaviors, components, parts, or a combination thereof disclosed in the present specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, behaviors, components, parts, or a combination thereof exist or are added.

[0069] It should also be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0070] Although current community detection technology has alleviated the efficiency problems faced by community detection to a certain extent, it often focuses on the topological structure of the network and pays less attention to the personalized preferences of users. It is difficult to accurately locate the target community based on user preferences for specific applications.

[0071] Taking the above-mentioned defects into consideration, the target community discovery method integrating content structure rules and time regularities provided by the present disclosure includes: obtaining the user's posting information and posting times; identifying the user's content structure rules from the posting information, and calculating the content structure rule similarity between users through the Jaccard similarity coefficient; constructing a user posting time regularity matrix based on the posting times, and calculating the time regularity similarity between users through the Pearson correlation coefficient; establishing a network undirected weighted graph based on the content structure rule similarity between users and the time regularity similarity between users; using the spectral clustering method to perform community discovery on the network undirected weighted graph to obtain community division results. In the above technical solution, by calculating the content structure rule similarity between users and the time regularity similarity between users, the potential connections and posting writing habits between users can be revealed, which helps to discover hidden communities and influence networks in social platforms, accurately locate target communities based on user preferences, provide valuable insights for researchers, and provide strong support for research and application in the field of community discovery.

[0072] Figure 1 A flow chart of a target community discovery method integrating content structure rules and time regularities according to an embodiment of the present disclosure is shown.

[0073] like Figure 1 As shown, the target community discovery method integrating content structure rules and time regularities includes the following steps S110-S150:

[0074] In step S110, the user's posting information and posting times are obtained;

[0075] In step S120, the user content structure rules are identified from the posting information, and the similarity of the content structure rules between users is calculated by using the Jaccard similarity coefficient;

[0076] In step S130, a user posting time regularity matrix is ​​constructed based on the posting times, and the time regularity similarity between users is calculated by the Pearson correlation coefficient;

[0077] In step S140, a network undirected weighted graph based on the content structure rule similarity between users and the time regularity similarity between users is established;

[0078] In step S150, a spectral clustering method is used to perform community discovery on the undirected weighted graph of the network to obtain a community division result.

[0079] In the disclosed method, the posting information of each user on the social platform is first collected, and then user feature mining is performed based on the posting information. Specifically, the posting information may include the posting time, posting content, and posting type such as text, picture, video, etc. When performing user feature mining, the content structure rules are extracted for the posting information of each user to obtain the user content structure rules of each user, and then the similarity of the content structure rules between users is calculated. In the disclosed method, the similarity of the content structure rules between users is calculated by the Jaccard similarity coefficient.

[0080] In the disclosed method, based on the posting records of user posting information, the number of posts by users in a preset time period, such as every hour or every two hours, can be counted, and then user feature mining can be performed based on the number of posts. Specifically, when performing user feature mining, the user posting information can be aggregated into a list according to the preset time period, such as hours, to represent the time pattern of the user in the corresponding hour, and then a user posting time pattern matrix is ​​constructed, in which the rows represent each user, the columns represent different time periods, and the elements in the matrix represent the number of posts or time patterns of a user in a certain time period. Then, based on the user posting time pattern matrix of each user, the time pattern similarity between users is calculated. In the present disclosure, the time pattern similarity between users is calculated by the Pearson correlation coefficient.

[0081] In the disclosed method, the similarity of content structure rules between users and the similarity of time rules between users are weighted and combined into a comprehensive weight, a network undirected weighted graph is constructed, and a spectral clustering method is used to perform community discovery on the network undirected weighted graph. Users with similar posting writing habits and similar posting time periods can be divided into the same community, so that the node similarity within the cluster is higher and more stable, the similar users that can be portrayed are more explanatory and convincing, and the quality of community discovery will be better.

[0082] In one implementation of the present disclosure, the step S120 of identifying the user content structure rule from the posting information includes:

[0083] Constructing a pattern set for content structure rule recognition; the pattern set includes preset pattern characters;

[0084] Recognizing preset pattern characters in the posted information;

[0085] All recognized preset pattern characters are used as user content structure rules.

[0086] In the disclosed method, when performing content structure rule recognition, it is necessary to construct a pattern set for content structure rule recognition. The pattern set in the disclosed method includes preset pattern characters, which can be dates such as year, month, day, etc., numbers such as 1, 2, 3, etc., symbols such as curly brackets {}, parentheses (), etc., special characters such as triangle ▲, circle ●, etc.

[0087] In a specific embodiment, the pattern set P = {year, month, day, release, name, http, #, \$, :, ;, ;, \+, ↓, ■, ▲, ◆, ●, ★, ☆, 〈, 〉, 《, 》, 「, 」, 『, 』,

,

[0088] Then, a regular expression or AC automaton algorithm is used to find all possible locations of the pattern set P in the text of the post information, and then all the preset pattern characters R={r 1 ,r 2 ,r 3 ,…,r n}, and use all the recognized preset pattern characters R as user content structure rules. 1 ,r 2 ,r 3 ,…,r n is a single recognized preset mode character, and n is the number of recognized preset mode characters.

[0089] In one implementation of the present disclosure, the Jaccard similarity coefficient (i.e., the Jaccard coefficient) is applied to the set operation of word frequency, and the ratio of the intersection and the union of two sets is calculated. The index related to the Jaccard coefficient is called the Jaccard distance, which is used to describe the dissimilarity between sets. The larger the Jaccard distance, the lower the sample similarity.

[0090] In step S120, the similarity of content structure rules between users is calculated by the Jaccard similarity coefficient, where the Jaccard coefficient is defined as the ratio of the size of the intersection of the preset pattern characters in the content structure rules of user a and user b to the size of the union of the preset pattern characters in the content structure rules of user a and user b, and is defined as follows:

[0091]

[0092] Among them, R a is the set of preset pattern characters in the content structure rule of user a, R b is the set of preset pattern characters in the content structure rule of user b. When the set R a , R b When both are empty, J(a,b) is defined as 1.

[0093] In an implementation of the present disclosure, constructing a user posting time regularity matrix based on the posting times in step S130 includes:

[0094] Count the number of posts in a preset time period;

[0095] Based on the preset time period, the number of posts in the preset time period is used to construct a user posting time regularity matrix.

[0096] In the disclosed method, a user posting time regularity feature is constructed to describe the user posting time habit. The preset time period in the disclosed method can be minutes, hours, etc., and the statistics of the number of user postings can only count the user's own posting information, and can also count the user's forwarding of other users' posting information as the number of user postings.

[0097] The preset time period for statistics is hourly. Taking 24 hours a day as an example, a matrix of user posting time patterns is constructed. For example Figure 3 As shown, the rows represent users V1, V2, V3...Vn, the columns represent time periods 0-23, and the elements in the matrix represent the number of posts by the user in the time period.

[0098] In an implementation of the present disclosure, constructing a user posting time regularity matrix based on the posting times in step S130 further includes:

[0099] If there is interaction between users, a bias is added to the corresponding element in the user posting time regularity matrix.

[0100] In the disclosed method, interactive behaviors between users, such as comments and likes, indicate that there is a relationship between the users. Therefore, a bias can be added to the matrix elements corresponding to the users with the interactive relationship. Subsequently, the similarity of the temporal patterns between users is calculated based on the biased matrix elements, and then the users with the interactive relationship can be divided into the same community as much as possible when community discovery is performed.

[0101] In an implementation of the present disclosure, before calculating the temporal regularity similarity between users by using the Pearson correlation coefficient in step S130, the method further includes:

[0102] The user posting time regularity matrix is ​​normalized.

[0103] In the disclosed method, since the number of posts by different users may vary greatly, the time regularity matrix needs to be normalized to eliminate the magnitude differences between users. Specifically, the maximum and minimum normalization method or the Z-score normalization method can be used to eliminate the magnitude differences.

[0104] In one implementation of the present disclosure, in step S130, the Pearson correlation coefficient is used to measure the linear correlation between two variables in calculating the temporal regularity similarity between users. The calculation formula is as follows:

[0105]

[0106] In the formula, x i Represents different preset time period lengths; y i Represents the time regularity score of user posting in each preset time period, which is set according to the number of posts by the user in the time period. For example, the more posts the user makes, the higher the score. m is the number of preset time periods in the statistical duration. For example, if the preset time period is counted in hours and the statistical duration is 1 day, then m is 24.

[0107] The posting and forwarding activities of users in the same region on social platforms generally have obvious periodicity. The posting time patterns of users in different time periods within 24 hours of the day vary greatly. The similarity of the heavy active posting time periods of different users every day is of reference significance for calculating the similarity of time patterns between users.

[0108] In an implementation of the present disclosure, the step S140 of establishing a network undirected weighted graph based on the content structure rule similarity between users and the time regularity similarity between users includes:

[0109] The social network is represented as an undirected unweighted graph; the undirected unweighted graph includes user nodes and edges connecting the user nodes;

[0110] Calculate the weight WJR of the edge in the undirected unweighted graph ij , the formula is as follows:

[0111] Wl ij =αJ(v i ,v j )+(1-α)R(v i ,v j );

[0112] Based on the weights WJR ij and undirected unweighted graphs to build network undirected weighted graphs;

[0113] Among them, α is an adjustable parameter, 0<α<1, J(v i ,v j ) represents user v i With user v j The similarity of content structure rules between users, R(v i ,v j ) represents user v i With user v j The temporal regularity similarity between users.

[0114] In an implementation of the present disclosure, in step S150, a spectral clustering method is used to perform community discovery on the undirected weighted graph of the network to obtain a community division result.

[0115] In one implementation of the present disclosure, it further includes:

[0116] The modularity of the community is calculated based on the community division result.

[0117] In the disclosed method, modularity is also called modularity measurement value, which is a commonly used method to measure the strength of network community structure. The size of modularity can be used to quantitatively measure the quality of network community division. The closer its value is to 1, the stronger the strength of the community structure divided by the network is, that is, the better the division quality is.

[0118] Figure 2 A flowchart of a target community discovery method integrating content structure rules and time regularities according to a specific embodiment of the present disclosure is shown.

[0119] like Figure 2 As shown, the target community discovery method integrating content structure rules and time regularities includes two stages: a feature mining comparison stage and a target community discovery stage. In the feature mining comparison stage, the posting information of each user on the social platform is collected, including posting time, content, type (such as text, picture, video), etc. On the one hand, content structure rules are extracted, and the specified pattern string is identified from the posting information of each user as the content structure rule of each user, and then the similarity of content structure rules between users is calculated by the Jaccard similarity coefficient. On the other hand, the posting information of each user is aggregated in time periods, for example, the user posting information is aggregated into a list according to hours, representing the user posting time regularity of the user in the corresponding hour, and then the similarity of time regularity between users is measured by the Pearson correlation coefficient. In the target community discovery stage, the similarity of content structure rules between users and the similarity of time regularity between users are combined into a comprehensive weight to construct a weighted undirected graph. Then, the target community is discovered based on the spectral clustering method (community discovery), and the modularity Q (Modularity) is used as the evaluation indicator for the community division results. The quality of the candidate target communities is comprehensively ranked by modularity comparison, and then the target communities with the highest ranking are selected.

[0120] Figure 4 The structure block diagram of the target community discovery device integrating content structure rules and time regularity according to the embodiment of the present disclosure is shown. The device can be implemented as part or all of the electronic device through software, hardware or a combination of both.

[0121] like Figure 4 As shown, the target community discovery device 400 integrating content structure rules and time regularities includes:

[0122] The acquisition module 410 is configured to acquire the user's posting information and posting times;

[0123] A first calculation module 420 is configured to identify user content structure rules from the posting information, and calculate the similarity of content structure rules between users through the Jaccard similarity coefficient;

[0124] A second calculation module 430 is configured to construct a user posting time regularity matrix based on the posting times, and calculate the time regularity similarity between users through the Pearson correlation coefficient;

[0125] A construction module 440 is configured to establish a network undirected weighted graph based on the content structure rule similarity between users and the time regularity similarity between users;

[0126] The community discovery module 450 is configured to use a spectral clustering method to perform community discovery on the network undirected weighted graph to obtain a community division result.

[0127] The target community discovery device that integrates content structure rules and time patterns provided by the present disclosure can reveal the potential connections and posting habits between users by calculating the similarity of content structure rules and time patterns between users. This helps to discover hidden communities and influence networks in social platforms, accurately locate target communities based on user preferences, provide researchers with valuable insights, and provide strong support for research and application in the field of community discovery.

[0128] In an implementation of the present disclosure, the part of the first calculation module 420 that identifies the user content structure rule from the posting information is configured as follows:

[0129] Constructing a pattern set for content structure rule recognition; the pattern set includes preset pattern characters;

[0130] Recognizing preset pattern characters in the posted information;

[0131] All recognized preset pattern characters are used as user content structure rules.

[0132] In an implementation of the present disclosure, the part of the second calculation module 430 that constructs the user posting time regularity matrix based on the posting times is configured as follows:

[0133] Count the number of posts in a preset time period;

[0134] Based on the preset time period, the number of posts in the preset time period is used to construct a user posting time regularity matrix.

[0135] In an implementation of the present disclosure, the part of the second calculation module 430 that constructs the user posting time regularity matrix based on the posting times is further configured as follows:

[0136] If there is interaction between users, a bias is added to the corresponding element in the user posting time regularity matrix.

[0137] In one implementation of the present disclosure, the device further includes:

[0138] The processing module is configured to normalize the user posting time regularity matrix.

[0139] In one implementation of the present disclosure, the building module 440 includes:

[0140] A construction unit is configured to represent a social network as an undirected unweighted graph; the undirected unweighted graph includes user nodes and edges connecting the user nodes;

[0141] A computing unit configured to calculate the weight WJR of the edge in the undirected unweighted graph ij , the formula is as follows:

[0142] Wl ij =αJ(v i ,v j )+(1-α)R(v i ,v j );

[0143] Establishing unit, configured to be based on the weight WJR ij and undirected unweighted graphs to build network undirected weighted graphs;

[0144] Among them, α is an adjustable parameter, 0<α<1, J(v i ,v j ) represents user v i With user v j The similarity of content structure rules between users, R(v i ,v j ) represents user v i With user v j The temporal regularity similarity between users.

[0145] In one implementation of the present disclosure, it further includes:

[0146] The third calculation module is configured to calculate the modularity of the community based on the community division result.

[0147] The present disclosure also discloses an electronic device, Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0148] like Figure 5As shown, the electronic device includes a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to an embodiment of the present disclosure.

[0149] The target community discovery method integrating content structure rules and time regularities includes:

[0150] Get the user's posting information and posting times;

[0151] Identifying user content structure rules from the posting information, and calculating the similarity of content structure rules between users by using the Jaccard similarity coefficient;

[0152] Based on the number of posts, a user posting time regularity matrix is ​​constructed, and the time regularity similarity between users is calculated by the Pearson correlation coefficient;

[0153] Establishing a network undirected weighted graph based on the similarity of content structure rules between users and the similarity of time rules between users;

[0154] The spectral clustering method is used to perform community discovery on the undirected weighted graph of the network to obtain a community division result.

[0155] In one implementation of the present disclosure, the identifying the user content structure rule from the posting information includes:

[0156] Constructing a pattern set for content structure rule recognition; the pattern set includes preset pattern characters;

[0157] Recognizing preset pattern characters in the posted information;

[0158] All recognized preset pattern characters are used as user content structure rules.

[0159] In an implementation of the present disclosure, constructing a user posting time regularity matrix based on the posting times includes:

[0160] Count the number of posts in a preset time period;

[0161] Based on the preset time period, the number of posts in the preset time period is used to construct a user posting time regularity matrix.

[0162] In an implementation of the present disclosure, constructing a user posting time regularity matrix based on the posting times further includes:

[0163] If there is interaction between users, a bias is added to the corresponding element in the user posting time regularity matrix.

[0164] In an implementation of the present disclosure, before calculating the temporal regularity similarity between users by using the Pearson correlation coefficient, the method further includes:

[0165] The user posting time regularity matrix is ​​normalized.

[0166] In an implementation of the present disclosure, the establishing of a network undirected weighted graph based on the content structure rule similarity between users and the time regularity similarity between users includes:

[0167] The social network is represented as an undirected unweighted graph; the undirected unweighted graph includes user nodes and edges connecting the user nodes;

[0168] Calculate the weight WJR of the edge in the undirected unweighted graph ij , the formula is as follows:

[0169] Wl ij =αJ(v i ,v j )+(1-α)R(v i ,v j );

[0170] Based on the weights WJR ij and undirected unweighted graphs to build network undirected weighted graphs;

[0171] Among them, α is an adjustable parameter, 0<α<1, J(v i ,v j ) represents user v i With user v j The similarity of content structure rules between users, R(v i ,v j ) represents user v i With user v j The temporal regularity similarity between users.

[0172] In one implementation of the present disclosure, it further includes:

[0173] The modularity of the community is calculated based on the community division result.

[0174] Figure 6 A schematic diagram showing the structure of a computer system suitable for implementing the method according to an embodiment of the present disclosure is shown.

[0175] like Figure 6As shown, the computer system includes a processing unit, which can perform the various methods in the above-mentioned embodiments according to the program stored in the read-only memory (ROM) or the program loaded from the storage part into the random access memory (RAM). In the RAM, various programs and data required for the operation of the computer system are also stored. The processing unit, ROM and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.

[0176] The following components are connected to the I / O interface: an input part including a keyboard, a mouse, etc.; an output part including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage part including a hard disk, etc.; and a communication part including a network interface card such as a LAN card, a modem, etc. The communication part performs a communication process via a network such as the Internet. The drive is also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that the computer program read therefrom is installed into the storage part as needed. Among them, the processing unit can be implemented as a processing unit such as a CPU, a GPU, a TPU, an FPGA, an NPU, etc.

[0177] In particular, according to an embodiment of the present disclosure, the method described above can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly contained on a machine-readable medium, and the computer program includes a program code for executing the above method. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium.

[0178] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0179] The units or modules involved in the embodiments described in the present disclosure may be implemented by software or programmable hardware. The units or modules described may also be set in a processor, and the names of these units or modules do not constitute limitations on the units or modules themselves in some cases.

[0180] As another aspect, the present disclosure further provides a computer-readable storage medium, which may be a computer-readable storage medium included in the electronic device or computer system in the above embodiment; or a computer-readable storage medium that exists independently and is not assembled into a device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the method described in the present disclosure.

[0181] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other.

Claims

1. A target community discovery method integrating content structure rules and time rules, characterized in that: include: Get the user's posting information and posting times; Identifying user content structure rules from the posting information, and calculating the similarity of content structure rules between users by using the Jaccard similarity coefficient; Based on the number of posts, a user posting time regularity matrix is ​​constructed, and the time regularity similarity between users is calculated by the Pearson correlation coefficient; Establishing a network undirected weighted graph based on the similarity of content structure rules between users and the similarity of time rules between users; The spectral clustering method is used to perform community discovery on the undirected weighted graph of the network to obtain a community division result.

2. The target community discovery method integrating content structure rules and time regularities according to claim 1 is characterized in that: The identifying user content structure rule from the posting information includes: Constructing a pattern set for content structure rule recognition; the pattern set includes preset pattern characters; Recognizing preset pattern characters in the posted information; All recognized preset pattern characters are used as user content structure rules.

3. The target community discovery method integrating content structure rules and time regularities according to claim 1 is characterized in that: The constructing of a user posting time regularity matrix based on the posting times includes: Count the number of posts in a preset time period; Based on the preset time period, the number of posts in the preset time period is used to construct a user posting time regularity matrix.

4. The method for discovering a target community by integrating content structure rules and time regularities according to claim 3, characterized in that: The constructing of a user posting time regularity matrix based on the posting times also includes: If there is interaction between users, a bias is added to the corresponding element in the user posting time regularity matrix.

5. The target community discovery method integrating content structure rules and time regularities according to claim 3 or 4, characterized in that: Before calculating the temporal regularity similarity between users by using the Pearson correlation coefficient, the method further includes: The user posting time regularity matrix is ​​normalized.

6. The method for discovering a target community by integrating content structure rules and time regularities according to claim 1, characterized in that: The establishing of a network undirected weighted graph based on the similarity of content structure rules between users and the similarity of time rules between users comprises: The social network is represented as an undirected unweighted graph; the undirected unweighted graph includes user nodes and edges connecting the user nodes; Calculate the weight WJR of the edge in the undirected unweighted graph ij , the formula is as follows: WJR ij =αJ(v i ,v j )+(1-α)R(v i ,v j ); Based on the weights WJR ij and undirected unweighted graphs to build network undirected weighted graphs; Among them, α is an adjustable parameter, 0<α<1, J(v i ,v j ) represents user v i With user v j The similarity of content structure rules between users, R(v i ,v j ) represents user v i With user v j The temporal regularity similarity between users.

7. The method for discovering a target community by integrating content structure rules and time regularities according to claim 1, characterized in that: Also includes: The modularity of the community is calculated based on the community division result.

8. A target community discovery device integrating content structure rules and time rules, characterized in that: include: An acquisition module is configured to acquire the user's posting information and posting times; A first calculation module is configured to identify user content structure rules from the posting information and calculate the similarity of content structure rules between users through the Jaccard similarity coefficient; A second calculation module is configured to construct a user posting time regularity matrix based on the posting times, and calculate the time regularity similarity between users through the Pearson correlation coefficient; A construction module configured to establish a network undirected weighted graph based on the similarity of content structure rules between users and the similarity of time rules between users; The community discovery module is configured to use a spectral clustering method to perform community discovery on the undirected weighted graph of the network to obtain a community division result.

9. An electronic device, characterized in that: The method comprises a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Interest and network structure double-cohesion social network community discovering method

    CN104268271A

  • A microblog social circle mining method and system based on an artificial immune network

    CN109597924A

  • Multi-target complex network community discovery method based on spectral clustering

    CN109859065A

  • Scholar name disambiguation method and device, storage medium and terminal

    CN111581949A

  • User preference recommendation method, device, electronic equipment and storage medium

    CN113610608A