Content-driven community search method and system on large-scale heterogeneous network, medium and program product

By building k-quad indexes and recursively generating k-quad communities in the content push network, the problem that existing community models are difficult to search for content-driven communities in the content push network is solved, and efficient and applicable content-driven community search is achieved.

CN120107005APending Publication Date: 2025-06-06HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510102358.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing community model is difficult to efficiently search for content-driven communities in the content push network, mainly because they rely on dense direct relationships, while there is less direct relationship between users and content in the content push network, and there is a huge difference in user browsing capabilities and content dissemination speed.

Method used

A content-driven community search method based on k-quad index is designed. By constructing a tree-shaped index structure, quadrilateral support for edges is calculated, and k-quad community is recursively generated to achieve efficient search of content-driven community.

Benefits of technology

It realizes an efficient search for content-driven community in large-scale heterogeneous networks, can handle indirect interactions, and has a search algorithm with sublinear time complexity, and is suitable for content push networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107005A_ABST
    Figure CN120107005A_ABST
Patent Text Reader

Abstract

The invention provides a content-driven community search method and system on a large-scale heterogeneous network, a medium and a program product. In order to solve the problem that a traditional community model is limited by direct relation dependence when processing a community search task on a content push network, the invention provides a content-driven community search model k-quad. The method comprises the steps of description of an indirect interaction mode, construction of a nested index and design of a content-driven community search algorithm. Wherein indirect interaction of a user through content comments is described by utilizing the shape of a quadrangle, a nested index is constructed on k-quads, and a search algorithm with sub-linear time complexity is designed. Validity and high efficiency of the model in the content push network are verified through a large number of experiments performed on the two common data sets and the two captured data sets. The method shows excellent performance in the aspect of processing the characteristics of the content push network, and the algorithm has efficient computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of community search on heterogeneous networks, and in particular relates to a content-driven community search method, system, medium and program product on a large-scale heterogeneous network. Background Art

[0002] Nowadays, various social networks emerge in an endless stream and play an important role in people's daily lives, such as Twitter, Bilibili, and Tiktok. Researchers always model social networks as graphs and mine valuable information through different techniques. Among them, community search is a widely studied technique that is widely used in many applications such as recommendation and advertising. Its goal is to find a subgraph that is densely connected and contains a given query vertex.

[0003] Community models are often used to impose structural constraints on communities to achieve effective community search. Take two popular community models as examples: k-truss restricts each edge to be contained in at least k-2 triangles, and k-core requires each vertex to have at least k neighbors. In this way, stable triangle relationships or friendship levels can be used to ensure the cohesion of the community. They both simulate the direct interaction mode between users and provide good performance for traditional social networks.

[0004] Unlike traditional social networks, more and more content-pushing social networks are more inclined to content dissemination rather than user socialization. For example, users watch videos on the Bilibili platform and share their opinions through video comments, but rarely make friends with others. Users who like specific content will gradually gather together because of the content they are interested in, rather than just forming a user community. It promotes the emergence of content-driven communities in which users interact with each other indirectly by using content as an intermediary. However, it is difficult to search for these content-driven communities due to two main characteristics of content-pushing networks. The first characteristic is that there are few direct relationships between content-pushing networks and users, and the second characteristic is that there is a huge gap between user degree and content degree due to the huge difference between users' browsing ability and content dissemination speed. Since existing community models such as k-core and k-truss have strict constraints on dense direct interaction relationships and treat vertices equally, they are not suitable for searching content-driven communities on specific content-pushing networks.

[0005] There are two typical strategies to enable the classic community model to handle indirect interaction situations. From a topological perspective, the first strategy aims to build a homogeneous graph by establishing direct relationships between users. From a semantic perspective, the second strategy adopts meta-paths to find different community models on heterogeneous graphs. Specifically, the first strategy establishes boundaries between users who browse and comment on the same content, and then searches for user communities on the constructed homogeneous graph. However, due to the large amount of content, the number of relationships between users has increased dramatically, which cannot meet the real-time search needs of users. From a semantic perspective, the second strategy uses meta-paths as edges to search for community models, and meta-paths are built on heterogeneous graphs. However, the space of all meta-paths of arbitrary length is too large to be enumerated, which also leads to an explosion of connection relationships in the graph. In summary, neither the classic community model nor its variants can handle content-driven community search tasks well because they rely heavily on dense direct relationships.

[0006] To solve this problem, it is necessary to design a content-driven community search method to better describe the indirect interaction patterns between users and express content-driven communities. Summary of the invention

[0007] The object of the present invention is to provide a content-driven community search method, system, medium and program product on a large-scale heterogeneous network to achieve efficient community search on a content push network.

[0008] The purpose of the present invention is achieved through the following technical solutions:

[0009] A content-driven community search method on a large-scale heterogeneous network, the specific steps are as follows:

[0010] Step 1: Data collection;

[0011] Obtain raw network data from DBLP, LMDB, Weibo, and Bilibili, and process and cleanse the data to ensure data integrity and accuracy;

[0012] Step 2: Index building;

[0013] First, a tree structure Tree is constructed with the original graph G as the root, and the nodes represent k-quad communities. Then, the k-quad communities are decomposed, and the tree hierarchy is built using the nested relationship of the decomposition to determine the leaf nodes. The quadrilateral support of each edge is calculated. Finally, the k-quad is generated by recursively decomposing the (k-1)-quad, and the edges that do not meet the constraints are deleted.

[0014] Step 3: Index maintenance;

[0015] When an edge e=(u, c) is added or deleted from the graph G, the index KQindex is updated according to different situations;

[0016] Step 4: Community search;

[0017] The QuadCS community search algorithm is based on the k-quad index. It quickly finds the target community that meets the k-level constraints based on the query vertex set and parameter k. The algorithm first finds the mapping community of the query vertex through the k-quad index, then determines the k-level ancestor communities of these communities through interpolation search in the index, and obtains the final community by finding the intersection. If the ancestor communities are consistent, the community is returned, otherwise an empty set is returned.

[0018] Further step 1 is specifically as follows:

[0019] Step 1.1: Collect raw network data from DBLP, IMDB, Weibo, and Bilibili, including post content and user information. All nodes are classified and stored according to their types, ensuring that each node type is uniquely identified as "user" or "content";

[0020] Step 1.2.: Based on the data obtained in step 1.1, construct a content push network G = (V, E), where the node set V includes user nodes and content nodes, and the edge set E represents the association between users and content; use the node mapping function φ: V→A to map each node to the corresponding type set A, thereby forming a bipartite graph;

[0021] Step 1.3: Generate adjacency matrix A for the constructed content push network G |U|×|C| , where U is the user node set and C is the content node set; if the edge (v i , v j )∈E, then A ij =1, otherwise A ij =0.

[0022] Furthermore, step 2 is specifically as follows:

[0023] Step 2.1: Build a tree structure; Build a tree structure Tree, where the root node is the graph G; the nodes of the tree represent k-quad communities, and the root node represents the original graph G;

[0024] Step 2.2: Decompose k-quad communities; decompose the graph G into several 1-quad communities, which constitute the second level of the tree; for each 1-quad community, recursively decompose it into (k-1)-quad communities, and so on, until the required k value is reached;

[0025] Step 2.3: Nested relationship formation: According to the nested relationship formed in the decomposition process, different levels of the tree structure are established to reflect the inclusion relationship of the community between different levels;

[0026] Step 2.4: Determine the leaf nodes; the leaf nodes of the tree are communities that cannot be further decomposed. The list of these leaf nodes is stored in hierarchical order and is called Clist(C);

[0027] Step 2.5: Quadrilateral support calculation; Quadrilateral support calculation refers to the number of quadrilaterals that contain a certain side; to calculate the quadrilateral support of each side, it is necessary to determine all quadrilaterals that contain the side;

[0028] Step 2.6: k-quad generation; k-quad generation is achieved by recursively decomposing the (k-1)-quad; first generate a 1-quad, and then recursively decompose each (k-1)-quad into a k-quad until no larger k-quad can be generated; during the decomposition process, delete the edges that do not meet the k-quad constraints, that is, those that do not participate in at least k quadrilaterals.

[0029] Furthermore, step 3 is specifically as follows:

[0030] Step 3.1: Adding edges; when an edge e = (u, c) is added to the graph G, the index KQindex is updated according to different situations; first, the support of the newly added edge is calculated. If the support is 0, the index remains unchanged; if the support is not 0, different treatments are performed according to the 1-quads where u and c are located; for edges in the same 1-quad, the support of the quadrilateral where the newly added edge is located is updated, and the index KQindex is updated through the k-truss generation algorithm; for edges in different 1-quads, the two 1-quads are merged, and the support of other affected edges is updated, and finally the index KQindex is updated through the k-truss generation algorithm;

[0031] Step 3.2: Deletion of edges; when an edge e = (u, c) is deleted from the graph G, the index KQindex also needs to be updated; if the support of the deleted edge is 0, the index remains unchanged; if the support is not 0, the support of other edges in the quadrilateral where the edge is located needs to be updated; the index KQindex is recalculated and updated through the k-truss generation algorithm to avoid the reduction, splitting or disappearance of the 1-quad.

[0032] Furthermore, the k-truss algorithm is used to extract a compact subgraph satisfying the k-truss condition from a graph, and ensure connectivity by removing edges that do not meet the condition; the k-truss generation algorithm: generates a complete subgraph satisfying the k-truss condition by iteratively deleting edges, and constructs a highly connected substructure in the graph.

[0033] Furthermore, step 4 is specifically as follows:

[0034] Step 4.1: Vertex mapping: For each vertex v in the query vertex set, use the k-Quad index to find its corresponding mapping community Cv; this step quickly locates the initial community where the vertex is located through the index structure;

[0035] Step 4.2: Ancestral community search; by mapping community C v Ancestor list (Clist (C v )) to perform interpolation search and determine the ancestral community C′ at level k v ;This step uses the index hierarchy to quickly find communities that meet k-level constraints;

[0036] Step 4.3: Find the intersection; find the intersection of the k-level ancestor communities C of all query vertices. If the intersection is not empty and all communities are the same, then the intersection is the target community, otherwise an empty set is returned.

[0037] A content-driven community search system on a large-scale heterogeneous network, comprising a data collection device, an index building device, an index maintenance device and a community search device;

[0038] The data collection device: The main task is to obtain original network data from DBLP, IMDB, Weibo, and Bilibili, so as to prepare for the subsequent construction of a content-driven community search system;

[0039] The index building device: The main task is to design a tree index structure, calculate the quadrilateral support of the edge and recursively generate k-quad communities, and build a nested k-quad index KQindex to maintain the community structure;

[0040] The index maintenance device: The main task is to update the KQindex in time according to the changes in the graph structure in the social network to avoid the degradation of community search performance and ensure that the index can accurately reflect the current network structure;

[0041] The community search device: The main task is to accurately and efficiently locate a specific community in the graph, which meets the requirements and structural constraints of the input query.

[0042] A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that when the computer program / instruction is executed by a processor, the steps of a content-driven community search method on a large-scale heterogeneous network are implemented.

[0043] A computer program product includes a computer program / instruction, characterized in that: when the computer program / instruction is executed by a processor, the steps of a content-driven community search method on a large-scale heterogeneous network are implemented.

[0044] The beneficial effects of the present invention are:

[0045] Aiming at the problem that traditional community models are limited by direct relationship dependence when dealing with community search tasks on content push networks, the present invention proposes a content-driven community search model k-quad. It includes the description of indirect interaction patterns, the construction of nested indexes, and the design of content-driven community search algorithms. Among them, the shape of quadrilaterals is used to describe the indirect interaction of users through content comments, and nested indexes are constructed on k-quads, and a search algorithm with sublinear time complexity is designed. Through a large number of experiments conducted on two public datasets and two crawled datasets, the effectiveness and efficiency of the model in content push networks are verified. Experimental results show that this method shows excellent performance in processing the characteristics of content push networks, and the algorithm has efficient computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 The classical community model pointed out for the present invention is not suitable for the illustration of content-driven networks;

[0047] Figure 2 is an illustration of the k-quad index structure of the present invention;

[0048] Figure 3 An illustration of calculating quadrilateral support for the present invention;

[0049] Figure 4 An illustration of the k-quad generation of the present invention;

[0050] Figure 5 Three diagrams illustrating the situation of adding an edge to the present invention;

[0051] Figure 6 This is a graph showing the experimental results of the index building time in the present invention;

[0052] Figure 7 This is a graph showing the experimental results of the time cost of community search in the present invention;

[0053] Figure 8 This is a graph showing the experimental results of the community search performance comparison in the present invention;

[0054] Fig. 9 This is an illustration of the present invention using the Bilibili dataset as an example;

[0055] Fig.10 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0056] The present invention is further described below in conjunction with the accompanying drawings.

[0057] according to Fig.10 The present invention provides a content-driven community search method on a large-scale heterogeneous network, and the specific steps are as follows:

[0058] Step 1: Data collection;

[0059] The original network data is obtained from DBLP, IMDB, Weibo, and Bilibili, and these data are processed and cleaned to ensure the integrity and accuracy of the data; the processed data includes attribute information such as the connection relationship between nodes and node degree.

[0060] Step 2: Index building;

[0061] The present invention introduces the construction process of k-quad index. In order to perform efficient community search, the present invention provides a tree index structure - k-quad. Based on the nested relationship of k-quad communities, k-quad communities in the graph can be effectively organized and retrieved. The construction process mainly includes two steps. The first step is to calculate the quadrilateral support calculation of the edge to provide a sufficient basis for further k-quad generation. The second step is to generate nested k-quad communities through recursive decomposition.

[0062] Step 3: Index maintenance;

[0063] Index maintenance reflects the dynamic changes of network data, including the operations of adding nodes and edges and removing nodes and edges, and the corresponding algorithms are designed to maintain the community structure based on nested indexes. When an edge is added to the graph, the quadrilateral support of the edge is first calculated. If the support is 0, the index does not need to be updated; otherwise, the 1-quad involved in the edge needs to be checked, the support of the edges in the newly formed quadrilateral needs to be updated, and the related 1-quads may need to be merged or updated. If an edge is deleted, if its support is 0, there is no need to update the index; if the support is not 0, it is necessary to update the support of other edges in the affected quadrilateral, and reprocess and update the index of the affected 1-quad. These steps ensure that the index remains accurate during the dynamic changes of the graph.

[0064] Step 4: Community search;

[0065] The QuadCS community search algorithm is based on the k-quad index and aims to quickly find the target community that meets the k-level constraints based on the query vertex set and the parameter k. The algorithm first finds the mapping community of the query vertex through the k-quad index, then determines the k-level ancestor communities of these communities through interpolation search in the index, and obtains the final community by finding the intersection. If the ancestor communities are consistent, the community is returned, otherwise an empty set is returned. The algorithm uses the k-quad index to achieve efficient community search with sublinear time complexity 0(l0glogm).

[0066] Figure 1 The classical community model proposed by the present invention is not suitable for the illustration of content-driven networks, where circles represent users and quadrilaterals represent content. Figure 1 The community inside the dotted circle in the traditional social network in (a) can be easily found through k-truss (this community is 4-truss) or k-core (this community is 3-core). Figure 1 In the content-driven network with sparse direct relationships in (b), it is difficult for the classic community model to find a reasonable community. In order to enable the classic community model to handle indirect interactions, there are two typical strategies, such as Figure 1 (c) and (d). The first strategy aims to build a homogeneous graph by establishing direct relationships between users. From a semantic point of view, the second strategy adopts meta-paths to find different community models on heterogeneous graphs. Specifically, the first strategy establishes boundaries between users who browse and comment on the same content, and then searches for user communities on the constructed homogeneous graph. Figure 1 (c) shows Figure 1 (b) is a homogeneous graph constructed by the content-driven network. Obviously, due to the large amount of content, the number of relationships between users has increased dramatically, which cannot meet the real-time search needs of users. The second strategy searches for community models by using meta-paths as edges, which are built on heterogeneous graphs. However, the space of all meta-paths is too large to be enumerated and will also lead to Figure 1 (d) shows the relationship explosion. In summary, Figure 1 It is shown that neither the classic community model nor its variants can handle content-driven community search tasks well because they heavily rely on dense direct relationships.

[0067] Figure 2 This is an illustration of the k-quad index structure of the present invention, specifically:

[0068] Given a graph G = (V, E), the k-quad index structure consists of a tree and a map, that is, KQindex = (Tree, Map), where Tree is the community of nodes in the tree structure, and It is a set of mappings from each vertex to its highest community (i.e., with the maximum value k). The root of the Tree is the graph G. Then, G is decomposed into several 1-quad communities of the second level. Similarly, the present invention recursively decomposes the (k-1)-quad communities to obtain their sub-communities k-quad, and increases k. The decomposition forms a nested relationship between communities at different levels. The leaves of the Tree are communities that cannot be decomposed any further. In particular, for each leaf node, the present invention stores a list of its ancestor nodes in the Tree, sorted by their levels, represented by Clist(C). Figure 2 An example of a k-quad index structure is given.

[0069] like Figure 2 As shown, the graph G contains 23 vertices and is a tree.G. G can be regarded as a 0-quad community. Then, G is decomposed into three 1-quads at the second level, marked in green, blue and orange respectively. Then, the 1-quad is further decomposed into 2-quads. As the k value increases, the decomposition process continues until there are no more k-quads to decompose. For example, the maximum value of the green branch is 2-quad, while the maximum value of the blue branch is 3-quad. Finally, the present invention constructs a mapping from each vertex to its highest community. For example, vertex 20 is mapped to the 1-quad community {20, 21, 22, 23}.

[0070] Figure 3 This is an illustration of the present invention's support for calculating quadrilaterals, specifically:

[0071] To complete the construction of the k-quad index, the first step is to calculate the quadrilateral support of the edge, which will be used as the basis for further generating the k-quad.

[0072] For each edge in a given graph G, the quadrilateral support refers to the number of quadrilaterals that contain the edge. Therefore, in order to calculate the quadrilateral support of each edge, it is necessary to calculate all the quadrilaterals it participates in. This invention explains the principle of calculating quadrilaterals through the adjacency matrix. Figure 3As shown, user vertices are represented by circular nodes and content vertices are represented by square nodes. The goal of the present invention is to calculate the quadrilateral support of edge (1, 5) in the graph, that is, point (1, 5) in the adjacent matrix. Obviously, the quadrilateral with four edges in the graph is actually a rectangle, and there are four coordinate points in the adjacent matrix with a value of 1. For example, the quadrilateral with four edges (1, 5), (2, 5), (2, 7), and (1, 7) marked with rectangular dotted lines corresponds to a rectangle with four points (1, 5), (2, 5), (2, 7), and (1, 7) in the adjacency matrix. In order to calculate the rectangle containing point (1, 5), the present invention fixes content 1 and finds another neighbor of user 5, that is, content 2, and then checks whether content 1 and content 2 have another common user other than 5, that is, user 6 or user 7. Therefore, point (1, 5) participates in two rectangles, which are marked with solid lines of different colors on the matrix respectively.

[0073] Figure 4 The illustration diagram generated for the k-quad of the present invention specifically includes:

[0074] In order to generate k-quads for different values ​​of k, the key idea is that k-quads are recursively extracted from (k-1)-quads with k ≥ 1, and their nested relationship forms a tree. In the process of decomposing (k-1)-quads to generate k-quads, the key is to find and delete the edges that do not satisfy the k-quad constraint, that is, the quadrilaterals that participate in less than k.

[0075] Figure 4 The bipartite graph of gives the generation process of k-truss. The present invention first finds and deletes the edges with support less than 1 in G. The three connected subgraphs built on the left edge are 1-quads, which are divided into left, middle and right sides according to the background color. In order to generate 2-quads, delete the edges with support less than 2 in 1-quads. If the support points of the dotted edge (2, 9) are less than 2, it should also be deleted. The remaining edges form three 2-quads in the 1-quads.

[0076] Figure 5 The following are three diagrams for explaining the situation of adding an edge to the present invention, specifically:

[0077] In view of the addition and deletion of content in social networks, the present invention considers the problem of adding and deleting edges. Take the way of adding edges as an example: Figure 5 The details of updating the index KQindex when adding an edge e = (u, c) to the graph G are explained. First, if qsup(e) = 0, that is, the addition of e does not form a new quadrilateral, then KQindex will not change. For example, adding an edge (4, 5) to Figure 5(a), but it does not consist of any quadrilaterals with support 0. Second, if qsup(e)≠0 and u,c appear in the same 1-quad, the present invention only needs to update the support of the edges in the newly formed quadrilateral, and then recursively decompose the 1-quad and update KQindex. For example, add edge (2, 5) to Figure 5 In the 1-quads in (b), it forms two new quadrilaterals {1, 2, 4, 5} and {2, 3, 4, 5}, resulting in support updates for the other edges in the quadrilaterals. Third, if qsup(e) ≠ 0 and u, c appear in two different 1-quads, the addition of e will result in the combination of the two 1-quads, also with support updates. Figure 5 In (c), edge (2, 5) is added and results in two 1-quads {1, 2, 3, 4} and {5, 6, 7, 8}. Note that the black dashed edge (3, 8) is removed when generating the 1-quads because its support is 0. Since edge (2, 5) is added, the support of (3, 8) becomes 2 and it is not removed in the 1-quad generation.

[0078] The following describes in detail the experimental example scenarios of the present invention, and analyzes the implementation results in combination with the advantages of the present invention.

[0079] The present invention uses four datasets. DBLP and IMDB are public datasets, while Weibo and Bilibili datasets are collected by crawlers written by the authors of the present invention. DBLP is a traditional network, which can be regarded as a content-driven network when the degree difference between two types of vertices is large. The other datasets are typical content-driven networks, whose content has a large degree and drives the formation of communities.

[0080] In WC-index, the present invention sets the edge weight to 1. In CSSH, the present invention uses its default parameters, for example: the maximum length of the meta-path is 4. In order to evaluate the community search performance on each dataset, the present invention randomly generates 200 queries for the query vertex numbers of 1, 2, 3, 4 and 5 respectively. Hardware settings. All experiments were conducted on a server equipped with a 2.20GHz Intel Xeon Gold 5320 processor. The processor has 52 cores (26 cores per socket, 2 sockets) and 78MiB L3 cache. The server has a 1.7TB disk and 755GB DDR4 DRAM. The operating system is 64-bit Ubuntu 22.04.4. The program is compiled using g++11.4.0.

[0081] In order to compare the difference in index building efficiency between the proposed method and the comparative method, the proposed method randomly selects 20%, 40%, 60%, and 80% of the edges from each data set. By changing the data set size from 20% to 100%, the index building time is Figure 6 shown. Figure 6 This is an experimental result diagram of the index building time in the present invention. Figure 6 As shown in Figure 2, in most cases, the index construction time of QuadCS is higher than that of other methods. Taking the 80% IMDB dataset as an example, QuadCS takes 1.39×10 3 s, which is 9.92, 1.71, 5.38 and 3.42 times that of CSRTI-LPA, CSRTI-Louvain, WCindex and CSSH. Since WC-index only calculates neighbors, the index construction time is shorter. Although CSRTI needs to build a nested index on k-truss, truss rarely exist in content-driven networks. Therefore, it only consumes the time of community detection. In terms of maximizing global complexity, CSRTI-Louvain takes more time, while CSRTI-LPA executes faster when the idea is simple. Obviously, the index complexity determines the construction time of QuadCS to build a nested index on k-quads, so it takes more time. By sacrificing the preprocessing time of building an index with valuable information, QuadCS will achieve efficient community search, as demonstrated later.

[0082] It is particularly important to emphasize that CSSH performs very poorly on Weibo and Bilibili datasets, which have larger data volumes and higher degrees. On 20% of the Weibo dataset, if the default maximum meta-path length of 4 is used, CSSH takes 1.32×104s on the Weibo dataset, which is 49.07 times that of QuadS, and occupies more than 100GB of memory. There is even insufficient memory on 40% of Weibo and 20% of Bilibili datasets. Therefore, the present invention reduces its maximum meta-path length on Weibo and Bilibili to 2, marked as CSSH-L2. Because its enumeration space is too large, CSSH-L2 still has the longest establishment time on 100% Weibo and 100% Bilibili.

[0083] Figure 7 This is a graph showing the experimental results of the time cost of community search in the present invention;

[0084] The present invention evaluates the efficiency of community search on different data sets. The time cost of community search is affected by the size of the data set and the size of the query set. The present invention evaluates the efficiency of community search of different methods on different data sets by changing the size of the data set and the size of the query set.

[0085] In terms of query data sets, the present invention randomly selects 20%, 40%, 60%, and 80% of the data from each data set. The query set size is set to 1 by default. By changing the data set size from 20% to 100%, the average community search time results on each data set are as follows: Figure 7 (a) to (d). First, it can be seen that QuadCS performs the fastest in community search, CSRTI-LPA and CSRTI-Louvain are also efficient, while CSSH and WC-index are quite slow. On the 20% Bilibili dataset, CSSH takes 4.55×103s, which is 4.79 times that of WC-index and 7 orders of magnitude slower than QuadCS, CSRTI-LPA and CSRTILouvain. CSSH needs to search the neighbors of the queried vertex layer by layer based on the meta-path. WC-index repeatedly calls the connected component algorithm for each deleted vertex. QuadCS is the most efficient because it has sublinear time complexity.

[0086] The QuadCS community search is as follows:

[0087]

[0088] Second, CSSH ran out of memory when searching for communities on 80% of Weibo and 60% of Bilibili datasets, but the present invention only adopted a maximum meta-path length of 2. This shows that the enumeration space of the meta-path-based method is too large, resulting in an explosive growth of relations. In addition, for many queries, it fails to find the target community, while other methods do not have such problems. Therefore, the meta-path-based method is not suitable for large-scale content-driven networks.

[0089] In terms of changing the query set size, by varying the query set size from 1, 2, 3, 4 to 5 on each 100% dataset, the community search time results are as follows Figure 7As shown in (e) to (h). On the four datasets, QuadCS achieved the shortest community search time under different query set sizes. When processing a query consisting of 3 vertices on the DBLP dataset, WC-index and CSSH took 138.43s and 28.81s respectively, while QuadCSS, CSRTI-LPA and CSRTILouvain took 4.02×10-7s, 5.55×10-7s and 9.63×10-7s respectively. Since QuadCS constructs nested indexes on k-quad communities, it can return the pointer to the target community by interpolating the index with a complexity of O(loglogm). Two experiments have shown that based on the constructed index, QuadCS achieves more efficient community search. The meta-path-based CSSH method not only takes a long search time, but also consumes a lot of memory because its enumeration space is large and difficult to traverse.

[0090] Figure 8 This is a graph showing the experimental results of the community search performance comparison in the present invention;

[0091] In order to prove the effectiveness of QuadCS, the present invention compares the community search performance through four indicators. The first indicator is the quadrilateral ratio, that is, the ratio between the quadrilaterals in the community and other polygons. The higher it is, the closer the indirect interaction between the communities is. The second indicator is the edge ratio, that is, the ratio of the internal degree to the external degree. It is usually used to reflect the degree of cohesion. Then, for the diameter and the jaccard similarity coefficient, the smaller the diameter means the higher the community cohesion, and the larger the jaccard similarity coefficient means the higher the neighborhood similarity. The results are as follows Figure 8 shown.

[0092] From the perspective of quadrilateral ratio, the quadrilateral ratio of QuadS is much higher than Figure 8 (a) The baseline. On the content-driven Weibo dataset, the quadrilateral ratio of QuadS is 980.29, which is 99.45 times, 105.73 times, and 4730.02 times of each variable, respectively. It proves that QuadS can fully mine the indirect interaction relationship of content-driven networks and connect users through the driving effect of content. Due to the absence of triangles, CSRTI degenerates into a community detection method and performs poorly on the crawled Weibo and Bilibili datasets. WC-index performs the worst because it only focuses on direct interactions of calculation degree.

[0093] Then, by combining Figure 8 (b) and Figure 8(d), it can be seen that QuadCS has lower jaccard and higher edge ratio than CSRTI-Louvain. Since CSRTI-Louvain has higher iaccard, the neighbors of the vertices in G are more similar for the CSRTI-Louvain community. Although the global neighbors are similar, the lower edge ratio of CSRTI-Louvain means that the community of CSRTI-Louvain misses a part of the neighboring vertices and is sparsely connected. In contrast, QuadCS can help find more complete content-driven communities. In addition, combined with Figure 8 (b) and Figure 8 (c), it is found that QuadCS has a lower edge ratio and a smaller diameter than CSRTI-LPA. The higher edge ratio of CSRTI-LPA means that the neighbors of the vertex are more likely to be in the community. At the same time, CSRTI-LPA has a larger diameter because CSRTI-LPA sequentially connects scattered star-shaped subgraphs. It ignores the relationship between boundary vertices and does not reflect the driving characteristics of the content. In contrast, QuadCS makes full use of the shape of the quadrilateral to find communities driven by the content that users follow.

[0094] In addition, combined Figure 8 (b) and Figure 8 (c), it is found that QuadCS has a lower edge ratio and a smaller diameter than CSRTI-LPA. The higher edge ratio of CSRTI-LPA means that the neighbors of the vertex are more likely to be in the community. At the same time, CSRTI-LPA has a larger diameter because CSRTI-LPA sequentially connects scattered star-shaped subgraphs. It ignores the relationship between boundary vertices and does not reflect the driving characteristics of the content. In contrast, QuadCS makes full use of the shape of the quadrilateral to find communities driven by the content that users follow.

[0095] Fig. 9 This is an illustration of the present invention using the Bilibili dataset as an example.

[0096] The present invention conducts a case study on the Bilibili dataset to demonstrate the effectiveness of QuadCS. Given a query vertex with user ID 35...38 and k=3, QuadC finds a cohesive community with the theme of buying a Huawei Mate60 mobile phone. When searching for query vertices with user ID 53...08 and k=6, a community with the theme of tank is returned. It can be seen that the content of each community has related topics, providing a bridge for users to communicate and promoting the formation of communities. These cases show that QuadCS can help find content-driven communities and meet users' search needs for interest communities in real scenarios.

[0097] The present invention discloses a content-driven community search system on a large-scale heterogeneous network, comprising a data collection device, an index building device, an index maintenance device, and a community search device;

[0098] The data collection device: The main task is to obtain original network data from DBLP, IMDB, Weibo, and Bilibili, so as to prepare for the subsequent construction of a content-driven community search system;

[0099] The index building device: The main task is to design a tree index structure, calculate the quadrilateral support of the edge and recursively generate k-quad communities, and build a nested k-quad index KQindex to maintain the community structure;

[0100] The index maintenance device: The main task is to update the KQindex in time according to the graph structure changes in the social network (such as the addition or deletion of edges) to avoid the degradation of community search performance and ensure that the index can accurately reflect the current network structure.

[0101] The community search device: The main task is to accurately and efficiently locate a specific community in the graph, which meets the requirements and structural constraints of the input query.

[0102] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A content-driven community search method on a large-scale heterogeneous network, characterized in that: The specific steps are as follows: Step 1: Data collection; Obtain raw network data from DBLP, LMDB, Weibo, and Bilibili, and process and cleanse the data to ensure data integrity and accuracy; Step 2: Index building; First, a tree structure Tree is constructed with the original graph G as the root, and the nodes represent k-quad communities. Then, the k-quad communities are decomposed, and the tree hierarchy is built using the nested relationship of the decomposition to determine the leaf nodes. The quadrilateral support of each edge is calculated. Finally, the k-quad is generated by recursively decomposing the (k-1)-quad, and the edges that do not meet the constraints are deleted. Step 3: Index maintenance; When an edge e = (u, c) is added or deleted from the graph G, the index KQindex is updated according to different situations; Step 4: Community search; The QuadCS community search algorithm is based on the k-quad index. It quickly finds the target community that meets the k-level constraints based on the query vertex set and parameter k. The algorithm first finds the mapping community of the query vertex through the k-quad index, then determines the k-level ancestor communities of these communities through interpolation search in the index, and obtains the final community by finding the intersection. If the ancestor communities are consistent, the community is returned, otherwise an empty set is returned.

2. The content-driven community search method on a large-scale heterogeneous network according to claim 1, characterized in that: Step 1 is as follows: Step 1.1: Collect raw network data from DBLP, IMDB, Weibo, and Bilibili, including post content and user information. All nodes are classified and stored according to their types, ensuring that each node type is uniquely identified as "user" or "content"; Step 1.2.: Based on the data obtained in step 1.1, construct a content push network G = (V, E), where the node set V includes user nodes and content nodes, and the edge set E represents the association between users and content; use the node mapping function φ:V→A to map each node to the corresponding type set A, thereby forming a bipartite graph; Step 1.3: Generate adjacency matrix A for the constructed content push network G |U|×|C| , where U is the user node set and C is the content node set; if the edge (v i ,v j )∈E, then A ij =1, otherwise A ij =0.

3. The content-driven community search method on a large-scale heterogeneous network according to claim 1, characterized in that: Step 2 is as follows: Step 2.1: Build a tree structure; Build a tree structure Tree, where the root node is the graph G; the nodes of the tree represent k-quad communities, and the root node represents the original graph G; Step 2.2: Decompose k-quad communities; decompose the graph G into several 1-quad communities, which constitute the second level of the tree; for each 1-quad community, recursively decompose it into (k-1)-quad communities, and so on, until the required k value is reached; Step 2.3: Nested relationship formation: According to the nested relationship formed in the decomposition process, different levels of the tree structure are established to reflect the inclusion relationship of the community between different levels; Step 2.4: Determine the leaf nodes; the leaf nodes of the tree are communities that cannot be further decomposed. The list of these leaf nodes is stored in hierarchical order and is called Clist(C); Step 2.5: Quadrilateral support calculation; Quadrilateral support calculation refers to the number of quadrilaterals that contain a certain side; to calculate the quadrilateral support of each side, it is necessary to determine all quadrilaterals that contain the side; Step 2.6: k-quad generation; k-quad generation is achieved by recursively decomposing the (k-1)-quad; first generate a 1-quad, and then recursively decompose each (k-1)-quad into a k-quad until no larger k-quad can be generated; during the decomposition process, delete the edges that do not meet the k-quad constraints, that is, those that do not participate in at least k quadrilaterals.

4. The content-driven community search method on a large-scale heterogeneous network according to claim 1, characterized in that: Step 3 is as follows: Step 3.1: Adding edges; when an edge e = (u, c) is added to the graph G, the index KQindex is updated according to different situations; first, the support of the newly added edge is calculated. If the support is 0, the index remains unchanged; if the support is not 0, different treatments are performed according to the 1-quads where u and c are located; for edges in the same 1-quad, the support of the quadrilateral where the newly added edge is located is updated, and the index KQindex is updated through the k-truss generation algorithm; for edges in different 1-quads, the two 1-quads are merged, and the support of other affected edges is updated, and finally the index KQindex is updated through the k-truss generation algorithm; Step 3.2: Deletion of edges; when an edge e = (u, c) is deleted from the graph G, the index KQindex also needs to be updated; if the support of the deleted edge is 0, the index remains unchanged; if the support is not 0, the support of other edges in the quadrilateral where the edge is located needs to be updated; the index KQindex is recalculated and updated through the k-truss generation algorithm to avoid the reduction, splitting or disappearance of the 1-quad.

5. The content-driven community search method on a large-scale heterogeneous network according to claim 4, characterized in that: The k-truss algorithm is used to extract a compact subgraph satisfying the k-truss condition from a graph, and ensure connectivity by removing edges that do not meet the condition; the k-truss generation algorithm generates a complete subgraph satisfying the k-truss condition by iteratively deleting edges, and constructs a highly connected substructure in the graph.

6. The content-driven community search method on a large-scale heterogeneous network according to claim 1, characterized in that: Step 4 is as follows: Step 4.1: Vertex mapping: For each vertex v in the query vertex set, use the k-Quad index to find its corresponding mapping community Cv; This step quickly locates the initial community where the vertex is located through the index structure; Step 4.2: Ancestral community search; by mapping community C v Ancestor list (Clist (C v )) to perform interpolation search and determine the ancestral community C at level k v '; This step uses the index hierarchy to quickly find communities that meet k-level constraints; Step 4.3: Find the intersection; for all query vertices, the k-level ancestral communities C v ′Find the intersection. If the intersection is not empty and all communities are the same, then the intersection is the target community, otherwise an empty set is returned.

7. A content-driven community search system on a large-scale heterogeneous network according to any one of claims 1 to 6, characterized in that: It includes a data collection device, an index building device, an index maintenance device and a community search device; The data collection device: The main task is to obtain original network data from DBLP, IMDB, Weibo, and Bilibili, so as to prepare for the subsequent construction of a content-driven community search system; The index building device: The main task is to design a tree index structure, calculate the quadrilateral support of the edge and recursively generate k-quad communities, and build a nested k-quad index KQindex to maintain the community structure; The index maintenance device: The main task is to update the KQindex in time according to the changes in the graph structure in the social network to avoid the degradation of community search performance and ensure that the index can accurately reflect the current network structure; The community search device: The main task is to accurately and efficiently locate a specific community in the graph, which meets the requirements and structural constraints of the input query.

8. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer program product comprising a computer program / instructions, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.