Scale-limited attribute graph community search method
By combining vertex attribute filtering and k-truss model structure on the attribute graph, the problem of failure to effectively consider vertex attributes and scale limitations in the existing technology is solved, and a fast and efficient community search effect is achieved.
Patent Information
- Application Number
- CN202510172422.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-13
AI Technical Summary
When the prior art conducts community search on the attribute graph, it fails to effectively consider the vertex attribute information and scale limitations, making it difficult to meet the complex query needs.
A scale-limited attribute graph community search method is proposed. By querying vertices and their attributes for filtering, combining the k-truss model structure, enumerating different attribute combinations, and quickly finding the result subgraph that conforms to scale limitations and maximum k value.
It achieves rapid execution speed and efficiency on real data graphs, and is stable especially when the data set is large in size, which is 1-3 orders of magnitude faster than traditional online search methods.
Smart Images

Figure CN119988756A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of graph data query, and relates to a method for querying a dense subgraph containing a given query vertex on a property graph, and specifically to a scale-limited property graph community search method. Background Art
[0002] In the real world, graphs are widely used in many fields, such as social networks, collaboration networks, and protein interaction networks, to represent complex relationships between objects. Unlike abstract graph models, graphs in the real world usually contain rich vertex attribute information, which makes the vertices in the network closely related to the attributes. For example, in social networks, each user has his or her own hobbies, such as photography, dancing, and playing the piano; in collaboration networks, authors have their own professional fields of expertise, such as machine learning, data mining, and natural language processing; and in protein interaction networks, proteins usually also have some characteristics, such as molecular functions, biological processes, and cellular components. Combining the vertices in the network with the attribute information can better characterize the relationship between entities, thereby discovering more meaningful community structures in community analysis.
[0003] In the field of community analysis, existing research can be divided into two categories: community detection and community search. The goal of community detection is to identify all communities in a graph, rather than targeting specific query requests (such as user-specified query vertices). In contrast, community search is a variant of community detection, which aims to find communities containing given query vertices and has been widely studied and applied in personalized community analysis. Community search is usually based on one or more model structures, among which commonly used models include k-core, k-truss, and k-clique. Specifically, k-core requires that each vertex has at least k neighbors, k-truss requires that any edge participates in the formation of at least k-2 triangles, and k-clique requires that k vertices are connected to each other to form a complete graph structure. These model structures provide a theoretical basis for the implementation of community search and are widely used in different community mining tasks.
[0004] Although existing community mining research on attribute graphs has made important progress in attribute calculation and model construction, it often ignores the problem of limited size of returned subgraphs. In practical applications, in addition to satisfying the relationship structure within the community, it is also necessary to consider the limitations of available resources. For example, suppose user q plans to organize a group trip and is deciding who to invite. When making a decision, the following factors are crucial. First, in order to ensure a friendly and pleasant atmosphere during the trip, user q hopes that all participants have more common hobbies. Second, the group size must not be less than 4 people to enjoy the group discount price of air tickets. Finally, due to the limitation of accommodation conditions, the maximum number of participants is 8. Therefore, user q can propose a query requiring the number of participants to be between 4 and 8. In the study of community search with size restrictions, existing work mainly explores the model structure in ordinary graphs or edge-weighted graphs, and adopts a variety of pruning strategies to effectively solve this problem. However, these methods do not consider the attribute information of vertices, making it difficult to meet the complex query requirements in practical scenarios. Summary of the invention
[0005] The primary purpose of the present invention is to propose a scale-limited property graph community search method, which aims to return a subgraph H by searching on a given property graph G, a query vertex q, a vertex lower limit l and a vertex upper limit h, and H must meet the following conditions: 1) the number of vertices in H is greater than or equal to l and less than or equal to h; 2) H is a subgraph that shares the largest number of attributes with the query vertex q; 3) H is a k-truss community that contains the query vertex q and has the largest k value.
[0006] A method for searching a community in a scale-constrained property graph comprises the following steps:
[0007] Step S1: Set a given query condition, which includes the attribute graph G, the query vertex q, the vertex lower limit l, and the vertex upper limit h;
[0008] Step S2: Start traversing from the query vertex q to obtain the vertex set corresponding to each attribute of vertex q and the subgraph formed by the set;
[0009] Step S3: enumerate all attributes of the query vertex q from the graph, obtain different types of attribute combinations, and sort them in descending order of the number of attributes;
[0010] Step S4: traverse all attribute combinations A in the result obtained in step S3, and for each combination, form a subgraph with multiple attributes;
[0011] Step S5: In all subgraphs formed by the vertex intersection, find a result subgraph that meets the scale limit and has the largest k value, and return the result.
[0012] The present invention proposes a scale-limited community search method on an attribute graph. When evaluating attribute cohesion, attention is paid to the number of common attributes of all vertices in the subgraph. In addition, when considering structural cohesion, the k-truss model structure is adopted because the model can identify communities with strong cohesion and is an enhanced version of the k-core model. The solution to this problem is: first, a preliminary screening is performed based on the attributes of the query vertex and its neighbors to narrow the search scope. Subsequently, nodes that are closely related to the query vertex attributes are included in the search scope. Next, the characteristics of the k-truss structure are used for expansion.
[0013] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0014] 1. The core of the present invention is that the execution speed of this method on real data graphs is faster than that of traditional online search methods, the performance effect is good, and it is more stable and efficient when the data set is large.
[0015] 2. The advantage of the present invention is that by obtaining vertex sets of single attributes and combining them to obtain vertex sets of attribute combinations, a result subgraph that meets the requirements can be quickly found.
[0016] 3. The method of the present invention has high execution efficiency in scalability experiments. As the number of graph vertices increases, the execution results of the method tend to be stable, which is 1-3 orders of magnitude faster than traditional online search methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is the property diagram of the present invention.
[0018] Figure 2 It is the query result subgraph of the present invention under given conditions.
[0019] Figure 3 for Figure 1 The subgraph that contains only attribute b.
[0020] Figure 4 for Figure 1 The subgraph that contains only the attribute c. Specific implementation plan
[0021] The accompanying drawings are only used for illustrative purposes and should not be construed as limiting the present patent.
[0022] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the drawings.
[0023] A method for searching a community in a limited-scale attribute graph, characterized by comprising the following steps:
[0024] Step S1: Set a given query condition, which includes the attribute graph G, the query vertex q, the vertex lower limit l, and the vertex upper limit h.
[0025] Step S2: Start traversing from the query vertex q to obtain the vertex set corresponding to each attribute of vertex q and the subgraph formed by the set.
[0026] The specific process is as follows:
[0027] Step S201: select an attribute a in vertex q, start from query vertex q, and traverse the graph;
[0028] Step S202: during traversal, if the attributes of the currently visited vertex include attribute a, the vertex is recorded;
[0029] Step S203: extract all recorded vertices from the graph to form a subgraph of the graph;
[0030] Step S204: Repeat the above process to obtain different subgraphs for each attribute.
[0031] Step S3: Enumerate all attributes of the query vertex q from the graph, obtain different types of attribute combinations, and sort them in descending order of the number of attributes.
[0032] Step S4: traverse all attribute combinations A in the result obtained in step S3, and for each combination, form a subgraph with multiple attributes.
[0033] The specific process is as follows:
[0034] Step S401: From these combinations, traverse from the combination with the most common attributes;
[0035] Step S402: Filter out the vertices with the greatest correlation with the query vertex attributes through attribute combinations, and find the intersection of the single attribute vertex sets that satisfy the attribute combination.
[0036] Step S5: In all subgraphs formed by the vertex intersection, find a result subgraph that meets the scale limit and has the largest k value, and return the result.
[0037] The specific process is as follows:
[0038] Step S501: perform truss decomposition on the subgraph formed by all vertex intersections H to obtain the decomposed subgraph and the corresponding k value;
[0039] Step S502: When the number of vertices of the decomposed subgraph is greater than or equal to the lower limit l and less than or equal to the upper limit h, that is, the number of vertices meets the scale limit and the subgraph k value is the largest, the subgraph is returned.
[0040] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0041] The implementation examples are as follows:
[0042] This implementation example provides a property subgraph method for performing a scale-limited property graph community search, where the property graph is as follows: Figure 1 As shown, the resulting subgraph is Figure 2 As shown, the query search process includes the following steps:
[0043] Step S1: Set a given query condition, where the attribute graph G is as follows Figure 1 As shown, the query vertex is vertex h in the graph, and the scale limit range is [l,h] = [4,8];
[0044] Step S2: Start traversing from the query vertex q to obtain the vertex set corresponding to each attribute, Figure 1 It can be seen that the query vertex q has three attributes, namely {a}, {b}, and {c}. Since attribute {a} only exists in the query vertex q, there is no need to consider the subsequent process, and only attributes {b} and {c} need to be considered.
[0045] Step S201: select one of the attributes, take attribute {b} as an example, start from vertex q, and traverse the graph;
[0046] Step S202: Record all vertices in the graph that contain the attribute {b}. It is known that the set of vertices in the graph that contain the attribute {b} is {q,v 1 ,v 2 ,v 4 ,v 5 ,v 8};
[0047] Step S203: Extract the vertex set in S202 and form a subgraph. The subgraph formed by the vertex set is as follows: Figure 3 As shown;
[0048] Step S204: Repeat the above process and perform the same operation on the attribute {c}. The vertex set corresponding to the attribute {c} is {q,v 1 ,v 2 ,v 3 ,v 4 ,v 5 ,v 7}, the corresponding subgraph is as follows Figure 4 As shown;
[0049] Step S3: Enumerate all attributes of the query vertex q to obtain different types of attribute combinations A, where the attribute combinations generated are {b,c}, {b}, {c} and And sort all possible different attribute combinations enumerated in descending order of the number of attributes, and the sorting results are {b,c}, {b}, {c}. Among these attribute combinations, the one with the most common attributes is {b,c};
[0050] Step S4: traverse all attribute combinations A in the result obtained in step S3, and for each combination, form a subgraph with multiple attributes.
[0051] Step S401: start traversing from the combination {b, c} with the most common attributes;
[0052] Step S402: Intersect the vertex sets in attribute {b} and attribute {c} to obtain the intersection, and form the vertex set {q,v} of the attribute combination {b,c}. 1 ,v 2 ,v 4 ,v 5}, and the subgraph formed by this attribute combination is as follows Figure 1 The same operation is performed for the other combinations, which will not be described here.
[0053] Step S5: In the subgraphs composed of different attribute combinations A, find the subgraph that meets the scale restriction and has the largest k value, and return the result.
[0054] Step S501: For each attribute combination {b,c}, {b}, {c} and The subgraphs formed are truss decomposed. After truss decomposition of the subgraphs formed by the above attribute combinations, the obtained subgraphs k are 5, 3, 3 respectively. Figure 2 , Figure 3 , Figure 4 As shown;
[0055] Step S502: Since the number of vertices of the above subgraphs meets the given size limit, the subgraph corresponding to the attribute combination with the largest k and the largest number of attributes is returned, that is, the subgraph corresponding to the attribute combination {b, c} is returned.
[0056] Selection of datasets: The experimental part uses 8 real-world datasets, including krogan, facebook, deezer, Email-EU, 144, Amazon, DBLP and YouTube. Among them, krogan is a PPI network in the BioGRID database; facebook contains an induced subgraph consisting of 10 given users and their neighbors; deezer represents a user social network, where vertices represent deezer users from various European countries and edges represent their mutual attention relationships; email-EU comes from EMAIL NETWORK; 144 is selected from DIMACS10 set; Amazon is a dataset collected by crawling the Amazon website, where if product i is often purchased together with product j, the graph contains an undirected edge from i to j. DBLP represents a co-author network, in which two authors are connected if they publish at least one paper together. Youtube represents a video sharing website containing a social network. Except for krogan and facebook, the vertices in other datasets have no attributes. Therefore, for the other six datasets, 3 attributes were randomly selected and each attribute was randomly assigned to 80% of the vertices in the graph.
[0057] 1. The experiment selected 8 graph datasets for scale-constrained attribute graph community search. When the scale constraint was small, the online search method performed well on all eight datasets. This is because when the breadth-first search was first performed, the number of vertices required to be enumerated was relatively small, allowing the online search method to quickly find and return the result subgraph that met the conditions. However, as the scale constraint increased, the number of layers required to be enumerated by the online search method also increased, resulting in a gradual increase in the running time. Therefore, in the case of larger scale, the running time of the online search method exceeded that of the attribute subgraph method. In contrast, the attribute subgraph method performed more stably and efficiently in the case of larger scale.
[0058] 2. In order to evaluate the scalability of the algorithm, 20%, 40%, 60%, 80% and 100% of the vertices were randomly selected in each data set for experiments. As the number of vertices increases, the running time of the method of the present invention shows an upward trend. However, the method of the present invention has high execution efficiency. As the number of graph vertices increases, the execution results of the method tend to be stable, which is 1-3 orders of magnitude faster than the traditional online search method.
[0059] In view of the problem that the existing community search methods on attribute graphs do not consider the community size limit, this paper proposes the problem of scale-limited attribute graph community search and designs a corresponding solution. The method obtains the vertex set of a single attribute and enumerates different attribute combinations at the same time. According to different attribute combinations, the intersection of different vertex sets is solved to quickly obtain the result subgraph. Experimental results show that when the scale limit is large, the attribute subgraph method takes less time than the traditional online search method, and the method is more stable.
Claims
1. A scale-constrained attribute graph community search method, characterized in that The following steps are involved: Step S1: Set a given query condition, which includes the attribute graph G, the query vertex q, the vertex lower limit l, and the vertex upper limit h; Step S2: Start traversing from the query vertex q to obtain the vertex set corresponding to each attribute of vertex q and the subgraph formed by the set; Step S3: enumerate all attributes of the query vertex q from the graph, obtain different types of attribute combinations, and sort them in descending order of the number of attributes; Step S4: traverse all attribute combinations A in the result obtained in step S3, and for each combination, form a subgraph with multiple attributes; Step S5: In all subgraphs formed by the vertex intersection, find the truss result subgraph that meets the scale limit and has the largest k value, and return the result.
2. The scale-constrained attribute graph community search method according to claim 1, characterized in that The specific process of the above step S2 is as follows: Step S201: select an attribute a in vertex q, start from query vertex q, and traverse the graph; Step S202: during traversal, if the attributes of the currently visited vertex include attribute a, the vertex is recorded; Step S203: extract all recorded vertices from the graph to form a subgraph of the graph; Step S204: Repeat the above process to obtain different subgraphs for each attribute.
3. The scale-constrained attribute graph community search method according to claim 2, characterized in that The specific process of the above step S4 is as follows: Step S401: From these combinations, traverse from the combination with the most common attributes; Step S402: Filter out the vertices with the greatest correlation with the query vertex attributes through attribute combinations, and find the intersection of the single attribute vertex sets that satisfy the attribute combination.
4. The scale-constrained attribute graph community search method according to claim 3, characterized in that The specific process of the above step S5 is as follows: Step S501: perform truss decomposition on the subgraph formed by all vertex intersections H to obtain the decomposed subgraph and the corresponding k value; Step S502: When the number of vertices of the decomposed subgraph is greater than or equal to the lower limit l and less than or equal to the upper limit h, that is, the number of vertices meets the scale limit and the subgraph k value is the largest, the subgraph is returned.