Method and apparatus for segmenting map data, electronic device and program product
By judging and segmenting the vertices in the graph data, forming virtual points and allocating them to the processing thread of the distributed cluster, the problem of load imbalance in graph data processing is solved and computing efficiency is improved.
Patent Information
- Application Number
- CN202510126192.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-23
AI Technical Summary
When processing large-scale graph data, the existing technology is difficult to achieve load balancing, resulting in some computing nodes taking on too many computing tasks and reducing computing efficiency.
By determining the vertex degree in the graph data, if the degree is greater than the threshold, the vertices are divided into multiple virtual points and these virtual points are evenly allocated to multiple processing threads of the distributed cluster.
Load balancing allocation of graph data is realized, and the overall performance of distributed clusters and the efficiency of graph data processing is improved.
Smart Images

Figure CN120029775A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present specification relate to the field of graph technology, and more specifically, to a method, apparatus, electronic device, and program product for segmenting graph data. Background Art
[0002] Graph data has a wide range of applications because it is easy to represent complex dependencies between different entities. For example, in social network analysis, graph data can be used to represent the relationship between users. In route planning, graph data can be used to represent the connection between road networks and cities. In the field of artificial intelligence, graph data can be used to represent knowledge graphs and applied to recommendation systems. For example, in an e-commerce recommendation system, graph data can be used to establish user nodes, product nodes, and edges between nodes. The edges between user nodes and product nodes are used to represent the interaction information between users and products, thereby predicting the user's intention for the product and determining the product recommendation list.
[0003] With the development of big data and the Internet, the scale of graph data has shown an exponential growth. There may be billions of points and trillions of edges in the graph data. Such a huge amount of data exceeds the processing capacity of a single machine. The processing of graph data usually requires the participation of multiple machines. Summary of the invention
[0004] Embodiments of the present specification provide a method, an apparatus, an electronic device, and a program product for segmenting graph data.
[0005] According to a first aspect of the present specification, a method for segmenting graph data is provided. The method includes determining the degree of each of a plurality of vertices in the graph data, wherein the degree indicates the number of a plurality of edges corresponding to the vertex. The method also includes, in response to the degree of the vertex being greater than a degree threshold, segmenting the vertex into a plurality of virtual points, wherein the plurality of virtual points are assigned edges corresponding to the vertex. In addition, the method also includes assigning the plurality of virtual points to a plurality of processing threads of a distributed cluster.
[0006] According to a second aspect of the present specification, there is provided an apparatus for partitioning graph data, the apparatus comprising a degree determination module configured to determine the degrees of each of a plurality of vertices in the graph data, wherein the degrees indicate the number of a plurality of edges corresponding to the vertices. The apparatus further comprises a vertex partitioning module configured to partition the vertices into a plurality of virtual points in response to the degree of the vertex being greater than a degree threshold, wherein the plurality of virtual points are assigned edges corresponding to the vertices. In addition, the apparatus further comprises a virtual point allocation module configured to allocate the plurality of virtual points to a plurality of processing threads of a distributed cluster.
[0007] According to a third aspect of the present specification, an electronic device is provided, which includes at least one processor and a memory coupled to the at least one processor and having instructions stored thereon, wherein when the instructions are executed by the processor, the electronic device executes the method according to the first aspect.
[0008] According to a fourth aspect of the present specification, there is provided a computer-readable storage medium on which a computer program product is stored. The computer program product includes machine-executable instructions, which, when executed, enable a machine to perform the steps of the method of the first aspect of the present specification.
[0009] According to a fifth aspect of the present specification, there is provided a computer program product, which is tangibly stored on a non-volatile computer-readable medium and includes machine-executable instructions, which when executed cause a machine to perform the steps of the method of the first aspect of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features and advantages of the present specification will become more apparent through a more detailed description of exemplary embodiments of the present specification in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components in the exemplary embodiments of the present specification.
[0011] Figure 1 A schematic diagram illustrating an example scenario in which the apparatus and / or method according to an embodiment of the present specification may be implemented;
[0012] Figure 2 A schematic flow chart of a method for segmenting graph data in an embodiment of the present specification is illustrated;
[0013] Figure 3 FIG. 1 is a schematic diagram showing the vertices of a segmentation graph in an embodiment of the present specification;
[0014] Figure 4 A schematic flow chart for updating hot spots in graph data in an embodiment of the present specification is illustrated;
[0015] Figure 5 A schematic block diagram of an apparatus for segmenting image data provided in an embodiment of this specification is illustrated; and
[0016] Figure 6 A schematic block diagram of an example device suitable for implementing embodiments of the present description is illustrated.
[0017] In the various drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION
[0018] The embodiments of the present specification will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present specification are shown in the accompanying drawings, it should be understood that the present specification can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present specification. It should be understood that the drawings and embodiments of the present specification are only for exemplary purposes and are not intended to limit the scope of protection of the present specification.
[0019] In the description of the embodiments of this specification, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0020] As mentioned above, the processing of graph data often requires the participation of multiple computing nodes. During the processing, the graph data needs to be distributed, and the graph tasks are assigned to different computing nodes for parallel computing. This distribution often has the problem of load imbalance. Some computing nodes may bear too many computing tasks, resulting in reduced computing efficiency and even node overload.
[0021] At least to solve the above and other potential problems, an embodiment of the present specification provides a method for segmenting graph data. In this method, the degree of each of multiple vertices in the graph data can be determined, where the degree represents the number of multiple edges corresponding to the vertex. When the degree of each vertex is greater than the degree threshold, the vertex is segmented into multiple virtual points, where the multiple virtual points are assigned edges corresponding to the vertex. The multiple virtual points are assigned to multiple processing threads of a distributed cluster. Through this method, a vertex in the graph data including a large number of edges can be segmented into multiple virtual points including a small number of edges, and the multiple virtual points are evenly assigned to the processing threads in the distributed cluster, so that the total number of vertices and edges in each processing thread is the same or similar, so that each processing thread achieves load balancing in the process of processing graph data. In this way, the overall performance of the distributed cluster can be improved and the efficiency of processing graph data can be improved.
[0022] To facilitate understanding, first, combine Figure 1 Describe the scenarios to which the embodiments of this specification are applicable. Figure 1The illustrated scenario 100 includes graph data 101, a control device 102, and a distributed cluster 103 composed of multiple machines. The graph data 101 is a data structure that uses vertices and edges to represent entities and relationships. The graph data 101 may include multiple vertices, each of which may represent an entity or object, such as a person, a place, or a transaction. The graph data 101 may include an edge, which is used to connect two vertices and represents a relationship or dependency between the two vertices. The edges in the graph data 101 may be undirected or directed.
[0023] In some embodiments, the graph data 101 may be represented as G=(V, E). Where G represents a graph, V represents the set of all vertices in the graph G, and E represents the set of all edges in the graph G. In the graph data 101, an edge e pointing from vertex u to vertex v may be represented by e=(u, v). The graph data 101 may be pre-stored in the memory of the control device 102, or may be obtained by the control device 102 from other devices, which is not limited in the present disclosure.
[0024] In some embodiments, the distributed cluster 103 may include multiple machines with the same configuration, and the machine refers to each independent computer or server in a distributed system, which is connected together through a network to complete a task or service together. These machines are generally referred to as computing nodes (Node), and each computing node is an independent computer system that can run applications and services. The machine may include various general and / or special processing components with processing and computing capabilities, such as but not limited to a processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc.
[0025] In some embodiments, the distributed cluster 103 may include a machine 111, a machine 112, and more machines not shown in the figure. Each machine may include multiple processing threads for processing training tasks, for example, the machine 111 may include a processing thread 121 and a processing thread 122, and the machine 112 may include a processing thread 123 and a processing thread 124. It should be understood that each of the multiple machines included in the distributed cluster 103 may have more processing threads. In some embodiments, the processing threads of the machines in the distributed cluster 103 may be used to process vertices and edges in the graph data 101, for example, to perform graph computing and / or graph learning.
[0026] The control device 102 may distribute the graph data 101 to the distributed cluster 103 for processing. In some embodiments, the control device 102 may distribute the graph data 101 to multiple machines in the distributed cluster 103 for processing. For example, the control device 102 may distribute the graph data 101 to machines 111 and 112 for processing. Further, the control device 102 may distribute the vertices in the graph data 101 to multiple threads of machines 111 and 112 for processing. For example, when there are 8 vertices in the graph data 101, the control device 102 may distribute the 8 vertices evenly to the processing threads 121, 122, 123, and 124 included in machines 111 and 112 for processing. In some embodiments, the edges associated with the vertices in the graph data 101 may be distributed to the same processing thread as the vertices. Exemplarily, the control device 102 may distribute the edges pointing to the vertices to the same processing thread together with the vertices.
[0027] In some embodiments, the control device 102 may determine the degree of each of the multiple vertices in the graph data 101, wherein the degree represents the number of multiple edges corresponding to the vertex. In some embodiments, the degree threshold represents the critical value of whether the vertex needs to be split into virtual points. When the degree of each vertex is greater than the degree threshold, the control device 102 may split the vertex into multiple virtual points, wherein the multiple virtual points are assigned edges corresponding to the vertex. The control device 102 may distribute the multiple virtual points to multiple processing threads of the distributed cluster 103. In some embodiments, the control device 102 may evenly distribute the multiple virtual points to the machine 111 and the machine 112 for processing. Further, the control device 102 may evenly distribute the multiple virtual points to the four processing threads included in the machine 111 and the machine 112 for processing. For example, the control device 102 may evenly distribute the multiple virtual points to the processing thread 121, the processing thread 122, the processing thread 123, and the processing thread 124 for processing. By dividing the vertices into multiple virtual points and allocating the virtual points to multiple processing threads of the distributed cluster, the amount of tasks processed by each processing thread is the same or similar, thereby improving the computing efficiency of the distributed cluster.
[0028] It should be understood that Figure 1 The illustrated scenario 100 is only an example of the present specification and cannot be construed as limiting the present specification. For example, in some embodiments, the control device 102 may be deployed in the distributed cluster 103, for example, it may be a module in a machine in the distributed cluster 103. It should also be understood that the control device 102 in the embodiments of the present specification is only illustrative, and the control device 102 may be any suitable device and may be implemented in software and / or hardware.
[0029] Combined with the above Figure 1Describes example scenarios in which the devices and / or methods of the embodiments of this specification may be implemented. Figure 2 The process of segmenting the graph data involved in the embodiments of this specification is described.
[0030] Figure 2 The method 200 for segmenting graph data in the embodiment of the present specification is exemplarily shown. The method 200 may be executed by the control device 102 in the environment 100, for example. For the convenience of explanation, the method 200 is schematically described below by taking the control device 102 as the execution subject. Figure 2 , method 200 may include blocks 202 to 206 .
[0031] In block 202, the degrees of each of the plurality of vertices in the graph data are determined, wherein the degrees indicate the number of the plurality of edges corresponding to the vertex. In some embodiments, the control device 102 may determine a vertex in the graph data 101, and based on the vertex, determine the number of all edges connected to the vertex, wherein the number of all edges connected to the vertex may be used as the degree of the vertex. In some embodiments, the control device 102 may determine the degrees of all vertices included in the graph data 101 based on the above-mentioned method of determining the degree of a vertex.
[0032] In block 204, in response to the degree of the vertex being greater than the degree threshold, the vertex is split into a plurality of virtual points, wherein the plurality of virtual points are assigned edges corresponding to the vertex. In some embodiments, a virtual point is a special point introduced for the convenience of analysis, calculation, or representation. When the degree of a vertex in the graph data 101 is greater than the degree threshold, it indicates that the vertex needs to be split; when a vertex in the graph data 101 is less than or equal to the degree threshold, it indicates that the vertex does not need to be split. After the vertices in the graph data 101 are split, a plurality of virtual points may be formed.
[0033] In some embodiments, the control device 102 may determine that a vertex in the graph data 101 whose degree exceeds a degree threshold is a hotspot, where a hotspot indicates that the vertex has a large degree. In some embodiments, the control device 102 may determine the degree threshold for determining a vertex as a hotspot according to actual needs, which is not limited here.
[0034] In some embodiments, when the degree of one or more vertices in the graph data 101 is greater than a degree threshold, the control device 102 may determine the one or more vertices as hot spots that need to be segmented. Further, the control device 102 may segment the one or more hot spots into multiple virtual points. In some embodiments, the control device 102 may assign the edges corresponding to the one or more vertices to multiple virtual points.
[0035] In block 206, the plurality of virtual points are assigned to the plurality of processing threads of the distributed cluster. In some embodiments, the control device 102 may divide one or more hot spots in the graph data 101 into a plurality of virtual points corresponding to the hot spots, and further, the control device 102 may assign the plurality of virtual points to the plurality of processing threads of the distributed cluster 103 for processing.
[0036] In this way, when using a distributed cluster to process graph data, hot spots with larger degrees can be divided into multiple virtual points, and the multiple virtual points can be assigned to multiple processing threads of the distributed cluster for processing, so that each processing thread can achieve load balancing in the process of processing graph data, avoiding the processing thread from being assigned to a hot spot with larger degree, thereby improving computing efficiency.
[0037] Combined with the above Figure 2 The process for segmenting the graph data involved in the embodiments of this specification is described. Figure 3 Describe the embodiments of this specification. Figure 3 A schematic diagram 300 of segmenting graph vertices in an embodiment of the present specification is shown. Schematic diagram 300 includes a hotspot 301, wherein hotspot 301 includes 8 adjacent edges connected thereto, and these 8 adjacent edges are respectively connected to 8 other vertices in the graph data, and these 8 other vertices can also be called neighboring points of the corresponding virtual point. Figure 3 The 8 adjacent edges shown are purely examples. In practice, the number of adjacent edges is much greater than 8.
[0038] In some embodiments, when the degree of the hotspot 301 is greater than the degree threshold, the control device 102 may split the hotspot 301 into multiple virtual points. Exemplarily, the control device 102 may split the hotspot 301 into multiple virtual points according to the total number of threads in the distributed cluster 103, wherein the total number of threads represents the sum of the number of threads included in the multiple machines in the distributed cluster. For example, the distributed cluster 103 includes 2 machines, each of which has 2 processing threads, so the total number of threads may be 4. In some embodiments, the control device 102 may split the hotspot 301 into 4 virtual points according to the total number of threads, including virtual point 302, virtual point 303, virtual point 304, and virtual point 305. It should be understood that the control device 102 may determine the number of virtual points according to the number of multiple machines in the distributed cluster and the sum of the number of threads included in the machine. For example, when the distributed cluster 103 includes m machines, each of which has n threads, the total number of threads may be m*n. The control device 102 may determine the number of virtual points as m*n based on the total number of threads.
[0039] In some embodiments, the control device 102 may allocate multiple edges connected to the hotspot 301 to the four virtual points. Exemplarily, the control device 102 may evenly allocate the eight edges included in the hotspot 301 to the four virtual points, and the virtual points 302, 303, 304, and 305 may be allocated two edges of the hotspot 301, respectively, and the neighboring points connected to the edges are also associated with the virtual points as the edges are allocated.
[0040] In some embodiments, the control device 102 may allocate multiple virtual points to multiple processing threads of the distributed cluster 103. In some embodiments, the control device 102 may determine the number of multiple virtual points based on the total number of threads of the multiple processing threads, and then allocate the multiple virtual points to the multiple processing threads according to the number. The total number of threads is the sum of the number of processing threads included in each of the multiple machines. In some embodiments, the number of multiple virtual points may be determined according to an integer multiple of the total number of threads, such as 1 times, 2 times, etc. of the total number of threads. In this way, it can be ensured that the multiple virtual points can be evenly distributed to the multiple processing threads.
[0041] In this determination mode, the number of virtual points assigned to each processing thread is determined based on the total number of threads and the number of multiple virtual points. In some embodiments, the multiple processing threads included in the distributed cluster 103 may include a target processing thread. In some embodiments, the control device 102 can take the processing thread 121 as the target processing thread, and assign the virtual point 302 and the two edges connected to the virtual point 302 to the processing thread 121. After the assignment is completed, the control device 102 can take the processing thread 122 as the target processing thread, and assign the virtual point 303 and the two edges connected to the virtual point 303 to the processing thread 122. After the assignment is completed, the control device 102 can take the processing thread 123 as the target processing thread, and assign the virtual point 304 and the two edges connected to the virtual point 304 to the processing thread 123. After the assignment is completed, the control device 102 can take the processing thread 124 as the target processing thread, and assign the virtual point 305 and the two edges connected to the virtual point 305 to the processing thread 124.
[0042] By executing the above steps, the hot spots in the graph data can be divided into multiple virtual points based on the total number of threads in the distributed cluster, and the virtual points can be allocated to the processing threads of the distributed cluster based on the total number of threads, thereby improving the efficiency of allocating virtual points and making the number of virtual points and edges allocated to each processing thread the same or similar, greatly improving the computational efficiency of the processing threads in processing hot spots.
[0043] In some embodiments, the control device 102 can establish a bidirectional edge between a virtual point and a vertex in a plurality of virtual points. The bidirectional edge represents a symmetrical connection relationship, in which the virtual point and the vertex have an equal status and are each other's neighbor points. By establishing a bidirectional edge, information can be transmitted between the virtual point and the hotspot, and the information may include identification information, attribute information, and structural information, wherein the identification information may include numbers and indexes to facilitate the positioning and operation of vertices in the graph data; wherein the attribute information may include numerical attributes, category attributes, and text attributes to describe the relevant features of the vertex; and the structural information may include the degree, adjacent points, and connectivity of the vertex. It should be understood that the above description of the information is only an example of the embodiments of this specification, and does not constitute a limitation on the scheme provided in this specification.
[0044] In some embodiments, after the control device 102 divides the hotspot 301 into four virtual points, the control device 102 may respectively establish bidirectional edges between the virtual point 302, the virtual point 303, the virtual point 304, and the virtual point 305 and the hotspot 301. Exemplarily, the control device 102 may establish a bidirectional edge 306 between the virtual point 302 and the hotspot 301, a bidirectional edge 307 between the virtual point 303 and the hotspot 301, a bidirectional edge 308 between the virtual point 304 and the hotspot 301, and a bidirectional edge 309 between the virtual point 305 and the hotspot 301. In some embodiments, the control device 102 may distribute the information included in the hotspot 301 to each virtual point through the established bidirectional edges.
[0045] By executing the above steps, the hotspot information can be distributed to multiple virtual points, so that virtual points with smaller degrees can replace the hotspots to transmit and update information, thereby improving the computing efficiency of the distributed cluster.
[0046] In some embodiments, the control device 102 can perform graph learning based on the virtual point, the edge assigned to the virtual point, and the neighbor point corresponding to the virtual point, so as to update the hot spots, virtual points, neighbor points, and edges. The hot spots are vertices with a degree greater than a degree threshold, for example Figure 3 The hotspot 301 in the figure; the edge assigned to the virtual point is the edge between the hotspot and its neighbor point, for example Figure 3 The eight edges between the hotspot 301 and the eight neighboring points in the figure are respectively assigned to virtual points 302, 303, 304, and 305 (each virtual point is assigned two edges); the virtual points are also assigned part or all of the hotspot information, for example Figure 3The information of hotspot 301 can be distributed to virtual points 302, 303, 304 and 305 through bidirectional edges 306, 307, 308 and 309, and each virtual point can inherit part or all of the information of hotspot 301. The specific process of learning in the figure in this embodiment will be described below. Figure 4 And the corresponding embodiments are described in detail.
[0047] Next, combine Figure 4 , Taking graph learning of graph data as an example, the updating and transmission process of virtual point and hotspot information is explained. For example, Figure 4 The diagram shows a schematic flow chart for updating hot spots in graph data in an embodiment of the present specification. The method 400 may be executed by the control device 102 in the environment 100, for example. For ease of description, the method 400 is schematically described below with the control device 102 as the execution subject. Figure 4 , method 400 may include blocks 402 to 410 .
[0048] In box 402, the degrees of each of the multiple vertices in the graph data are determined. In box 404, when the degrees of each vertex are greater than the degree threshold, the vertex is split into multiple virtual points, where multiple virtual points are assigned edges corresponding to the vertex. In box 406, a bidirectional edge is established between the virtual point and the hotspot. In box 408, the hotspot information is distributed to the virtual point, so that the virtual point transmits information through the adjacent edges and updates the edges and neighboring points. In box 410, when the hotspot needs to be updated, the information of the edges and neighboring points collected by all virtual points is aggregated to update the hotspot information.
[0049] In the aforementioned block 408, in some embodiments, the control device 102 may determine information assigned to a virtual point among the plurality of virtual points based on the edge to which the virtual point is assigned, wherein the information of the virtual point represents a subset or a full set of information included in the vertex. Figure 3 , the control device 102 may assign the attribute information of the hotspot 301 to the virtual points 302, 303, 304, and 305, and each virtual point may inherit a subset or the entire set of the attribute information of the hotspot 301. In addition, the control device 102 may assign the edge between the hotspot 301 and the neighboring points to each virtual point. Figure 3 , the control device 102 can assign the edges between the hotspot 301 and the eight neighboring points to the virtual point 302, the virtual point 303, the virtual point 304 and the virtual point 305 respectively, and each virtual point can inherit two edges.
[0050] In some embodiments, the control device 102 may determine neighboring points of multiple virtual points through the edges to which the virtual points are assigned. Figure 3 , the control device 102 can determine two neighboring points of the virtual point 302 based on the two edges assigned to the virtual point 302. By performing the above steps multiple times, the control device 102 can determine the neighboring points of the virtual point 303, the virtual point 304, and the virtual point 305 respectively.
[0051] In some embodiments, reference Figure 3 , the control device 102 can update the information of the edges and neighbor points corresponding to the neighbor points of the multiple virtual points based on the multiple virtual points, wherein the updated edges and updated information of the neighbor points are related to the virtual points. In some embodiments, the control device 102 can update the information of the edges and neighbor points corresponding to the neighbor points of the virtual point 302 based on the virtual point 302. Exemplarily, the control device 102 can update the information of the two edges connected to the virtual point 302 and the two neighbor points adjacent to the virtual point 302 based on the information of the hotspot 301 assigned to the virtual point 302. By performing the above steps multiple times, the control device 102 can update the information of the edges and neighbor points corresponding to the neighbor points of the virtual point 303, the virtual point 304 and the virtual point 305 based on the information assigned to the virtual point 303, the virtual point 304 and the virtual point 305 respectively.
[0052] In some embodiments, reference Figure 3 , the control device 102 can obtain the updated edges and neighbor point information of multiple neighbor points of the multiple virtual points. When multiple virtual points need to be updated, the control device 102 can obtain the updated two edges and the updated two neighbor point information of the two neighbor points of the virtual point 302. Similarly, the control device 102 can respectively obtain the updated two edges and the updated two neighbor point information of the two neighbor points of the virtual point 303, the virtual point 304 and the virtual point 305.
[0053] In the aforementioned block 410, in some embodiments, based on the updated edges of the plurality of neighbor points and the updated information of the plurality of neighbor points, the control device 102 may update the edges corresponding to the hotspot and the information of the hotspot. In some embodiments, the control device 102 may update the information of the virtual point 302, the virtual point 303, the virtual point 304, and the virtual point 305, respectively, based on the updated information and two updated edges of the two neighbor points of the virtual point 302, the updated information and two updated edges of the two neighbor points of the virtual point 303, the updated information and two updated edges of the two neighbor points of the virtual point 304, and the updated information and two updated edges of the two neighbor points of the virtual point 305.
[0054] In some embodiments, the control device 102 may aggregate the updated information of virtual point 302, virtual point 303, virtual point 304, and virtual point 305, thereby obtaining the updated information of hotspot 301. The aggregation may be represented as weighted summation of the features of the virtual points, and the updated information of the virtual points is integrated into the hotspot 301, so that in each layer of convolution operation, the feature information of the virtual points is integrated and the feature representation of the virtual points is continuously updated. The control device 102 may complete the update of the hotspot 301 based on the updated information of the aggregated multiple virtual points. In some embodiments, the control device 102 may complete the above-mentioned aggregation operation through bidirectional edges 306, bidirectional edges 307, bidirectional edges 308, and bidirectional edges 309.
[0055] By executing the above steps, the virtual points can be quickly updated, thereby updating the hot spots, avoiding directly updating the hot spots with larger degrees, and improving the training efficiency of the graph data.
[0056] In some embodiments, the control device 102 may assign the updated hotspot 301 information to multiple virtual points through a bidirectional edge. Exemplarily, the control device 102 may assign the updated hotspot 301 information to the virtual point 302 through a bidirectional edge 306; the control device 102 may assign the updated hotspot 301 information to the virtual point 303 through a bidirectional edge 307; the control device 102 may assign the updated hotspot 301 information to the virtual point 304 through a bidirectional edge 308; the control device 102 may assign the updated hotspot 301 information to the virtual point 305 through a bidirectional edge 309. By assigning the updated hotspot information to multiple virtual points, a new round of update steps is performed again, and by executing the above update steps multiple times, efficient training of graph data is achieved.
[0057] In some embodiments, the control device 102 may store multiple virtual points in the external memory of the distributed cluster 103. Exemplarily, the control device 102 may store virtual point 302, virtual point 303, virtual point 304, and virtual point 305 in the external memory of the distributed cluster 103, wherein the external memory of the distributed cluster refers to an external storage device or storage system for storing data in addition to the memory of each node in the distributed cluster architecture. By storing multiple virtual points in the external memory of the distributed cluster, the number of virtual points that can be stored in the distributed cluster is greatly increased.
[0058] In some embodiments, in response to a request to allocate multiple virtual points to multiple processing threads, the control device 102 may allocate the multiple virtual points stored in the external memory to the multiple processing threads.
[0059] It should be understood that the above steps of updating virtual points and hotspots are only examples, and more or fewer machines may be deployed in the distributed cluster 103. The machines may include more or fewer processing threads, for example, the machine 111 and / or the machine 112 may include more or fewer processing threads. In addition, the graph data 101 may include a greater number of hotspots, which is not limited here.
[0060] By using the method of this specification, in the process of processing graph data using a distributed cluster, hot spots with large degrees can be divided into multiple virtual points, and the multiple virtual points can be assigned to multiple processing threads of the distributed cluster for processing, so that the total number of vertices and edges in each processing thread is basically the same, so that each processing thread can achieve load balancing in the process of processing graph data. In this way, the overall performance of the processing thread can be improved and the efficiency of processing graph data can be increased.
[0061] Figure 5 is a schematic diagram of an apparatus 500 for dividing graph data provided in an embodiment of this specification. Figure 5 As shown, the apparatus 500 may include a degree determination module 502, configured to determine the degrees of each of the plurality of vertices in the graph data, wherein the degrees indicate the number of the plurality of edges corresponding to the vertices. The apparatus 500 also includes a vertex segmentation module 504, configured to segment the vertices into a plurality of virtual points in response to the respective degrees of the vertices being greater than a degree threshold, wherein the plurality of virtual points are assigned edges corresponding to the vertices. In addition, the apparatus 500 also includes a virtual point allocation module 506, configured to allocate the plurality of virtual points to a plurality of processing threads of the distributed cluster.
[0062] In some embodiments, the vertex segmentation module 504 includes: a virtual point number determination module, configured to determine the number of multiple virtual points based on the number of multiple machines in the distributed cluster and the number of processing threads in the machines; and a first allocation module, configured to allocate edges corresponding to the vertices to the multiple virtual points based on the number of multiple virtual points.
[0063] In some embodiments, wherein the target processing thread is included in the plurality of processing threads, the module for determining the number of virtual points comprises: a module for determining the total number of threads of the plurality of processing threads based on the number of the plurality of machines in the distributed cluster and the number of processing threads included in the machines; and a first number determination module configured to determine the number of the plurality of virtual points based on the total number of threads, wherein the number of virtual points allocated to the target processing thread is determined based on the total number of threads and the number of the plurality of virtual points.
[0064] In some embodiments, the device 500 also includes: a bidirectional edge establishment module configured to establish bidirectional edges between a virtual point in multiple virtual points and a vertex; and an information allocation module configured to allocate information included in the vertex to the virtual point through the bidirectional edge.
[0065] In some embodiments, the device 500 also includes: a graph learning module, configured to perform graph learning based on the virtual point and its assigned edges and corresponding neighbors; wherein the virtual point has partial or all information contained in its corresponding hotspot; the hotspot is a vertex with a degree greater than a degree threshold; the edges assigned to the virtual point are partial edges adjacent to the hotspot and its neighbors.
[0066] In some embodiments, the device 500 also includes a virtual point information acquisition module, which is configured to acquire information contained in multiple virtual points after graph learning; and an aggregation module, which is configured to aggregate the information contained in multiple virtual points after graph learning to obtain corresponding hotspot information.
[0067] In some embodiments, the virtual point allocation module 506 includes: a virtual point storage module, configured to store multiple virtual points in an external memory of a distributed cluster; and a second allocation module, configured to allocate the multiple virtual points stored in the external memory to multiple processing threads in response to a request to allocate multiple virtual points to multiple processing threads.
[0068] In some embodiments, multiple processing threads are used to perform graph learning on graph data.
[0069] Figure 6 A schematic block diagram of an example device 600 that may be used to implement embodiments of the present description is shown. Figure 1 The control device 102 in the embodiment can be implemented by using the device 600. Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 602 or loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604. Although not shown in FIG. Figure 6 As shown in FIG. 6 , device 600 may also include a co-processor.
[0070] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0071] The various methods or processes described above may be performed by the computing unit 601. For example, in some embodiments, the methods may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps or actions in the methods or processes described above may be performed.
[0072] The present specification may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions for executing various aspects of the present disclosure.
[0073] Computer readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. Computer readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of computer readable storage medium include: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device, such as a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination of the above. The computer readable storage medium used here is not interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.
[0074] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0075] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0076] Various aspects of this specification are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of this specification. It should be understood that each box of the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions. The various embodiments in this specification are described in a progressive manner, and the same and similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the hardware + program embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0077] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0078] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0079] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of this specification. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.
[0080] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0081] The embodiments of the present specification have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
[0082] The above are only optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for segmenting graph data, comprising: determining a degree of each of a plurality of vertices in the graph data, wherein the degree indicates a number of a plurality of edges corresponding to the vertex; In response to the degree of the vertex being greater than a degree threshold, splitting the vertex into a plurality of virtual points, wherein the plurality of virtual points are assigned edges corresponding to the vertex; as well as The plurality of virtual points are distributed to a plurality of processing threads of the distributed cluster.
2. The method according to claim 1, wherein splitting the vertex into a plurality of virtual points comprises: Determining the number of the plurality of virtual points based on the number of the plurality of processing threads in the distributed cluster; as well as Based on the number of the plurality of virtual points, the edges corresponding to the vertices are allocated to the plurality of virtual points.
3. The method according to claim 2, wherein the plurality of processing threads include a target processing thread, and based on the number of the plurality of processing threads in the distributed cluster, determining the number of the plurality of virtual points comprises: Determining the number of the plurality of processing threads in the distributed cluster based on the number of processing threads of each of the plurality of machines in the distributed cluster; as well as A product of the number of the plurality of processing threads and an integer multiple is determined as the number of the plurality of virtual points.
4. The method according to claim 1, further comprising: Establishing a bidirectional edge between a virtual point among the plurality of virtual points and the vertex; as well as The information included in the vertex is allocated to the virtual point through the bidirectional edge, and the information includes at least one of the identification information, attribute information, and structure information of the vertex.
5. The method according to claim 4, further comprising: Performing graph learning based on the virtual points and the edges assigned thereto and corresponding neighbors; The virtual point has part or all of the information contained in the corresponding hotspot; the hotspot is a vertex whose degree is greater than a degree threshold; and the edge assigned to the virtual point is a partial edge connecting the hotspot and its neighbors.
6. The method according to claim 5, further comprising: Obtain the information contained in multiple virtual points after graph learning; as well as The information contained in the multiple virtual points after graph learning is aggregated to obtain the corresponding hotspot information.
7. The method according to claim 1, wherein allocating the plurality of virtual points to a plurality of processing threads of a distributed cluster comprises: Storing the plurality of virtual points in an external memory of the distributed cluster; as well as In response to receiving a request to allocate the plurality of virtual points to the plurality of processing threads, the plurality of virtual points stored in the external memory are allocated to the plurality of processing threads.
8. The method according to claim 1, wherein the multiple processing threads are used to perform graph learning on the graph data.
9. A device for segmenting graph data, comprising: a degree determination module configured to determine a degree of each of a plurality of vertices in the graph data, wherein the degree indicates a number of a plurality of edges corresponding to the vertex; a vertex segmentation module, configured to segment the vertex into a plurality of virtual points in response to the degree of the vertex being greater than a degree threshold, wherein the plurality of virtual points are assigned edges corresponding to the vertex; as well as The virtual point allocation module is configured to allocate the plurality of virtual points to a plurality of processing threads of the distributed cluster.
10. An electronic device comprising: at least one processor; as well as A memory coupled to the at least one processor and having instructions stored thereon, the instructions causing the apparatus to perform the method according to any one of claims 1 to 8 when executed by the at least one processor.
11. A computer program product comprising machine executable instructions which, when executed, cause the method according to any one of claims 1 to 8 to be implemented.