Graph convolutional neural network training method suitable for network edge deployment
By dividing the training tasks of large-scale graph convolution neural networks into multiple subgraphs and collaborative training on edge devices, the problems of GPU resource dependence and computing resource consumption in the existing technology are solved, and more efficient training process and resource utilization are achieved.
Patent Information
- Application Number
- CN202510028615.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The existing technology faces huge challenges in computing and memory resources when training large-scale graph convolutional neural networks. It relies too much on expensive GPU resources, and the cost of CPU-GPU data transmission is high, and communication overhead and frequent synchronization problems cannot be solved simultaneously.
By dividing the entire map into multiple sub-graphs, and solving the joint optimization targets of the calculation frequency and transmission power of the edge device on the server side, the optimization results are sent to the edge device for training. The training task of graph convolutional neural network is coordinated by the exchange of sampling boundary node information between edge devices.
Reliance on expensive GPU resources is reduced, CPU resources are used for training, energy consumption and communication overhead are reduced, training efficiency and resource utilization are improved.
Smart Images

Figure CN120029763A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network edge computing technology, and in particular to a graph convolutional neural network training method suitable for network edge deployment. Background Art
[0002] With the popularity of artificial intelligence, graph convolutional neural networks have received increasing attention from academia and industry. Graph convolutional neural networks have shown superior performance in graph structured data tasks, such as link prediction, graph classification, and node classification tasks. The success of graph convolutional neural networks can be attributed to their ability to capture complex adjacency relationships by aggregating the features of neighbor nodes and updating the target node features using multi-layer perceptrons. This two-step process (neighbor feature aggregation and node feature update) enables graph convolutional neural networks to effectively learn from graph structures. With the rise of large-scale graph structured data in the real world, such as Sina Weibo, WeChat, Meta and other platforms in social networks, graph convolutional neural network training for large-scale graphs has become increasingly important. However, extending the training of graph convolutional neural networks to large-scale real-world graph structured data still faces huge challenges, mainly due to the huge computing and memory resources required. In order to cope with these resource challenges, various sampling-based techniques have been proposed, and neighbor sampling to create small batches of data techniques usually introduce significant feature approximation errors and sacrifice accuracy. Compared with the method of building small batches based on sampling, there is also a full-graph training method using multiple GPUs. One method stores data in the CPU and uses GPU resources for calculations. However, in actual training, there will be frequent communication between the CPU and the GPU. Another method stores and trains data in the GPU, but the proposed method cannot simultaneously solve the communication overhead of all boundary nodes during training and the frequent synchronization of boundary node features. In general, the existing solutions still have the following problems:
[0003] 1. Over-reliance on GPU resources. The deployment and rental costs of GPU servers are very high. Even if they are rented through a cloud platform, the cost is still quite expensive.
[0004] 2. Data transmission cost between CPU and GPU. Frequent data transmission between CPU and GPU during data processing leads to reduced efficiency.
[0005] 3. Communication overhead and frequent synchronization issues cannot be solved at the same time. Summary of the invention
[0006] In order to overcome the above-mentioned shortcomings and deficiencies of the prior art, an object of the present invention is to provide a graph convolutional neural network training method suitable for network edge deployment.
[0007] The purpose of the present invention is achieved through the following technical solutions:
[0008] A graph convolutional neural network training method suitable for network edge deployment includes the following steps:
[0009] The S1 server divides the full graph into multiple sub-graphs, solves the joint optimization target of the edge device regarding the computing frequency and transmission power to obtain the optimization result, and then sends the sub-graph and optimization result to the edge device;
[0010] After receiving the subgraph, the S2 edge device broadcasts the obtained subgraph boundary node set to other edge devices;
[0011] S3 calculates and samples the intersection of the subgraph internal nodes and the boundary node sets of other subgraphs, generates a sampled boundary node set, and sends the sampled boundary node set to other edge devices;
[0012] S4 each edge device reconstructs a subgraph according to the received sampled boundary node set, wherein the reconstructed subgraph is composed of internal nodes and some boundary nodes;
[0013] S5 Each edge device trains the reconstructed subgraph according to the optimization result and the forward features and feature gradient information of the sampling boundary node set obtained in the previous round, and transmits the forward features and feature gradient information obtained by the sampling boundary node set of the current subgraph to other edge devices;
[0014] S6: Based on the previous round of forward features and feature gradient information, each edge device calculates the local model gradient and sends it to the server;
[0015] The S7 server receives and aggregates the local model gradients sent by each edge device, updates the global model parameters, and sends them to the edge devices;
[0016] S8 loops through steps S5 to S7 until convergence.
[0017] Furthermore, there are several edge devices, and the communication between the server and the edge devices is realized through a downlink channel and an uplink channel, and the edge devices communicate with each other through a point-to-point channel;
[0018] The edge devices collaborate to complete the training task of the graph convolutional neural network by exchanging information of sampling boundary nodes, and the server is responsible for updating the global model parameters using the local model gradients uploaded by the edge devices.
[0019] Furthermore, the server stores the original graph structure data and the computing and communication resource information of each edge device, including the maximum CPU frequency, maximum transmission power, channel gain, etc.;
[0020] Furthermore, the server in S1 divides the whole graph into multiple sub-graphs, specifically:
[0021] The server divides the entire graph into K parts with equal numbers of internal nodes. Find the first-order neighbors corresponding to each internal node in the entire graph as the boundary nodes, that is, Construct a subgraph by combining the internal nodes of each part with the corresponding boundary nodes Where K represents the number of edge devices.
[0022] Furthermore, the joint optimization objective is specifically:
[0023]
[0024]
[0025] where {f i} refers to the CPU frequency used by the edge device in local computing; τ comp Refers to each round of local computation assigned to each edge device; It means that the edge device i sends the local model gradient G to the server every round i The energy consumed can be expressed as p i Refers to the total local model gradient G of the edge device i server to server method i the transmission power used; It refers to the energy consumed by edge device i to send the number of sampled boundary nodes to edge device j, which can be expressed as p i,j It refers to the transmission power used by edge device i to send sampled boundary node information to edge device j; It refers to the local model gradient G assigned to edge device i i Communication time to the server; It refers to the communication time allocated to edge device i to send the sampled boundary node information to edge device j; T refers to the total number of iterations set; K refers to the number of edge devices; |V| refers to the total number of nodes in the whole graph; a refers to the average number of floating-point operations required to update the nodes in the subgraph in one iteration; C i Refers to the number of floating-point operations that device i can perform in one CPU cycle; refers to the capacitance coefficient of device i; represents the computing energy consumption of edge device i; B w refers to the total bandwidth of the system; c refers to the total communication volume required for the edge device to communicate the weight parameters with the server in each iteration; q refers to the total communication volume required for sampling the boundary node information in each round of communication between the edge devices; h i refers to the uplink channel gain between the edge device i and the server; h i,jRefers to the point-to-point channel gain between the edge device i and the edge device j; σ 2 refers to the noise power; τ refers to the time limit of the total training; f max Refers to the maximum CPU frequency of the device; p max Refers to the maximum communication power of the device.
[0026] Furthermore, five constraints are included, specifically:
[0027] Constraint 1 indicates that the computing time allocated to each edge device in each round is greater than the time for any device to perform local calculations in each round; Constraint 2 indicates that the total amount of communication must be greater than the total amount of communication required for model gradient updates between the edge device and the server; Constraint 3 indicates that the time allocated to the edge device for communication with other edge devices must make the total amount of communication greater than the total amount of communication for the edge device to communicate with other edge devices to sample boundary node information; Constraint 4 indicates that the training time of all iterative rounds is less than the total time limit for training; Constraint 5 indicates that the CPU computing frequency used by each edge device cannot exceed its maximum CPU frequency, and the computing frequency is adjusted through dynamic voltage and frequency adjustment technology.
[0028] Furthermore, the joint optimization objective is solved by using convex optimization technology to obtain the optimization result {f i}、τ comp , and subgraph Sent to each edge device, where {f i} refers to the CPU frequency used by edge devices for local computing; τ comp Refers to each round of local computation assigned to each edge device; It means that the edge device i sends the local model gradient G to the server every round i Energy consumed; It refers to the energy consumed by edge device i sending the number of sampled boundary nodes to edge device j; It refers to the local model gradient G assigned to edge device i i Communication time to the server; It refers to the communication time allocated to edge device i to send the sampled boundary node information to the edge device.
[0029] Further, S3 calculates and samples the intersection of the subgraph internal nodes and the boundary node sets of other subgraphs to generate a sampled boundary node set, specifically:
[0030] The edge device receives the subgraph boundary node set, obtains the intersection of the edge device and other subgraph boundary node sets, constructs an initial sampling boundary node set, and samples the initial sampling boundary node set to obtain a sampling boundary node set.
[0031] Further, each edge device in S5 trains the reconstructed subgraph according to the optimization result, and transmits the forward features and feature gradient information obtained by the sampling boundary node set of the current subgraph to other edge devices, specifically:
[0032] According to the optimization target, the CPU frequency is obtained, and the internal nodes of the reconstructed subgraph are trained. For the sampled boundary node set Using the old information from the previous iteration, the specific forward process of the graph convolutional neural network is described as follows:
[0033]
[0034] Where v is an internal node of the edge device, u is an internal node of the device subgraph among the first-order neighbors of the v node, and m is a boundary node of the device subgraph after sampling among the first-order neighbors of the v node. Indicates that when calculating the l-th layer feature of node v during iteration t, the current round feature is used for the part belonging to the internal node and For the sampled boundary nodes, the features of the previous round are used. After this round of local iteration is completed, the device sends the node features and feature gradients of the boundary nodes of the subgraphs belonging to other devices and the internal nodes of its own subgraph to the corresponding devices for use in the next round of calculation.
[0035] Furthermore, the S7 server receives and aggregates the local model gradients sent by each edge device, updates the global model parameters and sends them to the edge device, specifically:
[0036]
[0037] Where W t represents the global model parameters of the tth iteration, η represents the learning rate, Represents the model gradient of the i-th subgraph at the t-th iteration.
[0038] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0039] (1) The present invention utilizes the CPU resources of edge devices to perform collaborative training of graph convolutional neural networks instead of relying on expensive and scarce GPU resources. This approach reduces the difficulty of implementing large-scale graph convolutional neural network training, making it more applicable and applicable in resource-limited environments, and improving the scalability of resource utilization.
[0040] (2) The present invention proposes a joint optimization problem of the computing frequency and transmission power of a device, and effectively reduces energy consumption through reasonable resource allocation.
[0041] (3) The present invention adopts the method of sampling boundary nodes and using stale boundary node information, which significantly reduces the need for frequent synchronization and communication overhead. This strategy enables collaborative graph neural network training between multiple devices to converge faster than single-device training while maintaining training accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flow chart of the method of the present invention;
[0043] Figure 2 is a schematic diagram of edge network deployment of the present invention;
[0044] Figure 3(a)-Figure 3(d) It is a schematic diagram of the partitioning of the server graph of the present invention;
[0045] Figure 4 It is a simulation diagram of the test set accuracy of the method of the present invention with respect to the training time under four edge devices and different numbers of sampling boundary nodes;
[0046] Figure 5 It is a simulation diagram of the system energy consumption with respect to the number of edge devices under different numbers of sampled boundary nodes of the method of the present invention. DETAILED DESCRIPTION
[0047] The present invention will be further described in detail below in conjunction with examples, but the embodiments of the present invention are not limited thereto.
[0048] Example
[0049] like Figure 1 As shown, a graph convolutional neural network training method suitable for network edge deployment includes the following steps:
[0050] The S1 server divides the entire graph into multiple sub-graphs, and solves the joint optimization problem of the computing frequency and transmission power of the device based on the computing and communication resources of the edge device, and then distributes the sub-graphs and optimization results to the edge device.
[0051] The edge device in this embodiment refers to a terminal device, including a computer, a mobile phone, etc., which is a device capable of performing graph neural network training.
[0052] In this embodiment, the full graph refers to the entire graph convolution training graph data.
[0053] like Figure 2As shown is a schematic diagram of the edge network deployment in this embodiment. The system architecture includes a server and several edge devices. The communication between the server and the edge devices is realized through downlink and uplink channels, and the edge devices communicate through point-to-point channels. The edge devices collaborate to complete the training task of the graph convolutional neural network by exchanging information between sampling boundary nodes. The server is responsible for updating the global model parameters using the local model gradients uploaded by the edge devices.
[0054] like Figure 3(a)-Figure 3(d) The schematic diagram of the server of the present invention dividing the whole graph into multiple sub-graphs is shown. Divide it into K parts with similar numbers of internal nodes Find all first-order neighbors of the internal node part from the whole graph as boundary nodes, that is, Construct a subgraph with the corresponding internal nodes and boundary nodes Wherein K represents the number of edge devices, and the number of the divided parts is the same as the number of edge devices.
[0055] The server establishes the following optimization objectives based on the number of internal nodes, the number of sampling boundary nodes, channel conditions, etc.:
[0056]
[0057]
[0058] where {f i} refers to the CPU frequency used by the edge device in local computing; τ comp Refers to each round of local computation assigned to each edge device; It means that the edge device i sends the local model gradient G to the server every round i The energy consumed can be expressed as p i Refers to the total local model gradient G of the edge device i server to server method i the transmission power used; It refers to the energy consumed by edge device i to send the number of sampled boundary nodes to edge device j, which can be expressed as p i,j It refers to the transmission power used by edge device i to send sampled boundary node information to edge device j; It refers to the local model gradient G assigned to edge device i i Communication time to the server; It refers to the communication time allocated to edge device i to send the sampled boundary node information to the edge device; T refers to the total number of iterations set; K refers to the total number of devices; |V| refers to the total number of nodes in the whole graph; a refers to the average number of floating-point operations required to update the nodes in the subgraph in one iteration; C i Refers to the number of floating-point operations that device i can perform in one CPU cycle; refers to the capacitance coefficient of device i; represents the computing energy consumption of edge device i; B w refers to the total bandwidth of the system; c refers to the total communication volume required for the edge device to communicate the weight parameters with the server in each iteration; q refers to the total communication volume required for sampling the boundary node information in each round of communication between the edge devices; h i refers to the uplink channel gain between the edge device i and the server; h i,j Refers to the point-to-point channel gain between the edge device i and the edge device j; σ 2 refers to the noise power; τ refers to the time limit of the total training; f max Refers to the maximum CPU frequency of the device; p max Refers to the maximum communication power of the device. The optimization problem only needs to be solved before formal training, and the optimization solution is sent to the device. The device adopts corresponding settings during the training process according to the optimization solution, which can reduce the total energy consumption of training.
[0059] The optimization target P1 reflects the energy consumption effect of graph neural network training deployed at the network edge. The smaller the objective function is, the less energy is consumed by graph neural network training and the higher the energy utilization efficiency is. To minimize the optimization target, it is necessary to find a compromise between computing and communication energy loss.
[0060] The optimization objective sets five constraints, specifically:
[0061] Among them, constraint one means that the computing time allocated to each edge device in each round is greater than the time for any device to perform local calculations in each round; constraint two means that the total amount of communication must be greater than the total amount of communication required for model gradient updates between the edge device and the server; constraint three means that the time allocated to the edge device for communication with other edge devices must make the total amount of communication greater than the total amount of communication between the edge device and other edge devices for sampling boundary node information; constraint four means that the training time of all iterative rounds is less than the total time limit for training; constraint five means that the CPU computing frequency used by each edge device cannot exceed its maximum CPU frequency, and the computing frequency can be adjusted through dynamic voltage and frequency scaling technology (Dynamic Voltage and Frequency Scaling, DVFS).
[0062] This embodiment uses convex optimization technology to solve the joint optimization problem P1. The server obtains the corresponding optimization result {f i}、τ comp , and subgraph Sent to the corresponding edge device.
[0063] where {f i} refers to the CPU frequency used by edge devices for local computing; τ comp Refers to each round of local computation assigned to each edge device; It means that the edge device i sends the local model gradient G to the server every round i Energy consumed; It refers to the energy consumed by edge device i sending the number of sampled boundary nodes to edge device j; It refers to the local model gradient G assigned to edge device i i Communication time to the server; It refers to the communication time allocated to edge device i to send the sampled boundary node information to the edge device.
[0064] S2. After receiving the subgraph, the edge device broadcasts the boundary node set of its subgraph to other edge devices in sequence to facilitate the subsequent exchange of node information.
[0065] In this embodiment, specifically:
[0066] The edge device receives the subgraph and optimization result sent by the server in step S1 and stores them in the storage unit. The device sequentially stores the boundary node set of the obtained subgraph. The broadcast is sent to other devices so that all devices have the boundary node sets of other devices for subsequent node information exchange.
[0067] S3, calculating and sampling the intersection of the subgraph internal nodes and the boundary node sets of other subgraphs, generating a sampled boundary node set, and sending the sampled boundary node set to other edge devices;
[0068] In this embodiment, specifically:
[0069] Each edge device receives a set of boundary nodes of the subgraph of edge devices other than itself in step S2.
[0070] Taking edge device 1 as an example, edge device 1 receives a set of boundary nodes. And store it in the storage unit, and then use the calculation unit to obtain the internal nodes of the device subgraph Boundary nodes of the subgraph with other edge devices The intersection of The initial intersection set is sampled by equal number of nodes to obtain the final intersection set It is called the sampling boundary node set, which means that the edge device 1 needs to send the sampling boundary node set to the device {2,3…,K} in the actual training, and the terminal point set after sampling is sent to the corresponding device, and all devices execute the above steps.
[0071] S4. Each edge device reconstructs a subgraph according to the received subgraph sampled boundary node set, and the reconstructed subgraph consists of internal nodes and some boundary nodes.
[0072] In this embodiment, specifically:
[0073] Each edge device receives the information block of step S3, and constructs a new subgraph containing only the device's own internal nodes and the sampled boundary node set according to the received subgraph's sampled boundary node set, and stores it in the storage unit, and deletes the original subgraph in the original storage unit. Taking edge device 1 as an example, it receives Merge to form new boundary nodes for Does not belong to The boundary nodes are removed to reconstruct the subgraph Right now All devices perform the above steps and can obtain their own reconstructed sub-graphs.
[0074] S5 Each edge device trains the reconstructed subgraph according to the optimization results, and at the same time obtains the forward features and feature gradient information from the sampled boundary node set of the current subgraph and transmits them to other edge devices.
[0075] In this embodiment, specifically:
[0076] The edge device adjusts the CPU frequency {f i}, for internal nodes For training, for the sampled boundary node set Using the old information from the previous iteration, the forward process of the specific graph convolutional neural network can be described as:
[0077]
[0078] Where v is an internal node of the edge device, u is an internal node of the device subgraph among the first-order neighbors of the v node, and m is a boundary node of the device subgraph after sampling among the first-order neighbors of the v node. Indicates that when calculating the l-th layer feature of node v during iteration t, the current round feature is used for the part belonging to the internal node and For the sampled boundary nodes, the features of the previous round are used. After this round of local iterations is completed, the device sends the node features and feature gradients of the boundary nodes of the subgraphs of other devices and the internal nodes of its own subgraph to the corresponding devices for the next round of calculations. Taking edge device 1 as an example, it sends the feature With feature gradient Give device K.
[0079] Furthermore, the obsolete information includes dielectric characteristics and characteristic gradients.
[0080] S6. Each edge device uploads the calculated local gradient to the server. In this embodiment, specifically:
[0081] After the calculation in step S5, each edge device can obtain the local model gradient in iteration round t The sending unit sends it to the server.
[0082] S7. The server receives and aggregates the model gradients uploaded by each edge device, updates the model parameters and sends them to the edge device. In this embodiment, specifically:
[0083] The server receives the information block containing the local model gradient, aggregates the model gradients of all edge devices, and then updates the model parameters W t , which can be expressed as Then W t Send it to the edge device to prepare parameters for the next iteration.
[0084] S8, looping steps S5 to S7 until convergence, in this embodiment, specifically:
[0085] Steps S5 to S7 are executed in a loop, and after each round of iteration, it is determined whether the training has converged. If it has converged, the training is terminated, and the training is terminated and the model parameters after the training are output; otherwise, steps S5 to S7 are continued.
[0086] Figure 4 This is a simulation effect diagram of the test set accuracy of the training time under four edge devices and different numbers of sampling boundary nodes. As shown in the figure, compared with the centralized training method, the method of the present invention not only maintains the accuracy of centralized training, but also improves the convergence speed, and the effect is more obvious when the number of sampling boundary nodes is smaller.
[0087] Figure 5 This is a simulation effect diagram of the system energy consumption of the method of the present invention with respect to the number of edge devices under different numbers of sampling boundary nodes. As shown in the figure, compared with the centralized training method, the method of the present invention consumes less energy than the centralized training method under different numbers of devices, and the effect is more obvious when the number of sampling boundary nodes is smaller.
[0088] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A graph convolutional neural network training method suitable for network edge deployment, characterized in that: The steps include: The S1 server divides the full graph into multiple sub-graphs, solves the joint optimization target of the edge device regarding the computing frequency and transmission power to obtain the optimization result, and then sends the sub-graph and optimization result to the edge device; After receiving the subgraph, the S2 edge device broadcasts the obtained subgraph boundary node set to other edge devices; S3 calculates and samples the intersection of the subgraph internal nodes and the boundary node sets of other subgraphs, generates a sampled boundary node set, and sends the sampled boundary node set to other edge devices; S4 each edge device reconstructs a subgraph according to the received sampled boundary node set, wherein the reconstructed subgraph is composed of internal nodes and some boundary nodes; S5 Each edge device trains the reconstructed subgraph according to the optimization result, and at the same time obtains the forward features and feature gradient information of the sampling boundary node set of the current subgraph and transmits them to other edge devices; S6: Based on the previous round of forward features and feature gradient information, each edge device calculates the local model gradient and sends it to the server; The S7 server receives and aggregates the local model gradients sent by each edge device, updates the global model parameters, and sends them to the edge devices; S8 loops through steps S5 to S7 until convergence.
2. The graph convolutional neural network training method according to claim 1, characterized in that: There are several edge devices, and the communication between the server and the edge devices is realized through a downlink channel and an uplink channel, and the edge devices communicate with each other through a point-to-point channel; Edge devices collaborate to complete the training task of graph convolutional neural networks through information exchange between border nodes, and the server is responsible for updating the global model parameters using the local model gradients uploaded by edge devices.
3. The graph convolutional neural network training method according to claim 1, characterized in that: The computing frequency and transmission power of the edge device include the maximum CPU frequency, maximum transmission power and channel gain.
4. The graph convolutional neural network training method according to claim 1, characterized in that: In S1, the server divides the whole graph into multiple sub-graphs, specifically: The server divides the entire graph into K parts with equal numbers of internal nodes. Find the first-order neighbors corresponding to each internal node in the entire graph as the boundary nodes, that is, Construct a subgraph by combining the internal nodes of each part with the corresponding boundary nodes Where K represents the number of edge devices.
5. The graph convolutional neural network training method according to claim 4, characterized in that: The joint optimization objectives are specifically: where {f i } refers to the CPU frequency used by the edge device in local computing; τ comp Refers to each round of local computation assigned to each edge device; It means that the edge device i sends the local model gradient G to the server every round i Energy consumed; It refers to the energy consumed by edge device i sending the number of sampled boundary nodes to edge device j; It refers to the local model gradient G assigned to edge device i i Communication time to the server; It refers to the communication time allocated to edge device i to send the sampled boundary node information to the edge device; T refers to the total number of iteration rounds set; K refers to the number of edge devices; |V| refers to the total number of nodes in the whole graph; a refers to the average number of floating-point operations required to update the nodes in the subgraph in one iteration; C i Refers to the number of floating-point operations that device i can perform in one CPU cycle; refers to the capacitance coefficient of device i; represents the computing energy consumption of edge device i; B w refers to the total bandwidth of the system; c refers to the total communication volume required for the edge device to communicate the weight parameters with the server in each iteration; q refers to the total communication volume required for sampling the boundary node information in each round of communication between the edge devices; h i refers to the uplink channel gain between the edge device i and the server; h i,j Refers to the point-to-point channel gain between the edge device i and the edge device j; σ 2 refers to the noise power; τ refers to the time limit of the total training; f max Refers to the maximum CPU frequency of the device; p max Refers to the maximum communication power of the device.
6. The graph convolutional neural network training method according to claim 5, characterized in that: There are also five constraints, specifically: Constraint 1 indicates that the computing time allocated to each edge device in each round is greater than the time for any device to perform local calculations in each round; Constraint 2 indicates that the total amount of communication must be greater than the total amount of communication required for model gradient updates between the edge device and the server; Constraint 3 indicates that the time allocated to the edge device for communication with other edge devices must make the total amount of communication greater than the total amount of communication for the edge device to communicate with other edge devices to sample boundary node information; Constraint 4 indicates that the training time of all iterative rounds is less than the total training time limit; Constraint 5 indicates that the CPU computing frequency used by each edge device cannot exceed its maximum CPU frequency, and the computing frequency is adjusted through dynamic voltage and frequency adjustment technology.
7. The graph convolutional neural network training method according to claim 6, characterized in that: The joint optimization objective is solved by convex optimization technology to obtain the optimization result and subgraph Sent to each edge device, where {f i } refers to the CPU frequency used by edge devices for local computing; τ comp Refers to each round of local computation assigned to each edge device; It means that the edge device i sends the local model gradient G to the server every round i Energy consumed; It refers to the energy consumed by edge device i sending the number of sampled boundary nodes to edge device j; It refers to the local model gradient G assigned to edge device i i Communication time to the server; It refers to the communication time allocated to edge device i to send the sampled boundary node information to the edge device.
8. The graph convolutional neural network training method according to claim 1, characterized in that: S3 calculates the intersection of the subgraph internal nodes and the boundary node sets of other subgraphs and samples them to generate a sampled boundary node set, specifically: The edge device receives the subgraph boundary node set, obtains the intersection of the edge device and other subgraph boundary node sets, constructs an initial sampling boundary node set, and samples the initial sampling boundary node set to obtain a sampling boundary node set.
9. The graph convolutional neural network training method according to claim 1, characterized in that: In S5, each edge device trains the reconstructed subgraph according to the optimization result, and at the same time obtains the forward features and feature gradient information of the sampling boundary node set of the current subgraph and transmits them to other edge devices, specifically: According to the optimization target, the CPU frequency is obtained, and the internal nodes of the reconstructed subgraph are trained. For the sampled boundary node set Using the old information from the previous iteration, the specific forward process of the graph convolutional neural network is described as follows: Among them, v is an internal node of the edge device, u is an internal node of the device subgraph among the first-order neighbors of the v node, and m is a boundary node of the first-order neighbors of the v node that belongs to the device subgraph after sampling. Indicates that when calculating the l-th layer feature of node v during iteration t, the current round feature is used for the part belonging to the internal node and For the sampled boundary nodes, the features of the previous round are used. After this round of local iteration is completed, the device sends the node features and feature gradients of the boundary nodes of the subgraphs belonging to other devices and the internal nodes of its own subgraph to the corresponding devices for use in the next round of calculation.
10. The graph convolutional neural network training method according to any one of claims 1 to 9, characterized in that: The S7 server receives and aggregates the local model gradients sent by each edge device, updates the global model parameters, and sends them to the edge device, specifically: Where W t represents the global model parameters of the tth iteration, η represents the learning rate, Represents the model gradient of the i-th subgraph at the t-th iteration.
Citation Information
Patent Citations
Federal map learning method based on network topology and graph neighbor sampling joint optimization
CN117875404A
Graph neural network training methods and systems
US11227190B1
Device and method for detecting anomalies in double-party interaction data
WO2024039294A1