Computer network performance evaluation method and system based on graph convolutional neural network
By constructing heterogeneous graph data based on graph convolution neural network and making predictions, the problems of large computing overhead and poor accuracy in the performance evaluation of smart computing center networks are solved, and fast and accurate performance evaluation is achieved.
Patent Information
- Application Number
- CN202510249828.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art has problems such as large computing overhead, long simulation time or program crash due to model complexity and scale in the network performance evaluation of intelligent computing centers.
Using a method based on graph convolution neural network, heterogeneous graph data is constructed by obtaining network feature data and inputting it into the trained graph convolution neural network model, the high-order hidden features of the intelligent computing center network are extracted and performance indicators are predicted.
It realizes fast and low overhead evaluation of the network performance of smart computing centers, improves the accuracy and efficiency of evaluation, and can detect potential problems in the network in advance.
Smart Images

Figure CN120200924A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent computing center network performance evaluation, and specifically relates to a method and system for computer network performance evaluation based on graph convolutional neural network. Background Art
[0002] As the core infrastructure for large-scale computing tasks, the network architecture design of an intelligent computing center is crucial for computing performance, data transmission efficiency, and management capabilities. Intelligent computing centers usually adopt multi-layer tree-like topologies, such as Fat-Tree, Spine-Leaf, and Dragonfly, etc., to provide efficient data transmission, load balancing, and high scalability to meet the high-concurrency communication requirements of ultra-large-scale computing tasks. Among them, the Spine-Leaf topology is widely adopted due to its high bandwidth, low latency, and good scalability.
[0003] Relational graph convolutional network (RGCN) is a technology specifically for processing heterogeneous graph structure data containing multiple relationships. By introducing independent parameterization methods for different types of edges and nodes, it can effectively capture the characteristics of different relationships and has been widely used in many fields such as knowledge graphs and recommendation systems. In computer network modeling, RGCN can more accurately capture the multiple relationships between different entities in the network and achieve efficient analysis and prediction of complex dependency structures.
[0004] However, in the aspect of intelligent computing center network performance evaluation, there are many challenges in the existing technologies. Traditional network performance evaluation methods include continuous time simulation, discrete event simulation, and data-driven artificial intelligence performance approximation methods. Continuous time simulation is difficult to model due to the complex network structure of the intelligent computing center; the method based on discrete event simulation has excessive computational overhead and even program crashes due to the huge amount of events; the data-driven artificial intelligence methods have problems such as too fine-grained modeling and inapplicable neural network models, resulting in slow speed and poor accuracy.
[0005] In addition, specific studies on data center networks, such as fully fine-grained packet-level simulation training of LSTM models, Infiniband interconnection network simulation systems based on OMNeT++, and methods for distributed execution of network simulation tasks based on discrete event simulators, all face problems such as large computational overhead, long simulation time, or program crashes due to the scale and complexity of the intelligent computing center network. Therefore, a more efficient, accurate, and applicable method for intelligent computing center network performance evaluation is needed. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a computer network performance evaluation method and system based on a graph convolutional neural network for solving the technical problems of large computational overhead, too long simulation time, or program crashes caused by the complexity and scale of existing evaluation models in the above-mentioned prior art.
[0007] The object of the present invention is achieved by the following technical solutions: In a first aspect, the present invention provides a computer network performance evaluation method based on a graph convolutional neural network, including: Obtain network feature data of the to-be-evaluated intelligent computing center network, where the network feature data is obtained according to network configuration data; Construct a heterogeneous graph data according to the network feature data; the nodes in the heterogeneous graph data are obtained according to the entities in the network configuration data; the edges of the heterogeneous graph data include the network topology relationship in the network feature data, the physical connection and logical connection between different entities; Input the heterogeneous graph data into a trained graph convolutional neural network model to extract the whole-graph high-order hidden features of the intelligent computing center network, and predict the performance indicators in the to-be-evaluated intelligent computing center network according to the high-order hidden features to obtain a performance prediction result; Evaluate the performance of the intelligent computing center network according to the performance prediction result to obtain an evaluation result.
[0008] As a further improvement of the present invention, the network configuration data includes network topology information, routing policies, and traffic distribution information.
[0009] As a further improvement of the present invention, after obtaining the network configuration data, perform identification processing on the network configuration data to obtain network feature data, specifically including: Analyze the network topology information: identify the entities in the intelligent computing center network and the physical or logical connection relationships between the entities, record the corresponding link attributes, and complete the analysis of the network topology information; Extract the routing policy: analyze the routing information of the intelligent computing center network to determine the data forwarding paths and routing selection algorithms between different entities; Obtain the traffic distribution information: determine the traffic scale, type, and distribution between different entities or different pairs according to the traffic configuration of the intelligent computing center network; Structurally extract the processed network topology information, routing policies, and traffic distribution information to obtain network feature data.
[0010] As a further improvement of the present invention, constructing a heterogeneous graph data according to the network feature data specifically includes: Regarding different types of entities and routing policies in the intelligent computing center network as corresponding types of nodes in a graph structure, where the entities at least include servers, switches, routers, and links; Determine the relationships between different entities according to the network topology information, and determine the traffic distribution information; Connect the corresponding nodes to form edges according to the relationships between different entities; connect the traffic nodes and the passing switch nodes according to the routing information and traffic information to form logical edges of the graph structure; Connect the two end nodes and the intermediate nodes where the data stream is located with directed or undirected edges to form initial graph data; Attach network feature data to each node and each edge in the initial graph structure to form the final heterogeneous graph data.
[0011] As a further improvement of the present invention, the nodes in the heterogeneous graph data are also used for feature initialization processing and normalization processing, specifically including: Wrap the network feature data into a Tensor and pad it with zeros to expand to the same dimension to obtain the initial feature vector of the node; Perform normalization processing on the initial feature vector, and represent the discrete attributes in the form of one-hot codes.
[0012] As a further improvement of the present invention, the training process of the graph convolutional neural network model is as follows: Obtain a training data set, where the data set includes a number of heterogeneous graph data and corresponding performance evaluation data; Input the heterogeneous graph data in the data set into the graph convolutional neural network model, and obtain the high-order hidden feature representation of the graph after several rounds of message passing processes; Input the high-order hidden features into the fully connected neural network in the graph convolutional neural network model to obtain the mapping relationship with the performance metrics; train for multiple rounds until convergence.
[0013] As a further improvement of the present invention, the obtaining of the training data set specifically includes: Use a network emulator to simulate the data packet transmission process of the intelligent computing center network to obtain performance evaluation data, where the performance evaluation data includes delay, jitter, and flow completion time; Correspond the performance evaluation data corresponding to the same network configuration to the constructed heterogeneous graph data in a mapped manner; Integrate the heterogeneous graph data and the corresponding performance evaluation data to obtain a training data set.
[0014] As a further improvement of the present invention, the graph convolutional neural network model uses a relational graph convolutional network or a heterogeneous graph attention network.
[0015] Second aspect, the present invention provides a computer network performance evaluation system based on a graph convolutional neural network for implementing the above-mentioned computer network performance evaluation method based on a graph convolutional neural network, including: A network configuration information module for obtaining network configuration data of the to-be-evaluated intelligent computing center network; A network information recognition module for parsing and recognizing the network configuration data and obtaining network feature data according to the parsed network configuration data; A graph data generation module for constructing a heterogeneous graph data according to the network feature data; the nodes in the heterogeneous graph data are obtained according to the entities in the network configuration data; the edges of the heterogeneous graph data include the network topology relationship in the network feature data and the connections between different entities; A graph convolutional neural network module, the graph convolutional neural network module includes a graph convolutional neural network model, and the graph convolutional neural network model is used for predicting performance indicators in the to-be-evaluated intelligent computing center network according to the input heterogeneous graph data to obtain a performance prediction result; An evaluation module for evaluating the performance prediction result to obtain an evaluation result.
[0016] Third aspect, the present invention provides a computer-readable storage medium storing one or more programs, the one or more programs include instructions, and when the instructions are executed by a computing device, the computing device is caused to execute the above-mentioned computer network performance evaluation method based on a graph convolutional neural network.
[0017] Fourth aspect, the present invention provides a computing device, including: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include steps for executing the above-mentioned computer network performance evaluation method based on a graph convolutional neural network.
[0018] The beneficial effects of the present invention are as follows: The present invention provides a method for evaluating the performance of a computer network, which is used for evaluating the performance of the network in an intelligent computing center. By comprehensively utilizing network feature data, the evaluation process can more accurately reflect the actual performance of the network, avoiding evaluation biases caused by incomplete data. The network structure and the relationships between entities are intuitively represented by graph data, making the complex network structure easy to analyze and process. Moreover, the graph data helps to capture local and global features in the network, thereby improving the accuracy of the evaluation. The constructed graph data is input into a trained graph convolutional neural network model to predict the performance metrics in the network of the intelligent computing center to be evaluated, obtaining a performance prediction result. Through prediction and evaluation, potential problems in the network can be discovered in advance. In view of the characteristics that the topology of the network in the intelligent computing center is fixed and regular, the present invention abstracts the network in the intelligent computing center into a heterogeneous graph, and a lightweight relational graph convolutional neural network can maximize the extraction of the parameter features of the network in the intelligent computing center. Modeling the network in the intelligent computing center from the granularity of the entire network, so the function of predicting the performance metrics of the network in the intelligent computing center can be realized quickly and with low overhead.
[0019] Furthermore, by identifying the network topologies between the networks in the intelligent computing center, the physical and logical structures of the network can be comprehensively understood to ensure that no key nodes or connections are missed during the evaluation process. By extracting the routing, the data forwarding paths and routing selection algorithms between different entities are determined. Then, potential bottleneck paths or inefficient paths are identified, providing a basis for path optimization. By understanding the routing selection algorithm, and then according to the network load and performance requirements, the routing strategy is dynamically adjusted to improve the overall efficiency of the network. According to the traffic configuration of the network in the intelligent computing center, the traffic scale, type and distribution between different entities or different pairs are determined. This provides data support for the load balancing strategy.
[0020] Furthermore, the heterogeneous graph data can comprehensively reflect the physical and logical structures of the network, including all key entities and their mutual relationships. The relationships between different entities are determined according to the network topology information, and the traffic nodes are connected to the passing switch nodes according to the traffic distribution information to form logical edges of the graph structure. The two end nodes and intermediate nodes where the data stream is located are connected by directed or undirected edges to form initial graph data, and network feature data is attached to each node and each edge to form the final graph data.
[0021] Furthermore, a network emulator is used for simulation, and the simulation results are integrated with the graph data to form an efficient network performance evaluation process. It can not only comprehensively obtain network performance data, but also establish the correlation between the graph data and the performance data to construct a rich training data set. The data obtained from the simulation provides strong support for network performance prediction and optimization, improving the efficiency of network simulation and evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a schematic flowchart of the computer network performance evaluation method of the graph convolutional neural network of the present invention; Figure 2 It is a framework description diagram of the graph convolutional network model of the present invention; Figure 3 It is a linear regression graph of the model accuracy test implemented by the performance evaluation method provided by the present invention; Figure 4 It is a comparison graph of the training time used by the performance evaluation method provided by the present invention; Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Specific Embodiments
[0024] In order to make the purpose and technical solutions of the present invention clearer and easier to understand. The following further details the present invention in conjunction with the accompanying drawings and embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0025] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings and specific embodiments. Among them, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0026] Embodiment 1 As Figures 1 to 4 shown, this embodiment provides a computer network performance evaluation method based on a graph convolutional neural network. Through a graph generation module, the entities and traffic in the network are used as nodes, and are connected together according to the actual physical connection and routing scheme to form graph data, which is directly input into the graph convolutional neural network for training and inference to predict network performance indicators. The following are the specific implementation manners.
[0027] Obtain the network configuration data of the intelligent computing center network to be evaluated, and obtain network feature data according to the network configuration data.
[0028] Specifically, the network configuration data of this embodiment includes network topology information, routing policies, and traffic distribution information. The network topology information in this embodiment refers to the connection relationships and structures of entities such as switches, routers, servers, and links; the routing policy is used for the routing method or policy of packet forwarding in the network. The traffic distribution information refers to the data traffic model or traffic scale (such as the bandwidth requirements of communication pairs, data sending frequencies, etc.) between each host or each end node in the network.
[0029] For example, in the spine-lead topology, it refers to the connection between the HCA and the leaf switch, and the connection between the leaf switch and the spine switch. Generally, it is an array, and the elements in the array are a pair of integers, representing a connection relationship:
[0030] Routing refers to static routing table entries, that is, an N*N matrix, ports, and port queues:
[0031] Traffic refers to an array representing all the traffic in the network, and a certain element in the array is the node number passed by a flow:
[0032] In this embodiment, after obtaining the network configuration data and performing identification processing on the network configuration data, network feature data is obtained. Network feature data refers to initial parameters such as physical hardware used in the network. Commonly used ones include link bandwidth, switch core frequency, collective communication algorithms used by traffic generation devices, sending rate, and the size of each traffic. The network feature data includes: Analyze network topology information: Identify the entities in the intelligent computing center network and the physical or logical connection relationships between the entities, record the corresponding link attributes, and complete the analysis of the network topology information. For example, identify the nodes in the network (such as switches, routers, servers, etc.) and the physical or logical connection relationships between these nodes, and record the corresponding link attributes (such as link bandwidth, delay, etc.).
[0033] Extract routing policies: Analyze the routing information of the intelligent computing center network to determine the data forwarding paths and routing selection algorithms between different entities. For example, analyze the routing information to understand the data forwarding paths between different nodes and possible routing selection algorithms.
[0034] Obtain traffic distribution information: According to the traffic configuration of the intelligent computing center network, determine the traffic scale, type, and distribution between different entities or different pairs.
[0035] Structurally extract the processed network topology information, routing policies, and traffic distribution information to obtain network feature data. In this embodiment, network feature data refers to initial parameters such as physical hardware used in the network. Commonly used ones include link bandwidth, switch core frequency, collective communication algorithms used by traffic generation devices, sending rate, and the size of each traffic flow.
[0036] Construct a heterogeneous graph data based on the network feature data; the nodes in the heterogeneous graph data are obtained from the entities in the network configuration data; the edges in the heterogeneous graph data include the network topology relationship in the network feature data, physical connections, and logical connections between different entities.
[0037] Specifically, different types of entities and routing policies in the intelligent computing center network are used as corresponding types of nodes in the graph structure. Entities at least include servers, switches, routers, and links; according to the network topology information, determine the relationships between different entities and the traffic distribution information; according to the relationships between different entities, connect the corresponding nodes to form edges; according to the routing information and traffic information, connect the traffic nodes with the passing switch nodes to form logical edges of the graph structure; connect the two end nodes and intermediate nodes where the data flow is located with directed or undirected edges to form initial graph data; attach network feature data to each node and each edge in the initial graph structure to form the final heterogeneous graph data.
[0038] Specifically, corresponding types of nodes are established respectively according to different entities existing in the network (such as servers, switches, routers, links, etc.). For example, servers can be regarded as one type of node, switches as another type of node, or different types of nodes can also be defined for links, routing policies, etc. For example, traffic generation devices, leaf switches, spine switches, traffic, etc.; using the topology information, connect the traffic generation device nodes, leaf switch nodes, and spine switch nodes together to form physical node connections; according to the routing and traffic information, connect the traffic nodes with the passing switch nodes to form logical connections.
[0039] Connect the nodes with "edges" according to the connection relationships in the network topology or the association relationships between different entities. Some common situations include: physical links between servers and switches, and between switches and switches. The forwarding paths specified by the routing policies connect the two end nodes and intermediate nodes where the data flow is located with directed or undirected edges. Integrate the feature parameters into the nodes and edges: on the basis of the constructed graph structure, attach the network feature data to each node and each edge. For example, the computing power of the node, link bandwidth, traffic demand, routing policy label, etc. can all be used as the feature attributes of the node or edge.
[0040] In this implementation, it also includes initializing and normalizing the features of the nodes in the heterogeneous graph data. Specifically, the purpose of this step is to provide compatible and effective input feature representations for each node during subsequent training. It mainly involves the process of initializing and encoding the features of each type of node in the heterogeneous graph data, specifically including: Wrap the feature parameter values into a Tensor and pad them with zeros to the same dimension.
[0041] For features describing types, such as the type of collective communication algorithm used, the type of port queue scheduling algorithm, etc., they need to be encoded into one-hot codes:
[0042] For the flow-sending device nodes, the features after initialization are as follows:
[0043] Among them represents the collective communication algorithm used by the flow-sending device, represents the initial sending rate of the flow-sending device, represents the data size sent by the flow-sending device.
[0044] For the switch nodes, the features after initialization are as follows:
[0045]
[0046] Among them represents the core frequency of the leaf switch, represents the core frequency of the spine switch. To align with the features of other nodes, some core frequency values are repeated and supplemented to form the features.
[0047] For the traffic nodes, the initialized features are as follows:
[0048] Among them represents the data volume size sent, etc. represent the nodes that a flow passes through.
[0049] Feature normalization and encoding: When using neural networks or other machine learning methods, different features often have different dimensions and numerical ranges. To make the training more stable and fast, it is usually necessary to normalize the features (such as min-max normalization or standardization), and represent some discrete attributes (such as node type, routing protocol type, etc.) in the form of one-hot encoding to ensure that the network can better learn this discrete information.
[0050] Input the heterogeneous graph data into the trained graph convolutional neural network model to extract the global high-order hidden features of the intelligent computing center network, predict the performance metrics in the intelligent computing center network to be evaluated based on the high-order hidden features, and obtain a performance prediction result; evaluate the performance of the intelligent computing center network based on the performance prediction result to obtain an evaluation result.
[0051] The graph convolutional neural network model adopts a relational graph convolutional network or a heterogeneous graph attention network. In this embodiment, a relational graph convolutional network is adopted. The relational graph convolutional network combines a fully connected neural network. The relational graph convolutional network (RGCN) uses different weight matrices to process different types of relationships, enhancing the flexibility of the model. The core idea of RGCN is to aggregate neighbor information and update node representations. The fully connected network is the basic form of a neural network, and it achieves the task goal by learning the mapping relationship between the input and the output. The model uses a fully connected neural network as the mapping between the high-order hidden features of the graph and the performance metrics.
[0052] In this embodiment, the training process of the graph convolutional neural network model is as follows: Obtain a training data set, which includes a number of heterogeneous graph data and corresponding performance evaluation data. The features in the data set are packaged into a dictionary type to improve the search efficiency.
[0053] Input the heterogeneous graph data in the data set into the graph convolutional neural network model, and obtain the high-order hidden feature representation of the graph after several rounds of message passing processes; Input the high-order hidden features into the fully connected neural network in the graph convolutional neural network model to obtain the mapping relationship with the performance metrics; go through multiple rounds of training until convergence.
[0054] Specifically, the method for obtaining performance evaluation data is as follows: Use a network emulator to simulate the data packet transmission process of the intelligent computing center network to obtain performance evaluation data, and the performance evaluation data includes delay, jitter, and flow completion time. Map the performance evaluation data corresponding to the same network configuration to the constructed graph data; integrate the heterogeneous graph data and the corresponding performance evaluation data to obtain a training data set.
[0055] Specifically, input the network configuration into the network emulator. The network emulator simulates the data packet transmission process in the network according to information such as topology, traffic, and routing, and statistically records the following key performance metrics during the simulation: delay, jitter, and flow completion time. Among them: Delay: Refers to the average time or maximum time experienced by a data packet from the source node to the destination node.
[0056] Jitter: The fluctuation of the packet transmission delay, which is crucial for delay-sensitive services (such as audio and video).
[0057] Flow completion time: The total time required for all packets of a specific traffic flow (such as a large file or a short flow) to successfully reach the destination.
[0058] Network simulation is specifically divided into three steps: First, build a simulation network topology in the network emulator according to the network topology information, configure the routing table entries for each router according to the routing information, set the traffic-sending device to send traffic according to the traffic information, and set the parameters of each router, link, and traffic-sending device using the parameter information.
[0059] Start the network emulator, and use the collection module to collect performance metrics such as flow completion time and end-to-end traffic delay. Write the configuration and performance metrics into a text file for temporary storage. By running a certain number of "virtual scenarios" in the emulator or conducting a global simulation within a fixed time, obtain the specific values or distributions of the above metrics, and finally form a set of relatively accurate performance evaluation data.
[0060] Correspond the simulation performance metrics (such as delay, jitter, flow completion time, etc.) corresponding to the same network configuration to the constructed heterogeneous graph data according to a certain mapping method. Integrate these heterogeneous graph data with the corresponding metrics into a complete training-validation dataset sample. If there are multiple sets of network configurations or multi-scenario simulation results, perform the same processing one by one. Finally, a training set and a validation set containing a large number of samples can be obtained, which is convenient for subsequent training and evaluation. Store the generated dataset.
[0061] When the graph convolutional neural network training is completed, the obtained model can be used to predict the performance of unknown network configurations or analyze the performance of existing networks. At this time, integrate or deploy the model into the "performance evaluation module": Given a new network topology or configuration, convert it into graph data in the same format as in the training stage, and use the trained neural network to infer performance metrics such as network delay, jitter, and flow completion time.
[0062] Evaluate the intelligent computing center network according to the prediction performance metrics. In the last step, the performance evaluation module combines the performance prediction results output by the neural network model to comprehensively evaluate the specific deployment plans, routing strategies, or resource allocation methods that may be used in the intelligent computing center network: if there are multiple feasible network configurations or different routing strategies, this evaluation module can be used for prediction comparison to select one or more better options in terms of metrics such as latency, jitter, and flow completion time. According to the evaluation results, make corresponding adjustments to the network topology improvement, traffic scheduling plan, or allocation strategy, and re-simulate and predict to form a sustainable optimization closed-loop process. The final evaluation result can be used as a reference for the actual deployment and operation and maintenance of the intelligent computing center network, helping managers balance resource utilization and performance metrics and improve the overall network efficiency.
[0063] In summary, compared with the existing technologies, the advantages of the present invention are that, aiming at the characteristics of the fixed and regular topology of the intelligent computing center network, the intelligent computing center network is abstracted into a heterogeneous graph, and the lightweight relational graph convolutional neural network can maximize the extraction of the parameter features of the intelligent computing center network. Modeling the intelligent computing center network from the granularity of the entire network, so the function of predicting the performance metrics of the intelligent computing center network can be realized quickly and with low overhead.
[0064] As Figure 3 Shows the results of a prediction accuracy test of the model (named FPRNet) and compares it with MimicNet: It can be seen that for FPRNet, the blue points in the figure show that most of the coordinates composed of the sample labels and prediction values are close to being proportional, indicating that the gap between the model prediction values and the sample true values is very small, and the final average relative error in the test set is about 3.79%. For MimicNet, which is represented by orange points in the figure, it can be seen that the difference between its prediction values and the sample labels is larger than that of FPRNet, and the final average relative error of testing MimicNet is 6.72% As Figure 4 Shows that due to the use of a lightweight design, FPRNet has a shorter training time compared to other artificial intelligence-based simulation solutions. From Figure 4 It can be seen that due to the differences in architecture design, there are significant differences in the inference time of the three models. Both MimicNet and RouteNet-E use a time series-based recurrent neural network (RNN), so their inference times are relatively long. Especially for RouteNet-E, due to its relatively complex model structure and large parameter scale, its inference time is particularly prominent. While FPRNet has the shortest inference time and shows relatively stable characteristics due to its relatively simple structure and direct processing of graph data.
[0065] Example 2 This embodiment provides a computer network performance evaluation system based on a graph convolutional neural network, which is used to implement the computer network performance evaluation method based on the graph convolutional neural network in Embodiment 1. The system specifically includes: A network configuration information module, which is used to obtain the network configuration data of the intelligent computing center network to be evaluated; A network information recognition module, which is used to parse and recognize the network configuration data, and obtain network feature data according to the parsed network configuration data.
[0066] After receiving the configuration information, the network recognition module sequentially performs the following operations: (a) Parse the network topology: Identify the nodes in the network (such as switches, routers, servers, etc.) and the physical or logical connection relationships between these nodes, and record the corresponding link attributes (such as link bandwidth, delay, etc.). (b) Extract routing policies: Analyze the routing information to understand the data forwarding paths between different nodes and the possible routing selection algorithms. (c) Obtain traffic characteristics: According to the traffic configuration, determine the traffic scale, type, and distribution of different nodes or different pairs in the network. (d) Collect characteristic parameters: Structurally extract the information parsed above and summarize it into a set of characteristic parameters to be used later, such as node types, link attributes, routing selection methods, traffic load levels, etc., to provide basic data support for the subsequent graph data construction and network performance evaluation.
[0067] A graph data generation module, which is used to construct a heterogeneous graph data according to the network feature data; the nodes in the heterogeneous graph data are obtained according to the entities in the network configuration data; the edges of the heterogeneous graph data include the network topology relationship in the network feature data and the connections between different entities.
[0068] This process generally includes: Constructing the node types of the graph: According to the different entities existing in the network (such as servers, switches, routers, links, etc.), set up corresponding types of nodes respectively. For example, servers can be regarded as one type of node, switches as another type of node, or different types of nodes can also be defined for links, routing policies, etc.
[0069] Constructing the edges of the graph: According to the connection relationships in the network topology or the association relationships between different entities, connect the nodes with "edges". Some common situations include: the physical links between servers and switches, and between switches and switches.
[0070] The forwarding path specified by the routing policy can connect the two end nodes and intermediate nodes where the data flow is located with directed or undirected edges. Integrate feature parameters into nodes and edges: Based on the constructed graph structure, attach network feature data to each node and each edge. For example, the computing power of nodes, link bandwidth, traffic demand, routing policy labels, etc. can all be used as the feature attributes of nodes or edges.
[0071] Finally, the graph data generation module will output a graph structure with multiple types of nodes, multiple types of edges, and corresponding feature information, providing a directly usable data representation form for subsequent graph neural network training.
[0072] A graph convolutional neural network module, the graph convolutional neural network module includes a graph convolutional neural network model, and the graph convolutional neural network model is used to predict the performance metrics in the to-be-evaluated intelligent computing center network according to the input heterogeneous graph data to obtain a performance prediction result; An evaluation module, used to evaluate the performance prediction result to obtain an evaluation result.
[0073] The system also includes a network emulator module. The input of the network emulator module is related to performance metric statistics. Input the network configuration (configuration information the same as or similar to that in step 1) into a dedicated "network emulator module". The network emulator usually simulates the transmission process of data packets in the network according to information such as topology, traffic, and routing, and statistically records the following key performance metrics during the simulation process: Delay: It refers to the average time or maximum time experienced by a data packet from the source node to the destination node.
[0074] Jitter: The fluctuation of the data packet transmission delay, which is very critical for delay-sensitive services (such as audio and video, etc.).
[0075] Flow completion time: It refers to the total time required for all data packets of a specific traffic (such as a large file or a short flow) to successfully reach the destination.
[0076] By running a certain number of "virtual scenarios" in the emulator or conducting a global simulation within a fixed time, the specific values or distributions of the above metrics can be obtained, and finally a relatively accurate set of performance evaluation data is formed.
[0077] The network emulator module specifically refers to a network simulation system built using the discrete event emulator OMNeT++. The network simulation is specifically divided into three steps: First, build the network topology in the network emulator according to the topology information in step 1, configure the routing table entries for each router according to the routing information, set the flow-sending device to send traffic according to the traffic information, and set the parameters of each router, link, and flow-sending device using the parameter information.
[0078] Start the network emulator, and use the collection module to collect performance metrics such as the traffic completion time and the end-to-end traffic delay.
[0079] Write the configuration and performance metrics into a text file for temporary storage.
[0080] It further includes a packaging module, which specifically refers to pairing the initialized graph data with the metrics generated by the network emulator to generate a format recognizable by the neural network, namely the dataloader.
[0081] Embodiment 3 In an embodiment of the present invention, a computer-readable storage medium is provided. This medium belongs to the memory device of the terminal device and is mainly used to store programs and data. The computer-readable storage medium includes both the storage medium built into the terminal and the extended storage medium supported by the terminal. Specifically, any tangible medium that can store a program and be used by an instruction execution system, device, or component belongs to this category. This storage medium provides storage space for storing the terminal operating system and instructions (including one or more computer programs and their codes) that can be loaded and executed by the processor. Examples include electrical connections, portable disks, hard disks, RAM, ROM, EPROM / flash memory, optical fibers, CD-ROMs, optical storage devices, magnetic storage devices, etc., and combinations thereof.
[0082] In addition, the computer-readable storage medium also relates to data signals propagated in the baseband or as a carrier wave, and these signals carry readable program codes, which can be in the form of electromagnetic signals, optical signals, etc. The readable storage medium is not limited to the above types, and also includes other media that can send, propagate, or transmit programs for use by an instruction execution system, device, or component. The program code can be transmitted via wireless, wired, optical cable, RF, etc.
[0083] The program code can be written in various programming languages, such as object-oriented languages (Python, Java, C++, etc.) and procedural languages (C language, etc.). The code can be executed completely or partially on the user device, or can be used as an independent software package, or can be executed partially / fully on a remote device. The remote device is connected to the user device via a LAN, WAN, or Internet service provider.
[0084] One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the computer network performance evaluation method based on the graph convolutional neural network in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps: Obtain the network feature data of the intelligent computing center network to be evaluated, where the network feature data is obtained according to the network configuration data; Construct graph data based on network feature data; the nodes in the graph data are obtained from the entities in the network configuration data; the edges of the graph data include the network topology relationship in the network feature data and the connections between different entities; Input the graph data into the trained graph convolutional neural network model to predict the performance metrics in the to-be-evaluated intelligent computing center network, and obtain a performance prediction result; Evaluate the performance of the intelligent computing center network according to the performance prediction result to obtain an evaluation result.
[0085] Embodiment 4 As Figure 5 shown, the terminal device in this embodiment is a computer device 60, which mainly includes a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and running on the processor 61. When the computer program 63 is executed by the processor 61, it can implement the computer network performance evaluation method based on the graph convolutional neural network, or implement the functions of each model / unit of the computing system of the graph convolutional neural network model in Embodiment 1, which will not be elaborated here.
[0086] The types of computer devices 60 are relatively diverse, including desktop computers, notebooks, palm computers, and cloud servers, etc. Its components are not limited to the processor 61 and the memory 62, and may also include input / output devices, network access devices, buses, etc. Figure 5 Only as an example, the number and types of components of the actual device may vary.
[0087] The processor 61 can be a central processing unit (CPU), or other general-purpose processors, graphics processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) and other programmable logic devices, discrete gates, transistor logic devices, or even data processing logic devices based on quantum computing, discrete hardware components, etc., or a conventional microprocessor.
[0088] Regarding the memory 62, it can be either an internal storage unit of the computer device 60, such as a hard disk or memory, or an external storage device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD card), flash card (FlashCard), etc. The memory 62 can also include both an internal storage unit and an external storage device, for storing computer programs, other required programs and data, and temporarily storing the data that has been output or to be output.
[0089] In various embodiments of the present application, the memory, database, or other media mentioned cover non-volatile and volatile memories. Non-volatile memories include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric memories (FRAM), phase change memories (PCM), graphene memories, etc.; volatile memories include random access memory (RAM) or external cache memories, etc., and RAM can be further divided into various forms such as static random access memory (SRAM) and dynamic random access memory (DRAM).
Claims
1. A computer network performance evaluation method based on graph convolutional neural network, characterized in that: include: Acquire network characteristic data of the intelligent computing center network to be evaluated, wherein the network characteristic data is obtained according to the network configuration data; Heterogeneous graph data is constructed according to network feature data; nodes in the heterogeneous graph data are obtained according to entities in the network configuration data; edges of the heterogeneous graph data include network topological relationships in the network feature data, physical connections between different entities, and logical connections; Input the heterogeneous graph data into a trained graph convolutional neural network model to extract high-order hidden features of the entire graph of the intelligent computing center network, and predict the performance indicators of the intelligent computing center network to be evaluated based on the high-order hidden features to obtain performance prediction results; Based on the performance prediction results, the performance of the intelligent computing center network is evaluated to obtain the evaluation results.
2. The computer network performance evaluation method based on graph convolutional neural network according to claim 1 is characterized in that: The network configuration data includes network topology information, routing strategy and traffic distribution information.
3. The computer network performance evaluation method based on graph convolutional neural network according to claim 2 is characterized in that: After the network configuration data is acquired, the network configuration data is identified and processed to obtain network feature data, which specifically includes: Analyze network topology information: Identify the entities in the intelligent computing center network and the physical or logical connection relationships between entities, record the corresponding link attributes, and complete the analysis of network topology information; Extract routing strategies: Analyze the routing information of the intelligent computing center network to determine the data forwarding paths and routing selection algorithms between different entities; Obtain traffic distribution information: Determine the traffic scale, type, and distribution between different entities or different pairs based on the traffic configuration of the intelligent computing center network; The processed network topology information, routing strategy and traffic distribution information are structured and extracted to obtain network feature data.
4. The computer network performance evaluation method based on graph convolutional neural network according to claim 3 is characterized in that: Construct heterogeneous graph data based on network feature data, including: Different types of entities and routing strategies in the intelligent computing center network are respectively used as corresponding types of nodes in the graph structure, and the entities include at least servers, switches, routers and links; Based on the network topology information, determine the relationship between different entities and confirm the traffic distribution information; According to the relationship between different entities, the corresponding nodes are connected to form edges; according to the routing information and traffic information, the traffic nodes are connected with the switch nodes passed through to form the logical edges of the graph structure; Connect the two end nodes and the middle node of the data flow with directed or undirected edges to form the initial graph data; Network feature data is added to each node and each edge in the initial graph structure to form the final heterogeneous graph data.
5. The computer network performance evaluation method based on graph convolutional neural network according to claim 4 is characterized in that: The nodes in the heterogeneous graph data are also used to perform feature initialization processing and normalization processing, which specifically includes: Pack the network feature data into Tensor and expand it to the same dimension with zero padding to obtain the initial feature vector of the node; The initial feature vector is normalized, and the discrete attributes are represented in a one-hot code manner.
6. The computer network performance evaluation method based on graph convolutional neural network according to claim 1, characterized in that: The training process of the graph convolutional neural network model is: Obtaining a training data set, wherein the data set includes a plurality of heterogeneous graph data and corresponding performance evaluation data; The heterogeneous graph data in the dataset is input into the graph convolutional neural network model, and after several rounds of message passing, the high-order hidden feature representation of the graph is obtained; The high-order hidden features are input into the fully connected neural network in the graph convolutional neural network model to obtain the mapping relationship between them and the performance indicators; multiple rounds of training are performed until convergence.
7. The computer network performance evaluation method based on graph convolutional neural network according to claim 6 is characterized in that: The obtaining of the training data set specifically includes: A network simulator is used to simulate the data packet transmission process of the intelligent computing center network to obtain performance evaluation data, wherein the performance evaluation data includes delay, jitter and flow completion time; The performance evaluation data corresponding to the same network configuration is mapped to the constructed heterogeneous graph data; The heterogeneous graph data is integrated with the corresponding performance evaluation data to obtain the training data set.
8. The computer network performance evaluation method based on graph convolutional neural network according to claim 1 is characterized in that: The graph convolutional neural network model adopts a relational graph convolutional neural network or a heterogeneous graph attention network.
9. A computer network performance evaluation system based on graph convolutional neural network, used to implement the computer network performance evaluation method based on graph convolutional neural network according to any one of claims 1 to 8, characterized in that: include: The network configuration information module is used to obtain the network configuration data of the intelligent computing center network to be evaluated; A network information identification module is used to parse and identify network configuration data, and obtain network feature data based on the parsed network configuration data; A graph data generation module is used to construct heterogeneous graph data according to network feature data; the nodes in the heterogeneous graph data are obtained according to the entities in the network configuration data; the edges of the heterogeneous graph data include the network topology relationship in the network feature data and the connection between different entities; A graph convolutional neural network module, wherein the graph convolutional neural network module includes a graph convolutional neural network model, and the graph convolutional neural network model is used to predict the performance indicators in the intelligent computing center network to be evaluated according to the input heterogeneous graph data to obtain a performance prediction result; The evaluation module is used to evaluate the performance prediction results and obtain evaluation results.
10. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions which, when executed by a computing device, cause the computing device to execute the computer network performance evaluation method based on a graph convolutional neural network as described in any one of claims 1 to 8.
Citation Information
Cited By
Power distribution station house intelligent gateway sensor equipment protocol automatic matching system
CN120658814A
An automatic protocol matching system for intelligent gateway sensor devices in power distribution rooms
CN120658814B