Multi-domain computing resource aggregation method and system based on virtualized user network
By collecting and analyzing real-time state data of multi-domain computing resources, generating multi-dimensional resource feature vectors, using intelligent aggregation policy network to match cross-domain resources, forming a virtualized resource pool, solving the problem of insufficient accuracy of resource aggregation in the existing technology, and achieving efficient resource utilization and business processing.
Patent Information
- Application Number
- CN202510687268.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing multi-domain computing resource aggregation technology cannot deeply explore the close correlation characteristics of computing resource nodes and power services, resulting in insufficient accuracy and effectiveness of resource aggregation, and cannot meet the demand for efficient resource utilization of power grid services.
Real-time state data of computing resource nodes in multi-domain environments is collected, multi-dimensional resource feature vectors are generated through virtualized resource mapping models, cross-domain resource association analysis is used to dynamically match weights, form a virtualized resource pool, and deploy resource scheduling topology to realize resource collaborative services.
It realizes organic integration and efficient collaboration of cross-domain resources, improves resource utilization and business processing efficiency, and meets the flexible scheduling needs of power grid services.
Smart Images

Figure CN120281776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing, and in particular, to a multi-domain computing resource aggregation method and system based on a virtualized user network. Background Art
[0002] In the power grid industry, with the continuous advancement of the construction of smart grids and the increasing complexity of power services, the efficient aggregation and flexible scheduling of resources in a multi-domain computing environment have become key issues to be urgently solved.
[0003] Currently, existing multi-domain computing resource aggregation technologies cannot deeply explore the characteristics of computing resource nodes closely related to power services from multiple dimensions. Generally, only simple resource classification or matching based on a single feature is performed, and multi-dimensional resource feature vectors including dynamic load fluctuation characteristics, resource compatibility characteristics, and task adaptability characteristics cannot be generated. This makes it difficult to accurately evaluate the adaptability and potential value of resources to different business scenarios in the power grid, affecting the accuracy and effectiveness of resource aggregation.
[0004] Moreover, existing technologies cannot perform dynamic correlation analysis and collaborative matching according to the actual situation of multi-domain resources in the power grid. Power grid services are characterized by high real-time and strong complexity, while existing technologies usually perform resource aggregation based on fixed rules or simple algorithms, and cannot fully consider the complex relationships and collaborative potential between cross-domain computing resources, resulting in the inability of the aggregated resource pool to fully utilize the advantages of each node, low resource utilization rate, and difficulty in meeting the requirements of power grid services for efficient resource utilization. Summary of the Invention
[0005] In view of the problems mentioned above, in combination with the first aspect of the present invention, embodiments of the present invention provide a multi-domain computing resource aggregation method based on a virtualized user network, and the method includes:
[0006] Collect a set of real-time resource status data of computing resource nodes dispersedly deployed in a multi-domain environment, and the set of real-time resource status data includes the computing power load rate, storage space occupancy rate, and network bandwidth utilization rate of each computing resource node;
[0007] Invoke a virtualized resource mapping model to perform cross-domain resource feature extraction processing on the set of real-time resource status data, and generate a multi-dimensional resource feature vector for each computing resource node, where the multi-dimensional resource feature vector includes dynamic load fluctuation characteristics, resource compatibility characteristics, and task adaptability characteristics;
[0008] Based on a preset intelligent aggregation strategy network, perform multi-domain resource correlation analysis processing on the multi-dimensional resource feature vector, determine the collaborative matching weight between cross-domain computing resources, and aggregate the dispersed computing resource nodes into a virtualized resource pool according to the collaborative matching weight;
[0009] Perform dynamic resource adaptation processing on the aggregated resources in the virtualized resource pool, generate a resource scheduling topology structure that matches the target service requirements, and deploy the resource scheduling topology structure to a multi-domain computing environment to trigger resource collaboration services.
[0010] In another aspect, an embodiment of the present invention further provides a multi-domain computing resource aggregation system based on a virtualized user network, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the machine-readable storage medium to implement the above method.
[0011] Based on the above aspects, the embodiment of the present invention invokes a virtualized resource mapping model to perform cross-domain resource feature extraction processing on a real-time resource status data set. The generated multi-dimensional resource feature vector not only covers dynamic load fluctuation features and resource compatibility features, but also includes task adaptation degree features, deeply depicting the characteristics of computing resource nodes from multiple dimensions. Based on a preset intelligent aggregation policy network, perform multi-domain resource association analysis processing on the multi-dimensional resource feature vector, and can accurately determine the collaborative matching weights between cross-domain computing resources. Furthermore, the scattered computing resource nodes are scientifically and reasonably aggregated into a virtualized resource pool, realizing the organic integration and efficient collaboration of cross-domain resources. Perform dynamic resource adaptation processing on the aggregated resources in the virtualized resource pool, and generate a resource scheduling topology structure that matches the target service requirements. It can flexibly and accurately allocate and schedule resources according to the demand characteristics of different services, significantly improving resource utilization and service processing efficiency. Deploy the resource scheduling topology structure to a multi-domain computing environment to trigger resource collaboration services, realizing the efficient collaboration and dynamic adaptation of multi-domain computing resources as a whole, and effectively improving the overall performance and service response ability of the multi-domain computing system. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a schematic execution flow diagram of a multi-domain computing resource aggregation method based on a virtualized user network provided by an embodiment of the present invention.
[0013] Figure 2 is a schematic diagram of exemplary hardware and software components of a multi-domain computing resource aggregation system based on a virtualized user network provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] The present invention will be specifically described below in conjunction with the accompanying drawings of the specification. Figure 1 is a schematic flow diagram of a multi-domain computing resource aggregation method based on a virtualized user network provided by an embodiment of the present invention. The multi-domain computing resource aggregation method based on a virtualized user network will be introduced in detail below.
[0015] Step S110: Collect the real-time resource status data set of computing resource nodes dispersed and deployed in a multi-domain environment. The real-time resource status data set includes the computing power load rate, storage space occupancy rate, and network bandwidth utilization rate of each computing resource node.
[0016] In this embodiment, in the multi-domain computing environment of the power grid field, there are a large number of computing resource nodes dispersed and deployed. These computing resource nodes are distributed in different substations, power generation stations, power control centers, and various monitoring stations, and undertake key tasks such as power data collection and processing, power grid fault diagnosis, and power load forecasting. In order to achieve efficient management and aggregation of these computing resources, the primary task is to collect real-time resource status data.
[0017] Taking a large power grid covering multiple regions as an example, there is a computing resource node A in a large substation, which is responsible for real-time processing and analysis of the operation data of many power equipment such as transformers and circuit breakers in the station; node B is located in the regional power control center and undertakes the power dispatching and distribution calculation tasks of the entire regional power grid; node C is set in a wind power station in a remote mountain area and mainly processes the operation status monitoring and fault warning data of wind turbines.
[0018] For node A, with the help of the monitoring software installed in its hardware system, the computing power load rate is collected at intervals of one minute. Suppose this node is equipped with 64 computing cores. At a certain specific moment, 40 computing cores are running at full load, another 10 computing cores are at a load level of 70%, and the remaining 14 computing cores are idle. Then the number of equivalent computing cores in use is 40 + 10×0.7 = 47, and its computing power load rate is 47 divided by 64, approximately 73.44%. At the same time, the storage space occupancy rate is collected. The total storage space of node A is 8TB (i.e., 8192GB), and the used storage space is 4.5TB (i.e., 4608GB), so the storage space occupancy rate is 4608 divided by 8192, approximately 56.25%. In terms of network bandwidth utilization rate, the total network bandwidth of node A is 2000Mbps, and the currently used network bandwidth is 1200Mbps, and the network bandwidth utilization rate is 1200 divided by 2000, that is, 60%.
[0019] For node B, the total number of computing cores is 128. At the current moment, 80 computing cores are in working state, among which 60 are running at full load and 20 are at 60% load state, and the remaining 48 are idle. The equivalent number of computing cores in use is 60 + 20×0.6 = 72. The computing power load rate is 72 divided by 128, approximately 56.25%. The total storage space is 16TB (i.e., 16384GB), 7TB (i.e., 7168GB) has been used, and the storage space occupancy rate is 7168 divided by 16384, approximately 43.75%. The total network bandwidth is 4000Mbps, and 2400Mbps is currently in use. The network bandwidth utilization rate is 2400 divided by 4000, which is 60%.
[0020] For node C, the total number of computing cores is 32. Currently, 20 computing cores are running, among which 15 are at full load and 5 are at 30% load state, and the remaining 12 are idle. The equivalent number of computing cores in use is 15 + 5×0.3 = 16.5. The computing power load rate is 16.5 divided by 32, approximately 51.56%. The total storage space is 2TB (i.e., 2048GB), 0.8TB (i.e., 819.2GB) has been used, and the storage space occupancy rate is 819.2 divided by 2048, approximately 40%. The total network bandwidth is 500Mbps, and 150Mbps is currently in use. The network bandwidth utilization rate is 150 divided by 500, which is 30%.
[0021] Thus, by collecting and summarizing the computing power load rate, storage space occupancy rate, and network bandwidth utilization rate data of all computing resource nodes such as nodes A, B, and C, a real-time resource status data set is formed.
[0022] Step S120: Invoke the virtualized resource mapping model to perform cross-domain resource feature extraction processing on the real-time resource status data set, and generate a multi-dimensional resource feature vector for each computing resource node. The multi-dimensional resource feature vector includes dynamic load fluctuation features, resource compatibility features, and task adaptability features.
[0023] In this embodiment, after obtaining the real-time resource status data set, it is necessary to invoke the virtualized resource mapping model to deeply process these data to extract the key features of each computing resource node. Specifically, dynamic load fluctuation features, resource compatibility features, and task adaptability features need to be extracted from the data, and then these three features are combined into a multi-dimensional resource feature vector. Taking nodes A, B, and C as examples, their real-time resource status data sets are input into the virtualized resource mapping model, and through a series of processes of the model, a multi-dimensional resource feature vector containing the above three features is generated for each node. The following elaborates in detail how to extract these features.
[0024] Step S121: Input the real-time resource status data set into the feature encoding layer of the virtualized resource mapping model, and extract the periodic fluctuation pattern of the computing power load rate through a temporal convolutional network to generate dynamic load fluctuation features.
[0025] To accurately extract the periodic fluctuation pattern of the computing power load rate, it is necessary to carefully process the collected computing power load rate data. Taking node A as an example, its computing power load rate data changes continuously over time. The computing power load rate of node A is periodically segmented according to a preset time window. Assume the preset time window is 2 hours. Then the computing power load rate data of node A within a day is segmented into 12 load rate time series segments, and each load rate time series segment contains load rate sampling data with 120 consecutive timestamps (one per minute).
[0026] For example, the first load rate time series segment is the computing power load rate data from 0:00 to 2:00 in the early morning, which contains the load rate sampling values for each minute within these 120 minutes. For example, the load rate at 0:00 is 70%, the load rate at 0:01 is 71%, and the load rate at 0:02 is 70.5%, etc.
[0027] Step S1211: Periodically segment the computing power load rate according to a preset time window to obtain multiple load rate time series segments, and each load rate time series segment contains load rate sampling data with consecutive timestamps.
[0028] Continuing with node A as an example, the preset time window is 2 hours, and the monitoring software collects the computing power load rate data once a minute. Starting from the beginning of the day, the first load rate time series segment is the computing power load rate data within 120 minutes from 0:00 to 2:00. Assume that at 0:00, the computing power load rate of node A is 70%, which is the first data point of this time series segment; at 0:01, the computing power load rate becomes 71%, becoming the second data point; and so on, until 2:00, obtaining the 120th data point, thus forming a complete load rate time series segment. In the same way, the computing power load rate data within a day is segmented into 12 such load rate time series segments, and these load rate time series segments contain the continuous changes in the computing power load rate of node A at different time periods.
[0029] Step S1212: Input the load rate time series segment into the dilated causal convolutional layer, and use convolutional kernels with different dilation coefficients to parallelly extract short-period fluctuation features and long-period trend features.
[0030] Input the segmented load rate time series segment of node A into the dilated causal convolutional layer. In this dilated causal convolutional layer, convolution operations are performed using convolutional kernels with different dilation coefficients. For example, a convolutional kernel with a dilation coefficient of 2 is used to extract short-term fluctuation features, and the load rate time series segment can be convolved at every 2 time steps. Suppose the first load rate time series segment is [70, 71, 70.5, 72, 71.5, 73, 72.5, 74, 73.5, 75, …, 78]. When the convolutional kernel with a dilation coefficient of 2 is convolved with this load rate time series segment, it will process the data points at odd positions such as the 1st, 3rd, 5th, 7th, 9th, etc. successively. Through the convolution operation, the local change relationships between these data points can be captured, thereby extracting short-term fluctuation features. For example, the changes from 70 to 70.5 and then to 71.5, etc.
[0031] Meanwhile, a convolutional kernel with a dilation coefficient of 10 is used to extract long-term trend features. For example, convolution can be performed at every 10 time steps to process data within a longer time range. For the above load rate time series segment, it can process the data points at positions such as the 1st, 11th, 21st, etc. From this, the change trend of the computing power load rate can be observed from a more macroscopic perspective, such as whether the overall trend is upward or downward. By using these convolutional kernels with different dilation coefficients in parallel, the short-term and long-term feature information can be obtained simultaneously.
[0032] Step S1213: Input the short-term fluctuation features and long-term trend features into the gated recurrent unit, and fuse the fluctuation patterns of different time scales through the time gating mechanism to generate a fused time series feature vector.
[0033] Input the short-term fluctuation features and long-term trend features obtained by node A through the dilated causal convolutional layer into the gated recurrent unit. The gated recurrent unit has a time gating mechanism, which can selectively update and retain information according to the importance of features at different time scales.
[0034] There are two key gating mechanisms in the gated recurrent unit, namely the update gate and the reset gate. The update gate determines how much information from the previous hidden state needs to be passed to the current moment, and the reset gate determines how much information from the previous hidden state needs to be reset.
[0035] For the short-term fluctuation features and long-term trend features, at each time step, the update gate will calculate an update weight based on the input features at the current moment and the previous hidden state. This update weight represents the proportion of the previous hidden state that needs to be retained. For example, if the short-term fluctuation features change significantly at the current moment, the update gate may reduce the retention proportion of the previous hidden state to pay more attention to the current short-term fluctuation situation.
[0036] Resetting the gate calculates a reset weight, which is used to determine how much information from the hidden state at the previous moment can participate in the calculation of the hidden state at the current moment. If there are significant changes in the long-term trend characteristics at the current moment, the reset gate may allow more hidden state information related to the long-term trend at the previous moment to participate in the calculation.
[0037] Through this time gating mechanism, the gated recurrent unit can adaptively fuse short-term fluctuation characteristics and long-term trend characteristics to generate a fused time series feature vector, which contains comprehensive information on the fluctuation patterns of the computing power load rate at different time scales and can more comprehensively reflect the changes in the computing power load of Node A.
[0038] Step S1214: Perform spectral analysis on the time series feature vector, identify the frequency domain components with significant periodicity, and retain the frequency domain components that match the service cycle through a band-pass filter to generate the filtered spectral features.
[0039] In this embodiment, spectral analysis can convert the signal in the time domain to the frequency domain to identify the components of different frequencies in the signal. In the power grid industry, different services may have different periodic characteristics. For example, the power load may have a daily cycle, a weekly cycle, etc.
[0040] Use methods such as the fast Fourier transform to convert the time series feature vector to the frequency domain to obtain the frequency domain signal. In the frequency domain, each frequency component corresponds to a different period. By analyzing the frequency domain signal, identify the frequency components with significant energy, and these components represent the parts with obvious periodicity in the time series feature vector.
[0041] For example, through spectral analysis, it is found that there is a frequency component in the frequency domain whose corresponding period is 24 hours, which may be related to the daily cycle of the power load; there is also a frequency component whose corresponding period is 7 days, which may be related to the weekly cycle of the power load.
[0042] Next, use a band-pass filter to filter the frequency domain signal. The band-pass filter can set a frequency range and only allow the frequency components within this frequency range to pass. According to the characteristics of the power grid service, determine the frequency range that matches the service cycle. For example, if the focus is on the fluctuations of the daily cycle and the weekly cycle, set the frequency range of the band-pass filter so that the frequency components corresponding to the 24-hour and 7-day cycles can pass, while other irrelevant frequency components are filtered out.
[0043] After being processed by the band-pass filter, the filtered spectral features are obtained, which only contain the periodic information that matches the service cycle and removes other interference information, enabling subsequent analysis to focus more on the fluctuation patterns related to the service.
[0044] Step S1215: Map the spectral features back to the time domain space, eliminate the dimensional differences of different computing resource nodes through adaptive normalization processing, and generate standardized dynamic load fluctuation features, where the dynamic load fluctuation features are used to characterize the load change law of computing resource nodes across time scales.
[0045] In this embodiment, this step can be implemented by methods such as inverse fast Fourier transform. For example, convert the spectral features in the frequency domain back to the time domain signal to obtain a time domain signal reflecting the periodic fluctuation of the computing power load rate.
[0046] However, different computing resource nodes may have different hardware configurations, workloads, etc., resulting in differences in the dimensionality and numerical range of their computing power load rate data. To eliminate these differences, adaptive normalization processing is required. Adaptive normalization processing will dynamically adjust the scale of the data according to the specific situation of each computing resource node. For example, for node A and node B, their computing power load rate data may have a relatively high overall value and a relatively low overall value. Adaptive normalization processing will calculate the mean and standard deviation of the data of each node, and then standardize the data according to these statistical information.
[0047] Specifically, for each data point in the time domain signal of node A, first subtract the mean of the node signal, and then divide by the standard deviation of the node signal. After such processing, the data of node A is standardized to a relatively unified scale. The same processing method is also applied to other computing resource nodes.
[0048] After adaptive normalization processing, standardized dynamic load fluctuation features are obtained. These dynamic load fluctuation features eliminate the dimensional differences between different computing resource nodes and can accurately characterize the load change law of each computing resource node across time scales. For example, through these dynamic load fluctuation features, it can be clearly seen the computing power load fluctuations of node A within the daily and weekly cycles, as well as the relative change trends compared with other nodes.
[0049] Step S122: Analyze the resource interaction and dependence relationships between each computing resource node through a graph attention network to generate resource compatibility features, where the resource compatibility features are used to quantify the hardware configuration differences and protocol interoperability of computing resource nodes between different domains.
[0050] In a multi-domain computing environment in the power grid industry, there are complex resource interaction and dependence relationships between different computing resource nodes. For example, the computing resource nodes in a substation may need to interact with the nodes in a power control center to achieve real-time power dispatching; the nodes in a power generation station need to share equipment operation data with the nodes in a monitoring station to detect faults in a timely manner. To quantify the resource compatibility between these nodes, analysis through a graph attention network is required.
[0051] First, construct a directed weighted graph based on the historical communication records between computing resource nodes. The historical communication records contain data transmission information between nodes, such as the frequency of data transmission, success rate, average latency, etc.
[0052] Step S1221: Construct a directed weighted graph based on the historical communication records between computing resource nodes. The edge weights in the directed weighted graph are jointly determined by the historical data transmission success rate and average latency between nodes.
[0053] Taking nodes A, B, and C as examples, assume that in the past period, node A transmitted data to node B 100 times, with 90 successes and an average latency of 20 milliseconds; node B transmitted data to node A 80 times, with 70 successes and an average latency of 25 milliseconds; node A transmitted data to node C 60 times, with 50 successes and an average latency of 30 milliseconds; node C transmitted data to node A 50 times, with 40 successes and an average latency of 35 milliseconds; node B transmitted data to node C 70 times, with 60 successes and an average latency of 28 milliseconds; node C transmitted data to node B 65 times, with 55 successes and an average latency of 32 milliseconds.
[0054] To determine the edge weights in the directed weighted graph, comprehensively consider the historical data transmission success rate and average latency. A simple calculation method can be adopted: edge weight = historical data transmission success rate / average latency.
[0055] For the edge from node A to node B, its edge weight = 90% / 20 milliseconds = 0.045. For the edge from node B to node A, the edge weight = 70% / 25 milliseconds = 0.028. Calculating in the same way, the edge weights of the edges from node A to node C, from node C to node A, from node B to node C, and from node C to node B are 0.0278, 0.0229, 0.0268, and 0.0234 respectively.
[0056] Construct a directed weighted graph based on these edge weights. The nodes in the graph represent computing resource nodes, the directed edges represent the data transmission directions between nodes, and the edge weights represent the efficiency and reliability of data transmission between nodes.
[0057] Step S1222: Perform random walk sampling on the directed weighted graph to generate a context node sequence for each computing resource node. The context node sequence contains neighbor nodes that have a strong dependence relationship with it.
[0058] After constructing the directed weighted graph, perform random walk sampling on the directed weighted graph. Random walk sampling is a method of randomly selecting paths in a graph. By performing multiple random walks, neighbor nodes that have a strong dependence relationship with each node can be found.
[0059] Taking node A as an example, start a random walk from node A. Each time when walking, according to the weights of the edges in the directed weighted graph, select the next node with a certain probability. For example, node A has two out-edges pointing to node B and node C respectively. The weight of the edge from node A to node B is 0.045, and the weight of the edge from node A to node C is 0.0278. Then when performing a random walk, the probability of selecting node B is 0.045 / (0.045 + 0.0278) ≈ 0.618, and the probability of selecting node C is 0.0278 / (0.045 + 0.0278) ≈ 0.382.
[0060] Perform multiple random walks (for example, 100 times), and record the sequence of nodes passed by each walk. After multiple walks, count the frequency of each node being visited. If a certain node is frequently visited, it indicates that it has a strong dependency relationship with node A. Sort the nodes that have a strong dependency relationship with node A according to the visit frequency to generate the context node sequence of node A. Suppose after random walk sampling, the context node sequence of node A is [node B, node C], which means that node B and node C have a strong resource interaction and dependency relationship with node A.
[0061] Step S1223: Input the context node sequence into the embedding layer of the graph attention network to generate a topological perception feature vector for each neighbor node, and the topological perception feature vector encodes the physical connection attributes and logical cooperation relationships between nodes.
[0062] Input the context node sequence [node B, node C] of node A into the embedding layer of the graph attention network. The role of the embedding layer is to convert node information into a low-dimensional vector representation for subsequent analysis and processing. In the embedding layer, a topological perception feature vector will be learned for each neighbor node (here are node B and node C).
[0063] For node B, the embedding layer will comprehensively consider information such as its position in the directed weighted graph, connection relationships with other nodes, and the attributes of the node itself. For example, the edge weight between node B and node A represents the data transmission efficiency and reliability between them, and this information will be incorporated into the learning of the topological perception feature vector. At the same time, the connection situation of node B with other nodes (such as node C) will also affect its topological perception feature vector. Suppose node B also has data interactions with multiple other nodes, and the weights and directions of these connections will be processed in the embedding layer to generate a topological perception feature vector that can encode the physical connection attributes and logical cooperation relationships of node B.
[0064] Specifically, the embedding layer uses a series of neural network operations to learn these feature vectors. First, a linear transformation is performed on the original features of node B (such as hardware configuration information, historical data transfer success rate, etc.) to obtain a preliminary feature representation. Then, combined with the structural information in the directed weighted graph, the importance of different connections is adjusted through an attention mechanism. The attention mechanism assigns an attention score to each connection based on information such as edge weights. The higher the score, the greater the impact of the connection on the topological perception feature vector of node B. For example, if the edge weight between node B and node A is large, then when generating the topological perception feature vector of node B, the information of node A will be given a higher weight.
[0065] After such processing, the embedding layer generates the topological perception feature vector of node B. Assume that the topological perception feature vector is a multi-dimensional vector with a length of 128, and the value of each dimension reflects some topological and attribute information of node B in the graph.
[0066] The same process is also applied to node C. The embedding layer generates the topological perception feature vector of node C according to the connection relationship between node C and other nodes (including node A and node B), edge weights, and the attributes of node C itself through linear transformation and the attention mechanism. Assume that the topological perception feature vector of node C is also a multi-dimensional vector with a length of 128, which also encodes the physical connection attributes and logical cooperation relationships of node C.
[0067] Step S1224: Perform bidirectional attention calculation on the feature vector of the current computing resource node and the topological perception feature vectors of neighbor nodes to generate a resource cooperation degree score between nodes, and the cooperation degree score reflects the hardware configuration alignment degree, protocol compatibility, and load complementarity between nodes.
[0068] Taking node A as an example, after obtaining the topological perception feature vectors of neighbor nodes B and C, it is necessary to perform bidirectional attention calculation on the feature vector of node A itself and the topological perception feature vectors of neighbor nodes to generate a resource cooperation degree score between nodes.
[0069] The feature vector of node A itself contains its hardware configuration information (such as the number of computing cores, storage space size, etc.), real-time resource status data (such as computing power load rate, storage space occupancy rate, network bandwidth utilization rate, etc.). First, for the case of node A and node B, the bidirectional attention calculation considers the information interaction in both directions.
[0070] In the forward attention calculation, the influence of the topological perception feature vector of node B on node A is concerned. By calculating the similarity between the feature vector of node A and the topological perception feature vector of node B, an attention score is obtained. For example, the dot product operation can be used to calculate the similarity, and then these similarities are converted into a probability distribution through the softmax function as the attention score, which represents the importance of the information of node B when generating the cooperation degree score between node A and node B.
[0071] In the reverse attention calculation, the influence of the feature vector of node A on node B is concerned. Similarly, the similarity between the topological perception feature vector of node B and the feature vector of node A is calculated to obtain the reverse attention score.
[0072] The forward and reverse attention scores are integrated to obtain a comprehensive attention score. Then, combined with the hardware configuration information, historical data transmission success rate, protocol type, etc. of node A and node B, the resource cooperation degree score between nodes is calculated. For example, if the hardware configurations of node A and node B are similar, the protocol types are the same, and the historical data transmission success rate is high, then their resource cooperation degree scores will be relatively high. Suppose through a series of calculations, the resource cooperation degree score between node A and node B is 0.8.
[0073] The same method is used to calculate the resource cooperation degree score between node A and node C. After calculation, suppose the resource cooperation degree score between node A and node C is 0.7. These cooperation degree scores reflect the situations in aspects such as the alignment degree of hardware configurations, protocol compatibility, and load complementarity between nodes. For example, if the computing power load of node A is high while the computing power load of node C is low, and their hardware configurations and protocols are compatible, then their load complementarity is good and the cooperation degree score will increase accordingly.
[0074] Step S1225: Dynamically weight and aggregate the topological perception feature vectors of neighbor nodes based on the cooperation degree score to generate the local compatibility feature of the current node.
[0075] After obtaining the resource cooperation degree scores between node A and neighbor nodes B and C, the topological perception feature vectors of neighbor nodes are dynamically weighted and aggregated based on these cooperation degree scores to generate the local compatibility feature of node A.
[0076] For the topological awareness feature vectors of node B and node C, weights are assigned according to their cooperation degree scores with node A. The cooperation degree score of node B with node A is 0.8, and the cooperation degree score of node C with node A is 0.7. Then, during aggregation, the weight of the topological awareness feature vector of node B is 0.8 / (0.8 + 0.7) ≈ 0.533, and the weight of the topological awareness feature vector of node C is 0.7 / (0.8 + 0.7) ≈ 0.467.
[0077] Multiply the topological awareness feature vector of node B by its weight, multiply the topological awareness feature vector of node C by its weight, and then concatenate these two weighted vectors. Assume that the topological awareness feature vectors of node B and node C are both multi-dimensional vectors of length 128. After concatenation, a vector of length 256 is obtained, which contains the information of neighbor nodes B and C, and is weighted according to the cooperation degree score, and can reflect the local compatibility situation between node A and its neighbor nodes, which is the local compatibility feature of node A.
[0078] Step S1226: Cross the local compatibility feature with the cross-domain resource distribution feature extracted by the global graph pooling layer, and suppress redundant features and enhance the cross-domain cooperation signal through a gating fusion mechanism to generate the final resource compatibility feature.
[0079] In this embodiment, the global graph pooling layer can process the entire directed weighted graph and extract features reflecting the cross-domain resource distribution. For example, the global graph pooling layer can comprehensively consider the information of all nodes, and through aggregating and statistically analyzing the feature vectors of the nodes in the directed weighted graph, obtain a cross-domain resource distribution feature vector. For example, the average computing power load rate, average storage space occupancy rate, etc. of all nodes can be calculated to reflect the resource distribution of the entire multi-domain computing environment. Assume that the cross-domain resource distribution feature vector extracted by the global graph pooling layer is a multi-dimensional vector of length 128.
[0080] Cross the local compatibility feature of node A (a vector of length 256) with the cross-domain resource distribution feature vector. This can be achieved by concatenating the two vectors in the dimension to obtain a vector of length 384.
[0081] Then, the concatenated vector is processed through a gating fusion mechanism. The gating fusion mechanism uses a gating unit that calculates a gating signal based on the values of each dimension of the concatenated vector. The gating signal is a vector of the same length as the concatenated vector, and the value of each dimension is between 0 and 1. For the values of certain dimensions in the concatenated vector, if the value of the corresponding dimension of the gating signal is close to 0, it indicates that the feature of this dimension is redundant and will be suppressed; if the value of the corresponding dimension of the gating signal is close to 1, it indicates that the feature of this dimension is important for enhancing the cross-domain collaborative signal and will be retained and enhanced.
[0082] After being processed by the gating fusion mechanism, the resource compatibility feature of node A is finally generated. This resource compatibility feature comprehensively considers the local compatibility between node A and its neighbor nodes as well as the distribution of the entire cross-domain resources, and can accurately quantify the resource compatibility of node A in a multi-domain computing environment.
[0083] Step S123: Perform a non-linear coupling analysis on the storage space occupancy rate and the network bandwidth utilization rate through a multi-layer perceptron network to generate a task adaptability feature, which is used to predict the computing throughput support ability of a computing resource node for a target service.
[0084] In the power grid industry, different target services have different requirements for the storage space and network bandwidth of computing resource nodes. To predict the computing throughput support ability of a computing resource node for a target service, it is necessary to perform a non-linear coupling analysis on the storage space occupancy rate and the network bandwidth utilization rate through a multi-layer perceptron network.
[0085] Taking node A as an example, its storage space occupancy rate is 56.25% and its network bandwidth utilization rate is 60%. These two metrics are used as inputs and fed into the multi-layer perceptron network. The multi-layer perceptron network is a feedforward neural network composed of an input layer, a hidden layer, and an output layer.
[0086] The input layer receives these two input values of the storage space occupancy rate and the network bandwidth utilization rate. Assume that the input layer has 2 neurons, corresponding to the storage space occupancy rate and the network bandwidth utilization rate respectively. The hidden layer has multiple neurons, for example, 10 neurons. In the hidden layer, each neuron performs a weighted sum on the output of the input layer and processes it through a non-linear activation function (such as the ReLU function).
[0087] For the first neuron in the hidden layer, the two output values of the input layer can be multiplied by their corresponding weights and then summed. Suppose the weight corresponding to the storage space occupancy rate is 0.6 and the weight corresponding to the network bandwidth utilization rate is 0.4. Then the result of the weighted sum is 56.25%×0.6 + 60%×0.4 = 57.75%. Then, the result of this weighted sum is processed through the ReLU function. If the result is greater than 0, it remains unchanged; if the result is less than 0, it becomes 0.
[0088] Similar calculations are also performed on other neurons in the hidden layer, and each neuron uses different weights to capture different features of the input data. After being processed by the hidden layer, a vector of length 10 is obtained, which contains non-linear combination information of the storage space occupancy rate and the network bandwidth utilization rate.
[0089] The output layer receives the output of the hidden layer and converts it into a task fitness feature. Suppose there is 1 neuron in the output layer, the output of the hidden layer can be weighted and summed, and processed through a linear activation function (such as the identity function). The finally obtained output value is the task fitness feature of node A, which can be used to predict the computing throughput support ability of node A for the target service. For example, if the task fitness feature value is high, it means that node A can better support the computing throughput requirements of the target service under the current storage space occupancy rate and network bandwidth utilization rate.
[0090] Step S124: Concatenate the dynamic load fluctuation feature, the resource compatibility feature, and the task fitness feature in tensors to generate a multi-dimensional resource feature vector for each computing resource node.
[0091] After obtaining the dynamic load fluctuation feature, the resource compatibility feature, and the task fitness feature of node A, they are concatenated in tensors to generate a multi-dimensional resource feature vector of node A.
[0092] In this embodiment, suppose the dynamic load fluctuation feature is a multi-dimensional vector of length 128, the resource compatibility feature is a multi-dimensional vector of length 384, and the task fitness feature is a vector of length 1. Then, they are concatenated in dimensions to obtain a multi-dimensional resource feature vector of length 128 + 384 + 1 = 513.
[0093] This multi-dimensional resource feature vector contains information on various aspects such as the dynamic load fluctuation situation, resource compatibility, and computing throughput support ability of node A for the target service. The same method is also applied to node B and node C to generate their respective multi-dimensional resource feature vectors.
[0094] Step S130: Based on a preset intelligent aggregation policy network, perform multi-domain resource association analysis on the multi-dimensional resource feature vectors, determine the collaborative matching weights between cross-domain computing resources, and aggregate the scattered computing resource nodes into a virtualized resource pool according to the collaborative matching weights.
[0095] After obtaining the multi-dimensional resource feature vectors of each computing resource node, it is necessary to perform multi-domain resource association analysis on these vectors based on a preset intelligent aggregation policy network to determine the collaborative matching weights between cross-domain computing resources, and aggregate the scattered computing resource nodes into a virtualized resource pool.
[0096] Step S131: Input the multi-dimensional resource feature vectors into the cross-attention encoding layer of the intelligent aggregation policy network, perform cross-domain bidirectional attention calculation on the multi-dimensional resource feature vectors of each computing resource node, and generate a similarity association matrix containing the inter-domain node feature similarities.
[0097] Input the multi-dimensional resource feature vectors of nodes A, B, C, etc. into the cross-attention encoding layer of the intelligent aggregation policy network. In this cross-attention encoding layer, cross-domain bidirectional attention calculation can be performed on the multi-dimensional resource feature vectors of each computing resource node.
[0098] Taking nodes A and B as an example, first perform forward attention calculation. Calculate the similarity between the multi-dimensional resource feature vector of node A and the multi-dimensional resource feature vector of node B. The dot product operation can be used to calculate the similarity, and then these similarities are converted into a probability distribution through the softmax function to obtain the forward attention score, which represents the importance of the information of node B when generating the association between node A and node B.
[0099] Then perform reverse attention calculation, calculate the similarity between the multi-dimensional resource feature vector of node B and the multi-dimensional resource feature vector of node A, and also obtain the reverse attention score through the softmax function.
[0100] Integrate the forward and reverse attention scores to obtain the cross-domain bidirectional attention score between node A and node B. Calculate the cross-domain bidirectional attention scores between all node pairs in the same way.
[0101] Organize the above cross-domain bidirectional attention scores into a matrix, which is the similarity association matrix containing the inter-domain node feature similarities. Suppose there are 3 nodes A, B, and C in total, then the similarity association matrix is a 3×3 matrix. The element in the i-th row and j-th column of the matrix represents the cross-domain bidirectional attention score between node i and node j, reflecting their feature similarity.
[0102] Step S132: Perform a sparsification process based on threshold filtering on the similarity association matrix, retain the association edges with similarity higher than the dynamic threshold in each computing resource node and cross-domain nodes, generate a sparsified association relationship graph, input the sparsified association relationship graph into a differentiable graph sorting network, and perform end-to-end gradient backpropagation training based on the multi-dimensional resource feature vectors of the nodes and the weights of the association edges to generate the collaborative matching weights of each cross-domain association edge.
[0103] After obtaining the similarity association matrix, perform a sparsification process based on threshold filtering on it. The dynamic threshold is determined dynamically according to the element values in the similarity association matrix. For example, the average value of all elements in the similarity association matrix can be calculated and used as the dynamic threshold.
[0104] For each element in the similarity association matrix, if its value is higher than the dynamic threshold, retain the corresponding association edge; if its value is lower than the dynamic threshold, delete the corresponding association edge. Taking the similarity association matrix of three nodes A, B, and C as an example, assume the similarity association matrix is:
[0105] |1.0 0.8 0.3|
[0106] |0.8 1.0 0.7|
[0107] |0.3 0.7 1.0|
[0108] If the dynamic threshold is 0.6, then retain the association edges between node A and node B, and between node B and node C, and delete the association edge between node A and node C to generate a sparsified association relationship graph.
[0109] Input the sparsified association relationship graph into a differentiable graph sorting network. The differentiable graph sorting network can perform end-to-end gradient backpropagation training based on the multi-dimensional resource feature vectors of the nodes and the weights of the association edges. During the training process, the differentiable graph sorting network continuously adjusts the weights of the association edges to make the associations between nodes more reasonable.
[0110] Specifically, the differentiable graph sorting network calculates a loss function based on the multi-dimensional resource feature vectors of the nodes and the current weights of the association edges. The loss function represents the difference between the current weights of the association edges and the ideal association situation. Through the gradient backpropagation algorithm, the gradient of the loss function with respect to the weights of the association edges is calculated, and then the weights of the association edges are updated according to the gradient. After multiple training iterations, the collaborative matching weights of each cross-domain association edge are finally generated.
[0111] Step S133: Update the edge weights of the sparsified association graph according to the collaborative matching weights to generate a weighted cross-domain resource association graph. Input the cross-domain resource association graph into an overlapping community discovery algorithm, and identify node clusters with stable collaboration relationships based on the principle of maximizing modularity to generate multiple candidate resource clusters.
[0112] For example, in the sparsified association graph, the original weight of the association edge between node A and node B is 0.8, and the collaborative matching weight obtained through training is 0.9. Then, update the weight of this association edge to 0.9.
[0113] After updating the edge weights, a weighted cross-domain resource association graph is generated. For example, input this cross-domain resource association graph into an overlapping community discovery algorithm. The goal of the overlapping community discovery algorithm is to identify node clusters with stable collaboration relationships based on the principle of maximizing modularity.
[0114] Modularity is an index to measure the quality of the community structure in a graph, which represents the difference between the sum of the weights of the edges within the community and the sum of the weights of the edges in the random case. The overlapping community discovery algorithm will continuously try to partition nodes into different communities to maximize the modularity.
[0115] For example, through the calculation of the algorithm, it is found that node A and node B often conduct data interactions and their collaborative matching weights are relatively high, so they are partitioned into one community; node C has a strong collaboration relationship with some nodes in another community, so it is partitioned into another community. In this way, multiple candidate resource clusters are generated.
[0116] Step S134: Perform mean pooling operations on the multi-dimensional resource feature vectors of the nodes within each candidate resource cluster to generate cluster-level resource feature vectors. Screen the target clusters according to the cosine similarity between the cluster-level resource feature vectors and the preset global resource demand template to generate a set of target resource clusters.
[0117] For each candidate resource cluster, perform mean pooling operations on the nodes within the cluster. Take the candidate resource cluster containing node A and node B as an example. The multi-dimensional resource feature vector of node A is a vector with a length of 513, and the multi-dimensional resource feature vector of node B is also a vector with a length of 513.
[0118] Add the values of the corresponding dimensions of the multi-dimensional resource feature vectors of node A and node B, and then divide by the number of nodes (here it is 2) to obtain the cluster-level resource feature vector. For example, for the first dimension, the feature vector value of node A is 0.5, and the feature vector value of node B is 0.6. Then, the value of the first dimension of the cluster-level resource feature vector is (0.5 + 0.6) / 2 = 0.55.
[0119] After obtaining the cluster-level resource feature vectors of each candidate resource cluster, calculate their cosine similarity with a preset global resource demand template. The global resource demand template is a multi-dimensional vector preset according to the business requirements of the entire power grid industry.
[0120] Cosine similarity is an index to measure the cosine value of the angle between two vectors. The closer the value is to 1, the more similar the two vectors are. For example, the cosine similarity between the cluster-level resource feature vector of candidate resource cluster 1 and the global resource demand template is 0.8, and the cosine similarity of candidate resource cluster 2 is 0.6.
[0121] Screen the target clusters according to the cosine similarity. Set a similarity threshold, such as 0.7. Screen out the candidate resource clusters with a cosine similarity higher than 0.7 to form a set of target resource clusters. Suppose among all candidate resource clusters, the cosine similarities of candidate resource cluster 1, candidate resource cluster 3, and candidate resource cluster 5 are 0.8, 0.85, and 0.72 respectively, all higher than the set threshold of 0.7, then these three candidate resource clusters are included in the set of target resource clusters.
[0122] Step S135: Perform virtualized resource encapsulation operations on each cluster in the set of target resource clusters to generate logical resource units with a unified interface protocol, and establish two-way resource redundancy channels between the logical resource units based on the collaborative matching weights to generate a virtualized resource pool with fault tolerance capabilities.
[0123] For each cluster in the set of target resource clusters, perform virtualized resource encapsulation operations. Taking candidate resource cluster 1 as an example, this cluster contains node A and node B. First, abstract and encapsulate the computing resources (such as computing power, storage space, network bandwidth, etc.) of node A and node B.
[0124] For the computing power resources of node A, integrate information such as the available number of computing cores and computing capabilities, and convert them into a standardized representation form according to certain rules. For example, the 32 computing cores of node A are abstracted into a resource unit with a specific computing power value according to its different performance indicators. Similarly, perform similar encapsulation on the storage space and network bandwidth of node A.
[0125] For node B, perform the same operations. Then, integrate the encapsulated resources of node A and node B to generate a logical resource unit with a unified interface protocol. Among them, this unified interface protocol stipulates how an external system interacts with this logical resource unit, such as how to request computing resources and how to read data in the storage space.
[0126] Next, establish bidirectional resource redundancy channels between logical resource units based on the collaborative matching weights. Extract the collaborative matching weights between each pair of logical resource units from the target resource cluster set to generate a weight connection relationship table between units. Suppose there are three logical resource units in the target resource cluster set, coming from candidate resource cluster 1, candidate resource cluster 3, and candidate resource cluster 5 respectively. After calculation and arrangement of their collaborative matching weights, the following weight connection relationship table between units is formed:
[0127] The collaborative matching weight between logical resource unit 1 and logical resource unit 2 is 0.8; the collaborative matching weight between logical resource unit 1 and logical resource unit 3 is 0.7; the collaborative matching weight between logical resource unit 2 and logical resource unit 3 is 0.6.
[0128] Perform the top-K weight screening for each logical resource unit according to this weight connection relationship table. Suppose K = 2. For logical resource unit 1, its collaborative matching weights with logical resource unit 2 and logical resource unit 3 are relatively high. Take logical resource unit 2 and logical resource unit 3 as its redundant associated units to generate a redundant associated unit list corresponding to logical resource unit 1. Similarly, generate redundant associated unit lists for logical resource unit 2 and logical resource unit 3 respectively.
[0129] Verify the protocol handshake between the units in the redundant associated unit list and the current logical resource unit. Take logical resource unit 1 and its redundant associated unit logical resource unit 2 as an example. By sending a specific verification message, check whether they can communicate following a unified interface protocol. If the verification is successful, generate an interoperable redundant unit pair. Perform resource mirroring synchronization operations on each redundant unit pair, copy the status information, data, etc. of the current logical resource unit to the redundant associated unit to generate redundant resource copies with consistent real-time states.
[0130] Configure priority routing identifiers for each redundant resource copy based on a preset failover strategy. For example, for the redundant resource copies logical resource unit 2 and logical resource unit 3 of logical resource unit 1, according to their collaborative matching weights with logical resource unit 1 and other factors, configure a priority of 1 for logical resource unit 2 and a priority of 2 for logical resource unit 3. Generate a redundant channel configuration table with priority markings and inject this table into the load balancer of the virtualized resource pool. The load balancer generates dynamic traffic distribution rules according to this redundant channel configuration table. When there is a task request, first assign the task to the normally working logical resource unit. If this unit fails, switch the task to the redundant resource copy in the order of priority.
[0131] During the operation of the virtualized resource pool, the heartbeat signals of each logical resource unit are monitored in real time. The heartbeat signal is a signal regularly sent by the logical resource unit to indicate its normal operation. By monitoring these signals, a unit health status matrix is generated. For example, each element in the matrix represents the health status of a logical resource unit, where 1 indicates normal and 0 indicates a fault. When a heartbeat signal timeout anomaly is detected, it means that a certain logical resource unit has failed. According to the redundant channel configuration table, the task flow of the faulty unit is switched to the redundant resource copy with the highest priority, and an updated resource scheduling path is generated. Suppose logical resource unit 1 fails. The load balancer, based on the redundant channel configuration table, switches the tasks originally assigned to logical resource unit 1 to logical resource unit 2 with a priority of 1 and updates the resource scheduling path to ensure that the tasks can continue to execute normally. In this way, a virtualized resource pool with fault tolerance is generated, improving the reliability and stability of the entire power grid computing environment.
[0132] Step S140: Perform dynamic resource adaptation processing on the aggregated resources in the virtualized resource pool to generate a resource scheduling topology structure that matches the target service requirements, and deploy the resource scheduling topology structure to the multi-domain computing environment to trigger resource collaboration services.
[0133] Step S141: Input the real-time task characteristics of the target service requirements into the third candidate subset of the virtualized resource pool, and perform transmission path traversal simulation based on the network bandwidth utilization characteristics of the logical resource units in the third candidate subset to generate the end-to-end delay prediction value from each logical resource unit to the service access point.
[0134] In the power grid industry, different target services have different real-time task characteristics. For example, the power failure repair service needs to obtain relevant data of the fault point in real time, and has extremely high requirements for the real-time data transmission; while for the regular inspection data upload service of power equipment, the requirements for real-time are relatively low. Suppose the current target service is power failure repair, and its real-time task characteristics include that the data transmission delay is required to be within 100 milliseconds.
[0135] The third candidate subset of the virtualized resource pool is a set of a part of the logical resource units obtained after being screened and processed in the previous steps. For each logical resource unit in the third candidate subset, perform transmission path traversal simulation based on its network bandwidth utilization characteristics.
[0136] Taking the logical resource unit A as an example, it is located at a computing node of a substation, and the service access point is the server of the fault repair command center. The network bandwidth utilization rate of the logical resource unit A is 60%, indicating its current network bandwidth usage. In the transmission path traversal simulation, all possible transmission paths from the logical resource unit A to the service access point are considered. These paths may pass through multiple network nodes and links, and each node and link has its own transmission delay and bandwidth limitation.
[0137] During the simulation process, according to the network bandwidth utilization rate of the logical resource unit A and the status of nodes and links on each transmission path, the transmission delay of data on each path is calculated. For example, path 1 passes through 3 network nodes, and the processing delays of each node are 10 milliseconds, 15 milliseconds, and 20 milliseconds respectively, and the transmission delay of the link is 30 milliseconds. Since the network bandwidth utilization rate of the logical resource unit A is relatively high, it may cause a reduction in the effective bandwidth of the link, thereby increasing the transmission delay. Assuming that an additional 10 milliseconds of delay is added, then the total delay of path 1 is 10 + 15 + 20 + 30 + 10 = 85 milliseconds. In the same way, the delays of all possible paths are calculated, and the minimum value among them is taken as the end-to-end delay prediction value from the logical resource unit A to the service access point. Such calculations are performed for all logical resource units in the third candidate subset to obtain the end-to-end delay prediction values from each logical resource unit to the service access point.
[0138] Step S142: Compare the end-to-end delay prediction values with the delay upper limit in the real-time task characteristics node by node, and eliminate the logical resource units whose delay prediction values continuously exceed the upper limit to generate a fourth candidate subset with compliant delays.
[0139] Compare the end-to-end delay prediction value of each logical resource unit with the delay upper limit in the target service real-time task characteristics node by node. Taking the power fault repair service as an example, the delay upper limit is 100 milliseconds.
[0140] For the logical resource unit A in the third candidate subset, its end-to-end delay prediction value is 85 milliseconds, which is less than the delay upper limit and meets the requirements; the end-to-end delay prediction value of the logical resource unit B is 110 milliseconds, which exceeds the delay upper limit. To ensure the accuracy of the screening, multiple consecutive comparisons are made. Assuming that 3 consecutive delay predictions and comparisons are made, and the delay prediction values of the logical resource unit B exceed 100 milliseconds in all these 3 comparisons, then the logical resource unit B is eliminated from the third candidate subset.
[0141] After such a screening process, all the logical resource units whose delay prediction values continuously exceed the upper limit are eliminated, and the remaining logical resource units form a fourth candidate subset with compliant delays, and the logical resource units included in it can meet the real-time requirements of the target service in terms of data transmission delay.
[0142] Step S143: Perform a sliding window matching between the computing power load rate fluctuation characteristics of the logical resource units in the fourth candidate subset and the floating-point operation requirements of the compute-intensive task characteristics, calculate the cumulative distribution function of the computing power supply margin within each window, and generate a fifth candidate subset that meets the task peak demand.
[0143] In the power grid industry, some services belong to compute-intensive tasks, such as power flow calculations and fault analysis in power systems. The above tasks require a large number of floating-point operations. Suppose the target service is to perform a large-scale power flow calculation of the power system. The floating-point operation requirements of its compute-intensive task characteristics have different requirements at different time periods. For example, at the beginning stage, 10 million floating-point operations per second are required, at the middle stage, 15 million floating-point operations per second are required, and at the end stage, 8 million floating-point operations per second are required.
[0144] For each logical resource unit in the fourth candidate subset, perform a sliding window matching between its computing power load rate fluctuation characteristics and the floating-point operation requirements of the compute-intensive task characteristics. Taking logical resource unit C as an example, its computing power load rate fluctuation characteristics record the change of the computing power load rate over time within a period of time. Set the size of a sliding window, for example, 10 minutes, and slide this sliding window on the time axis, each time sliding by one time step, for example, 1 minute.
[0145] Within each window, calculate its current computing power supply capacity according to the computing power load rate fluctuation characteristics of logical resource unit C, and then compare it with the floating-point operation requirements of the compute-intensive task characteristics corresponding to this window to calculate the computing power supply margin. For example, within a certain 10-minute window, the computing power supply capacity of logical resource unit C is 12 million floating-point operations per second, and the floating-point operation requirements of the compute-intensive task characteristics corresponding to this window are 10 million floating-point operations per second. Then the computing power supply margin is 2 million floating-point operations per second.
[0146] Statistically analyze the computing power supply margin within each window and calculate its cumulative distribution function. The cumulative distribution function represents the probability that the computing power supply margin is less than or equal to a certain value. For example, the cumulative distribution function shows that the probability that the computing power supply margin is less than or equal to 1 million floating-point operations per second is 0.2, and the probability that it is less than or equal to 2 million floating-point operations per second is 0.5, etc.
[0147] Set a threshold for the peak demand of a task. For example, it is required that the computing power supply margin can meet the maximum floating-point operation demand of the task for at least 90% of the time. According to the cumulative distribution function, filter out the logical resource units that meet this threshold requirement. These logical resource units form a fifth candidate subset that meets the peak demand of the task. The logical resource units in this fifth candidate subset can better handle the computing-intensive task requirements of the target service in terms of computing power.
[0148] Step S144: Perform spatial continuity analysis on the storage space occupancy rate characteristics of the logical resource units in the fifth candidate subset, identify the storage fragmentation areas, and perform block remapping operations on the virtual storage volume to generate a sixth candidate subset with optimized storage continuity.
[0149] For each logical resource unit in the fifth candidate subset, perform spatial continuity analysis on its storage space occupancy rate characteristics. Taking the logical resource unit D as an example, there may be fragmentation problems in its storage space occupancy. Storage fragmentation means that stored data is scattered in different storage blocks, resulting in reduced data read and write efficiency.
[0150] By analyzing the storage space occupancy rate characteristics of the logical resource unit D, identify the storage fragmentation areas. It can be identified by checking the usage of storage blocks and the distribution of data. For example, it is found that there are multiple small free storage blocks scattered throughout the storage space, and these small free storage blocks cannot meet the storage requirements of a larger data block, indicating the existence of storage fragmentation areas.
[0151] For the identified storage fragmentation areas, perform block remapping operations on the virtual storage volume. First, organize and migrate the data in the scattered storage blocks. For example, merge some small and scattered data blocks into a larger continuous storage area. Then, update the mapping table of the virtual storage volume to record the new location information of the data blocks. In this way, the originally scattered data blocks are reorganized into continuous storage areas, improving the utilization rate of the storage space and the data read and write efficiency.
[0152] Perform such spatial continuity analysis and block remapping operations on all logical resource units in the fifth candidate subset. After processing, the logical resource units that meet the storage continuity requirements form a sixth candidate subset with optimized storage continuity. The logical resource units in this sixth candidate subset are optimized in terms of storage space and can better support the data storage requirements of the target service.
[0153] Step S145: Construct an initial resource topology graph based on the sixth candidate subset. The nodes of the initial resource topology graph represent the logical resource units with optimized storage continuity, and the edges represent the historical average bandwidth of the cross-domain communication links.
[0154] Construct an initial resource topology graph based on the logical resource units in the sixth candidate subset. The nodes in the graph are the logical resource units storing continuous optimizations, such as logical resource unit E, logical resource unit F, etc. The edges represent the historical average bandwidth of cross-domain communication links.
[0155] For the cross-domain communication link between logical resource unit E and logical resource unit F, by collecting the historical communication data between them, calculate the historical average bandwidth of this link. Suppose after statistics, the historical average bandwidth between logical resource unit E and logical resource unit F is 500 Mbps, then in the initial resource topology graph, the edge connecting logical resource unit E and logical resource unit F is marked as 500 Mbps.
[0156] In the same way, determine the historical average bandwidth of the cross-domain communication links between all logical resource units in the sixth candidate subset and mark them in the graph, thus constructing the initial resource topology graph, which intuitively shows the communication relationship and bandwidth situation between the logical resource units storing continuous optimizations.
[0157] Step S146: Input the initial resource topology graph into the multi-objective optimization module to synchronously optimize the computing resource utilization rate, storage balance degree, and network load variance, and iteratively adjust the edge weights and node clustering centers through the gradient descent algorithm to generate an intermediate topology graph with optimized edge weights.
[0158] Input the initial resource topology graph into the multi-objective optimization module. The goal of this module is to synchronously optimize the computing resource utilization rate, storage balance degree, and network load variance.
[0159] The computing resource utilization rate refers to the proportion of the computing resources of the logical resource unit that are actually used. The storage balance degree refers to the balance degree of the storage space usage among various logical resource units. The network load variance refers to the degree of dispersion of the network loads of the cross-domain communication links.
[0160] In the multi-objective optimization module, use the gradient descent algorithm to iteratively adjust the edge weights and node clustering centers. The gradient descent algorithm is an optimization algorithm for finding the minimum value of a function. First, define an objective function that comprehensively considers the computing resource utilization rate, storage balance degree, and network load variance. For example, the objective function can be expressed as: objective function value = computing resource utilization rate penalty term + storage balance degree penalty term + network load variance penalty term.
[0161] In each iteration, calculate the gradients of the objective function with respect to the edge weights and the node clustering centers. The gradient represents the rate of change and the direction of the objective function at the current point. According to the direction of the gradient, adjust the values of the edge weights and the node clustering centers so that the value of the objective function gradually decreases. For example, if the penalty term for computing resource utilization is large, it indicates that the computing resource utilization is low. By adjusting the edge weights and the node clustering centers, try to allocate tasks to the logical resource units with low computing resource utilization to improve the overall computing resource utilization.
[0162] After multiple iterations, when the value of the objective function converges to a small value, stop the iteration and generate an intermediate topology graph with optimized edge weights. In this intermediate topology graph, the edge weights and the node clustering centers are optimized and adjusted, so that the computing resource utilization, the storage balance degree, and the network load variance are all better optimized.
[0163] Step S147: Hierarchically cluster the nodes of the intermediate topology graph according to the fluctuation characteristics of the computing power load rate, generate a hierarchical cascade structure including a core computing layer, a buffer storage layer, and an edge access layer, and allocate a priority routing protocol for the bidirectional data channels between the layers.
[0164] Hierarchically cluster the nodes of the intermediate topology graph with optimized edge weights according to the fluctuation characteristics of the computing power load rate. Hierarchical clustering is a clustering method that gradually merges data points into different levels.
[0165] According to the fluctuation characteristics of the computing power load rate of the nodes, classify the nodes into different categories. For example, classify the nodes with small fluctuations in the computing power load rate and strong computing capabilities into one category, and these nodes can form the core computing layer; classify the nodes with large storage space and the ability to buffer data into one category to form the buffer storage layer; classify the nodes located at the network edge and responsible for connecting to external service access points into one category to form the edge access layer.
[0166] Generate a hierarchical cascade structure including a core computing layer, a buffer storage layer, and an edge access layer. In this hierarchical cascade structure, the core computing layer is mainly responsible for processing compute-intensive tasks, the buffer storage layer is responsible for storing intermediate data and temporary data, and the edge access layer is responsible for data interaction with external service access points.
[0167] Allocate a priority routing protocol for the bidirectional data channels between the layers. For example, for the data channel between the core computing layer and the buffer storage layer, since the tasks of the core computing layer have high requirements for data real-time, allocate a high-priority routing protocol for this channel to ensure fast data transmission. For the data channel between the buffer storage layer and the edge access layer, allocate an appropriate priority routing protocol according to the type and importance of the data. In this way, by reasonably allocating the priority routing protocol, the efficiency and reliability of data transmission are improved.
[0168] Step S148: Embed a resource status tracking agent in the hierarchical cascade structure to collect the task execution latency and resource consumption rate of each layer node in real time, and generate a topology health status indicator.
[0169] Embed a resource status tracking agent in the hierarchical cascade structure. The resource status tracking agent is a software module that can monitor the running status of each layer node in real time.
[0170] For the nodes in the core computing layer, the resource status tracking agent will collect their task execution latency, that is, the time taken from the start of the task to the completion of the task, and the resource consumption rate, such as CPU utilization, memory usage, etc. For the nodes in the buffer storage layer, collect their data read / write latency and storage utilization rate. For the nodes in the edge access layer, collect the end-to-end delay between them and the external service access point and the network bandwidth utilization rate.
[0171] Based on the collected data, generate a topology health status indicator. For example, define a comprehensive indicator that comprehensively considers factors such as the task execution latency and resource consumption rate of each layer node. This indicator can be calculated by weighted summation. For example, topology health status indicator = weight of core computing layer task execution latency × core computing layer task execution latency + weight of buffer storage layer storage utilization rate × buffer storage layer storage utilization rate + weight of edge access layer end-to-end delay × edge access layer end-to-end delay, etc.
[0172] By monitoring and calculating the topology health status indicator in real time, the running status of the hierarchical cascade structure can be understood in a timely manner.
[0173] Step S149: When the latency or consumption rate of any layer in the topology health status indicator exceeds the dynamic threshold, trigger a local topology reconstruction, re-execute the edge weight adjustment and node clustering according to the resource status of the current sixth candidate subset, and generate an updated resource scheduling topology structure.
[0174] Set a dynamic threshold to judge whether the topology health status indicator is normal. The dynamic threshold is dynamically adjusted according to the historical operation data of the system and the current business requirements. For example, for the task execution latency of the core computing layer, according to the average latency in the past period of time and the real-time requirement of the business, set a dynamic threshold of 500 milliseconds; for the storage utilization rate of the buffer storage layer, set a dynamic threshold of 80%; for the end-to-end delay of the edge access layer, set a dynamic threshold of 200 milliseconds.
[0175] When the latency or consumption rate of any layer in the topology health status indicator exceeds the dynamic threshold, local topology reconstruction is triggered. Suppose the task execution latency of a node in the core computing layer exceeds the dynamic threshold of 500 milliseconds, indicating that there may be a bottleneck in the computing resources of this node and it cannot meet the business requirements. At this time, it is necessary to re - execute edge weight adjustment and node clustering according to the resource status of the current sixth candidate subset.
[0176] First, re - collect the real - time resource status data of the logical resource units in the sixth candidate subset, including computing power load rate, storage space occupancy rate, network bandwidth utilization rate, etc. These data reflect the actual operation of each current logical resource unit.
[0177] Next, reconstruct the initial resource topology graph again. Since the resource status has changed, the historical average bandwidth of the cross - domain communication links between logical resource units may also be different, so it is necessary to recalculate and update the weights of the edges. For example, the historical average bandwidth between logical resource unit A and logical resource unit B was originally 500 Mbps, but due to the change in the current network load, after recalculation, it becomes 450 Mbps. Then, in the new initial resource topology graph, the weight of the edge connecting these two nodes is updated to 450 Mbps.
[0178] Then, input the new initial resource topology graph into the multi - objective optimization module. Similar to before, synchronously optimize the computing resource utilization rate, storage balance degree, and network load variance. Iteratively adjust the edge weights and node clustering centers through the gradient descent algorithm. During the iteration process, according to the new resource status data, recalculate the gradients of the objective function with respect to the edge weights and node clustering centers. For example, if the computing power load rate of a certain node suddenly increases, resulting in an increase in the penalty term of the computing resource utilization rate, then the gradient descent algorithm will try to adjust the edge weights and node clustering centers to allocate tasks to other nodes with lower computing power load rates to improve the overall computing resource utilization rate.
[0179] After multiple iterations, when the value of the objective function converges to a smaller value, stop the iteration and generate a new intermediate topology graph with optimized edge weights.
[0180] After that, hierarchically cluster the nodes of the new intermediate topology graph according to the characteristics of the computing power load rate fluctuations. Due to the change in the resource status, the classification of nodes may change. For example, a certain node that originally belonged to the core computing layer may be re - classified into the buffer storage layer or the edge access layer due to the increase in its computing power load rate fluctuations. Regenerate the hierarchical cascade structure including the core computing layer, buffer storage layer, and edge access layer, and re - allocate the priority routing protocol for the bidirectional data channels between each layer.
[0181] Finally, an updated resource scheduling topology structure is generated, which takes into account the real-time resource status of each current logical resource unit, can better meet business requirements, and improve the performance and reliability of the system.
[0182] Step S150: Deploy the resource scheduling topology structure to a multi-domain computing environment to trigger resource collaboration services.
[0183] Step S151: Convert the hierarchical cascading structure of the resource scheduling topology structure into a cross-domain resource orchestration instruction set, where the instruction set includes the virtual machine deployment coordinates of each layer node, the storage volume mapping relationship, and the priority routing table.
[0184] Convert the hierarchical cascading structure of the updated resource scheduling topology structure to generate a cross-domain resource orchestration instruction set. For the nodes in the core computing layer, determine their virtual machine deployment coordinates. The virtual machine deployment coordinates specify on which physical computing node the virtual machine should be deployed and its specific location on that node. For example, if a virtual machine in the core computing layer needs to be deployed on the 3rd computing core of physical node P1 and use a specific memory area on that node, then the virtual machine deployment coordinates record these details.
[0185] At the same time, determine the storage volume mapping relationship. The storage volume mapping relationship defines how the virtual machine connects to and uses the storage resources. For example, if a virtual machine in the core computing layer needs to access the storage volume V3 stored on the buffer storage layer node S2, then the storage volume mapping relationship clarifies this mapping.
[0186] For the bidirectional data channels between each layer, generate a priority routing table according to the previously assigned priority routing protocol. The priority routing table records the priority information of different data channels and is used to guide data transmission. For example, if the data channel priority between the core computing layer and the buffer storage layer is high, then this information will be clearly recorded in the priority routing table.
[0187] Integrate these virtual machine deployment coordinates, storage volume mapping relationship, and priority routing table together to form a cross-domain resource orchestration instruction set.
[0188] Step S152: Inject the priority routing table into the forwarding tables of each domain border router through the policy distribution interface of the software-defined network controller to generate a cross-domain bandwidth guarantee channel.
[0189] Through the policy distribution interface of the software-defined network controller, inject the priority routing table into the forwarding tables of each domain border router. The software-defined network controller is a centralized network management device that can send configuration information to each network device through the policy distribution interface.
[0190] Send the information in the priority routing table, such as the priorities of different data channels, bandwidth allocation, etc., to each domain border router. After receiving this information, each domain border router updates it to its own forwarding table. For example, after a certain domain border router receives the information that the data channel priority between the core computing layer and the buffer storage layer is high, it will allocate higher priority and more bandwidth resources for this channel in the forwarding table.
[0191] In this way, cross-domain bandwidth guarantee channels are generated. These channels can ensure that high-priority data is preferentially processed during transmission, avoiding data transmission delays caused by network congestion.
[0192] Step S153: Distribute the virtual machine deployment coordinates and storage volume mapping relationship to the target computing resource nodes through the resource scheduling engine of the cloud management platform, and generate virtualized resource instances distributed according to the topological structure.
[0193] Through the resource scheduling engine of the cloud management platform, distribute the virtual machine deployment coordinates and storage volume mapping relationship to the target computing resource nodes. The cloud management platform is a platform for managing cloud computing resources, and its resource scheduling engine is responsible for sending resource allocation and deployment information to each computing resource node.
[0194] For the virtual machine deployment coordinates, the resource scheduling engine sends them to the corresponding physical computing nodes. After receiving the information, the physical computing nodes create virtual machine instances at the specified locations according to the instructions of the virtual machine deployment coordinates. For example, after receiving the information that a virtual machine needs to be created on the 3rd computing core of physical node P1, physical node P1 will create the virtual machine as required.
[0195] For the storage volume mapping relationship, the resource scheduling engine sends it to the relevant storage nodes and virtual machine nodes. After receiving the information, the storage nodes configure the access permissions and mapping relationships of the storage volumes; after receiving the information, the virtual machine nodes access the corresponding storage volumes according to the storage volume mapping relationship.
[0196] In this way, virtualized resource instances distributed according to the topological structure are generated. These instances are distributed according to the hierarchical cascading structure of the resource scheduling topology and can work better together to meet business requirements.
[0197] Step S154: During the operation of the virtualized resource instances, continuously collect the CPU utilization rate of the core computing layer, the IO throughput of the buffer storage layer, and the end-to-end delay data of the edge access layer through the resource status tracking agent, and generate a real-time resource load matrix.
[0198] During the operation of virtualized resource instances, the resource status tracking agent continuously collects relevant data of nodes at each layer. For the core computing layer, CPU utilization data is collected. The CPU utilization reflects the usage of computing resources of nodes in the core computing layer. For example, the CPU utilization of a certain node in the core computing layer is collected every 1 minute and recorded.
[0199] For the buffer storage layer, IO throughput data is collected. The IO throughput represents the data read / write speed of nodes in the buffer storage layer. Similarly, the IO throughput of a certain node in the buffer storage layer is collected every 1 minute.
[0200] For the edge access layer, end-to-end delay data is collected. The end-to-end delay reflects the data transmission delay between nodes in the edge access layer and external service access points. Also, the end-to-end delay of a certain node in the edge access layer is collected every 1 minute.
[0201] The collected data is sorted to generate a real-time resource load matrix. Each row of the matrix represents a node, each column represents a time point, and the elements in the matrix are data such as the CPU utilization, IO throughput, or end-to-end delay of the node at the corresponding time point. For example, the element in the first row and first column of the matrix represents the CPU utilization of node 1 in the core computing layer at the 1st minute.
[0202] Step S155: Input the real-time resource load matrix into the overload determination model. When the CPU utilization of the core computing layer continuously exceeds the first threshold and the IO throughput of the buffer storage layer is lower than the second threshold, a calculation-storage resource imbalance alarm is generated.
[0203] Input the real-time resource load matrix into the overload determination model. The overload determination model is a pre-trained model that can determine whether there is a resource imbalance in the system based on the real-time resource load matrix.
[0204] Set the first threshold and the second threshold. For example, the first threshold is 80%, and the second threshold is 100 MB / s. When the CPU utilization of the core computing layer continuously exceeds 80%, it indicates that the computing resources of the core computing layer may be overloaded; at the same time, when the IO throughput of the buffer storage layer is lower than 100 MB / s, it indicates that the data read / write ability of the buffer storage layer is not fully utilized, and there may be a problem of calculation-storage resource imbalance.
[0205] When the overload determination model detects that the CPU utilization of the core computing layer continuously exceeds the first threshold and the IO throughput of the buffer storage layer is lower than the second threshold, a calculation-storage resource imbalance alarm is generated. The alarm information will be promptly notified to the system administrator so that corresponding measures can be taken for adjustment.
[0206] Step S156: Trigger a resource rescheduling request based on the calculated storage resource imbalance alarm, and re-enter the current real-time resource load matrix and the target service requirements into the dynamic resource adaptation processing module to generate a rescheduled resource scheduling topology structure.
[0207] Trigger a resource rescheduling request based on the calculated storage resource imbalance alarm. Re-enter the current real-time resource load matrix and the target service requirements into the dynamic resource adaptation processing module.
[0208] The dynamic resource adaptation processing module will perform operations such as resource screening, topology graph construction, optimization, and clustering again according to the previous steps. For example, according to the data in the real-time resource load matrix, re-screen the logical resource units that meet the requirements of latency, computing power, and storage, construct a new initial resource topology graph, and then perform multi-objective optimization and hierarchical clustering.
[0209] After a series of processes, a rescheduled resource scheduling topology structure is generated, which takes into account the current resource imbalance situation and the target service requirements, and can better balance the use of computing and storage resources.
[0210] Step S157: Use the hot migration engine to perform a lossless migration of the task status and data cache of the virtualized resource instance according to the rescheduled topology structure, and update the priority routing table of the cross-domain bandwidth guarantee channel after the migration is completed to generate an updated version of the resource collaboration service.
[0211] Step S157-1: Extract the list of edge access layer nodes in the hierarchical cascade structure of the rescheduled resource scheduling topology structure to generate a new set of access point addresses.
[0212] Extract the list of edge access layer nodes from the hierarchical cascade structure of the rescheduled resource scheduling topology structure. Each edge access layer node has its corresponding access point address. Organize these access point addresses into a set to generate a new set of access point addresses. For example, after rescheduling, there are nodes E1, E2, and E3 in the edge access layer, and their access point addresses are IP1, IP2, and IP3 respectively. Then the new set of access point addresses is {IP1, IP2, IP3}.
[0213] Step S157-2: Compare the differences between the new set of access point addresses and the addresses in the historical routing table to identify newly added access points and failed access points, and generate a routing change instruction set.
[0214] Compare the new set of access point addresses with the addresses in the historical routing table. The historical routing table records the previous access point addresses and routing information. By comparison, find out the newly added access point addresses and the failed access point addresses.
[0215] For example, if the set of access point addresses in the historical routing table is {IP0, IP1, IP4} and the set of new access point addresses is {IP1, IP2, IP3}, then the newly added access point addresses are {IP2, IP3}, and the failed access point addresses are {IP0, IP4}.
[0216] Based on the identified newly added access points and failed access points, a routing change instruction set is generated. The routing change instruction set includes instructions to delete the routing entries of the failed access points and add the routing entries of the newly added access points.
[0217] Step S157-3: Through the routing update interface of the software-defined network controller, the routing change instruction set is sent to each domain border router, the routing entries corresponding to the failed access points are deleted, and guaranteed bandwidth is allocated for the newly added access points.
[0218] Through the routing update interface of the software-defined network controller, the routing change instruction set is sent to each domain border router. After receiving the instruction, each domain border router first deletes the routing entries corresponding to the failed access points. For example, the routing entries related to IP0 and IP4 in the historical routing table are deleted.
[0219] Then, guaranteed bandwidth is allocated for the newly added access points. According to the service requirements and network resource conditions, a certain amount of bandwidth resources is allocated for each newly added access point to ensure its normal operation. For example, 200Mbps of guaranteed bandwidth is allocated for IP2 and IP3 respectively.
[0220] Step S157-4: After the routing update is completed, through the resource status tracking agent in the edge access layer, verify whether the end-to-end delay of the newly added access points meets the requirements of the real-time task characteristics, and generate a delay verification result.
[0221] After the routing update is completed, through the resource status tracking agent in the edge access layer, verify the end-to-end delay of the newly added access points. The resource status tracking agent will collect the end-to-end delay data between the newly added access points and the external service access points.
[0222] Compare the collected delay data with the requirements of the real-time task characteristics. For example, if the real-time task characteristics require the end-to-end delay to be within 200 milliseconds, then check whether the end-to-end delay of each newly added access point is less than 200 milliseconds.
[0223] Based on the comparison result, generate a delay verification result. The delay verification result records the information on whether the delay of each newly added access point meets the requirements.
[0224] Step S157-5: When there are abnormal access points in the delay verification result, re-trigger the dynamic resource adaptation processing module, replace the abnormal nodes from the sixth candidate subset, and generate a secondary rescheduling topology structure.
[0225] When there are abnormal access points in the delay verification results, it indicates that the end-to-end delays of these access points do not meet the requirements of real-time task characteristics and need to be adjusted. The dynamic resource adaptation processing module is re-triggered.
[0226] Select a suitable node from the sixth candidate subset to replace the abnormal node. For example, if the end-to-end delay of the newly added access point IP2 exceeds 200 milliseconds, then select a node from the sixth candidate subset that can meet the delay requirement to replace IP2.
[0227] The dynamic resource adaptation processing module will re-perform operations such as resource screening, topology graph construction, optimization, and clustering according to the new node situation, and generate a secondary rescheduling topology structure.
[0228] Step S157-6: Perform partial migration on the edge access layer nodes of the secondary rescheduling topology structure again through the live migration engine until the delay verification results of all access points meet the standards, and generate a final stable resource collaboration service version.
[0229] Perform partial migration on the edge access layer nodes of the secondary rescheduling topology structure again through the live migration engine. The live migration engine can migrate virtual machines and data from one node to another without interrupting the service.
[0230] Perform delay verification on the migrated access points again. If there are still abnormal access points, continue to repeat the above steps, replace the abnormal nodes from the sixth candidate subset, generate a new rescheduling topology structure, and perform migration and verification.
[0231] Until the delay verification results of all access points meet the standards, generate a final stable resource collaboration service version. The resource collaboration service of this resource collaboration service version can better meet the business requirements and improve the performance and reliability of the system.
[0232] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of a multi-domain computing resource aggregation system 100 based on a virtualized user network that can implement the inventive concept provided by some embodiments of the present invention. For example, the processor 120 can be used on the multi-domain computing resource aggregation system 100 based on a virtualized user network and is used to execute the functions in the present invention.
[0233] The multi-domain computing resource aggregation system 100 based on a virtualized user network can be a general-purpose server or a special-purpose server, and both can be used to implement the multi-domain computing resource aggregation method based on a virtualized user network of the present invention. Although only one server is shown in the present invention, for convenience, the functions described in the present invention can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0234] For example, the multi-domain computing resource aggregation system 100 based on a virtualized user network may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROMs, or RAMs, or any combination thereof. Exemplarily, the multi-domain computing resource aggregation system 100 based on a virtualized user network may further include program instructions stored in a ROM, a RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present invention may be implemented according to these program instructions. The multi-domain computing resource aggregation system 100 based on a virtualized user network further includes an input / output (I / O) interface 150 between the computer and other input / output devices.
[0235] For ease of illustration, only one processor is described in the multi-domain computing resource aggregation system 100 based on a virtualized user network. However, it should be noted that the multi-domain computing resource aggregation system 100 based on a virtualized user network in the present invention may further include multiple processors. Therefore, the steps performed by one processor described in the present invention may also be jointly performed or separately performed by multiple processors. For example, if the processor of the multi-domain computing resource aggregation system 100 based on a virtualized user network performs step A and step B, it should be understood that step A and step B may also be jointly performed by two different processors or separately performed in one processor. For example, the first processor performs step A, the second processor performs step B, or the first processor and the second processor jointly perform steps A and B.
[0236] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above multi-domain computing resource aggregation method based on a virtualized user network is implemented.
[0237] It should be noted that, in order to simplify the presentation of the disclosure of the present invention and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are incorporated into one embodiment, drawing, or description thereof.
Claims
1. A multi-domain computing resource aggregation method based on a virtualized user network, characterized in that The method includes: Collecting a real-time resource status data set of computing resource nodes deployed dispersedly in a multi-domain environment, where the real-time resource status data set includes the computing power load rate, storage space occupancy rate, and network bandwidth utilization rate of each computing resource node; Invoking a virtualized resource mapping model to perform cross-domain resource feature extraction processing on the real-time resource status data set, generating a multi-dimensional resource feature vector for each computing resource node, where the multi-dimensional resource feature vector includes dynamic load fluctuation features, resource compatibility features, and task adaptability features; Based on a preset intelligent aggregation policy network, performing multi-domain resource correlation analysis processing on the multi-dimensional resource feature vector, determining the collaborative matching weights between cross-domain computing resources, and aggregating the dispersed computing resource nodes into a virtualized resource pool according to the collaborative matching weights; Performing dynamic resource adaptation processing on the aggregated resources in the virtualized resource pool, generating a resource scheduling topology structure that matches the target service requirements, and deploying the resource scheduling topology structure to a multi-domain computing environment to trigger resource collaborative services.
2. The multi-domain computing resource aggregation method based on a virtualized user network according to claim 1, wherein The invoking a virtualized resource mapping model to perform cross-domain resource feature extraction processing on the real-time resource status data set and generating a multi-dimensional resource feature vector for each computing resource node includes: Inputting the real-time resource status data set into the feature encoding layer of the virtualized resource mapping model, and extracting the periodic fluctuation pattern of the computing power load rate through a temporal convolutional network to generate dynamic load fluctuation features; Analyzing the resource interaction dependence relationship between each computing resource node through a graph attention network to generate resource compatibility features, where the resource compatibility features are used to quantify the hardware configuration differences and protocol interoperability between computing resource nodes in different domains; Performing non-linear coupling analysis on the storage space occupancy rate and the network bandwidth utilization rate through a multi-layer perceptron network to generate task adaptability features, where the task adaptability features are used to predict the computing throughput support ability of the computing resource node for the target service; Performing tensor splicing on the dynamic load fluctuation features, resource compatibility features, and task adaptability features to generate a multi-dimensional resource feature vector for each computing resource node.
3. The multi-domain computing resource aggregation method based on a virtualized user network according to claim 2, wherein The extracting the periodic fluctuation pattern of the computing power load rate through a temporal convolutional network and generating dynamic load fluctuation features includes: Performing periodic segmentation on the computing power load rate according to a preset time window to obtain multiple load rate time series segments, where each load rate time series segment contains load rate sampling data with continuous timestamps; Inputting the load rate time series segment into a dilated causal convolutional layer, and using convolutional kernels with different dilation coefficients to parallelly extract short-period fluctuation features and long-period trend features; Inputting the short-period fluctuation features and long-period trend features into a gated recurrent unit, and fusing the fluctuation patterns of different time scales through a time gating mechanism to generate a fused temporal feature vector; Performing spectral analysis on the temporal feature vector, identifying the frequency domain components with significant periodicity, and retaining the frequency domain components that match the service cycle through a band-pass filter to generate filtered spectral features; Map the spectral features back to the time-domain space, eliminate the dimensional differences of different computing resource nodes through adaptive normalization processing, and generate standardized dynamic load fluctuation features, where the dynamic load fluctuation features are used to characterize the load change law of computing resource nodes across time scales.
4. The multi-domain computing resource aggregation method based on a virtualized user network according to claim 2, wherein, Analyze the resource interaction and dependence relationships between each computing resource node through a graph attention network to generate resource compatibility features, including: Construct a directed weighted graph based on the historical communication records between computing resource nodes, where the edge weights in the directed weighted graph are jointly determined by the historical data transmission success rate and average delay between nodes; Perform random walk sampling on the directed weighted graph to generate a context node sequence for each computing resource node, where the context node sequence contains neighbor nodes with strong dependence relationships with it; Input the context node sequence into the embedding layer of the graph attention network to generate a topological perception feature vector for each neighbor node, where the topological perception feature vector encodes the physical connection attributes and logical collaboration relationships between nodes; Perform two-way attention calculation on the feature vector of the current computing resource node and the topological perception feature vectors of neighbor nodes to generate a resource collaboration degree score between nodes, where the collaboration degree score reflects the hardware configuration alignment degree, protocol compatibility, and load complementarity between nodes; Perform dynamic weighted aggregation on the topological perception feature vectors of neighbor nodes based on the collaboration degree score to generate the local compatibility feature of the current node; Perform feature cross between the local compatibility feature and the cross-domain resource distribution feature extracted by the global graph pooling layer, and suppress redundant features and enhance cross-domain collaboration signals through a gated fusion mechanism to generate the final resource compatibility feature.
5. The multi-domain computing resource aggregation method based on a virtualized user network according to claim 1, characterized in that Perform multi-domain resource association analysis processing on the multi-dimensional resource feature vector based on a preset intelligent aggregation strategy network to determine the collaborative matching weights between cross-domain computing resources, and aggregate the scattered computing resource nodes into a virtualized resource pool according to the collaborative matching weights, including: Input the multi-dimensional resource feature vector into the cross-attention encoding layer of the intelligent aggregation strategy network, perform cross-domain two-way attention calculation on the multi-dimensional resource feature vectors of each computing resource node, and generate a similarity association matrix containing the similarity between inter-domain nodes; Perform sparsification processing based on threshold filtering on the similarity association matrix, retain the association edges with similarity higher than the dynamic threshold between each computing resource node and cross-domain nodes, generate a sparsified association relationship graph, input the sparsified association relationship graph into a differentiable graph sorting network, and perform end-to-end gradient backpropagation training based on the multi-dimensional resource feature vectors of nodes and the association edge weights to generate the collaborative matching weights of each cross-domain association edge; Update the edge weights of the sparsified association relationship graph according to the collaborative matching weights to generate a weighted cross-domain resource association graph, input the cross-domain resource association graph into an overlapping community discovery algorithm, and identify node clusters with stable collaboration relationships based on the principle of modularity maximization to generate multiple candidate resource clusters; Perform mean pooling operation on the nodes within each candidate resource cluster to generate a cluster-level resource feature vector, and screen the target cluster according to the cosine similarity between the cluster-level resource feature vector and a preset global resource demand template to generate a set of target resource clusters; Perform virtualized resource encapsulation operation on each cluster in the set of target resource clusters to generate logical resource units with a unified interface protocol, and establish bidirectional resource redundancy channels between the logical resource units based on the collaborative matching weights to generate a virtualized resource pool with fault tolerance capabilities.
6. The multi-domain computing resource aggregation method based on a virtualized user network according to claim 5, characterized in that The establishing bidirectional resource redundancy channels between the logical resource units based on the collaborative matching weights to generate a virtualized resource pool with fault tolerance capabilities includes: Extract the collaborative matching weights between the logical resource units to generate a unit-interconnection weight relationship table, and perform the top-K weight screening on each logical resource unit according to the weight connection relationship table to generate a redundant association unit list corresponding to each unit; Perform protocol handshake verification on the units in the redundant association unit list with the current logical resource unit to generate interoperable redundant unit pairs, and perform resource mirror synchronization operation on each redundant unit pair to generate redundant resource copies with consistent real-time states; Configure priority routing identifiers for each redundant resource copy based on a preset failover strategy to generate a redundant channel configuration table with priority markings, and inject the redundant channel configuration table into the load balancer of the virtualized resource pool to generate dynamic traffic distribution rules; During the operation of the virtualized resource pool, monitor the heartbeat signals of each logical resource unit in real time to generate a unit health status matrix. When a heartbeat signal timeout anomaly is detected, switch the task flow of the faulty unit to the redundant resource copy with the highest priority according to the redundant channel configuration table to generate an updated resource scheduling path.
7. The multi-domain computing resource aggregation method based on a virtualized user network according to claim 1, wherein The performing dynamic resource adaptation processing on the aggregated resources in the virtualized resource pool to generate a resource scheduling topology structure matching the target service requirements includes: Input the real-time task characteristics of the target service requirements into the third candidate subset of the virtualized resource pool, and perform transmission path traversal simulation based on the network bandwidth utilization characteristics of the logical resource units in the third candidate subset to generate end-to-end delay prediction values from each logical resource unit to the service access point; Perform node-by-node comparison of the end-to-end delay prediction values with the delay upper limit in the real-time task characteristics, and eliminate the logical resource units whose delay prediction values continuously exceed the upper limit to generate a fourth candidate subset with compliant delays; Perform sliding window matching on the computing power load rate fluctuation characteristics of the logical resource units in the fourth candidate subset with the floating-point operation demand of the compute-intensive task characteristics, and calculate the cumulative distribution function of the computing power supply margin within each window to generate a fifth candidate subset that meets the task peak demand; Perform spatial continuity analysis on the storage space occupancy rate characteristics of the logical resource units in the fifth candidate subset, identify storage fragmentation areas and perform block remapping operations on virtual storage volumes to generate a sixth candidate subset with optimized storage continuity; Construct an initial resource topology graph based on the sixth candidate subset. The nodes of the initial resource topology graph represent the logical resource units for storing continuous optimization, and the edges represent the historical average bandwidth of cross-domain communication links. Input the initial resource topology graph into the multi-objective optimization module to synchronously optimize the computing resource utilization rate, storage balance degree, and network load variance. Iteratively adjust the edge weights and node clustering centers through the gradient descent algorithm to generate an intermediate topology graph with optimized edge weights. Hierarchically cluster the nodes of the intermediate topology graph according to the fluctuation characteristics of the computing power load rate to generate a hierarchical cascade structure including a core computing layer, a buffer storage layer, and an edge access layer, and assign a priority routing protocol to the bidirectional data channels between the layers. Embed a resource status tracking agent in the hierarchical cascade structure to collect the task execution delay and resource consumption rate of each layer node in real time and generate a topology health status indicator. When the delay or consumption rate of any layer in the topology health status indicator exceeds the dynamic threshold, trigger local topology reconstruction, and re-execute edge weight adjustment and node clustering according to the resource status of the current sixth candidate subset to generate an updated resource scheduling topology structure.
8. The multi-domain computing resource aggregation method based on a virtualized user network according to claim 7, characterized in that The deployment of the resource scheduling topology structure to the multi-domain computing environment to trigger resource collaborative services includes: Convert the hierarchical cascade structure of the resource scheduling topology structure into a cross-domain resource orchestration instruction set, which includes the virtual machine deployment coordinates, storage volume mapping relationships, and priority routing tables of each layer node. Inject the priority routing table into the forwarding table of each domain boundary router through the policy distribution interface of the software-defined network controller to generate a cross-domain bandwidth guarantee channel. Distribute the virtual machine deployment coordinates and storage volume mapping relationships to the target computing resource nodes through the resource scheduling engine of the cloud management platform to generate virtualized resource instances distributed according to the topology structure. During the operation of the virtualized resource instances, continuously collect the CPU utilization rate of the core computing layer, the IO throughput of the buffer storage layer, and the end-to-end delay data of the edge access layer through the resource status tracking agent to generate a real-time resource load matrix. Input the real-time resource load matrix into the overload determination model. When the CPU utilization rate of the core computing layer continuously exceeds the first threshold and the IO throughput of the buffer storage layer is lower than the second threshold, generate a computing and storage resource imbalance alarm. Trigger a resource rescheduling request according to the computing and storage resource imbalance alarm, re-input the current real-time resource load matrix and the target service requirements into the dynamic resource adaptation processing module, and generate a rescheduled resource scheduling topology structure. Use the hot migration engine to perform lossless migration of the task status and data cache of the virtualized resource instances according to the rescheduled topology structure, and update the priority routing table of the cross-domain bandwidth guarantee channel after the migration is completed to generate an updated version of the resource collaborative service.
9. The multi-domain computing resource aggregation method based on a virtualized user network according to claim 8, wherein The use of the hot migration engine to perform lossless migration of the task status and data cache of the virtualized resource instances according to the rescheduled topology structure, and update the priority routing table of the cross-domain bandwidth guarantee channel after the migration is completed to generate an updated version of the resource collaborative service includes: Extract the list of edge access layer nodes in the hierarchical cascaded structure of the resource scheduling topology after rescheduling to generate a new set of access point addresses; Compare the new set of access point addresses with the addresses in the historical routing table to identify newly added access points and failed access points, and generate a routing change instruction set; Send the routing change instruction set to each domain border router through the routing update interface of the software-defined network controller, delete the routing entries corresponding to the failed access points, and allocate guaranteed bandwidth for the newly added access points; After the routing update is completed, verify whether the end-to-end delay of the newly added access points meets the requirements of real-time task characteristics through the resource status tracking agent in the edge access layer to generate a delay verification result; When there are abnormal access points in the delay verification result, re-trigger the dynamic resource adaptation processing module to replace the abnormal nodes from the sixth candidate subset and generate a secondary rescheduling topology; Perform partial migration on the edge access layer nodes of the secondary rescheduling topology again through the hot migration engine until the delay verification results of all access points meet the standards, generating a final stable resource collaboration service version.
10. A multi-domain computing resource aggregation system based on a virtualized user network, characterized in that, It includes a processor and a memory. The memory is connected to the processor. The memory is used to store programs, instructions, or codes. The processor is used to execute the programs, instructions, or codes in the memory to implement the multi-domain computing resource aggregation method based on the virtualized user network according to any one of claims 1-9 above.
Citation Information
Patent Citations
An SDN-based container network resource scheduling method
CN109743261A
Edge computing collaborative franchisee discovery method based on comprehensive trust evaluation
CN112132202A
Container cluster service dynamic management method and system
CN115665158A
Cloud edge computing power resource adaptive computing system
CN116389491A
Container cloud resource intelligent scheduling method oriented to LVC simulation
CN117687760A
Cited By
Audio and video system data comprehensive analysis processing method based on virtual distributed architecture
CN120475226A
Hyper-converged server resource pooling method and system
CN120547063A
Intelligent electric energy meter power consumption data analysis method based on cloud computing
CN120596281A
Unmanned aerial vehicle detection method and system based on visible light polarization imaging
CN120652460A
A method and system for detecting unmanned aerial vehicles based on visible light polarization imaging
CN120652460B