Method and system for multi-domain computing resource aggregation based on virtualized user network
By collecting and analyzing real-time status data of multi-domain computing resources, multi-dimensional resource feature vectors are generated. Cross-domain resource association analysis is performed using an intelligent aggregation strategy network, which solves the problem of insufficient accuracy in resource aggregation in existing technologies and achieves efficient resource integration and collaborative services.
Patent Information
- Application Number
- CN202510687268.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Existing multi-domain computing resource aggregation technologies cannot deeply explore the characteristics of computing resource nodes that are closely related to power business, resulting in insufficient accuracy and effectiveness of resource aggregation, and failing to meet the power grid business's demand for efficient resource utilization.
Real-time status data of computing resource nodes in a multi-domain environment is collected, multi-dimensional resource feature vectors are generated through a virtualized resource mapping model, cross-domain resource association analysis is performed using an intelligent aggregation strategy network, collaborative matching weights are determined, and resource nodes are aggregated into a virtualized resource pool for dynamic resource adaptation.
It has achieved the organic integration and efficient collaboration of cross-domain resources, improved resource utilization and business processing efficiency, and enhanced the overall performance and business response capability of the multi-domain computing system.
Smart Images

Figure CN120281776B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of cloud computing, in particular to a multi-domain computing resource aggregation method and system based on a virtualized user network. BACKGROUND
[0002] In the power grid industry, with the continuous advancement of smart grid construction and the increasing complexity of power business, efficient aggregation and flexible scheduling of resources in a multi-domain computing environment have become key issues that need to be addressed.
[0003] Currently, existing multi-domain computing resource aggregation technologies cannot deeply mine the characteristics of computing resource nodes closely related to power business from multiple dimensions. Generally, only simple resource classification or matching based on a single feature is performed, and a multi-dimensional resource feature vector containing dynamic load fluctuation characteristics, resource compatibility characteristics and task adaptation degree characteristics cannot be generated. This makes it difficult to accurately assess the adaptability and potential value of resources to different business scenarios of the power grid, affecting the accuracy and effectiveness of resource aggregation.
[0004] Moreover, existing technologies cannot perform dynamic correlation analysis and collaborative matching according to the actual situation of multi-domain resources in the power grid. Power grid business has the characteristics of high real-time and strong complexity, while existing technologies usually perform resource aggregation based on fixed rules or simple algorithms, and cannot fully consider the complex relationships and collaborative potential between cross-domain computing resources, resulting in a resource pool that cannot fully utilize the advantages of each node, low resource utilization, and difficulty in meeting the demand for efficient use of resources by power grid business. SUMMARY
[0005] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide a multi-domain computing resource aggregation method based on a virtualized user network, which comprises:
[0006] Collecting a set of real-time resource state data of computing resource nodes dispersedly deployed in a multi-domain environment, the set of real-time resource state data comprising the computing power load rate, storage space occupancy rate and network bandwidth utilization rate of each computing resource node;
[0007] Calling a virtualized resource mapping model to perform cross-domain resource feature extraction processing on the set of real-time resource state data, to generate a multi-dimensional resource feature vector of each computing resource node, the multi-dimensional resource feature vector containing dynamic load fluctuation characteristics, resource compatibility characteristics and task adaptation degree characteristics;
[0008] Performing multi-domain resource correlation analysis processing on the multi-dimensional resource feature vector based on a preset intelligent aggregation strategy network, determining the collaborative matching weight between cross-domain computing resources, and aggregating the dispersed computing resource nodes into a virtualized resource pool according to the collaborative matching weight;
[0009] performing dynamic resource adaptation processing on the aggregated resources in the virtualized resource pool, generating a resource scheduling topology structure matched with a target service requirement, and deploying the resource scheduling topology structure to a multi-domain computing environment to trigger a resource coordination service.
[0010] In yet another aspect, an embodiment of the present application also provides a multi-domain computing resource aggregation system based on a virtualized user network, comprising a processor, a machine readable storage medium, the machine readable storage medium being connected with the processor, the machine readable storage medium being used to store programs, instructions or codes, and the processor being used to execute the programs, instructions or codes in the machine readable storage medium to implement the above method.
[0011] Based on the above aspects, an embodiment of the present application calls a virtualized resource mapping model to perform cross-domain resource feature extraction processing on a real-time resource state data set, a generated multi-dimensional resource feature vector not only covers dynamic load fluctuation features, resource compatibility features, but also contains task adaptation degree features, deeply characterizing the characteristics of computing resource nodes from multiple dimensions, and based on a preset intelligent aggregation strategy network, performing multi-domain resource correlation analysis processing on the multi-dimensional resource feature vector can accurately determine the collaborative matching weight between cross-domain computing resources, and then scientifically and reasonably aggregate the dispersed computing resource nodes into a virtualized resource pool, realizing the organic integration and efficient collaboration of cross-domain resources. Performing dynamic resource adaptation processing on the aggregated resources in the virtualized resource pool generates a resource scheduling topology structure matched with a target service requirement, which can flexibly and accurately allocate and schedule resources according to the demand characteristics of different services, significantly improving the resource utilization rate and service processing efficiency. Deploying the resource scheduling topology structure to a multi-domain computing environment to trigger a resource coordination service realizes efficient collaboration and dynamic adaptation of multi-domain computing resources as a whole, effectively improving the overall performance and service response capability of the multi-domain computing system. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is a flowchart of the multi-domain computing resource aggregation method based on a virtualized user network provided by an embodiment of the present application.
[0013] Figure 2 is a schematic diagram of exemplary hardware and software components of the multi-domain computing resource aggregation system based on a virtualized user network provided by an embodiment of the present application. DETAILED DESCRIPTION
[0014] The present application will be described in detail below with reference to the accompanying drawings, Figure 1 is a flowchart of the multi-domain computing resource aggregation method based on a virtualized user network provided by an embodiment of the present application, and the multi-domain computing resource aggregation method based on a virtualized user network will be described in detail below.
[0015] Step S110: Collecting a real-time resource state data set of the computing resource nodes dispersedly deployed in the multi-domain environment, the real-time resource state data set including a computing power load rate, a storage space occupancy rate and a network bandwidth utilization rate of each computing resource node.
[0016] In this embodiment, in the multi-domain computing environment in the field of power grid, there are a large number of computing resource nodes dispersedly deployed, which are distributed in different transformer substations, power generation stations, power regulation centers and various monitoring sites, and undertake key tasks such as power data acquisition and processing, power grid fault diagnosis, power load prediction. In order to realize efficient management and aggregation of these computing resources, the primary task is to collect real-time resource state data.
[0017] Taking a large power grid covering multiple areas as an example, there is a computing resource node A in a large transformer substation, which is responsible for real-time processing and analysis of operation data of numerous power equipment such as transformers and circuit breakers in the substation; node B is located in a regional power regulation center, which undertakes power dispatching and distribution calculation tasks of the entire regional power grid; node C is set in a wind power station in a remote mountainous area, which mainly processes operation state monitoring and fault early warning data of wind turbine generators.
[0018] For node A, the computing power load rate is collected by the monitoring software installed in its hardware system at an interval of one minute. Assuming that the node is equipped with 64 computing cores, at a certain time, 40 computing cores are in full load operation state, 10 computing cores are in 70% load level, and the remaining 14 computing cores are in idle state. Then the number of equivalent computing cores in use is 40+10×0.7=47, and the computing power load rate is 47 divided by 64, which is about 73.44%. At the same time, the storage space occupancy rate is collected. The total storage space of node A is 8TB (i.e. 8192GB), and the used storage space is 4.5TB (i.e. 4608GB), so the storage space occupancy rate is 4608 divided by 8192, which is about 56.25%. In terms of network bandwidth utilization rate, the total network bandwidth of node A is 2000Mbps, and the current used network bandwidth is 1200Mbps, so the network bandwidth utilization rate is 1200 divided by 2000, which is 60%.
[0019] For node B, the total number of computing cores is 128, at the current time, 80 computing cores are in working state, of which 60 are full load and 20 are in 60% load state, the remaining 48 are idle. The equivalent number of computing cores in use is 60+20x0.6=72, and the computing power load rate is 72 divided by 128, about 56.25%. The total storage space is 16TB (i.e. 16384GB), and 7TB (i.e. 7168GB) has been used, the storage space occupancy rate is 7168 divided by 16384, about 43.75%. The total network bandwidth is 4000Mbps, and the current use is 2400Mbps, the network bandwidth utilization rate is 2400 divided by 4000, i.e. 60%.
[0020] The total number of computing cores of node C is 32, and currently 20 computing cores are running, of which 15 are full load and 5 are in 30% load state, and the remaining 12 are idle. The equivalent number of computing cores in use is 15+5x0.3=16.5, and the computing power load rate is 16.5 divided by 32, about 51.56%. The total storage space is 2TB (i.e. 2048GB), and 0.8TB (i.e. 819.2GB) has been used, the storage space occupancy rate is 819.2 divided by 2048, about 40%. The total network bandwidth is 500Mbps, and the current use is 150Mbps, the network bandwidth utilization rate is 150 divided by 500, i.e. 30%.
[0021] Thus, the computing power load rate, storage space occupancy rate and network bandwidth utilization rate data of all computing resource nodes such as nodes A, B and C are collected and summarized, forming a real-time resource state data set.
[0022] Step S120: calling a virtualized resource mapping model to perform cross-domain resource feature extraction processing on the real-time resource state data set, generating a multi-dimensional resource feature vector for each computing resource node, the multi-dimensional resource feature vector containing dynamic load fluctuation features, resource compatibility features and task adaptation degree features.
[0023] In this embodiment, after obtaining the real-time resource state data set, the virtualized resource mapping model needs to be called to deeply process these data to extract the key features of each computing resource node. Specifically, dynamic load fluctuation features, resource compatibility features and task adaptation degree features are extracted from the data, and then the three features are combined into a multi-dimensional resource feature vector. Taking nodes A, B and C as an example, input their real-time resource state data set into the virtualized resource mapping model, and through a series of processing of the model, generate a multi-dimensional resource feature vector containing the above three features for each node. How to extract these features is described in detail below.
[0024] Step S121: input the real-time resource state data set into the feature encoding layer of the virtualized resource mapping model, extract the periodic fluctuation mode of the computing power load rate through a time series convolution network, and generate dynamic load fluctuation features.
[0025] In order to accurately extract the periodic fluctuation mode of the computing power load rate, the collected computing power load rate data needs to be processed in detail. Taking node A as an example, the computing power load rate data of node A changes over time. The computing power load rate of node A is periodically divided according to a preset time window. Assuming that the preset time window is 2 hours. Then the computing power load rate data of node A in a day is divided into 12 load rate time sequence segments, each of which contains 120 consecutive time stamp (one per minute) load rate sampling data.
[0026] For example, the first load rate time sequence segment is the computing power load rate data from 0 o'clock to 2 o'clock, which contains the computing power load rate sampling value every minute in these 120 minutes, such as 0 o'clock 0 minute load rate 70%, 0 o'clock 1 minute load rate 71%, 0 o'clock 2 minute load rate 70.5%, etc.
[0027] Step S1211: periodically divide the computing power load rate according to a preset time window to obtain a plurality of load rate time sequence segments, each of which contains consecutive time stamp load rate sampling data.
[0028] Continuing to take node A as an example, the preset time window is 2 hours, and the monitoring software collects computing power load rate data every minute. From the start of the day, the first load rate time sequence segment is the computing power load rate data in the 120 minutes from 0 o'clock to 2 o'clock. Assuming that at 0 o'clock, the computing power load rate of node A is 70%, which is the first data point of this time sequence segment; at 0 o'clock 1 minute, the computing power load rate becomes 71%, which becomes the second data point; and so on, until 2 o'clock, the 120th data point is obtained, thus forming a complete load rate time sequence segment. In the same way, the computing power load rate data in a day is divided into 12 such load rate time sequence segments, which contain the continuous changes of the computing power load rate of node A in different time periods.
[0029] Step S1212: input the load rate time sequence segment into the dilated causal convolution layer, and extract short-period fluctuation features and long-period trend features in parallel using convolution kernels with different dilation coefficients.
[0030] The segmented load rate time series segment of node A is input into an expanded causal convolution layer, in which convolution kernels with different dilation coefficients are used for convolution operations. For example, a convolution kernel with a dilation coefficient of 2 is used to extract short-period fluctuation features. The load rate time series segment can be convolved every 2 time steps. Assuming that the first load rate time series segment is [70, 71, 70.5, 72, 71.5, 73, 72.5, 74, 73.5, 75, …, 78], the convolution kernel with a dilation coefficient of 2 is convolved with the load rate time series segment, and the data points at positions 1, 3, 5, 7, 9, etc. are processed in turn. Through convolution operations, the local change relationship between these data points can be captured, thereby extracting short-period fluctuation features. For example, from 70 to 70.5 to 71.5, etc.
[0031] At the same time, a convolution kernel with a dilation coefficient of 10 is used to extract long-period trend features. For example, convolution can be performed every 10 time steps to process data over a longer time range. For the above load rate time series segment, data points at positions 1, 11, 21, etc. can be processed, thereby observing the change trend of the computing power load rate from a more macro perspective, such as the overall upward or downward trend. By using these convolution kernels with different dilation coefficients in parallel, short-period and long-period feature information can be obtained at the same time.
[0032] Step S1213: input the short-period fluctuation features and long-period trend features into a gated recurrent unit, and fuse the fluctuation patterns of different time scales through a time gating mechanism to generate a fused time sequence feature vector.
[0033] The short-period fluctuation features and long-period trend features of node A obtained through the expanded causal convolution layer are input into the gated recurrent unit. The gated recurrent unit has a time gating mechanism and can selectively update and retain information according to the importance of features of different time scales.
[0034] There are two key gating mechanisms in the gated recurrent unit, namely the update gate and the reset gate. The update gate determines how much information of the hidden state at the previous time needs to be passed to the current time, and the reset gate determines how much information of the hidden state at the previous time needs to be reset.
[0035] For short-period fluctuation features and long-period trend features, at each time step, the update gate calculates an update weight according to the input features at the current time and the hidden state at the previous time. The update weight represents the proportion of the hidden state at the previous time that needs to be retained. For example, if the short-period fluctuation features at the current time change greatly, the update gate can reduce the retention proportion of the hidden state at the previous time, so as to pay more attention to the current short-period fluctuation.
[0036] The reset gate calculates a reset weight to determine how much information of the hidden state at the previous time can participate in the calculation of the hidden state at the current time. If there is a significant change in the long-period trend feature at the current time, the reset gate can allow more hidden state information related to the long-period trend at the previous time to participate in the calculation.
[0037] Through this time-gating mechanism, the gated recurrent unit can adaptively fuse the short-period fluctuation feature and the long-period trend feature to generate a fused time-series feature vector that contains comprehensive information of the fluctuation patterns of the computing power load rate at different time scales, and can more comprehensively reflect the computing power load change of the node A.
[0038] Step S1214: performing spectrum analysis on the time-series feature vector, identifying frequency domain components with significant periodicity, retaining frequency domain components matching the business period through a band-pass filter, and generating filtered spectral features.
[0039] In this embodiment, spectrum analysis can convert signals in the time domain to the frequency domain, thereby identifying components of different frequencies in the signal. In the power grid industry, different businesses can have different periodic characteristics, for example, power load can have daily and weekly cycles.
[0040] The time-series feature vector is converted to the frequency domain using methods such as fast Fourier transform to obtain a frequency domain signal. In the frequency domain, each frequency component corresponds to a different period. By analyzing the frequency domain signal, frequency components with significant energy are identified, which represent the parts of the time-series feature vector with significant periodicity.
[0041] For example, through spectrum analysis, it is found that there is a frequency component in the frequency domain corresponding to a period of 24 hours, which may be related to the daily cycle of power load; there is also a frequency component corresponding to a period of 7 days, which may be related to the weekly cycle of power load.
[0042] Next, a band-pass filter is used to filter the frequency domain signal. The band-pass filter can set a frequency range, allowing only frequency components within the frequency range to pass. According to the characteristics of the power grid business, determine the frequency range matching the business period. For example, if the daily and weekly fluctuations are concerned, set the frequency range of the band-pass filter so that the frequency components corresponding to the 24-hour and 7-day periods can pass, while other irrelevant frequency components are filtered out.
[0043] After the band-pass filter processing, the filtered spectral features are obtained, which only contain the periodic information matching the business period, and remove other interference information, so that the subsequent analysis can focus more on the fluctuation patterns related to the business.
[0044] Step S1215: Map the spectrum features back to the time domain space, eliminate the dimensional differences between different computing resource nodes through adaptive normalization processing, and generate standardized dynamic load fluctuation features, which are used to represent the load change law of the computing resource nodes across the time scale.
[0045] In this embodiment, this step can be implemented by methods such as inverse fast Fourier transform. For example, the spectrum features in the frequency domain are converted back to the time domain signal to obtain a time domain signal reflecting the periodic fluctuation of the computing power load rate.
[0046] However, different computing resource nodes may have different hardware configurations, workloads, etc., resulting in differences in the dimension and numerical range of their computing power load rate data. In order to eliminate these differences, adaptive normalization processing is needed. Adaptive normalization processing dynamically adjusts the scale of the data according to the specific circumstances of each computing resource node. For example, for node A and node B, their computing power load rate data may have a higher overall value and a lower overall value. Adaptive normalization processing calculates the mean and standard deviation of each node data, and then normalizes the data according to these statistical information.
[0047] Specifically, for each data point in the time domain signal of node A, the mean of the node signal is first subtracted, and then divided by the standard deviation of the node signal. After such processing, the data of node A is standardized to a relatively uniform scale. The same processing method is also applied to other computing resource nodes.
[0048] After adaptive normalization processing, the standardized dynamic load fluctuation features are obtained. The dynamic load fluctuation features eliminate the dimensional differences between different computing resource nodes and can accurately represent the load change law of each computing resource node across the time scale. For example, through the dynamic load fluctuation features, the computing power load fluctuation of node A in the daily and weekly cycles can be clearly seen, as well as the relative change trend compared with other nodes.
[0049] Step S122: Analyze the resource interaction dependency relationship between the computing resource nodes through the graph attention network, generate resource compatibility features, and the resource compatibility features are used to quantify the hardware configuration differences and protocol interoperability of the computing resource nodes in different domains.
[0050] In the multi-domain computing environment of the power grid industry, there are complex resource interaction dependency relationships between different computing resource nodes. For example, the computing resource nodes of the substation may need to interact with the nodes of the power control center to realize real-time scheduling of power; the nodes of the power station need to share equipment operation data with the nodes of the monitoring site to timely detect faults. In order to quantify the resource compatibility between these nodes, the graph attention network needs to be analyzed.
[0051] First, a directed weighted graph is constructed according to the historical communication records between the computing resource nodes. The historical communication records contain the data transmission information between the nodes, such as the frequency of data transmission, success rate, average delay, etc.
[0052] Step S1221: Construct a directed weighted graph according to the historical communication records between the computing resource nodes, and the edge weight in the directed weighted graph is determined by the historical data transmission success rate and average delay between the nodes.
[0053] Taking nodes A, B, and C as an example, assume that in the past period of time, node A transmits data to node B 100 times, with 90 successes and an average delay of 20 milliseconds; node B transmits data to node A 80 times, with 70 successes and an average delay of 25 milliseconds; node A transmits data to node C 60 times, with 50 successes and an average delay of 30 milliseconds; node C transmits data to node A 50 times, with 40 successes and an average delay of 35 milliseconds; node B transmits data to node C 70 times, with 60 successes and an average delay of 28 milliseconds; node C transmits data to node B 65 times, with 55 successes and an average delay of 32 milliseconds.
[0054] In order to determine the edge weight in the directed weighted graph, the historical data transmission success rate and average delay are considered comprehensively. A simple calculation method can be used, edge weight = historical data transmission success rate / average delay.
[0055] For the edge from node A to node B, the edge weight = 90% / 20 milliseconds = 0.045. For the edge from node B to node A, the edge weight = 70% / 25 milliseconds = 0.028. In the same way, the edge weights of the edges from node A to node C, from node C to node A, from node B to node C, and from node C to node B are calculated as 0.0278, 0.0229, 0.0268, and 0.0234, respectively.
[0056] According to these edge weights, a directed weighted graph is constructed. The nodes in the graph represent computing resource nodes, and the directed edges represent the data transmission direction between the nodes, and the weight of the edge represents the efficiency and reliability of data transmission between the nodes.
[0057] Step S1222: Perform random walk sampling on the directed weighted graph to generate a context node sequence for each computing resource node, which contains neighbor nodes with strong dependency relationships.
[0058] After the directed weighted graph is constructed, random walk sampling is performed on the directed weighted graph. Random walk sampling is a method of randomly selecting paths in a graph, and through multiple random walks, neighbor nodes with strong dependency relationships with each node can be found.
[0059] Taking node A as an example, start random walk from node A. At each walk, the next node is selected with a certain probability according to the weight of the edge in the directed weighted graph. For example, node A has two out-edges pointing to node B and node C, the weight of the edge from node A to node B is 0.045, and the weight of the edge from node A to node C is 0.0278. Then in the random walk, the probability of selecting node B is 0.045 / (0.045+0.0278)≈0.618, and the probability of selecting node C is 0.0278 / (0.045+0.0278)≈0.382.
[0060] Perform multiple random walks (e.g. 100 times), and record the node sequence passed through each time. After multiple walks, count the frequency of each node being visited. If a node is frequently visited, it indicates that it has a strong dependency relationship with node A. Sort these nodes with strong dependency relationship with node A according to the visit frequency to generate the context node sequence of node A. Suppose after random walk sampling, the context node sequence of node A is [node B, node C], which indicates that node B and node C have a strong resource interaction dependency relationship with node A.
[0061] Step S1223: input the context node sequence into the embedding layer of the graph attention network to generate a topological awareness feature vector of each neighbor node, the topological awareness feature vector encoding the physical connection attributes and logical collaboration relationships between nodes.
[0062] Input the context node sequence [node B, node C] of node A into the embedding layer of the graph attention network. The role of the embedding layer is to convert node information into low-dimensional vector representation for subsequent analysis and processing. In the embedding layer, a topological awareness feature vector will be learned for each neighbor node (here, node B and node C).
[0063] For node B, the embedding layer will consider its position in the directed weighted graph, connection relationship with other nodes, and its own attributes, etc. For example, the edge weight between node B and node A represents the data transmission efficiency and reliability between them, which will be included in the learning of the topological awareness feature vector. At the same time, the connection of node B with other nodes (such as node C) will also affect its topological awareness feature vector. Suppose node B also has data interaction with other multiple nodes, the edge weights and directions of these connections will be processed in the embedding layer to generate a topological awareness feature vector that can encode the physical connection attributes and logical collaboration relationships of node B.
[0064] Specifically, the embedding layer will use a series of neural network operations to learn these feature vectors. First, a linear transformation will be performed on the original features of node B (such as hardware configuration information, historical data transmission success rate, etc.) to obtain a preliminary feature representation. Then, the structural information in the directed weighted graph will be combined to adjust the importance of different connections through an attention mechanism. The attention mechanism will assign an attention score to each connection based on edge weights and other information, with a higher score indicating a greater impact on the topology-aware feature vector of node B. For example, if the edge weight between node B and node A is large, the information of node A will be given a higher weight when generating the topology-aware feature vector of node B.
[0065] After such processing, the embedding layer will generate the topology-aware feature vector of node B, assuming that this topology-aware feature vector is a multi-dimensional vector of length 128, where each dimension reflects some topological and attribute information of node B in the graph.
[0066] The same process will also be applied to node C. The embedding layer will generate the topology-aware feature vector of node C based on its connection relationships with other nodes (including nodes A and B), edge weights, and the attributes of node C itself, through linear transformation and attention mechanism. Assuming that the topology-aware feature vector of node C is also a multi-dimensional vector of length 128, it also encodes the physical connection attributes and logical collaboration relationships of node C.
[0067] Step S1224: Perform bidirectional attention calculation on the feature vector of the current computing resource node and the topology-aware feature vectors of the neighbor nodes to generate an inter-node resource collaboration score, which reflects the alignment of hardware configurations, protocol compatibility, and load complementarity between nodes.
[0068] Taking node A as an example, after obtaining the topology-aware feature vectors of neighbor nodes B and C, bidirectional attention calculation needs to be performed on the feature vector of node A itself and the topology-aware feature vectors of the neighbor nodes to generate an inter-node resource collaboration score.
[0069] The feature vector of node A itself contains its hardware configuration information (such as the number of computing cores, the size of storage space, etc.), real-time resource state data (such as computing power load rate, storage space occupancy rate, network bandwidth utilization rate, etc.), etc. First, for the case of node A and node B, bidirectional attention calculation will consider information exchange in both directions.
[0070] In the forward attention calculation, the influence of the topology-aware feature vector of node B on node A is focused on. By calculating the similarity between the feature vector of node A and the topology-aware feature vector of node B, an attention score is obtained. For example, the dot product operation can be used to calculate the similarity, and then the similarity is converted into a probability distribution by the softmax function as the attention score, which represents the importance of the information of node B in generating the collaborative degree score of node A and node B.
[0071] In the reverse attention calculation, the influence of the feature vector of node A on node B is focused on. Similarly, the similarity between the topology-aware feature vector of node B and the feature vector of node A is calculated to obtain a reverse attention score.
[0072] The forward and reverse attention scores are integrated to obtain a comprehensive attention score. Then, combined with the hardware configuration information, historical data transmission success rate, protocol type and other information of node A and node B, the inter-node resource collaboration degree score is calculated. For example, if the hardware configurations of node A and node B are similar, the protocol types are the same, and the historical data transmission success rate is high, their resource collaboration degree score will be higher. Assuming that through a series of calculations, the resource collaboration degree score of node A and node B is 0.8.
[0073] The same method is used to calculate the resource collaboration degree score of node A and node C. After calculation, it is assumed that the resource collaboration degree score of node A and node C is 0.7. These collaboration degree scores reflect the alignment degree of hardware configuration, protocol compatibility and load complementarity between nodes. For example, if the computing power load of node A is high and the computing power load of node C is low, and their hardware configurations and protocol are compatible, the load complementarity between them is good, and the collaboration degree score will also be improved accordingly.
[0074] Step S1225: dynamically weighting and aggregating the topology-aware feature vectors of neighbor nodes based on the collaboration degree scores to generate the local compatibility feature of the current node.
[0075] After obtaining the resource collaboration degree scores of node A and neighbor nodes B and C, the topology-aware feature vectors of neighbor nodes are dynamically weighted and aggregated based on these collaboration degree scores to generate the local compatibility feature of node A.
[0076] For the topology-aware feature vector of node B and the topology-aware feature vector of node C, weights are assigned according to their collaboration scores with node A. The collaboration score of node B with node A is 0.8, and the collaboration score of node C with node A is 0.7. Then when aggregating, the weight of the topology-aware feature vector of node B is 0.8 / (0.8+0.7)≈0.533, and the weight of the topology-aware feature vector of node C is 0.7 / (0.8+0.7)≈0.467.
[0077] The topology-aware feature vector of node B is multiplied by its weight, and the topology-aware feature vector of node C is multiplied by its weight, and then the two weighted vectors are spliced. Assuming that the topology-aware feature vectors of node B and node C are both multi-dimensional vectors with a length of 128, after splicing, a vector with a length of 256 is obtained, which contains the information of neighbor nodes B and C, and is weighted according to the collaboration scores, and can reflect the local compatibility between node A and neighbor nodes, that is, the local compatibility feature of node A.
[0078] Step S1226: Cross the local compatibility feature with the cross-domain resource distribution feature extracted by the global graph pooling layer, suppress redundant features and enhance cross-domain collaboration signals through a gating fusion mechanism, and generate the final resource compatibility feature.
[0079] In this embodiment, the global graph pooling layer can process the entire directed weighted graph and extract a feature reflecting the cross-domain resource distribution. For example, the global graph pooling layer can comprehensively consider the information of all nodes, aggregate and count the feature vectors of the nodes in the directed weighted graph, and obtain a cross-domain resource distribution feature vector. For example, the average computing power load rate, average storage space occupancy rate and other information of all nodes can be calculated to reflect the resource distribution of the entire multi-domain computing environment. Assuming that the cross-domain resource distribution feature vector extracted by the global graph pooling layer is a multi-dimensional vector with a length of 128.
[0080] Cross the local compatibility feature of node A (a vector with a length of 256) with the cross-domain resource distribution feature vector. The two vectors can be spliced in dimension to obtain a vector with a length of 384.
[0081] Then, the concatenated vector is processed by a gating fusion mechanism. The gating fusion mechanism uses a gating unit that computes a gating signal based on the values of each dimension of the concatenated vector. The gating signal is a vector of the same length as the concatenated vector, with values between 0 and 1 for each dimension. For some dimensions of the concatenated vector, if the corresponding dimension of the gating signal is close to 0, it means that the feature of that dimension is redundant and will be suppressed. If the corresponding dimension of the gating signal is close to 1, it means that the feature of that dimension is important for enhancing the cross-domain collaborative signal and will be preserved and enhanced.
[0082] After processing by the gating fusion mechanism, the resource compatibility feature of node A is finally generated. This feature takes into account both the local compatibility of node A with its neighbor nodes and the distribution of resources across the entire cross-domain, allowing for accurate quantification of the resource compatibility of node A in a multi-domain computing environment.
[0083] Step S123: Perform nonlinear coupling analysis of the storage space occupancy rate and network bandwidth utilization through a multilayer perceptron network to generate a task adaptation degree feature, which is used to predict the computing throughput support capability of the computing resource node for the target service.
[0084] In the power grid industry, different target services have different requirements for the storage space and network bandwidth of computing resource nodes. In order to predict the computing throughput support capability of the computing resource node for the target service, nonlinear coupling analysis of the storage space occupancy rate and network bandwidth utilization needs to be performed through a multilayer perceptron network.
[0085] Taking node A as an example, its storage space occupancy rate is 56.25%, and its network bandwidth utilization is 60%. These two indicators are input into the multilayer perceptron network. The multilayer perceptron network is a type of feedforward neural network composed of an input layer, a hidden layer, and an output layer.
[0086] The input layer receives the two input values of storage space occupancy rate and network bandwidth utilization. Suppose the input layer has 2 neurons, corresponding to the storage space occupancy rate and network bandwidth utilization. The hidden layer has multiple neurons, such as 10 neurons. In the hidden layer, each neuron performs weighted summation on the output of the input layer and processes it through a nonlinear activation function (such as the ReLU function).
[0087] For the first neuron of the hidden layer, the two output values of the input layer can be multiplied by the corresponding weights respectively, and then summed. Assuming that the weight corresponding to the storage space occupancy is 0.6 and the weight corresponding to the network bandwidth utilization is 0.4, the weighted sum result is 56.25% x 0.6 + 60% x 0.4 = 57.75%. Then, the weighted sum result is processed through the ReLU function. If the result is greater than 0, it remains unchanged; if the result is less than 0, it becomes 0.
[0088] The other neurons of the hidden layer also perform similar calculations, and each neuron uses different weights to capture different features of the input data. After processing by the hidden layer, a vector of length 10 is obtained, which contains the nonlinear combination information of the storage space occupancy and the network bandwidth utilization.
[0089] The output layer receives the output of the hidden layer and converts it into a task adaptation feature. Assuming that the output layer has 1 neuron, the output of the hidden layer can be weighted and summed, and processed through a linear activation function (such as an identity function). The final output value is the task adaptation feature of node A, which can be used to predict the computing throughput support capability of node A for the target business. For example, if the task adaptation feature value is high, it means that node A can better support the computing throughput demand of the target business under the current storage space occupancy and network bandwidth utilization.
[0090] Step S124: Tensor splicing the dynamic load fluctuation feature, resource compatibility feature and task adaptation feature to generate a multi-dimensional resource feature vector of each computing resource node.
[0091] After obtaining the dynamic load fluctuation feature, resource compatibility feature and task adaptation feature of node A, they are tensor spliced to generate a multi-dimensional resource feature vector of node A.
[0092] In this embodiment, it is assumed that the dynamic load fluctuation feature is a multi-dimensional vector of length 128, the resource compatibility feature is a multi-dimensional vector of length 384, and the task adaptation feature is a vector of length 1. Then, they are spliced in dimension to obtain a multi-dimensional resource feature vector of length 128 + 384 + 1 = 513.
[0093] The multi-dimensional resource feature vector contains information such as the dynamic load fluctuation, resource compatibility and computing throughput support capability of node A for the target business. The same method is applied to node B and node C to generate their respective multi-dimensional resource feature vectors.
[0094] Step S130: Based on the preset intelligent aggregation strategy network, the multi-dimensional resource feature vectors are processed for multi-domain resource correlation analysis, the collaborative matching weight between cross-domain computing resources is determined, and the dispersed computing resource nodes are aggregated into a virtualized resource pool according to the collaborative matching weight.
[0095] After obtaining the multi-dimensional resource feature vectors of each computing resource node, it is necessary to process these vectors for multi-domain resource correlation analysis based on the preset intelligent aggregation strategy network, so as to determine the collaborative matching weight between cross-domain computing resources, and aggregate the dispersed computing resource nodes into a virtualized resource pool.
[0096] Step S131: The multi-dimensional resource feature vectors are input into the cross-attention encoding layer of the intelligent aggregation strategy network, the multi-dimensional resource feature vectors of each computing resource node are executed for cross-domain bidirectional attention calculation, and a similarity correlation matrix containing inter-domain node feature similarity is generated.
[0097] The multi-dimensional resource feature vectors of nodes A, B, C, etc. are input into the cross-attention encoding layer of the intelligent aggregation strategy network. In the cross-attention encoding layer, cross-domain bidirectional attention calculation can be performed on the multi-dimensional resource feature vectors of each computing resource node.
[0098] Taking nodes A and B as an example, first, forward attention calculation is performed. The similarity between the multi-dimensional resource feature vector of node A and the multi-dimensional resource feature vector of node B is calculated. Dot product operation can be used to calculate the similarity, and then the similarity is converted into a probability distribution through a softmax function to obtain a forward attention score, which represents the importance of the information of node B in generating the association between node A and node B.
[0099] Then, reverse attention calculation is performed. The similarity between the multi-dimensional resource feature vector of node B and the multi-dimensional resource feature vector of node A is calculated, and a reverse attention score is obtained through a softmax function.
[0100] The forward and reverse attention scores are integrated to obtain the cross-domain bidirectional attention score of node A and node B. In the same way, the cross-domain bidirectional attention scores between all node pairs are calculated.
[0101] The above cross-domain bidirectional attention scores are arranged into a matrix, which is the similarity correlation matrix containing inter-domain node feature similarity. Assuming that there are three nodes A, B, and C, the similarity correlation matrix is a 3x3 matrix. The element in the ith row and jth column of the matrix represents the cross-domain bidirectional attention score of node i and node j, reflecting the feature similarity between them.
[0102] Step S132: performing a sparsification processing based on threshold filtering on the similarity correlation matrix, retaining an associated edge between each computing resource node and a cross-domain node with a similarity higher than a dynamic threshold, generating a sparsified correlation graph, inputting the sparsified correlation graph into the differentiable graph ranking network, performing end-to-end gradient backpropagation training based on the multi-dimensional resource feature vector of the node and the associated edge weight, and generating a collaborative matching weight of each cross-domain associated edge.
[0103] After obtaining the similarity correlation matrix, a sparsification processing based on threshold filtering is performed thereon. The dynamic threshold is dynamically determined according to the element values in the similarity correlation matrix. For example, the average value of all elements in the similarity correlation matrix can be calculated, and the average value is taken as the dynamic threshold.
[0104] For each element in the similarity correlation matrix, if the value thereof is higher than the dynamic threshold, the corresponding associated edge is retained; if the value thereof is lower than the dynamic threshold, the corresponding associated edge is deleted. Taking a similarity correlation matrix of 3 nodes A, B and C as an example, it is assumed that the similarity correlation matrix is as follows:
[0105] |1.0 0.8 0.3|
[0106] |0.8 1.0 0.7|
[0107] |0.3 0.7 1.0|
[0108] The dynamic threshold is 0.6, so the associated edges between node A and node B, and between node B and node C are retained, and the associated edge between node A and node C is deleted, generating a sparsified correlation graph.
[0109] The sparsified correlation graph is input into the differentiable graph ranking network. The differentiable graph ranking network can perform end-to-end gradient backpropagation training based on the multi-dimensional resource feature vector of the node and the associated edge weight. During the training process, the differentiable graph ranking network continuously adjusts the weight of the associated edge, so that the association between nodes is more reasonable.
[0110] Specifically, the differentiable graph ranking network calculates a loss function according to the multi-dimensional resource feature vector of the node and the current associated edge weight. The loss function represents the difference between the current associated edge weight and the ideal association. Through the gradient backpropagation algorithm, the gradient of the loss function with respect to the associated edge weight is calculated, and then the weight of the associated edge is updated according to the gradient. After multiple training iterations, the collaborative matching weight of each cross-domain associated edge is finally generated.
[0111] Step S133: updating the edge weight of the sparse association graph according to the collaborative matching weight, generating a weighted cross-domain resource association graph, inputting the cross-domain resource association graph into an overlapping community discovery algorithm, identifying a node cluster with stable cooperative relationship based on the principle of modularity maximization, and generating a plurality of candidate resource clusters.
[0112] For example, in the sparse association graph, the original weight of the association edge between node A and node B is 0.8, and the collaborative matching weight obtained through training is 0.9. Then the weight of the association edge is updated to 0.9.
[0113] After updating the edge weight, a weighted cross-domain resource association graph is generated. For example, the cross-domain resource association graph is input into an overlapping community discovery algorithm. The goal of the overlapping community discovery algorithm is to identify a node cluster with stable cooperative relationship based on the principle of modularity maximization.
[0114] Modularity is an index for measuring the quality of community structure in a graph, which represents the difference between the sum of weights of edges within the community and the sum of weights of edges in a random case. The overlapping community discovery algorithm will constantly try to divide nodes into different communities to maximize modularity.
[0115] For example, through the calculation of the algorithm, it is found that node A and node B often interact with data, and their collaborative matching weight is high, so they are divided into a community; node C has a strong cooperative relationship with some nodes in another community, and is divided into another community. In this way, a plurality of candidate resource clusters are generated.
[0116] Step S134: performing mean pooling operation of multi-dimensional resource feature vector on nodes in each candidate resource cluster to generate a cluster-level resource feature vector, screening a target cluster according to the cosine similarity between the cluster-level resource feature vector and a preset global resource demand template, and generating a target resource cluster set.
[0117] For each candidate resource cluster, mean pooling operation of multi-dimensional resource feature vector is performed on nodes in the cluster. Taking a candidate resource cluster containing node A and node B as an example, the multi-dimensional resource feature vector of node A is a vector of length 513, and the multi-dimensional resource feature vector of node B is also a vector of length 513.
[0118] The values of the corresponding dimensions of the multi-dimensional resource feature vectors of node A and node B are added, and then divided by the number of nodes (here, 2), to obtain a cluster-level resource feature vector. For example, for the first dimension, the feature vector value of node A is 0.5, and the feature vector value of node B is 0.6, so the value of the first dimension of the cluster-level resource feature vector is (0.5+0.6) / 2=0.55.
[0119] After obtaining the cluster-level resource feature vector of each candidate resource cluster, the cosine similarity between them and the preset global resource demand template is calculated. The global resource demand template is a multidimensional vector preset according to the business demand of the entire power grid industry.
[0120] The cosine similarity is an index for measuring the cosine value of the included angle between two vectors. The closer the value is to 1, the more similar the two vectors are. For example, the cosine similarity of the cluster-level resource feature vector of candidate resource cluster 1 to the global resource demand template is 0.8, and the cosine similarity of candidate resource cluster 2 is 0.6.
[0121] According to the cosine similarity, the target cluster is screened, and a similarity threshold is set, for example, 0.7. The candidate resource clusters with a cosine similarity higher than 0.7 are screened out to form a target resource cluster set. Assuming that among all candidate resource clusters, the cosine similarities of candidate resource cluster 1, candidate resource cluster 3, and candidate resource cluster 5 are 0.8, 0.85, and 0.72 respectively, all of which are higher than the set threshold 0.7, then these three candidate resource clusters are included in the target resource cluster set.
[0122] Step S135: performing a virtualization resource packaging operation on each cluster in the target resource cluster set to generate a logical resource unit with a unified interface protocol, establishing a bidirectional resource redundancy channel between the logical resource units based on the collaborative matching weight, and generating a virtualization resource pool with fault tolerance capability.
[0123] For each cluster in the target resource cluster set, a virtualization resource packaging operation is performed. Taking candidate resource cluster 1 as an example, the cluster includes node A and node B. First, the computing resources (such as computing power, storage space, network bandwidth, etc.) of node A and node B are abstracted and packaged.
[0124] For the computing power resources of node A, the available computing core number, computing power, and other information are integrated and converted into a standardized representation according to certain rules. For example, the 32 computing cores of node A are abstracted into a resource unit with a specific computing power value according to their different performance indicators. Similarly, the storage space and network bandwidth of node A are also packaged in a similar manner.
[0125] For node B, the same operation is also performed. Then, the packaged resources of node A and node B are integrated to generate a logical resource unit with a unified interface protocol. The unified interface protocol specifies how external systems interact with the logical resource unit, such as how to request computing resources, how to read data in the storage space, etc.
[0126] Next, based on the collaborative matching weight, a bidirectional resource redundancy channel is established between the logical resource units. The collaborative matching weight between each logical resource unit in the target resource cluster set is extracted to generate an inter-unit weight connection relationship table. Assuming that there are three logical resource units in the target resource cluster set, which are from candidate resource cluster 1, candidate resource cluster 3 and candidate resource cluster 5, the collaborative matching weight between them is calculated and arranged to form the following inter-unit weight connection relationship table:
[0127] The collaborative matching weight between logical resource unit 1 and logical resource unit 2 is 0.8; the collaborative matching weight between logical resource unit 1 and logical resource unit 3 is 0.7; and the collaborative matching weight between logical resource unit 2 and logical resource unit 3 is 0.6.
[0128] According to the weight connection relationship table, the top K weight screening is performed on each logical resource unit. Assuming that K = 2, for logical resource unit 1, the collaborative matching weight with logical resource unit 2 and logical resource unit 3 is higher, so logical resource unit 2 and logical resource unit 3 are taken as the redundancy associated units of logical resource unit 1 to generate a redundancy associated unit list corresponding to logical resource unit 1. Similarly, a redundancy associated unit list is generated for logical resource unit 2 and logical resource unit 3 respectively.
[0129] The units in the redundancy associated unit list are verified by protocol handshake with the current logical resource unit. Taking logical resource unit 1 and its redundancy associated unit logical resource unit 2 as an example, a specific verification message is sent to check whether they can communicate in accordance with a unified interface protocol. If the verification is successful, an interoperable redundancy unit pair is generated. The resource mirror synchronization operation is performed on each redundancy unit pair to copy the state information and data of the current logical resource unit to the redundancy associated unit to generate a real-time state consistent redundancy resource copy.
[0130] Based on the preset fault switching strategy, a priority routing identifier is configured for each redundancy resource copy. For example, for the redundancy resource copies of logical resource unit 1, logical resource unit 2 and logical resource unit 3, according to their collaborative matching weight with logical resource unit 1 and other factors, logical resource unit 2 is configured with a priority of 1 and logical resource unit 3 is configured with a priority of 2. A redundancy channel configuration table with priority markers is generated and injected into the load balancer of the virtualization resource pool. The load balancer generates a dynamic traffic distribution rule according to the redundancy channel configuration table, and when a task request occurs, the task is preferentially assigned to the normally working logical resource unit, and if the unit fails, the task is switched to the redundancy resource copy according to the priority order.
[0131] During the running of the virtualized resource pool, the heartbeat signals of each logical resource unit are monitored in real time. The heartbeat signal is a signal periodically sent by the logical resource unit to indicate its normal operation. By monitoring these signals, a unit health state matrix is generated. For example, each element in the matrix represents the health state of a logical resource unit, 1 indicating normal and 0 indicating failure. When a heartbeat signal timeout exception is detected, it indicates that a certain logical resource unit has failed, and the task flow of the failed unit is switched to the redundant resource copy with the highest priority according to the redundant channel configuration table, and an updated resource scheduling path is generated. Assuming that logical resource unit 1 fails, the load balancer switches the tasks originally assigned to logical resource unit 1 to logical resource unit 2 with a priority of 1 according to the redundant channel configuration table, and updates the resource scheduling path to ensure that the tasks can continue to be executed normally. In this way, a fault-tolerant virtualized resource pool is generated, improving the reliability and stability of the entire power grid computing environment.
[0132] Step S140: Perform dynamic resource adaptation processing on the aggregated resources in the virtualized resource pool, generate a resource scheduling topology structure that matches the target business demand, and deploy the resource scheduling topology structure to the multi-domain computing environment to trigger resource coordination services.
[0133] Step S141: Input the real-time task characteristics of the target business demand into the third candidate subset of the virtualized resource pool, perform transmission path traversal simulation based on the network bandwidth utilization characteristics of the logical resource units in the third candidate subset, and generate an end-to-end delay prediction value for each logical resource unit to the business access point.
[0134] In the power grid industry, different target businesses have different real-time task characteristics. For example, the power failure repair business needs to obtain the relevant data of the fault point in real time, and the real-time requirement for data transmission is extremely high; while the periodic inspection data uploading business of power equipment has relatively low real-time requirements. Assuming that the current target business is power failure repair, its real-time task characteristics include the requirement that the data transmission delay be within 100 milliseconds.
[0135] The third candidate subset of the virtualized resource pool is a collection of logical resource units obtained after the previous steps of screening and processing. For each logical resource unit in the third candidate subset, transmission path traversal simulation is performed based on its network bandwidth utilization characteristics.
[0136] Take logical resource unit A as an example, which is located in a computing node of a substation, and the service access point is the server of the fault repair command center. The network bandwidth utilization of logical resource unit A is 60%, which represents the current network bandwidth usage. In the transmission path traversal simulation, all possible transmission paths from logical resource unit A to the service access point are considered. These paths may pass through multiple network nodes and links, each of which has its own transmission delay and bandwidth limit.
[0137] During the simulation, the transmission delay of data on each path is calculated according to the network bandwidth utilization of logical resource unit A and the state of nodes and links on each transmission path. For example, path 1 passes through 3 network nodes, each with a processing delay of 10 ms, 15 ms, and 20 ms, and the link transmission delay is 30 ms. Due to the high network bandwidth utilization of logical resource unit A, the effective bandwidth of the link may be reduced, thereby increasing the transmission delay. Assuming that the delay is thus increased by 10 ms, the total delay of path 1 is 10+15+20+30+10=85 ms. In the same way, the delays of all possible paths are calculated, and the minimum value is taken as the end-to-end delay prediction value of logical resource unit A to the service access point. The same calculation is performed for all logical resource units in the third candidate subset to obtain the end-to-end delay prediction value of each logical resource unit to the service access point.
[0138] Step S142: According to the node-by-node comparison between the end-to-end delay prediction value and the delay upper limit in the real-time task characteristics, logical resource units whose delay prediction values continuously exceed the upper limit are removed, and a fourth candidate subset that complies with the delay is generated.
[0139] The end-to-end delay prediction value of each logical resource unit is compared with the delay upper limit in the real-time task characteristics of the target service. Taking the power fault repair service as an example, the delay upper limit is 100 ms.
[0140] For logical resource unit A in the third candidate subset, the end-to-end delay prediction value is 85 ms, which is less than the delay upper limit and meets the requirements; the end-to-end delay prediction value of logical resource unit B is 110 ms, which exceeds the delay upper limit. To ensure the accuracy of the screening, multiple consecutive comparisons are performed. Assuming that 3 consecutive delay predictions and comparisons are performed, the delay prediction value of logical resource unit B exceeds 100 ms in these 3 comparisons, and then logical resource unit B is removed from the third candidate subset.
[0141] After such a screening process, all logical resource units whose delay prediction values continuously exceed the upper limit are removed, and the remaining logical resource units form a fourth candidate subset that complies with the delay, which contains logical resource units that can meet the real-time requirements of the target service in terms of data transmission delay.
[0142] Step S143: Perform sliding window matching between the fluctuation feature of the computing power load rate of the logical resource units in the fourth candidate subset and the floating point operation demand of the computing intensive task feature, calculate the cumulative distribution function of the computing power supply margin in each window, and generate a fifth candidate subset that meets the peak demand of the task.
[0143] In the power grid industry, some businesses belong to computing intensive tasks, such as power system flow calculation, fault analysis, etc., which require a large number of floating point operations. Assuming that the target business is to perform a large-scale power system flow calculation, the floating point operation demand of the computing intensive task feature has different requirements in different time periods, for example, 10 million floating point operations per second are required at the beginning, 15 million floating point operations per second are required in the middle, and 8 million floating point operations per second are required at the end.
[0144] For each logical resource unit in the fourth candidate subset, perform sliding window matching between its computing power load rate fluctuation feature and the floating point operation demand of the computing intensive task feature. Take logical resource unit C as an example, its computing power load rate fluctuation feature records the change of computing power load rate over time within a period of time. Set the size of a sliding window, for example, 10 minutes, and slide the sliding window on the time axis, each time by a time step, for example, 1 minute.
[0145] In each window, calculate the current computing power supply capacity of logical resource unit C according to its computing power load rate fluctuation feature, and then compare it with the floating point operation demand of the computing intensive task feature corresponding to the window to calculate the computing power supply margin. For example, in a certain 10-minute window, the computing power supply capacity of logical resource unit C is 12 million floating point operations per second, and the floating point operation demand of the computing intensive task feature corresponding to the window is 10 million floating point operations per second, so the computing power supply margin is 2 million floating point operations per second.
[0146] Statistically analyze the computing power supply margin in each window to calculate its cumulative distribution function. The cumulative distribution function represents the probability that the computing power supply margin is less than or equal to a certain value. For example, the cumulative distribution function shows that the probability that the computing power supply margin is less than or equal to 1 million floating point operations per second is 0.2, the probability that the computing power supply margin is less than or equal to 2 million floating point operations per second is 0.5, and so on.
[0147] A threshold of peak demand of a task is set, for example, requiring that the computing power supply margin can meet the maximum floating point operation demand of the task at least 90% of the time. According to the cumulative distribution function, logical resource units that meet the threshold requirement are screened out, and these logical resource units form a fifth candidate subset that meets the peak demand of the task. The logical resource units in the fifth candidate subset can better cope with the computing-intensive task demand of the target business in terms of computing power.
[0148] Step S144: Perform spatial continuity analysis on the storage space occupancy characteristics of the logical resource units in the fifth candidate subset, identify storage fragmentation areas, and perform block remapping operations of virtual storage volumes to generate a sixth candidate subset optimized for storage continuity.
[0149] For each logical resource unit in the fifth candidate subset, perform spatial continuity analysis on its storage space occupancy characteristics. Take logical resource unit D as an example, its storage space occupancy may have fragmentation problems. Storage space fragmentation refers to the fact that storage data is scattered in different storage blocks, resulting in reduced data read-write efficiency.
[0150] By analyzing the storage space occupancy characteristics of logical resource unit D, storage fragmentation areas are identified. This can be identified by checking the usage of storage blocks and the distribution of data. For example, it is found that there are multiple small idle storage blocks scattered throughout the storage space, and these small idle storage blocks cannot meet the storage needs of a larger data block, indicating that there are storage fragmentation areas.
[0151] For the identified storage fragmentation areas, perform block remapping operations of virtual storage volumes. First, organize and migrate the data in the scattered storage blocks. For example, some small, scattered data blocks are combined into a larger continuous storage area. Then, update the mapping table of the virtual storage volume to record the new location information of the data blocks. In this way, the originally scattered data blocks are reorganized into continuous storage areas, improving the utilization of storage space and data read-write efficiency.
[0152] All logical resource units in the fifth candidate subset are subjected to such spatial continuity analysis and block remapping operations. After processing, the logical resource units that meet the storage continuity requirement form a sixth candidate subset optimized for storage continuity. The logical resource units in the sixth candidate subset have been optimized in terms of storage space and can better support the data storage needs of the target business.
[0153] Step S145: Construct an initial resource topology graph based on the sixth candidate subset. The nodes of the initial resource topology graph represent the logical resource units optimized for storage continuity, and the edges represent the historical average bandwidth of cross-domain communication links.
[0154] The initial resource topology graph is built according to the logical resource units in the sixth candidate subset. The nodes in the graph are the logical resource units that store continuous optimization, such as logical resource unit E, logical resource unit F, etc. The edges represent the historical average bandwidth of the cross-domain communication link.
[0155] For the cross-domain communication link between logical resource unit E and logical resource unit F, the historical average bandwidth of the link is calculated by collecting the historical communication data between them. Assuming that the historical average bandwidth between logical resource unit E and logical resource unit F is 500 Mbps after statistics, the edge connecting logical resource unit E and logical resource unit F is marked as 500 Mbps in the initial resource topology graph.
[0156] In the same way, the historical average bandwidth of the cross-domain communication link between all logical resource units in the sixth candidate subset is determined and marked in the graph, thereby building the initial resource topology graph, which intuitively shows the communication relationship and bandwidth between the logical resource units that store continuous optimization.
[0157] Step S146: input the initial resource topology graph into the multi-objective optimization module, and simultaneously optimize the computing resource utilization, storage balance, and network load variance, and iteratively adjust the edge weight and node clustering center by gradient descent algorithm to generate an intermediate topology graph after optimization of the edge weight.
[0158] The initial resource topology graph is input into the multi-objective optimization module. The goal of this module is to simultaneously optimize the computing resource utilization, storage balance, and network load variance.
[0159] The computing resource utilization refers to the proportion of the computing resources of the logical resource unit that are actually used, the storage balance refers to the balance degree of the storage space usage between the logical resource units, and the network load variance refers to the dispersion degree of the network load of each cross-domain communication link.
[0160] In the multi-objective optimization module, the gradient descent algorithm is used to iteratively adjust the edge weight and node clustering center. The gradient descent algorithm is an optimization algorithm for finding the minimum value of a function. First, a target function is defined, which considers the computing resource utilization, storage balance, and network load variance. For example, the target function can be expressed as: target function value = computing resource utilization penalty term + storage balance penalty term + network load variance penalty term.
[0161] In each iteration, the gradient of the objective function with respect to the edge weights and node cluster centers is calculated. The gradient represents the rate and direction of change of the objective function at the current point. Based on the direction of the gradient, the values of the edge weights and node cluster centers are adjusted so that the value of the objective function gradually decreases. For example, if the computational resource utilization penalty term is large, it indicates that the computational resource utilization is low. By adjusting the edge weights and node cluster centers, the task is attempted to be assigned to the logical resource unit with lower computational resource utilization, thereby improving the overall computational resource utilization.
[0162] After multiple iterations, when the value of the objective function converges to a small value, the iteration is stopped, and the intermediate topology graph after edge weight optimization is generated. In this intermediate topology graph, the edge weights and node cluster centers are optimized and adjusted, so that the computational resource utilization, storage balance, and network load variance are all well optimized.
[0163] Step S147: Hierarchical clustering of the nodes of the intermediate topology graph according to the computational power load rate fluctuation characteristics to generate a hierarchical cascading structure containing a core computing layer, a buffer storage layer, and an edge access layer, and assign a priority routing protocol to the bidirectional data channel between the layers.
[0164] The nodes of the intermediate topology graph after edge weight optimization are hierarchically clustered according to the computational power load rate fluctuation characteristics. Hierarchical clustering is a method of gradually merging data points into different levels of clusters.
[0165] According to the computational power load rate fluctuation characteristics of the nodes, the nodes are divided into different categories. For example, nodes with small computational power load rate fluctuations and strong computing capabilities are classified into one category, which can form the core computing layer; nodes with large storage space and capable of buffering data are classified into one category, forming the buffer storage layer; nodes located at the edge of the network and responsible for connecting with external business access points are classified into one category, forming the edge access layer.
[0166] A hierarchical cascading structure containing a core computing layer, a buffer storage layer, and an edge access layer is generated. In this hierarchical cascading structure, the core computing layer is mainly responsible for processing computationally intensive tasks, the buffer storage layer is responsible for storing intermediate and temporary data, and the edge access layer is responsible for data interaction with external business access points.
[0167] A priority routing protocol is assigned to the bidirectional data channel between the layers. For example, for the data channel between the core computing layer and the buffer storage layer, since the tasks of the core computing layer have high real-time requirements for data, a high-priority routing protocol is assigned to this channel to ensure fast data transmission. For the data channel between the buffer storage layer and the edge access layer, appropriate priority routing protocols are assigned according to the type and importance of the data. In this way, by reasonably assigning priority routing protocols, the efficiency and reliability of data transmission are improved.
[0168] Step S148: Embedding resource state tracking agents in the hierarchical cascading structure, collecting task execution time delay and resource consumption rate of each layer node in real time, and generating topology health status indicators.
[0169] Resource state tracking agents are embedded in the hierarchical cascading structure. A resource state tracking agent is a software module that can monitor the running state of each layer node in real time.
[0170] For the nodes of the core computing layer, the resource state tracking agent collects the task execution time delay, i.e. the time taken from the start of the task to the completion of the task, and the resource consumption rate, such as CPU utilization, memory usage, etc. For the nodes of the buffer storage layer, the data read / write time delay and storage utilization are collected. For the nodes of the edge access layer, the end-to-end delay and network bandwidth utilization between the node and the external service access point are collected.
[0171] According to the collected data, topology health status indicators are generated. For example, a comprehensive indicator is defined, which takes into account factors such as task execution time delay, resource consumption rate, etc. of each layer node. The indicator can be calculated by weighted summation, for example, topology health status indicator = core computing layer task execution time delay weight x core computing layer task execution time delay + buffer storage layer storage utilization weight x buffer storage layer storage utilization + edge access layer end-to-end delay weight x edge access layer end-to-end delay, etc.
[0172] By monitoring and calculating the topology health status indicators in real time, the running status of the hierarchical cascading structure can be understood in a timely manner.
[0173] Step S149: When the time delay or consumption rate of any layer in the topology health status indicator exceeds the dynamic threshold, trigger local topology reconstruction, re-execute edge weight adjustment and node clustering according to the resource state of the current sixth candidate subset, and generate an updated resource scheduling topology structure.
[0174] A dynamic threshold is set to determine whether the topology health status indicator is normal. The dynamic threshold is dynamically adjusted according to the historical running data of the system and the current business requirements. For example, for the task execution time delay of the core computing layer, a dynamic threshold of 500 milliseconds is set according to the average time delay in the past period of time and the real-time requirements of the business; for the storage utilization of the buffer storage layer, a dynamic threshold of 80% is set; for the end-to-end delay of the edge access layer, a dynamic threshold of 200 milliseconds is set.
[0175] When the latency or consumption rate of any layer in the topology health status indicator exceeds the dynamic threshold, local topology reconstruction is triggered. Assuming that the task execution latency of a node in the core computing layer exceeds the dynamic threshold of 500 milliseconds, it indicates that the computing resources of the node may have bottlenecks and cannot meet business demands. At this time, the edge weight adjustment and node clustering need to be re-executed according to the current resource status of the sixth candidate subset.
[0176] Firstly, real-time resource status data of logical resource units in the sixth candidate subset is re-collected, including computing power load rate, storage space occupancy rate, network bandwidth utilization rate, etc. These data reflect the actual running conditions of the current logical resource units.
[0177] Then, the initial resource topology graph is constructed again. Since the resource status has changed, the historical average bandwidth of the cross-domain communication link between the logical resource units may also be different, so the weight of the edge needs to be recalculated and updated. For example, the historical average bandwidth between logical resource unit A and logical resource unit B was originally 500 Mbps, but due to the change of current network load, it changes to 450 Mbps after recalculation, so the weight of the edge connecting the two nodes in the new initial resource topology graph is updated to 450 Mbps.
[0178] Then, the new initial resource topology graph is input into the multi-objective optimization module. As before, the computing resource utilization rate, storage balance degree, and network load variance are simultaneously optimized. The edge weight and node clustering center are iteratively adjusted by the gradient descent algorithm. During the iteration process, the gradient of the objective function with respect to the edge weight and node clustering center is recalculated according to the new resource status data. For example, if the computing power load rate of a node suddenly increases, resulting in an increase in the computing resource utilization rate penalty term, the gradient descent algorithm will try to adjust the edge weight and node clustering center to distribute tasks to other nodes with lower computing power load rate, in order to improve the overall computing resource utilization rate.
[0179] After several iterations, when the value of the objective function converges to a small value, the iteration is stopped, and a new intermediate topology graph with optimized edge weight is generated.
[0180] After that, the nodes of the new intermediate topology graph are hierarchically clustered according to the computing power load rate fluctuation characteristics. Due to the change of resource status, the classification of nodes may change. For example, a node originally belonging to the core computing layer may be reclassified to the buffer storage layer or the edge access layer due to the increase in its computing power load rate fluctuation. The hierarchical cascading structure including the core computing layer, the buffer storage layer, and the edge access layer is regenerated, and the priority routing protocol for the bidirectional data channel between the layers is re-assigned.
[0181] Finally, an updated resource scheduling topology is generated, which takes into account the real-time resource status of each logical resource unit and can better meet the business requirements, improving the performance and reliability of the system.
[0182] Step S150: deploying the resource scheduling topology to the multi-domain computing environment to trigger resource coordination services.
[0183] Step S151: converting the hierarchical cascade structure of the resource scheduling topology into a cross-domain resource orchestration instruction set, which includes virtual machine deployment coordinates, storage volume mapping relationships, and priority routing tables for each layer node.
[0184] The hierarchical cascade structure of the updated resource scheduling topology is converted to generate a cross-domain resource orchestration instruction set. For the nodes of the core computing layer, their virtual machine deployment coordinates are determined. Virtual machine deployment coordinates specify where the virtual machine should be deployed on a physical computing node and the specific location on that node. For example, a virtual machine in the core computing layer needs to be deployed on the 3rd computing core of physical node P1 and use a specific memory area on that node, so the virtual machine deployment coordinates record these detailed information.
[0185] At the same time, the storage volume mapping relationship is determined. The storage volume mapping relationship defines how the virtual machine connects and uses the storage resource. For example, a virtual machine in the core computing layer needs to access the storage volume V3 stored on the buffer storage layer node S2, so the storage volume mapping relationship clearly defines this mapping.
[0186] For the bidirectional data channels between layers, a priority routing table is generated according to the previously assigned priority routing protocol. The priority routing table records the priority information of different data channels to guide the transmission of data. For example, the data channel priority between the core computing layer and the buffer storage layer is high, so this information is clearly recorded in the priority routing table.
[0187] These virtual machine deployment coordinates, storage volume mapping relationships, and priority routing tables are integrated to form a cross-domain resource orchestration instruction set.
[0188] Step S152: injecting the priority routing table into the forwarding table of each domain border router through the policy delivery interface of the software-defined network controller to generate cross-domain bandwidth guarantee channels.
[0189] The priority routing table is injected into the forwarding table of each domain border router through the policy delivery interface of the software-defined network controller. The software-defined network controller is a centralized network management device that can send configuration information to each network device through the policy delivery interface.
[0190] The information in the priority routing table, such as the priority of different data channels, bandwidth allocation, etc., is sent to each domain border router. After each domain border router receives this information, it updates its own forwarding table. For example, after a domain border router receives information that the data channel between the core computing layer and the buffer storage layer has high priority, it allocates higher priority and more bandwidth resources to this channel in the forwarding table.
[0191] In this way, cross-domain bandwidth guarantee channels are generated. These channels can ensure that high-priority data is given priority in the transmission process, avoiding data transmission delays due to network congestion.
[0192] Step S153: Distribute the virtual machine deployment coordinates and storage volume mapping relationship to the target computing resource node through the resource scheduling engine of the cloud management platform, and generate virtualized resource instances distributed according to the topology structure.
[0193] The virtual machine deployment coordinates and storage volume mapping relationship are distributed to the target computing resource node through the resource scheduling engine of the cloud management platform. The cloud management platform is a platform for managing cloud computing resources, and its resource scheduling engine is responsible for sending resource allocation and deployment information to each computing resource node.
[0194] For virtual machine deployment coordinates, the resource scheduling engine sends them to the corresponding physical computing node. After the physical computing node receives the information, it creates a virtual machine instance at the specified location according to the virtual machine deployment coordinates. For example, after receiving information that a virtual machine needs to be created on the 3rd computing core of physical node P1, physical node P1 will create the virtual machine as required.
[0195] For storage volume mapping relationships, the resource scheduling engine sends them to related storage nodes and virtual machine nodes. After the storage node receives the information, it configures the access permissions and mapping relationship of the storage volume; after the virtual machine node receives the information, it accesses the corresponding storage volume according to the storage volume mapping relationship.
[0196] In this way, virtualized resource instances distributed according to the topology structure are generated. These instances are distributed according to the hierarchical cascading structure of the resource scheduling topology structure, and can work better in coordination to meet business needs.
[0197] Step S154: During the operation of the virtualized resource instance, continuously collect CPU utilization of the core computing layer, IO throughput of the buffer storage layer, and end-to-end delay data of the edge access layer through the resource state tracking agent, and generate a real-time resource load matrix.
[0198] During the running of the virtualized resource instance, the resource state tracking agent continuously collects relevant data of each layer node. For the core computing layer, CPU utilization data is collected. CPU utilization reflects the computing resource usage of the core computing layer node. For example, the CPU utilization of a certain node of the core computing layer is collected every 1 minute, and recorded.
[0199] For the buffer storage layer, IO throughput data is collected. IO throughput represents the data read / write speed of the buffer storage layer node. Similarly, the IO throughput of a certain node of the buffer storage layer is collected every 1 minute.
[0200] For the edge access layer, end-to-end delay data is collected. End-to-end delay reflects the data transmission delay between the edge access layer node and the external service access point. Also, the end-to-end delay of a certain node of the edge access layer is collected every 1 minute.
[0201] The collected data is sorted to generate a real-time resource load matrix. Each row of the matrix represents a node, and each column represents a time point. The elements in the matrix are the CPU utilization, IO throughput, or end-to-end delay data of the node at the corresponding time point. For example, the first row and first column element of the matrix represents the CPU utilization of the core computing layer node 1 at the 1st minute.
[0202] Step S155: input the real-time resource load matrix into the overload determination model, and generate a computing and storage resource imbalance alarm when the CPU utilization of the core computing layer is continuously higher than the first threshold and the IO throughput of the buffer storage layer is lower than the second threshold.
[0203] The real-time resource load matrix is input into the overload determination model. The overload determination model is a pre-trained model that can determine whether the system has a resource imbalance condition according to the real-time resource load matrix.
[0204] The first threshold and the second threshold are set. For example, the first threshold is 80%, and the second threshold is 100 MB / s. When the CPU utilization of the core computing layer is continuously higher than 80%, it indicates that the computing resource of the core computing layer may be in an overload state; at the same time, the IO throughput of the buffer storage layer is lower than 100 MB / s, indicating that the data read / write capability of the buffer storage layer is not fully utilized, and there may be a problem of computing and storage resource imbalance.
[0205] When the overload determination model detects that the CPU utilization of the core computing layer is continuously higher than the first threshold and the IO throughput of the buffer storage layer is lower than the second threshold, a computing and storage resource imbalance alarm is generated. The alarm information will timely notify the system administrator so as to take corresponding measures for adjustment.
[0206] Step S156: According to the calculation of the storage resource imbalance alarm, trigger the resource rescheduling request, input the current real-time resource load matrix and the target business demand into the dynamic resource adaptation processing module again, and generate the resource scheduling topology structure after re-scheduling.
[0207] According to the calculation of the storage resource imbalance alarm, trigger the resource rescheduling request. Input the current real-time resource load matrix and the target business demand into the dynamic resource adaptation processing module again.
[0208] The dynamic resource adaptation processing module will perform resource screening, topology graph construction, optimization and clustering operations again according to the previous steps. For example, according to the data in the real-time resource load matrix, re-screen the logical resource units that meet the delay, computing power and storage requirements, construct a new initial resource topology graph, and then perform multi-objective optimization and hierarchical clustering.
[0209] After a series of processing, the resource scheduling topology structure after re-scheduling is generated, which takes into account the current resource imbalance and target business demand, and can better balance the use of computing and storage resources.
[0210] Step S157: Through the live migration engine, the task state and data cache of the virtualized resource instance are migrated losslessly according to the re-scheduled topology structure, and the priority routing table of the cross-domain bandwidth guarantee channel is updated after migration is completed, generating a resource collaborative service update version.
[0211] Step S157-1: Extract the edge access layer node list in the hierarchical cascading structure of the resource scheduling topology structure after re-scheduling, and generate a new access point address set.
[0212] Extract the edge access layer node list from the hierarchical cascading structure of the resource scheduling topology structure after re-scheduling. Each edge access layer node has its corresponding access point address. Organize these access point addresses into a set to generate a new access point address set. For example, after re-scheduling, the edge access layer has nodes E1, E2, E3, and their access point addresses are IP1, IP2, IP3, respectively. Then the new access point address set is {IP1, IP2, IP3}.
[0213] Step S157-2: Compare the new access point address set with the addresses in the historical routing table to identify new access points and invalid access points, and generate a routing change instruction set.
[0214] Compare the new access point address set with the addresses in the historical routing table. The historical routing table records the previous access point addresses and routing information. By comparing, find out the new access point addresses and invalid access point addresses.
[0215] For example, the set of access point addresses in the historical routing table is {IP0, IP1, IP4}, and the new set of access point addresses is {IP1, IP2, IP3}. Then, the new access point addresses are {IP2, IP3}, and the invalid access point addresses are {IP0, IP4}.
[0216] According to the identified new access points and invalid access points, a routing change instruction set is generated. The routing change instruction set contains instructions for deleting routing entries of invalid access points and adding routing entries of new access points.
[0217] Step S157-3: The routing update interface of the software-defined network controller is used to distribute the routing change instruction set to each domain border router, delete the routing entries corresponding to the invalid access points, and allocate guaranteed bandwidth for the new access points.
[0218] The routing update interface of the software-defined network controller is used to distribute the routing change instruction set to each domain border router. After receiving the instructions, each domain border router first deletes the routing entries corresponding to the invalid access points. For example, the routing entries related to IP0 and IP4 in the historical routing table are deleted.
[0219] Then, guaranteed bandwidth is allocated for the new access points. According to the business requirements and network resource conditions, a certain amount of bandwidth resources is allocated to each new access point to ensure its normal operation. For example, 200Mbps of guaranteed bandwidth is allocated to IP2 and IP3, respectively.
[0220] Step S157-4: After the routing update is completed, the resource state tracking agent of the edge access layer is used to verify whether the end-to-end delay of the new access points meets the real-time task feature requirements, and a delay verification result is generated.
[0221] After the routing update is completed, the resource state tracking agent of the edge access layer is used to verify the end-to-end delay of the new access points. The resource state tracking agent collects end-to-end delay data between the new access points and external business access points.
[0222] The collected delay data is compared with the real-time task feature requirements. For example, if the real-time task feature requirement is that the end-to-end delay is within 200 milliseconds, then it is checked whether the end-to-end delay of each new access point is less than 200 milliseconds.
[0223] According to the comparison result, a delay verification result is generated. The delay verification result records information about whether the delay of each new access point meets the requirements.
[0224] Step S157-5: When there is an abnormal access point in the delay verification result, the dynamic resource adaptation processing module is retriggered to replace the abnormal node from the sixth candidate subset and generate a secondary rescheduling topology structure.
[0225] When there are abnormal access points in the delay verification result, it indicates that the end-to-end delay of these access points does not meet the real-time task feature requirements and needs to be adjusted. The dynamic resource adaptation processing module is retriggered.
[0226] A suitable node is selected from the sixth candidate subset to replace the abnormal node. For example, if the end-to-end delay of the newly added access point IP2 exceeds 200 milliseconds, a node that can meet the delay requirement is selected from the sixth candidate subset to replace IP2.
[0227] The dynamic resource adaptation processing module re-performs resource screening, topology graph construction, optimization, and clustering operations according to the new node conditions to generate a secondary rescheduling topology structure.
[0228] Step S157-6: The edge access layer nodes of the secondary rescheduling topology structure are again partially migrated by the live migration engine until the delay verification results of all access points meet the requirements, and a final stable resource collaborative service version is generated.
[0229] The edge access layer nodes of the secondary rescheduling topology structure are again partially migrated by the live migration engine. The live migration engine can migrate virtual machines and data from one node to another without interrupting services.
[0230] The migrated access points are again subjected to delay verification. If there are still abnormal access points, the above steps are repeated to replace the abnormal nodes from the sixth candidate subset, generate a new rescheduling topology structure, and perform migration and verification.
[0231] Until the delay verification results of all access points meet the requirements, a final stable resource collaborative service version is generated, and the resource collaborative service of the resource collaborative service version can better meet the business requirements and improve the performance and reliability of the system.
[0232] Figure 2 A schematic diagram of exemplary hardware and software components of a multi-domain computing resource aggregation system 100 based on a virtualized user network that can implement the inventive concept according to some embodiments of the present application is shown. For example, a processor 120 can be used in the multi-domain computing resource aggregation system 100 based on a virtualized user network and used to perform the functions in the present application.
[0233] The multi-domain computing resource aggregation system 100 based on a virtualized user network can be a general-purpose server or a special-purpose server, both of which can be used to implement the multi-domain computing resource aggregation method based on a virtualized user network of the present application. Although only one server is shown in the present application, for the sake of convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0234] For example, the virtualized user network based multi-domain computing resource aggregation system 100 can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, a ROM, or a RAM, or any combination thereof. The virtualized user network based multi-domain computing resource aggregation system 100 can also include program instructions stored in a ROM, a RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The virtualized user network based multi-domain computing resource aggregation system 100 also includes an input / output (I / O) interface 150 between the computer and other input / output devices.
[0235] For ease of illustration, only one processor is described in the virtualized user network based multi-domain computing resource aggregation system 100. However, it should be noted that the virtualized user network based multi-domain computing resource aggregation system 100 in the present application can also include multiple processors, so the steps performed by one processor described in the present application can also be jointly performed or individually performed by multiple processors. For example, if the processor of the virtualized user network based multi-domain computing resource aggregation system 100 performs steps A and B, it should be understood that steps A and B can also be jointly performed by two different processors or individually performed in one processor. For example, a first processor performs step A, a second processor performs step B, or the first processor and the second processor jointly perform steps A and B.
[0236] In addition, the embodiment of the present application also provides a readable storage medium, wherein computer executable instructions are preset in the readable storage medium, and when the processor executes the computer executable instructions, the virtualized user network based multi-domain computing resource aggregation method is realized.
[0237] It should be noted that, in order to simplify the description of the present application and to help understand one or more embodiments of the present application, in the foregoing description of the embodiments of the present application, various features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A multi-domain computing resource aggregation method based on virtualized user network, characterized in that, The method comprises: Collecting a real-time resource state data set of the computing resource nodes scattered in a multi-domain environment, the real-time resource state data set comprising a computing power load rate, a storage space occupancy rate and a network bandwidth utilization rate of each computing resource node; Calling a virtualization resource mapping model to perform cross-domain resource feature extraction processing on the real-time resource state data set, to generate a multi-dimensional resource feature vector of each computing resource node, the multi-dimensional resource feature vector comprising a dynamic load fluctuation feature, a resource compatibility feature and a task adaptation degree feature; Performing multi-domain resource correlation analysis processing on the multi-dimensional resource feature vector based on a preset intelligent aggregation strategy network, to determine a collaborative matching weight between cross-domain computing resources, and aggregating the scattered computing resource nodes into a virtualization resource pool according to the collaborative matching weight; Performing dynamic resource adaptation processing on the aggregated resources in the virtualization resource pool, to generate a resource scheduling topology structure matched with a target business requirement, and deploying the resource scheduling topology structure to a multi-domain computing environment to trigger resource collaboration services; The method comprises: inputting the multi-dimensional resource feature vector into a cross-attention encoding layer of the intelligent aggregation strategy network, performing cross-domain bidirectional attention calculation on the multi-dimensional resource feature vector of each computing resource node, and generating a similarity correlation matrix comprising an inter-domain node feature similarity; performing threshold filtering-based sparsification processing on the similarity correlation matrix, retaining an associated edge between each computing resource node and a cross-domain node with a similarity higher than a dynamic threshold, generating a sparsified association relationship graph, inputting the sparsified association relationship graph into a differentiable graph sorting network, performing end-to-end gradient backpropagation training based on the multi-dimensional resource feature vector of the node and the associated edge weight, and generating a collaborative matching weight of each cross-domain associated edge; updating the edge weight of the sparsified association relationship graph according to the collaborative matching weight, generating a weighted cross-domain resource association graph, inputting the cross-domain resource association graph into an overlapping community discovery algorithm, identifying a node cluster with a stable collaboration relationship based on a modularity maximization principle, and generating a plurality of candidate resource clusters; performing mean pooling operation on the multi-dimensional resource feature vector of the nodes in each candidate resource cluster, generating a cluster-level resource feature vector, selecting a target cluster according to a cosine similarity between the cluster-level resource feature vector and a preset global resource demand template, and generating a target resource cluster set; performing virtualization resource encapsulation operation on each cluster in the target resource cluster set, generating a logical resource unit with a unified interface protocol, establishing a bidirectional resource redundancy channel between the logical resource units based on the collaborative matching weight, and generating a virtualization resource pool with fault tolerance capability. 2.The method of claim 1, wherein, The calling virtualization resource mapping model performs cross-domain resource feature extraction processing on the real-time resource state data set, generates a multi-dimensional resource feature vector of each computing resource node, including: Input the real-time resource state data set into the feature encoding layer of the virtualization resource mapping model, extract the periodic fluctuation mode of the computing power load rate through a time series convolution network, and generate dynamic load fluctuation features; Analyze the resource interaction dependency relationship between each computing resource node through a graph attention network, generate resource compatibility features, and the resource compatibility features are used to quantify the hardware configuration difference and protocol interoperability of different domain computing resource nodes; Perform nonlinear coupling analysis on the storage space occupancy rate and network bandwidth utilization rate through a multilayer perceptron network, generate task adaptation degree features, and the task adaptation degree features are used to predict the computing throughput support capability of the computing resource node to the target business; The dynamic load fluctuation features, resource compatibility features and task adaptation degree features are spliced to generate a multi-dimensional resource feature vector of each computing resource node. 3.The method of claim 2, wherein, The periodic fluctuation mode of the computing power load rate is extracted through a time series convolution network to generate dynamic load fluctuation features, including: Periodically divide the computing power load rate according to a preset time window to obtain a plurality of load rate time series segments, each load rate time series segment containing load rate sampling data of consecutive time stamps; Input the load rate time series segment into the dilated causal convolution layer, and extract short-period fluctuation features and long-period trend features in parallel using convolution kernels with different dilation coefficients; Input the short-period fluctuation features and long-period trend features into the gated recurrent unit, fuse the fluctuation modes of different time scales through the time gating mechanism, and generate a fused time series feature vector; Perform frequency spectrum analysis on the time series feature vector, identify frequency domain components with significant periodicity, retain frequency domain components matching the business period through a band-pass filter, and generate filtered frequency spectrum features; Map the frequency spectrum features back to the time domain space, eliminate the dimensional differences of different computing resource nodes through adaptive normalization processing, generate standardized dynamic load fluctuation features, and the dynamic load fluctuation features are used to represent the load change law of the computing resource node in the cross-time scale. 4.The method of claim 2, wherein, The resource interaction dependency relationship between each computing resource node is analyzed through a graph attention network to generate resource compatibility features, including: A directed weighted graph is constructed according to the historical communication records between the computing resource nodes, and the edge weight in the directed weighted graph is determined by the historical data transmission success rate and the average delay between the nodes; Random walk sampling is performed on the directed weighted graph to generate a context node sequence of each computing resource node, and the context node sequence contains neighbor nodes having a strong dependency relationship with it; Input the context node sequence into the embedding layer of the graph attention network to generate a topology awareness feature vector of each neighbor node, and the topology awareness feature vector encodes the physical connection attribute and logical cooperation relationship between the nodes; Perform bidirectional attention calculation on the feature vector of the current computing resource node and the topology-aware feature vector of the neighbor node, generate an inter-node resource collaboration score, which reflects the alignment of hardware configuration, protocol compatibility and load complementarity between nodes; Based on the collaboration score, the topology-aware feature vector of the neighbor node is dynamically weighted and aggregated to generate the local compatibility feature of the current node; Cross the local compatibility feature with the cross-domain resource distribution feature extracted by the global graph pooling layer, suppress redundant features and enhance cross-domain collaboration signals through a gating fusion mechanism to generate the final resource compatibility feature.
5. The method of claim 1, wherein, The establishment of a bidirectional resource redundancy channel between the logical resource units based on the collaboration matching weight generates a virtualized resource pool with fault tolerance capability, including: Extract the collaboration matching weight between the logical resource units to generate a weight connection relationship table between units, and perform a top-K weight screening on each logical resource unit according to the weight connection relationship table to generate a redundant associated unit list corresponding to each unit; Perform protocol handshake verification between the units in the redundant associated unit list and the current logical resource unit to generate an interoperable redundant unit pair, and perform resource mirror synchronization operation on each redundant unit pair to generate a real-time state consistent redundant resource copy; Based on the preset fault switching strategy, configure a priority routing identifier for each redundant resource copy to generate a redundant channel configuration table with priority markers, and inject the redundant channel configuration table into the load balancer of the virtualized resource pool to generate dynamic traffic distribution rules; Real-time monitoring of the heartbeat signal of each logical resource unit during the operation of the virtualized resource pool generates a unit health status matrix. When a heartbeat signal timeout exception is detected, switch the task flow of the faulty unit to the redundant resource copy with the highest priority according to the redundant channel configuration table to generate an updated resource scheduling path. 6.The method of claim 1, wherein, The dynamic resource adaptation processing of the aggregated resources in the virtualized resource pool generates a resource scheduling topology structure that matches the target business demand, including: Input the real-time task features of the target business demand into the third candidate subset of the virtualized resource pool, perform transmission path traversal simulation based on the network bandwidth utilization features of the logical resource units in the third candidate subset to generate an end-to-end delay prediction value from each logical resource unit to the business access point; According to the end-to-end delay prediction value and the delay upper limit in the real-time task feature, eliminate the logical resource units whose delay prediction value continuously exceeds the upper limit to generate a fourth candidate subset that complies with the delay; Slide window matching of the computing power load rate fluctuation feature of the logical resource units in the fourth candidate subset with the floating point operation demand of the compute-intensive task feature calculates the cumulative distribution function of the computing power supply margin in each window to generate a fifth candidate subset that meets the task peak demand; Perform spatial continuity analysis on the storage space occupancy rate feature of the logical resource units of the fifth candidate subset, identify storage fragmentation areas and perform block remapping operations of virtual storage volumes, and generate a storage-continuous optimized sixth candidate subset; Construct an initial resource topology graph based on the sixth candidate subset, with nodes representing storage-continuous optimized logical resource units and edges representing the historical average bandwidth of cross-domain communication links; Input the initial resource topology graph into a multi-objective optimization module to simultaneously optimize resource utilization, storage balance, and network load variance, iteratively adjust edge weights and node cluster centers through a gradient descent algorithm, and generate an intermediate topology graph with optimized edge weights; Hierarchical cluster the nodes of the intermediate topology graph according to the computational load rate fluctuation feature to generate a hierarchical cascading structure containing a core computing layer, a buffer storage layer, and an edge access layer, and assign priority routing protocols to bidirectional data channels between layers; Embed resource state tracking agents in the hierarchical cascading structure to collect task execution delays and resource consumption rates of nodes in each layer in real time, and generate topology health status indicators; When the delay or consumption rate of any layer in the topology health status indicators exceeds the dynamic threshold, trigger local topology reconstruction, re-execute edge weight adjustment and node clustering based on the current resource state of the sixth candidate subset, and generate an updated resource scheduling topology structure. 7.The method of claim 6, wherein, The deployment of the resource scheduling topology structure to a multi-domain computing environment to trigger resource coordination services includes: Convert the hierarchical cascading structure of the resource scheduling topology structure into a cross-domain resource orchestration instruction set, which contains virtual machine deployment coordinates, storage volume mapping relationships, and priority routing tables for each layer node; Inject the priority routing table into the forwarding table of each domain boundary router through the policy delivery interface of the software-defined network controller to generate cross-domain bandwidth guarantee channels; Distribute the virtual machine deployment coordinates and storage volume mapping relationships to target computing resource nodes through the resource scheduling engine of the cloud management platform to generate virtualized resource instances distributed according to the topology structure; During the operation of the virtualized resource instances, continuously collect CPU utilization of the core computing layer, IO throughput of the buffer storage layer, and end-to-end delay data of the edge access layer through the resource state tracking agent to generate a real-time resource load matrix; Input the real-time resource load matrix into an overload determination model, and when the CPU utilization of the core computing layer is continuously higher than the first threshold and the IO throughput of the buffer storage layer is lower than the second threshold, generate a computing and storage resource imbalance alarm; Trigger a resource rescheduling request according to the computing and storage resource imbalance alarm, re-input the current real-time resource load matrix and target business demand into the dynamic resource adaptation processing module, and generate a rescheduled resource scheduling topology structure; Migrate the task state and data cache of the virtualized resource instances according to the rescheduled topology structure through the live migration engine, and update the priority routing table of the cross-domain bandwidth guarantee channel after migration is complete to generate an updated version of the resource coordination service. 8.The method of claim 7, wherein, The task state and data cache of the virtualized resource instance are live-migrated by the hot migration engine according to the re-scheduled topology structure, and the priority routing table of the cross-domain bandwidth guarantee channel is updated after the migration is completed, a resource coordination service update version is generated, including: Extracting the edge access layer node list in the hierarchical cascade structure of the re-scheduled resource scheduling topology structure, generating a new access point address set; Differentially comparing the new access point address set with the addresses in the historical routing table, identifying new and invalid access points, and generating a routing change instruction set; Through the routing update interface of the software-defined network controller, the routing change instruction set is sent to each domain border router, the routing entries corresponding to the invalid access points are deleted, and the new access points are allocated with guaranteed bandwidth; After the routing update is completed, the resource state tracking agent of the edge access layer verifies whether the end-to-end delay of the new access points meets the real-time task feature requirements, and generates a delay verification result; When there is an abnormal access point in the delay verification result, the dynamic resource adaptation processing module is re-triggered to replace the abnormal node from the sixth candidate subset and generate a secondary re-scheduled topology structure; The edge access layer nodes of the secondary re-scheduled topology structure are again executed by the hot migration engine to perform partial migration until the delay verification results of all access points meet the requirements, and a final stable resource coordination service version is generated.
9. A multi-domain computing resource aggregation system based on virtualized user networks, characterized in that, A processor and a memory, the memory and the processor are connected, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to realize the multi-domain computing resource aggregation method based on virtualized user network in any one of claims 1-8.
Citation Information
Patent Citations
Real-time image processing method and system based on edge calculation
CN118467181A
Cloud platform computing power resource performance monitoring and real-time scheduling optimization method
CN119883651A