A deep learning-based computing power performance dynamic allocation optimization method
By utilizing deep learning technology, task phase partitioning, an improved StemGNN model, and the CBBA algorithm, the problem of insufficient accuracy and stability in the dynamic allocation of computing resources in existing technologies has been solved, achieving high-precision resource demand matching and rapid response.
Patent Information
- Application Number
- CN202610782411.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-25
AI Technical Summary
Existing technologies struggle to accurately represent the differences in resource requirements for the same task during data loading, computation bursts, memory swapping, network backhaul, and waiting/blocking phases, resulting in insufficient accuracy and scheduling stability in the dynamic allocation of computing resources.
A deep learning-based approach is adopted, utilizing task phase partitioning, an improved StemGNN model, and the CBBA algorithm. By collecting and preprocessing data, a standardized computing power operation dataset is generated. A phase semantic spectrogram embedding layer and a degradation spectrum state filtering layer are constructed, a ternary influence hypergraph is built, resource bundle bidding and consensus conflict resolution are performed, and a dynamic allocation scheme is generated.
It improves the accuracy and response speed of computing resource allocation, enhances resource utilization and scheduling stability, and reduces resource occupation due to expired bidding and task migration damage.
Smart Images

Figure CN122633394A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning technology, and in particular to a method for dynamic allocation and optimization of computing power performance based on deep learning. Background Technology
[0002] As cloud computing, edge computing, and heterogeneous computing resource pools expand in scale, online inference, batch processing analysis, image rendering, database synchronization, log retrieval, and report generation tasks run in parallel within the same computing platform. Existing technologies typically collect data on CPU utilization, GPU utilization, memory utilization, bandwidth utilization, task queue length, response latency, and container running status, and generate computing resource allocation results based on fixed thresholds, load balancing rules, resource quota adjustment strategies, and ordinary time-series prediction models.
[0003] Existing methods still use complete tasks or fixed scheduling time windows as allocation objects, making it difficult to express the differences in resource requirements of the same task during data loading, computation bursts, memory swapping, network backhaul, and waiting / blocking phases. Ordinary prediction models often output load trends based on the correlation of resource variables, failing to express the structural changes of resource variables, node degradation states, and phase resource bundle information within adjacent scheduling time windows. Existing scheduling methods mostly establish binary matching relationships between tasks and nodes, failing to uniformly incorporate link congestion, container migration impairment, and non-migratable states into the allocation process. This easily leads to problems such as accurate candidate node selection but congested transmission links, excessive task migration impairment, and expired bid records still occupying resources, resulting in insufficient accuracy and scheduling stability in dynamic allocation of computing resources.
[0004] Therefore, how to provide a method for dynamic allocation and optimization of computing power performance based on deep learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a dynamic allocation optimization method for computing power performance based on deep learning. This invention utilizes task phase partitioning, an improved StemGNN model, and the CBBA algorithm to achieve phase identification, node matching, and dynamic scheduling of computing power resources, which has the advantages of high allocation accuracy, fast response speed, and high resource utilization.
[0006] A method for dynamically allocating and optimizing computing power performance based on deep learning according to an embodiment of the present invention includes:
[0007] Collect task execution data, node resource data, link status data, container execution data, and task execution feedback data from the computing power resource pool, perform preprocessing, and generate a standardized computing power execution dataset;
[0008] Based on the standardized computing power operation dataset, resource occupancy curves, task queue length curves, and response latency curves are extracted. The Toeplitz inverse covariance structure of resource variables within adjacent scheduling time windows is calculated. Synchronous segmentation is performed and phase semantic calibration is carried out to generate task computing power phase segments.
[0009] An improved StemGNN model is constructed, which includes a phase semantic spectrogram embedding layer, a degenerate spectral state filtering layer, and a phase resource bundle readout head. The phase fragment of the task computing power is input into the improved StemGNN model to generate phase resource bundle information.
[0010] A ternary influence hypergraph is constructed based on task computing power phase segments, phase resource bundle information, node resource data, link status data, and container operation data;
[0011] The CBBA algorithm, which incorporates a Bid Time-Stamp refresh rule, is used to read the ternary influence hypergraph, perform resource bundle bidding and consensus conflict resolution, and generate candidate computing power dynamic allocation schemes.
[0012] Perform Bid Time-Stamp refresh processing and local threshold-triggered rebidding on candidate computing power dynamic allocation schemes to generate target computing power dynamic allocation schemes;
[0013] The computing power scheduling instructions are generated based on the target computing power dynamic allocation scheme, and the scheduling feedback data is collected and written into the standardized computing power operation dataset.
[0014] Optionally, the task execution data includes task number, task type, task submission time, task start time, task end time, task log identifier, task queue length, and response latency. The task log identifier includes data loading identifier, computation start identifier, video memory exchange identifier, network backhaul identifier, and waiting / blocking identifier. The node resource data includes node number, remaining resources of the same type, node running status, and node degradation status data. The link status data includes link number, link endpoint, and link congestion status data. The container execution data includes container number, container-borne task number, container node number, container migration status, and container congestion status.
[0015] Optionally, the preprocessing includes time alignment, field unification, unit unification, missing value completion, outlier removal, normalization, generation of scheduling time window markers, and generation of associated indexes.
[0016] Optionally, the generation of task computing power phase segments includes:
[0017] Read the task number, task type, task submission time, task start time, task end time, scheduling time window marker, resource usage curve, task queue length curve, response latency curve, and task log identifier from the standardized computing power operation dataset, and arrange them into a resource variable sequence according to the task number and scheduling time window marker;
[0018] Resource variable sequences are extracted in units of continuous scheduling time windows. The inverse covariance relationship between resource variables within the same continuous scheduling time window is calculated, and the Toeplitz inverse covariance structure is generated in chronological order.
[0019] Based on the Toeplitz inverse covariance structure corresponding to adjacent consecutive scheduling time windows, segmentation cost and segment continuity cost are generated.
[0020] Based on the segmentation cost and segment continuity cost, synchronous segmentation is performed on the resource variable sequence to generate candidate segments for computing power phase.
[0021] Based on the task type and task log identifier, phase semantic labeling is performed on the candidate segments of computing power phase, and the task number, phase type, phase start and end time, phase duration, resource variable sequence within the phase and Toeplitz inverse covariance structure within the phase are written to generate task computing power phase segments.
[0022] Optionally, generating phase resource bundle information includes:
[0023] An improved StemGNN model is constructed, which includes a phase semantic spectrogram embedding layer, a degenerate spectral state filtering layer, and a phase resource bundle readout head.
[0024] The phase semantic spectrum embedding layer reads the phase segment of task computing power, sets the resource variables corresponding to the resource occupancy curve, task queue length curve and response latency curve as resource variable nodes, and writes the task number, phase type, scheduling time window mark and resource type into the resource variable nodes to generate a phase semantic resource graph.
[0025] The phase semantic spectrum embedding layer performs graph Fourier transform on the phase semantic resource graph and discrete Fourier transform on the sequence of resource variables within the phase to generate a temporal representation of the phase spectrum.
[0026] The degradation spectrum state filtering layer reads node degradation state data, link congestion state data, and container congestion state to generate node degradation spectrum state.
[0027] The phase resource bundle readout head reads the node degradation spectrum status, phase semantic resource graph and node resource data, and generates phase node compatibility value, resource elastic boundary, node performance margin and non-migratable flag with associated task number, phase type, scheduling time window mark and node number. The resource elastic boundary includes the compressible lower limit and scalable upper limit of the corresponding resource.
[0028] The phase node compatibility value, resource elasticity boundary, node performance margin, and non-migration identifier are combined into phase resource bundle information;
[0029] A training sample set is constructed, which includes task computing power phase segments, phase semantic resource graphs, node degradation state data, ternary influence hypergraphs, and corresponding historical scheduling feedback data. The historical scheduling feedback data includes actual resource usage, actual node performance margin, actual link congestion value, actual migration impairment value, and actual migration result. The training sample set is input into the improved StemGNN model, and the improved StemGNN model is trained using a joint loss function that includes phase resource bundle prediction constraints, node degradation spectrum state reconstruction constraints, phase node compatibility value discrimination constraints, resource elastic boundary constraints, and non-migratable identifier discrimination constraints, resulting in the trained improved StemGNN model.
[0030] Optionally, the construction of the ternary influence hypergraph includes:
[0031] Read task computing power phase segments and phase resource bundle information, generate a candidate computing power node set based on node resource data, generate a candidate transmission link set and link congestion value based on link status data, and generate migration impairment value based on container running data;
[0032] The intersection of the resource elastic boundary and the remaining resources of the same type in the candidate computing power node set is truncated to generate the resource elastic boundary truncation result, and candidate computing power nodes with empty resource elastic boundary truncation results are deleted.
[0033] Connect the task computing power phase segment, candidate computing power nodes, and candidate transmission links into a ternary influence hyperedge, and write the task number, phase type, phase start and end time, phase node compatibility value, resource elastic boundary truncation result, node performance margin, migration impairment value, link congestion value, and non-migration flag into the ternary influence hyperedge.
[0034] The ternary influence hyperedges are grouped according to task number and phase type, and ternary influence hyperedges that do not have a connection between candidate computing power nodes and candidate transmission links are deleted to generate a ternary influence hypergraph.
[0035] Optionally, the dynamic allocation scheme for generating candidate computing power includes:
[0036] Read the ternary influence hyperedges in the ternary influence hypergraph, set the candidate computing power nodes as node agents, and generate phase resource bundles according to the combination of task number, phase type, candidate computing power nodes and candidate transmission links in the ternary influence hyperedges. Write the phase resource bundles that have not entered the task bundles of the node agents into the common candidate pool.
[0037] The node agent reads the phase node compatibility value, node performance margin, migration impairment value and link congestion value in the phase resource bundle, uses the phase node compatibility value and node performance margin as positive bidding items, and uses the migration impairment value and link congestion value as negative deduction items, and generates the resource bundle bidding value based on the positive bidding items minus the negative deduction items.
[0038] Write the resource bundle bidding value, node agent number, candidate computing power node, task number, phase type, candidate transmission link, phase start and end time, resource elastic boundary truncation result, node performance margin, non-migratable identifier, link congestion value and Bid Time-Stamp field into the bidding record, and arrange the bidding records according to the resource bundle bidding value to generate the node agent task bundle.
[0039] When task bundles of intelligent agents at different nodes contain the same task number and the same phase type, the bidding record with the higher resource bundle bidding value is retained. When the resource bundle bidding values are the same, the bidding record with the earlier time corresponding to the Bid Time-Stamp field is retained, and a consensus bidding record is generated.
[0040] Based on the consensus bidding record, extract the task number, phase type, phase start and end time, candidate computing power nodes, candidate transmission links, resource elastic boundary truncation results, node performance margin, non-migratable identifier, link congestion value and bidding timestamp to generate a candidate computing power dynamic allocation scheme.
[0041] Optionally, the dynamic allocation scheme for generating target computing power includes:
[0042] Read the consensus bidding record and Bid Time-Stamp field in the candidate computing power dynamic allocation scheme, perform BidTime-Stamp refresh processing, and record the duration between the start time of the current scheduling time window and the end time of the current scheduling time window as the effective bidding time window;
[0043] The auction survival time is calculated based on the Bid Time-Stamp field and the current scheduling time window termination time. Consensus auction records with a survival time longer than the effective auction time window are marked as expired auction records, and the phase resource bundles corresponding to the expired auction records are rolled back to the public candidate pool.
[0044] Resource security thresholds are generated based on the compressible lower bound of resource elasticity boundaries, and link congestion thresholds are generated based on the link congestion values in the candidate transmission link set.
[0045] When one of the following conditions is met: the node performance margin is less than the resource safety threshold, there are similar resources below the resource safety threshold in the resource elastic boundary truncation results, or the link congestion value is greater than the link congestion threshold, the corresponding candidate computing power node, candidate transmission link, and task computing power phase segment are marked as affected objects.
[0046] Based on the affected objects, the affected phase resource bundles are extracted from the ternary influence hypergraph. Resource bundle bidding and consensus conflict resolution are re-executed on the affected phase resource bundles to generate updated consensus bidding records. Consensus bidding records that are not marked as expired are merged with the updated consensus bidding records to generate a target computing power dynamic allocation scheme.
[0047] Optionally, the generation of computing power scheduling instructions includes:
[0048] Read the task number, phase type, candidate computing power nodes, candidate transmission links, resource elastic boundary truncation results, non-migratable flag, node performance margin and bidding timestamp from the target computing power dynamic allocation scheme, and generate scheduling execution records according to task number and phase type;
[0049] Based on candidate computing power nodes, candidate transmission links, resource elastic boundary truncation results, and node performance margin, task deployment instructions, resource expansion instructions, resource reduction instructions, bandwidth adjustment instructions, and node isolation instructions are generated, and task migration instructions are generated when the non-migration mark is changed to allow migration.
[0050] Collect the actual resource usage and task completion time after the scheduling is executed, generate the actual link congestion value based on link status data, generate the actual migration damage value based on container operation data, generate the actual node degradation status based on node resource data, and combine them to generate scheduling feedback data.
[0051] The scheduling feedback data is written into the standardized computing power operation dataset according to the task number, phase type, candidate computing power node, candidate transmission link, and scheduling time window.
[0052] The beneficial effects of this invention are:
[0053] The present invention provides a deep learning-based method for dynamic allocation and optimization of computing power performance. By collecting task running data, node resource data, link status data, container running data, and task execution feedback data, a standardized computing power running dataset is generated. Based on resource occupancy curves, task queue length curves, and response latency curves, a Toeplitz inverse covariance structure is calculated to form task computing power phase segments. This allows data loading, computation bursts, memory swapping, network backhaul, and waiting / blocking phases to enter the computing power scheduling link respectively, improving the accuracy of task running phase identification, resource demand matching accuracy, and the reliability of dynamically allocated input data.
[0054] This invention generates phase resource bundle information by improving the StemGNN model, and combines it with a ternary influence hypergraph to uniformly associate task computing power phase segments, candidate computing power nodes, and candidate transmission links. It also employs the CBBA algorithm to perform resource bundle bidding, Bid Time-Stamp refresh processing, and local threshold-triggered rebidding, so that node degradation state, link congestion state, migration impairment value, non-migration flag, and resource elasticity boundary jointly participate in the allocation decision. This reduces the resource occupation caused by expired bidding, excessive task migration impairment, and scheduling imbalance caused by link congestion, thereby improving computing power resource utilization, task response speed, and computing platform operation stability. Attached Figure Description
[0055] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They explain the invention together with the embodiments of the invention and do not constitute a limitation thereof. In the drawings:
[0056] Figure 1 This is an overall flowchart of a deep learning-based dynamic allocation and optimization method for computing power performance proposed in this invention.
[0057] Figure 2 This is a schematic diagram of the improved StemGNN model structure, which is a deep learning-based method for dynamic allocation and optimization of computing power performance.
[0058] Figure 3 This is a schematic diagram of the CBBA algorithm, which is a method for dynamically allocating and optimizing computing power performance based on deep learning proposed in this invention. Detailed Implementation
[0059] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0060] refer to Figure 1 , Figure 2 and Figure 3 A deep learning-based method for dynamic allocation and optimization of computing power performance includes:
[0061] Collect task execution data, node resource data, link status data, container execution data, and task execution feedback data from the computing power resource pool, perform preprocessing, and generate a standardized computing power execution dataset;
[0062] Based on the standardized computing power operation dataset, resource occupancy curves, task queue length curves, and response latency curves are extracted. The Toeplitz inverse covariance structure of resource variables within adjacent scheduling time windows is calculated. Synchronous segmentation is performed and phase semantic calibration is carried out to generate task computing power phase segments.
[0063] An improved StemGNN model is constructed, which includes a phase semantic spectrogram embedding layer, a degenerate spectral state filtering layer, and a phase resource bundle readout head. The phase fragment of the task computing power is input into the improved StemGNN model to generate phase resource bundle information.
[0064] A ternary influence hypergraph is constructed based on task computing power phase segments, phase resource bundle information, node resource data, link status data, and container operation data;
[0065] The CBBA algorithm, which incorporates a Bid Time-Stamp refresh rule, is used to read the ternary influence hypergraph, perform resource bundle bidding and consensus conflict resolution, and generate candidate computing power dynamic allocation schemes.
[0066] Perform Bid Time-Stamp refresh processing and local threshold-triggered rebidding on candidate computing power dynamic allocation schemes to generate target computing power dynamic allocation schemes;
[0067] The computing power scheduling instructions are generated based on the target computing power dynamic allocation scheme, and the scheduling feedback data is collected and written into the standardized computing power operation dataset.
[0068] In this embodiment, the task operation data includes task number, task type, task submission time, task start time, task end time, task log identifier, task queue length, and response latency. The task log identifier includes data loading identifier, computation start identifier, video memory exchange identifier, network backhaul identifier, and waiting / blocking identifier. The node resource data includes node number, remaining resources of the same type, node operating status, and node degradation status data. The link status data includes link number, link endpoint, and link congestion status data. The container operation data includes container number, container-borne task number, node number where the container is located, container migration status, and container congestion status.
[0069] In this embodiment, the preprocessing includes time alignment, field unification, unit unification, missing value completion, outlier removal, normalization, scheduling time window marker generation, and associated index generation. Specifically, the scheduling time window marker generation involves:
[0070] The system reads the collection timestamps from the time-aligned task execution data, node resource data, link status data, container execution data, and task execution feedback data. The earliest collection timestamp is used as the scheduling baseline start time. All collection timestamps are segmented according to a 60-second scheduling time window. The time offset is obtained by subtracting the scheduling baseline start time from the collection timestamp of each record. The time offset is divided by 60 seconds and the integer part is taken to obtain the scheduling time window number. Based on the scheduling time window number, the corresponding scheduling time window start time, scheduling time window end time, and scheduling time window center time are generated. The task number, node number, link number, container number, and scheduling time window number are associated with the scheduling time window number to generate a scheduling time window marker written into the standardized computing power execution dataset.
[0071] In this embodiment, generating the task computing power phase segment includes:
[0072] Read the task number, task type, task submission time, task start time, task end time, scheduling time window marker, resource usage curve, task queue length curve, response latency curve, and task log identifier from the standardized computing power operation dataset, and arrange them into a resource variable sequence according to the task number and scheduling time window marker;
[0073] Resource variable sequences are extracted in units of continuous scheduling time windows. The inverse covariance relationship between resource variables within the same continuous scheduling time window is calculated, and a Toeplitz inverse covariance structure is generated in chronological order. Specifically, the generation of the Toeplitz inverse covariance structure in chronological order is as follows:
[0074] Using five consecutive scheduling time windows as one computation segment, five 9-dimensional resource variable vectors are extracted from the resource variable sequence to form a 5x9 resource variable matrix. The covariance values between each pair of the nine resource variables are calculated to generate a 9x9 covariance matrix. 0.001 is added to the main diagonal of the covariance matrix as a numerical stabilizing term. The covariance matrix after adding the numerical stabilizing term is inverted to generate a 9x9 inverse covariance matrix. A Toeplitz arrangement is constructed according to the adjacent position differences of the resource variables within the time window. The inverse covariance elements corresponding to the same time position difference are written into the same diagonal direction to generate a Toeplitz inverse covariance structure.
[0075] Based on the Toeplitz inverse covariance structure corresponding to adjacent consecutive scheduling time windows, segmentation cost and segment continuity cost are generated, where:
[0076] The cost of generating segment segmentation is as follows:
[0077] Read the Toeplitz inverse covariance structure corresponding to two consecutive scheduling time windows, subtract each item with the same row number and column number to obtain the structural difference matrix, take the absolute value of all elements in the structural difference matrix and sum them to generate the resource variable connection relationship change. Read the 9-dimensional resource variable vector corresponding to two consecutive scheduling time windows, subtract each item with the same resource variable position, take the absolute value of the difference and sum them to generate the resource variable state change. Add the resource variable connection relationship change and the resource variable state change to generate the segmentation cost.
[0078] The cost of generating continuous segments is as follows:
[0079] Read the time interval between two consecutive scheduling time windows. The scheduling time window length is 60s. When the time interval is equal to 60s, write 0 to the time continuity flag. When the time interval is greater than 60s, write 1 to the time continuity flag. Read the task log identifiers corresponding to two consecutive scheduling time windows. If the task log identifiers are the same, write 0 to the log continuity flag. If the task log identifiers are different, write 1 to the log continuity flag. Add the time continuity flag and the log continuity flag to generate the segment continuity cost.
[0080] Synchronous segmentation is performed on the resource variable sequence based on the segmentation cost and segment continuity cost to generate candidate segments for computing power phases. Specifically, generating candidate segments for computing power phases involves:
[0081] Read the segmentation cost and segment continuity cost corresponding to each continuous scheduling time window in the resource variable sequence. Mark the position where the segmentation cost is greater than 1.5 times the average segmentation cost of the first 3 continuous scheduling time windows as a candidate segmentation point. Mark the position where the segment continuity cost is equal to 1 as a forced segmentation point. Merge the candidate segmentation points and forced segmentation points in chronological order. Extract the resource variable sequence according to adjacent segmentation points to obtain multiple continuous segments. Read the Toeplitz inverse covariance structure in each continuous segment. Delete continuous segments with a duration of less than 2 scheduling time windows. Write the remaining continuous segments into the computing power phase candidate segments.
[0082] Based on the task type and task log identifier, phase semantic labeling is performed on the candidate computing power phase segments, and the task number, phase type, phase start and end time, phase duration, resource variable sequence within the phase, and Toeplitz inverse covariance structure within the phase are written to generate task computing power phase segments, where:
[0083] Phase semantic labeling is performed on candidate computing phase segments based on task type and task log identifier, specifically as follows:
[0084] Read the task type and task log identifier corresponding to the candidate fragment of computing power phase. When the task log identifier contains a data loading identifier, it is marked as a data loading phase. When the task log identifier contains a computing start identifier and the GPU usage value or NPU usage value is located at the maximum value position of the 9-dimensional resource variable vector within the fragment, it is marked as a computing burst phase. When the task log identifier contains a memory swap identifier and the change in memory usage value is greater than 0.3 times the average memory usage value of the fragment, it is marked as a memory swap phase. When the task log identifier contains a network backhaul identifier and the bandwidth usage value is located at the maximum value position of the 9-dimensional resource variable vector within the fragment, it is marked as a network backhaul phase. When the task log identifier contains a waiting block identifier and the task queue length value is greater than the average task queue length within the fragment, it is marked as a waiting block phase.
[0085] The task computing power phase segment is generated as follows:
[0086] Read the candidate computing power phase segments after phase semantic calibration, write the start time of the first scheduling time window of the segment into the phase start time, write the end time of the last scheduling time window of the segment into the phase end time, subtract the phase start time from the phase end time to obtain the phase duration, write all 9-dimensional resource variable vectors in the segment into the phase resource variable sequence in chronological order, write all Toeplitz inverse covariance structures in the segment into the phase Toeplitz inverse covariance structure in chronological order, and combine the task number, phase type, phase start time, phase end time, phase duration, phase resource variable sequence, and phase Toeplitz inverse covariance structure to generate the task computing power phase segment.
[0087] In this embodiment, generating phase resource bundle information includes:
[0088] An improved StemGNN model is constructed, comprising a phase semantic spectrogram embedding layer, a degenerate spectral state filtering layer, and a phase resource bundle readout head. Specifically, the construction of the improved StemGNN model is as follows:
[0089] This paper retains the graph Fourier transform, discrete Fourier transform, spectral domain graph convolution, time series prediction, residual connection, and output mapping structures of the traditional StemGNN spectral time series graph neural network model. A task computing power phase segment interface is added to the traditional graph input position. The original input structure, which directly reads multivariate time series data, is transformed into a phase input structure that reads the task number, phase type, scheduling time window marker, in-phase resource variable sequence, and in-phase Toeplitz inverse covariance structure. A phase semantic spectral graph embedding layer is added before the traditional graph learning structure, transforming the original structure based on variable correlation to generate implicit adjacency relationships into a structure based on in-phase Toeplitz inverse covariance. The structure establishes a phase semantic resource graph generation structure for resource variable nodes and resource connection edges. A degradation spectral state filtering layer is set within the traditional spectral domain graph convolution structure. The original structure that only performs convolution calculations on the time-series frequency domain features of resource variables is transformed into a structure that synchronously reads node degradation state data, link congestion state data, and container congestion state, and writes the degradation state into the degradation spectral filtering structure of the corresponding resource variable node. A phase resource bundle readout head is set after the traditional output mapping structure. The original structure that outputs load prediction values is transformed into a resource bundle output structure that outputs phase node compatibility values, resource elastic boundaries, node performance margins, and non-migratable identifiers according to task number, phase type, scheduling time window mark, and node number.
[0090] Phase semantic spectrum embedding layer, including:
[0091] Resource Variable Node Reading Unit: Reads the sequence of resource variables within the phase of the task computing power phase segment, sets the CPU usage value, GPU usage value, NPU usage value, memory usage value, video memory usage value, bandwidth usage value, storage read / write value, task queue length value, and response latency value into 9 resource variable nodes in a fixed order, numbers the 9 resource variable nodes from 1 to 9, and writes the task number, phase type, scheduling time window flag, and resource type into the corresponding resource variable node fields;
[0092] Phase edge structure writing unit: Read the 9 rows and 9 columns of the Toeplitz inverse covariance structure in the phase, delete the main diagonal elements, calculate the absolute value of each off-diagonal element, establish resource connection edges between two resource variable nodes with an absolute value greater than 0.05, and do not establish resource connection edges between two resource variable nodes with an absolute value not greater than 0.05. The setting of 0.05 is based on the weak correlation perturbation filtering requirements between normalized resource variables, and is used to remove accidental covariance connections close to 0.
[0093] Adjacency Matrix Generation Unit: Generates a 9-row, 9-column adjacency matrix according to the order of the 9 resource variable nodes. Writes 1 to the positions where there are resource connection edges and 0 to the positions where there are no resource connection edges. Writes the absolute value of the corresponding off-diagonal element in the Toeplitz inverse covariance structure into the resource connection edge strength field to generate a phase semantic resource graph with edge strength.
[0094] The spectrogram embedding computation unit reads the adjacency matrix of the phase semantic resource graph, counts the number of connection edges for each resource variable node and generates a degree matrix, subtracts the adjacency matrix from the degree matrix to obtain the graph Laplacian matrix, performs eigenvalue decomposition on the graph Laplacian matrix to obtain the eigenvector matrix, and multiplies the phase-internal resource variable sequence with the eigenvector matrix to generate phase semantic spectrogram embedding features.
[0095] The degraded spectral state filtering layer includes:
[0096] Degradation field reading unit: Reads memory fragmentation rate, GPU queue wait time, CPU cache miss rate and node temperature from node degradation state data; reads link queue length, packet loss rate and retransmission rate from link congestion state data; reads container queue length and container restart count from container congestion state; and aligns the above fields according to task number, node number and scheduling time window mark.
[0097] Degradation Node Mapping Unit: Writes the CPU cache miss rate to the CPU resource variable node, the GPU queue wait time to the GPU resource variable node, the video memory fragmentation rate to the video memory resource variable node, the average of the link queue length, packet loss rate, and retransmission rate to the bandwidth resource variable node, the average of the container queue length and container restart count to the task queue length resource variable node, and the node temperature to the CPU, GPU, NPU, and video memory resource variable nodes. For cases where there are more than two degradation fields in the same resource variable node, the arithmetic mean is taken to generate a 9-dimensional degradation state vector.
[0098] Degradation Spectrum Filtering Unit: Reads the phase semantic spectrum embedding features and the 9-dimensional degradation state vector, expands the 9-dimensional degradation state vector to the same dimension as the phase semantic spectrum embedding features according to the resource variable node order, multiplies the values of the same resource variable node positions one by one to obtain the degradation modulation spectrum features, and adds the degradation modulation spectrum features to the phase semantic spectrum embedding features one by one to generate the node degradation spectrum state.
[0099] Phase resource bundle readout head, including:
[0100] Phase node compatibility value output unit: Reads the node degradation spectrum state and the remaining resources of the same type in the node resource data, subtracts the resource requirement representation corresponding to the task computing power phase segment from the degradation representation corresponding to the candidate computing power node in the order of nodes with the same resource variable, takes the absolute value of the difference and sums them to generate the phase node difference value, inputs the phase node difference value into the fully connected mapping unit, and outputs the phase node compatibility value in the range of 0 to 1 after Sigmoid mapping;
[0101] Resource elastic boundary output unit: Read the minimum, average and maximum values of each resource variable in the phase segment of the task computing power within the phase duration, take the minimum value as the basic value of the compressible lower limit and the maximum value as the basic value of the scalable upper limit, read the remaining resources of the same type in the node resource data, take the smaller value between the basic value of the compressible lower limit and the remaining resources of the same type as the compressible lower limit, take the smaller value between the basic value of the scalable upper limit and the remaining resources of the same type as the scalable upper limit, and generate the resource elastic boundary;
[0102] Node performance margin output unit: Reads the remaining resources of the same type and the scalability limit in the resource elastic boundary from the node resource data, subtracts the scalability limit of the corresponding resource from the remaining resources of the same type to obtain the resource remaining difference, arranges all resource remaining differences in the order of resource variable nodes, inputs them into the margin mapping unit, and outputs the node performance margin.
[0103] The non-migratable flag output unit reads the task status snapshot size, container rebuild time, cache loss, video memory reload, and service interruption time from the container migration status. After normalizing the five values to the range of 0 to 1, it sums them to obtain the total migration damage. If the total migration damage is greater than 3, it outputs "migration prohibited"; if the total migration damage is not greater than 3, it outputs "migration allowed". The threshold of 3 is set based on the medium-high damage boundary after normalizing and summing the five migration damage fields, indicating that migration will no longer be performed when the average migration damage of each field reaches 0.6 or higher.
[0104] The phase semantic spectrum embedding layer reads the task computing power phase segment, sets the resource variables corresponding to the resource occupancy curve, task queue length curve, and response latency curve as resource variable nodes, and writes the task number, phase type, scheduling time window marker, and resource type into the resource variable nodes to generate a phase semantic resource graph. Specifically, generating the phase semantic resource graph involves:
[0105] Read the task number, phase type, scheduling time window marker, resource variable sequence within the phase, and Toeplitz inverse covariance structure within the phase from the task computing power phase segment. Set CPU, GPU, NPU, memory, video memory, bandwidth, storage read / write, task queue length, and response latency as 9 resource variable nodes respectively. Number the resource variable nodes from 1 to 9 in a fixed order. Write the task number, phase type, scheduling time window marker, and resource type into the corresponding resource variable nodes. Read the off-diagonal elements in the Toeplitz inverse covariance structure within the phase. Establish resource connection edges between two resource variable nodes whose absolute values of off-diagonal elements are greater than 0.05. Combine the resource variable nodes and resource connection edges to generate a phase semantic resource graph.
[0106] The phase semantic spectrum embedding layer performs a graph Fourier transform on the phase semantic resource graph and a discrete Fourier transform on the sequence of resource variables within the phase, generating a temporal representation of the phase spectrum. Specifically, generating the temporal representation of the phase spectrum involves:
[0107] Read the resource connection edges in the phase semantic resource graph, generate a 9x9 adjacency matrix according to the order of the 9 resource variable nodes, count the number of connection edges of each resource variable node and write it into a 9x9 degree matrix, subtract the adjacency matrix from the degree matrix to generate the graph Laplace matrix, perform eigenvalue decomposition on the graph Laplace matrix to obtain the eigenvector matrix, multiply the phase resource variable sequence with the eigenvector matrix to generate the graph Fourier feature, read the numerical sequence of each resource variable in the phase resource variable sequence arranged along the scheduling time window, perform discrete Fourier transform on each numerical sequence, extract the real part, imaginary part and magnitude, and concatenate the graph Fourier feature, real part, imaginary part and magnitude according to the order of resource variable nodes to generate the phase spectrum time series representation;
[0108] The degradation spectrum state filtering layer reads node degradation state data, link congestion state data, and container congestion state to generate node degradation spectrum states. Specifically, generating node degradation spectrum states involves:
[0109] Read memory fragmentation rate, GPU queue wait time, CPU cache miss rate, and node temperature from node degradation state data; read link queue length from link congestion state data; read container congestion state value from container congestion state data; write CPU cache miss rate to CPU resource variable node location; write GPU queue wait time to GPU resource variable node location; write memory fragmentation rate to memory resource variable node location; write link queue length to bandwidth resource variable node location; write container congestion state value to task queue length resource variable node location; write node temperature to CPU, GPU, NPU, and memory resource variable node locations; take the arithmetic mean for the case where there are two degradation data for the same resource variable node; generate a 9-dimensional degradation state vector; multiply the 9-dimensional degradation state vector with the features of the corresponding resource variable node in the phase spectrum temporal representation item by item to generate node degradation spectrum state;
[0110] The phase resource bundle readout head reads the node degradation spectrum state, phase semantic resource graph, and node resource data, generating a phase node compatibility value, resource elasticity boundary, node performance margin, and non-migratable flag, associated with the task number, phase type, scheduling time window marker, and node number. The resource elasticity boundary includes the lower limit of compressibility and the upper limit of scalability for the corresponding resource, where:
[0111] Generate phase node compatibility values that associate task number, phase type, scheduling time window flag, and node number, specifically as follows:
[0112] Read the degradation representation corresponding to the candidate computing power node in the node degradation spectrum state, read the resource demand representation corresponding to the task computing power phase segment in the phase semantic resource graph, subtract the resource demand representation from the degradation representation in the order of nodes with the same resource variables, take the absolute value of the difference and sum them to generate the phase node difference value, input the phase node difference value into the fully connected mapping unit in the phase resource bundle readout head, output the phase node compatibility value in the range of 0 to 1, and write the task number, phase type, scheduling time window mark and node number into the phase node compatibility value index;
[0113] The resource elastic boundary is generated as follows:
[0114] Read the minimum, average, and maximum values of each resource variable in the task computing power phase segment during the phase duration. Use the minimum value as the base value of the compressible lower limit and the maximum value as the base value of the scalable upper limit. Read the remaining resources of the same type in the node resource data. Use the smaller value between the base value of the compressible lower limit and the remaining resources of the same type as the compressible lower limit. Use the smaller value between the base value of the scalable upper limit and the remaining resources of the same type as the scalable upper limit. Arrange the compressible lower limit and scalable upper limit corresponding to CPU, GPU, NPU, memory, video memory, bandwidth, storage read / write, task queue length, and response latency according to the node order of resource variables to generate resource elastic boundaries.
[0115] Generate node performance margin and non-migration flags, specifically as follows:
[0116] Read the remaining resources of the same type in the node resource data, read the scalability limit in the resource elastic boundary, subtract the scalability limit of the corresponding resource from the remaining resources of the same type to obtain the resource remaining difference, arrange all resource remaining differences in the order of resource variable nodes and input them into the margin mapping unit in the phase resource bundle readout head to generate node performance margin, read the task status snapshot size, container reconstruction time, cache loss, video memory reload and service interruption time in the container migration status, normalize the five values to the range of 0 to 1 and add them together to generate the total migration damage. If the total migration damage is greater than 3, write the non-migration mark to prohibit migration, and if the total migration damage is not greater than 3, write the non-migration mark to allow migration.
[0117] The phase node compatibility value, resource elasticity boundary, node performance margin, and non-migration identifier are combined into phase resource bundle information;
[0118] A training sample set is constructed, comprising task computing power phase segments, phase semantic resource graphs, node degradation state data, ternary influence hypergraphs, and corresponding historical scheduling feedback data. The historical scheduling feedback data includes actual resource usage, actual node performance margin, actual link congestion value, actual migration impairment value, and actual migration result. This training sample set is input into the improved StemGNN model. A joint loss function, including phase resource bundle prediction constraints, node degradation spectrum state reconstruction constraints, phase node compatibility value discrimination constraints, resource elastic boundary constraints, and non-migratable identifier discrimination constraints, is used to train the improved StemGNN model, resulting in the trained improved StemGNN model.
[0119] The training sample set is constructed as follows:
[0120] Read the task computing power phase segments, phase semantic resource graph, node degradation state data and ternary influence hypergraph generated within the historical scheduling period, read the historical scheduling feedback data corresponding to the same historical scheduling period, use the actual resource occupation, actual node performance margin, actual link congestion value, actual migration impairment value and actual migration result as supervision labels, use the task number, phase type, scheduling time window mark and node number as sample index, pair the input data and supervision labels according to the sample index to generate a training sample set;
[0121] The improved StemGNN model is trained as follows:
[0122] The training sample set is input into the improved StemGNN model, which outputs predicted phase resource bundle information, predicted node degradation spectrum state, predicted phase node compatibility value, predicted resource elasticity boundary, and predicted non-migrating indicator. The average absolute difference between the predicted phase resource bundle information and the actual resource occupancy is calculated to generate phase resource bundle prediction constraints. The average absolute difference between the predicted node degradation spectrum state and the actual node degradation state is calculated to generate node degradation spectrum state reconstruction constraints. The average absolute difference between the predicted phase node compatibility value and the actual node performance margin normalized value is calculated to generate phase node compatibility value discrimination constraints. The average absolute difference between the predicted resource elasticity boundary and the portion exceeding the boundary in the actual resource occupancy is calculated to generate... Resource elastic boundary constraints are applied, and the cross-entropy value between the predicted non-transferable identifier and the actual migration result is calculated to generate non-transferable identifier discrimination constraints. The weights of the five constraints are all set to 1. The reason for setting the weights of the five constraints to 1 is that the five supervision objectives correspond to resource prediction, degradation reconstruction, node matching, boundary output and migration discrimination, respectively. The five supervision objectives jointly determine the phase resource bundle information. The five constraint values are directly added to generate a joint loss value. Based on the joint loss value, the network parameters of the phase semantic spectrum embedding layer, degradation spectrum state filtering layer and phase resource bundle readout head are updated in reverse. When the absolute value of the difference between the joint loss values of two adjacent rounds in five consecutive training rounds is less than 0.001, the improved StemGNN model is obtained after training.
[0123] In this embodiment, the construction of the ternary influence hypergraph includes:
[0124] Read the task's computing power phase segment and phase resource bundle information; generate a candidate computing power node set based on node resource data; generate a candidate transmission link set and link congestion value based on link status data; and generate a migration impairment value based on container runtime data, wherein:
[0125] A candidate computing power node set is generated based on node resource data, specifically as follows:
[0126] Read the task number, phase type, phase start and end time and resource variable sequence within the phase from the task computing power phase segment; read the resource elastic boundary and non-migratable identifier from the phase resource bundle information; read the node number, node running status, CPU remaining cores, GPU remaining cards, NPU remaining cards, memory remaining capacity, video memory remaining capacity, bandwidth remaining capacity and storage read / write remaining capacity from the node resource data; write nodes with online running status and non-isolated node isolation status into the initial candidate node table; filter nodes with remaining capacity of the same type of resources from the initial candidate node table according to the resource type required by the task phase; write the filtered node number, node resource type and remaining resources of the same type into the candidate computing power node set.
[0127] A candidate transmission link set is generated based on link state data, specifically as follows:
[0128] Read the link number, link start node, link end node, link bandwidth capacity, link occupied bandwidth, link queue length, packet loss rate and retransmission rate from the link status data; read the node number where the container is located from the container running data; use the node number where the container is located as the current bearing node of the task; perform endpoint matching between the current bearing node of the task and the candidate computing power nodes in the candidate computing power node set; write the links whose link start node and link end node contain the current bearing node of the task and the candidate computing power nodes into the candidate transmission link set; search for link paths within 2 hops for candidate node pairs that do not have direct links according to the continuous relationship of link endpoints; and write the searched link numbers into the candidate transmission link set in path order.
[0129] The link congestion value is generated as follows: read the link bandwidth capacity, occupied bandwidth, link queue length, packet loss rate, and retransmission rate of each link in the candidate transmission link set; divide the occupied bandwidth by the link bandwidth capacity to obtain the bandwidth occupancy ratio; divide the link queue length by the maximum queue length of all candidate links within the same scheduling time window to obtain the queue normalization value; normalize the packet loss rate and retransmission rate to the range of 0 to 1 respectively; and calculate the arithmetic mean of the bandwidth occupancy ratio, queue normalization value, packet loss rate normalization value, and retransmission rate normalization value to generate the link congestion value.
[0130] Migration impairment values are generated based on container runtime data, specifically:
[0131] Read the task status snapshot size, number of container image layers, container rebuild time, cache misses, memory reloads, and service interruption duration from the container runtime data. Divide the task status snapshot size by the maximum task status snapshot size within the same scheduling time window to obtain the snapshot normalization value. Divide the number of container image layers by the maximum number of container image layers within the same scheduling time window to obtain the image normalization value. Normalize the container rebuild time, cache misses, memory reloads, and service interruption duration to the range of 0 to 1. Calculate the arithmetic mean of the 6 normalized values to generate the migration impairment value.
[0132] The resource elastic boundary is truncated by intersecting with the remaining resources of the same type in the candidate computing power node set, generating a resource elastic boundary truncation result. Candidate computing power nodes with empty resource elastic boundary truncation results are then deleted. Specifically, generating the resource elastic boundary truncation result involves:
[0133] Read the resource elastic boundary from the phase resource bundle information. The resource elastic boundary includes the compressible lower limit and scalable upper limit corresponding to CPU, GPU, NPU, memory, video memory, bandwidth, and storage read / write. Read the remaining resources of the same type for each candidate computing power node in the candidate computing power node set. Set the node allocable range for each type of resource to 0 to the remaining resources of the same type. Calculate the intersection between the resource elastic boundary and the node allocable range. The lower limit of the intersection is the larger value between the compressible lower limit and 0. The upper limit of the intersection is the smaller value between the scalable upper limit and the remaining resources of the same type. If the upper limit of the intersection is not less than the lower limit of the intersection, write the resource elastic boundary truncation result. If the upper limit of the intersection is less than the lower limit of the intersection, mark the corresponding candidate computing power node as an empty truncation node and delete it from the candidate computing power node set.
[0134] The task computing power phase segment, candidate computing power nodes, and candidate transmission links are connected to form a ternary influence hyperedge. The task number, phase type, phase start and end time, phase node compatibility value, resource elastic boundary truncation result, node performance margin, migration impairment value, link congestion value, and non-migration flag are written to the ternary influence hyperedge. Specifically, writing the task number, phase type, phase start and end time, phase node compatibility value, resource elastic boundary truncation result, node performance margin, migration impairment value, link congestion value, and non-migration flag to the ternary influence hyperedge is as follows:
[0135] Read the task number, phase type, and phase start and end time from the task computing power phase segment; read the candidate computing power nodes from the candidate computing power node set; read the candidate transmission links from the candidate transmission link set; connect one task computing power phase segment, one candidate computing power node, and one candidate transmission link under the same task number and phase type into one ternary influence hyperedge; read the phase node compatibility value, resource elastic boundary truncation result, node performance margin, and non-migration flag corresponding to the same task number, same phase type, and same candidate computing power node from the phase resource bundle information; read the link congestion value corresponding to the candidate transmission link; read the migration impairment value corresponding to the container running data; and write the task number, phase type, phase start and end time, phase node compatibility value, resource elastic boundary truncation result, node performance margin, migration impairment value, link congestion value, and non-migration flag into the ternary influence hyperedge field;
[0136] The ternary influence hyperedges are grouped according to task number and phase type, and ternary influence hyperedges with no connectivity between candidate computing power nodes and candidate transmission links are deleted, generating a ternary influence hypergraph, where:
[0137] The ternary influence on the superedge grouping is as follows:
[0138] Read all ternary influence hyperedges, establish Level 1 groups by task number, establish Level 2 groups by phase type, write ternary influence hyperedges with the same task number and phase type into the same hyperedge group, arrange the ternary influence hyperedges in each hyperedge group in ascending order by candidate computing power node number and candidate transmission link number, and generate a grouped ternary influence hyperedge table.
[0139] The connection deletion and generation of ternary influence hypergraphs are as follows:
[0140] Read the candidate computing power nodes and candidate transmission links from the grouped ternary influence hyperedge table, and read the link start node and link end node from the link status data. Mark the ternary influence hyperedges where the candidate computing power node and any endpoint of the candidate transmission link are consistent as directly connected edges. Mark the ternary influence hyperedges where the candidate computing power node and the endpoint of the candidate transmission link are inconsistent and the link path within 2 hops can reach the candidate computing power node as path connected edges. Delete the ternary influence hyperedges that are neither directly connected edges nor path connected edges. Combine the remaining ternary influence hyperedges, task computing power phase segment nodes, candidate computing power nodes, candidate transmission link nodes and hyperedge fields to generate a ternary influence hypergraph.
[0141] In this embodiment, the generation of a candidate computing power dynamic allocation scheme includes:
[0142] Read the ternary influence hyperedges in the ternary influence hypergraph, set candidate computing power nodes as node agents, and generate phase resource bundles according to the combination of task number, phase type, candidate computing power node, and candidate transmission link ternary influence hyperedges. Write the phase resource bundles that have not entered the node agent task bundles into the common candidate pool. Specifically, setting candidate computing power nodes as node agents is as follows:
[0143] Read all ternary influence hyperedges in the ternary influence hypergraph, extract task number, phase type, candidate computing power node, candidate transmission link, phase start and end time, phase node compatibility value, resource elastic boundary truncation result, node performance margin, migration impairment value, link congestion value, and non-migration flag. Establish each candidate computing power node as a node agent. Merge the ternary influence hyperedges corresponding to the same task number, phase type, candidate computing power node, and candidate transmission link into a phase resource bundle record. Write the phase resource bundle record into the bidable resource bundle table of the corresponding node agent. Write the phase resource bundle record that has not yet been written into the task bundle of any node agent into the common candidate pool. The fields of the common candidate pool include resource bundle number, task number, phase type, candidate computing power node, candidate transmission link, and entry timestamp.
[0144] The node agent reads the phase node compatibility value, node performance margin, migration impairment value, and link congestion value from the phase resource bundle. It uses the phase node compatibility value and node performance margin as positive bidding terms, and the migration impairment value and link congestion value as negative deduction terms. The resource bundle bidding value is generated by subtracting the negative deduction terms from the positive bidding terms. Specifically, the generation of the resource bundle bidding value by subtracting the negative deduction terms from the positive bidding terms is as follows:
[0145] The node agent reads the phase node compatibility value, node performance margin, migration impairment value, and link congestion value from the list of available bidable resource bundles. It keeps the phase node compatibility value between 0 and 1, divides the node performance margin by the maximum node performance margin of all candidate computing power nodes within the same scheduling time window to generate a margin normalization value, divides the migration impairment value by the maximum migration impairment value within the same scheduling time window to generate a migration impairment normalization value, and keeps the link congestion value between 0 and 1. The phase node compatibility value and margin normalization value are used as positive bidding items, while the migration impairment normalization value and link congestion value are used as negative deduction items. The weight of each of the four bidding items is set to 1. The phase node compatibility value, node performance margin, migration impairment value, and link congestion value correspond to task matching capability, resource carrying capacity, migration cost, and link transmission pressure, respectively. All four fields directly affect the executability of the resource bundle. The resource bundle bidding value is generated by adding the margin normalization value to the phase node compatibility value and then subtracting the migration impairment normalization value and link congestion value.
[0146] The resource bundle bidding value, node agent ID, candidate computing power node, task ID, phase type, candidate transmission link, phase start and end time, resource elastic boundary truncation result, node performance margin, non-migratable flag, link congestion value, and Bid Time-Stamp field are written into the bidding record. The bidding records are then arranged according to the resource bundle bidding value to generate the node agent task bundle, where:
[0147] The Bid Time-Stamp field, specifically:
[0148] After the node agent generates the resource bundle bidding value, it reads the current scheduling time window number, node agent number, resource bundle number, task number, phase type and candidate transmission link number, reads the bidding generation time under the unified clock, converts the bidding generation time into a millisecond-level timestamp, combines the scheduling time window number, bidding generation time, node agent number and resource bundle number into a Bid Time-Stamp field, and writes the Bid Time-Stamp field into the corresponding bidding record, so that each bidding record has both resource bundle bidding value and bidding generation time order;
[0149] The task bundle for generating node intelligent agents is as follows:
[0150] Read the resource bundle bidding value and the corresponding phase resource bundle record. Write the node agent number, candidate computing power node, task number, phase type, candidate transmission link, phase start and end time, resource elastic boundary truncation result, node performance margin, non-migratable flag, link congestion value, and system time within the current scheduling time window into the bidding record. Use the system time within the current scheduling time window as the bidding timestamp. Group the bidding records by node agent number. Within each node agent group, arrange the bidding records in descending order of resource bundle bidding value. Delete the bidding records with empty resource elastic boundary truncation result, non-migratable flag marked as prohibited from migration, and candidate computing power nodes that are different from the current task carrying node. Write the remaining bidding records into the node agent task bundle in the sorted order.
[0151] When task bundles of agents at different nodes contain the same task number and the same phase type, the bidding record with the higher resource bundle bidding value is retained. When the resource bundle bidding values are the same, the bidding record with the earlier time corresponding to the Bid Time-Stamp field is retained to generate a consensus bidding record, wherein:
[0152] When resource bundle values are equal, the bidding record with the earlier time corresponding to the Bid Time-Stamp field is retained, specifically:
[0153] Read bidding records with the same task number and phase type from the task bundles of different node agents. First, compare the bidding value of the resource bundle and retain the bidding record with the larger resource bundle bidding value. If the bidding value of the resource bundle is the same, read the bidding generation time in the BidTime-Stamp field and retain the bidding record with the earlier bidding generation time. Then, back the phase resource bundle corresponding to the bidding record that was not retained to the public candidate pool to generate consensus bidding records.
[0154] The consensus bidding record is generated as follows:
[0155] Read the bidding records in all node agent task bundles, establish a conflict index according to task number and phase type, write multiple bidding records corresponding to the same task number and phase type into the same conflict group, compare the resource bundle bidding value in each conflict group, retain the bidding record with the highest resource bundle bidding value, and back the bidding record with the lower resource bundle bidding value to the public candidate pool, compare the bidding timestamps when the resource bundle bidding values are the same, retain the bidding record with the earliest bidding timestamp, and back the bidding record with the later bidding timestamp to the public candidate pool, sort all the retained bidding records according to task number, phase type and phase start time, and generate consensus bidding records;
[0156] Based on the consensus bidding records, the task number, phase type, phase start and end time, candidate computing power nodes, candidate transmission links, resource elastic boundary truncation results, node performance margin, non-migratable flag, link congestion value, and bidding timestamp are extracted to generate a candidate computing power dynamic allocation scheme. Specifically, generating the candidate computing power dynamic allocation scheme involves:
[0157] Read the task number, phase type, phase start and end time, candidate computing power nodes, candidate transmission links, resource elastic boundary truncation result, node performance margin, non-migratable flag, link congestion value, and bidding timestamp from the consensus bidding record. Establish a task-level allocation table by task number and a phase-level allocation table by phase type and phase start and end time. Write each consensus bidding record into the corresponding task-level and phase-level allocation tables. Use candidate computing power nodes as phase execution nodes, candidate transmission links as phase transmission links, resource elastic boundary truncation result as the phase allocable resource range, and use node performance margin, link congestion value, non-migratable flag, and bidding timestamp as scheme verification fields to generate a dynamic candidate computing power allocation scheme.
[0158] In this embodiment, the dynamic allocation scheme for generating target computing power includes:
[0159] Read the consensus bidding record and Bid Time-Stamp field from the candidate computing power dynamic allocation scheme, perform BidTime-Stamp refresh processing, and record the duration from the start time to the end time of the current scheduling time window as the effective bidding time window, where:
[0160] Perform a Bid Time-Stamp refresh, specifically as follows:
[0161] Read the consensus bidding records in the candidate computing power dynamic allocation scheme, extract the BidTime-Stamp field in each consensus bidding record, read the bidding generation time in the BidTime-Stamp field, read the current scheduling time window end time, subtract the bidding generation time from the current scheduling time window end time to obtain the bidding survival time, take the time between the current scheduling time window start time and the current scheduling time window end time as the bidding effective time window, when the scheduling time window length is 60s, the bidding effective time window is 60s, mark the consensus bidding records with a bidding survival time greater than 60s as expired bidding records;
[0162] The duration from the start time to the end time of the current scheduling time window is recorded as the effective bidding time window, specifically:
[0163] Read the consensus bidding records in the candidate computing power dynamic allocation scheme, extract the task number, phase type, candidate computing power node, candidate transmission link, bidding timestamp, start time and end time of the current scheduling time window, subtract the start time of the current scheduling time window from the end time of the current scheduling time window to obtain the bidding effective time window. When the length of the current scheduling time window is 60s, the bidding effective time window is 60s. Write the bidding effective time window into the verification field of each consensus bidding record.
[0164] The auction survival time is calculated based on the Bid Time-Stamp field and the current scheduling time window termination time. Consensus auction records with a survival time longer than the effective auction time window are marked as expired auction records, and the phase resource bundles corresponding to the expired auction records are rolled back to the public candidate pool. Specifically, rolling back the phase resource bundles corresponding to the expired auction records to the public candidate pool is as follows:
[0165] Read the bidding timestamp from each consensus bidding record, subtract the bidding timestamp from the current scheduling time window termination time to obtain the bidding survival time, mark consensus bidding records with a bidding survival time greater than 60s as expired bidding records, extract the task number, phase type, candidate computing power node and candidate transmission link from the expired bidding records, read the corresponding phase resource bundle from the ternary influence hypergraph according to the same fields, write the corresponding phase resource bundle into the public candidate pool, and delete the expired bidding records from the candidate computing power dynamic allocation scheme;
[0166] Resource security thresholds are generated based on the compressibility lower bound of resource elasticity boundaries, and link congestion thresholds are generated based on the link congestion values in the candidate transmission link set, where:
[0167] The resource safety threshold is generated based on the compressibility lower bound of the resource elastic boundary, specifically as follows:
[0168] Read the CPU compressible lower bound, GPU compressible lower bound, NPU compressible lower bound, memory compressible lower bound, video memory compressible lower bound, bandwidth compressible lower bound, and storage read / write compressible lower bound from the resource elastic boundary. Read the total number of CPU cores, total number of GPUs, total number of NPUs, total memory capacity, total video memory capacity, total bandwidth capacity, and total storage read / write capacity from the node resource data. Divide the compressible lower bound of each type of resource by the corresponding total resource amount to obtain the resource lower bound percentage. Calculate the arithmetic mean of all resource lower bound percentages to obtain the node-level resource safety threshold. Use the compressible lower bound of each type of resource as the safety threshold of the same type of resource. The node-level resource safety threshold is used to determine the node performance margin, and the same type of resource safety threshold is used to determine the resource elastic boundary truncation result.
[0169] The link congestion threshold is generated based on the link congestion values in the candidate transmission link set, specifically as follows:
[0170] Read all link congestion values in the candidate transmission link set, sort the link congestion values in ascending order, calculate the arithmetic mean of all link congestion values, read the link congestion value in the middle position after sorting as the median congestion value, add the arithmetic mean and the median congestion value and divide by 2 to obtain the link congestion threshold. The reason for using the arithmetic mean and the median congestion value to generate the link congestion threshold is that the arithmetic mean reflects the overall link pressure, and the median congestion value reduces the impact of a single abnormal link on the threshold.
[0171] When any of the following conditions are met: node performance margin is less than the resource safety threshold, there are similar resources below the resource safety threshold in the resource elastic boundary truncation results, or link congestion value is greater than the link congestion threshold, the corresponding candidate computing power node, candidate transmission link, and task computing power phase segment are marked as affected objects. Specifically, marking the corresponding candidate computing power node, candidate transmission link, and task computing power phase segment as affected objects is as follows:
[0172] Read the node performance margin, resource elastic boundary truncation result, and link congestion value from each consensus bidding record. Mark candidate computing power nodes whose node performance margin is less than the node-level resource security threshold as affected candidate computing power nodes. Mark candidate computing power nodes whose resource allocability value of any similar resource in the resource elastic boundary truncation result is less than the security threshold of the similar resource as affected candidate computing power nodes. Mark candidate transmission links whose link congestion value is greater than the link congestion threshold as affected candidate transmission links. Mark task computing power phase segments connected to affected candidate computing power nodes or affected candidate transmission links as affected task computing power phase segments. Combine affected candidate computing power nodes, affected candidate transmission links, and affected task computing power phase segments into affected objects.
[0173] Based on the affected objects, the affected phase resource bundles are extracted from the ternary influence hypergraph. Resource bundle bidding and consensus conflict resolution are re-executed on the affected phase resource bundles to generate updated consensus bidding records. Consensus bidding records not marked as expired are merged with the updated consensus bidding records to generate a target computing power dynamic allocation scheme, wherein:
[0174] The resource bundle bidding and consensus conflict resolution processes for the affected phase resource bundles will be re-executed, specifically as follows:
[0175] Read the task number, phase type, candidate computing power node and candidate transmission link from the affected object, retrieve the ternary influence hyperedge from the ternary influence hypergraph according to the same field, recombine the retrieved ternary influence hyperedge into the affected phase resource bundle, write the affected phase resource bundle into the public candidate pool, the node agent rereads the affected phase resource bundle in the public candidate pool, recalculates the resource bundle bidding value, regenerates the bidding record and node agent task bundle, re-executes consensus conflict resolution on the bidding record with the same task number and the same phase type, and generates an updated consensus bidding record;
[0176] The target computing power dynamic allocation scheme is generated as follows:
[0177] Read consensus bidding records from candidate computing power dynamic allocation schemes that are not marked as expired and are not associated with affected objects. Read updated consensus bidding records. Build a merge index by task number and phase type. Overwrite the original consensus bidding records with updated consensus bidding records under the same task number and phase type. Retain the original consensus bidding records that do not have updated records. Rearrange the merged consensus bidding records by task number, phase type, and phase start and end time. Write the results to candidate computing power nodes, candidate transmission links, resource elastic boundary truncation results, node performance margin, non-migration flag, link congestion value, and bidding timestamp to generate the target computing power dynamic allocation scheme.
[0178] In this embodiment, the generation of computing power scheduling instructions includes:
[0179] Read the task number, phase type, candidate computing power nodes, candidate transmission links, resource elastic boundary truncation results, non-migratable flag, node performance margin and bidding timestamp from the target computing power dynamic allocation scheme, and generate scheduling execution records according to task number and phase type;
[0180] Based on candidate computing power nodes, candidate transmission links, resource elastic boundary truncation results, and node performance margin, task deployment instructions, resource expansion instructions, resource reduction instructions, bandwidth adjustment instructions, and node isolation instructions are generated, and task migration instructions are generated when the non-migration mark is changed to allow migration.
[0181] Collect the actual resource usage and task completion time after the scheduling is executed, generate the actual link congestion value based on link status data, generate the actual migration damage value based on container operation data, generate the actual node degradation status based on node resource data, and combine them to generate scheduling feedback data.
[0182] The scheduling feedback data is written into the standardized computing power operation dataset according to the task number, phase type, candidate computing power node, candidate transmission link, and scheduling time window.
[0183] Example 1: In a continuous operating cycle of an enterprise-level hybrid computing resource pool, the pool comprises 32 computing nodes: 18 general-purpose CPU nodes, 8 GPU nodes, 4 high-memory nodes, and 2 edge access nodes. The general-purpose CPU nodes are configured with 32 CPU cores, 128GB of memory, 4TB of storage, and 10Gbps bandwidth; the GPU nodes are configured with 32 CPU cores, 4 GPU cards, 256GB of memory, 96GB of video memory, and 25Gbps bandwidth; the high-memory nodes are configured with 48 CPU cores, 512GB of memory, 8TB of storage, and 10Gbps bandwidth; and the edge access nodes are configured with 16 CPU cores, 2 NPU cards, 64GB of memory, and 10Gbps bandwidth. Within 24 hours, the resource pool receives 21,600 tasks, including online inference, batch analysis, database synchronization, log retrieval, image rendering, and report generation, and receives 21,600 task execution feedback data. Node resource data is sampled once every 10 seconds, resulting in 276,480 node sampling records; link status data comes from 48 transmission links, sampled once every 10 seconds, resulting in 414,720 link sampling records; container operation data comes from an average of 680 active containers, summarized once every 60 seconds, resulting in 979,200 container window records.
[0184] Traditional methods employ fixed thresholds and minimum load node selection strategies, using complete tasks as scheduling objects. Task migration or scaling is triggered when CPU utilization exceeds 80% or GPU utilization exceeds 85%. This method fails to anticipate sudden increases in GPU demand when a task transitions from data loading to a computational surge, causing GPU queue waiting times to rise from 1.8 seconds to 7.4 seconds. Some online inference tasks remain on general-purpose CPU nodes. During network backhaul, the traditional method only assesses the remaining resources of the target node, failing to simultaneously evaluate link congestion and container migration impairments. This resulted in task backlog across 12 nodes for three consecutive scheduling time windows, with congestion values exceeding 0.7 on links L18, L23, and L31.
[0185] In this embodiment, the system first standardizes the fields of task number, node number, link number, container number, and collection timestamp, and arranges CPU utilization, GPU utilization, NPU utilization, memory utilization, video memory utilization, bandwidth utilization, storage read / write rate, task queue length, and response latency into a 9-dimensional resource variable vector. The scheduling time window length is set to 60 seconds, based on the original sampling period of nodes and links being 10 seconds. One scheduling time window contains 6 sampling points, which can cover short-term resource fluctuations of online tasks without causing excessive segmentation of batch processing tasks. 1440 scheduling time windows are formed within 24 hours. The system associates task records, node sampling records, link sampling records, and container window records according to the scheduling time window number to generate a standardized computing power operation dataset.
[0186] The system extracts the resource utilization curve, task queue length curve, and response latency curve for each task. Taking task T0856 as an example, the task type is an online inference task. Within six adjacent scheduling time windows, the GPU utilization rate is 0.18, 0.21, 0.72, 0.81, 0.64, and 0.32, respectively; the task queue length is 3, 4, 18, 22, 11, and 5, respectively; and the response latency is 42ms, 47ms, 126ms, 138ms, 91ms, and 58ms, respectively. The system forms a 5x9 resource variable matrix for every five consecutive scheduling time windows, calculates the covariance matrix among the nine resource variables, adds a 0.001 stabilizing term to the main diagonal, and then inverts it to generate a 9x9 Toeplitz inverse covariance structure. The segmentation cost is obtained by summing the absolute differences between elements of adjacent structures. For task T0856, the segmentation cost at the third scheduling window is 2.41, which is 1.5 times higher than the average cost of 1.32 for the previous three scheduling windows. The system marks this location as a candidate segmentation point and, combined with the computation start identifier, identifies the corresponding segment as a computation burst phase. After full processing, 21,600 tasks are divided into 58,320 task computation phase segments, including 14,820 data loading phases, 12,640 computation burst phases, 7,920 memory swapping phases, 10,350 network backhaul phases, and 12,590 waiting / blocking phases.
[0187] The training sample set for the improved StemGNN model is constructed using historical runtime data, containing 126,000 samples and 18,000 validation samples. Each training sample includes a task computational phase segment, a phase semantic resource graph, node degradation state data, a ternary influence hypergraph, and historical scheduling feedback data. In sample A, task T0312 is in a computational burst phase, with resource requirements of CPU 0.42, GPU 0.86, NPU 0, memory 0.51, GPU memory 0.73, bandwidth 0.24, storage read / write 0.31, task queue length 16, and response latency 112ms. Node N07 has a GPU memory fragmentation rate of 0.28, a GPU queue wait time of 5.4s, a node temperature of 68℃, and scheduling feedback showing an actual GPU utilization of 0.78, a node performance margin of 0.19, a link congestion value of 0.22, a migration impairment value of 0.16, and a migration result of allowed migration. In sample B, task T1448 is in the network backhaul phase, with a bandwidth utilization of 0.82, a link queue length of 43, a packet loss rate of 0.013, a retransmission rate of 0.018, and scheduling feedback showing a link congestion value of 0.76 and a migration impairment value of 0.33. During training, a joint loss function is composed of five constraints: phase resource bundle prediction constraints, node degradation spectrum state reconstruction constraints, phase node compatibility value discrimination constraints, resource elastic boundary constraints, and non-migrating identifier discrimination constraints. The weights of all five constraints are set to 1. The basis for this setting is that the five constraints correspond to resource prediction, degradation reconstruction, node matching, boundary output, and migration discrimination, respectively, and any output error will affect the executability of the phase resource bundle. After 30 rounds of training, the absolute value of the difference between the joint loss values of five consecutive adjacent rounds is less than 0.001, resulting in the completed improved StemGNN model.
[0188] During real-time allocation, the system inputs the task's computational power phase segments into the trained improved StemGNN model, generating phase node compatibility values, resource elasticity boundaries, node performance margins, and non-migratable flags. For the computational burst phase of task T0856, node N07 has a phase node compatibility value of 0.86, resource elasticity boundaries of 1 to 2 GPUs, 22GB to 38GB of VRAM, and 6 to 12 CPU cores, a node performance margin of 0.23, and a non-migratable flag indicating migration is allowed. Node N12 has a phase node compatibility value of 0.79, but a VRAM fragmentation rate of 0.41 and a node performance margin of 0.08. The system connects task T0856, node N07, and link L18 as a ternary influence hyperedge, writing a migration impairment value of 0.14 and a link congestion value of 0.21. The CBBA algorithm sets candidate computing power nodes as node agents and generates resource bundle competition value by adding the phase node compatibility value and margin normalization value and subtracting the migration damage normalization value and link congestion value. The competition value corresponding to node N07 is 1.02 and the competition value corresponding to node N12 is 0.64. After the conflict is resolved, task T0856 is assigned to node N07.
[0189] After candidate schemes are generated, the system performs a Bid Time-Stamp refresh process. The effective bidding time window is 60 seconds, and 17 expired bidding records exceeding 60 seconds are rolled back to the public candidate pool. The system calculates the resource safety threshold based on the compressible lower bound of the resource elastic boundary. Based on the congestion values of 48 candidate links, a link congestion threshold of 0.57 is calculated, and 6 candidate links are found to have congestion values exceeding the link congestion threshold, involving 312 phase resource bundles. The system only rebids for the affected phase resource bundles, with local rebid taking 1.9 seconds, compared to 8.6 seconds for traditional full rescheduling. After the target computing power dynamic allocation scheme is generated, the system outputs 3420 task deployment instructions, 816 resource expansion instructions, 653 resource reduction instructions, 428 task migration instructions, 536 bandwidth adjustment instructions, and 11 node isolation instructions, and writes the scheduling feedback data back to the standardized computing power operation dataset.
[0190] Comparative tests were conducted on the same batch of 21,600 tasks. The average task response latency of the traditional method was 146ms, while that of the method of this invention was 104ms; the P95 response latency of the traditional method was 412ms, while that of the method of this invention was 276ms; the average CPU utilization of the traditional method was 61.8%, while that of the method of this invention was 72.6%; the average GPU utilization of the traditional method was 58.4%, while that of the method of this invention was 71.3%; the peak number of high-load nodes of the traditional method was 14, while that of the method of this invention was 8; the number of link congestion events of the traditional method was 126, while that of the method of this invention was 72; the number of invalid migrations of the traditional method was 97, while that of the method of this invention was 41; the task timeout rate of the traditional method was 5.8%, while that of the method of this invention was 2.7%; and the average scheduling decision time of the traditional method was 6.4s, while that of the method of this invention was 3.2s. As can be seen from this embodiment, the present invention subdivides tasks according to computing power phases and incorporates node degradation, link congestion, migration damage and resource elasticity boundary into the allocation process. This can solve the problems of node mismatch, link congestion and expired bidding occupying resources caused by coarse-grained scheduling of complete tasks, and verify the feasibility and effectiveness of the present invention in a hybrid computing power resource pool.
[0191] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for dynamically allocating and optimizing computing power performance based on deep learning, characterized in that, include: Collect task execution data, node resource data, link status data, container execution data, and task execution feedback data from the computing power resource pool, perform preprocessing, and generate a standardized computing power execution dataset; Based on the standardized computing power operation dataset, resource occupancy curves, task queue length curves, and response latency curves are extracted. The Toeplitz inverse covariance structure of resource variables within adjacent scheduling time windows is calculated. Synchronous segmentation is performed and phase semantic calibration is carried out to generate task computing power phase segments. An improved StemGNN model is constructed, which includes a phase semantic spectrogram embedding layer, a degenerate spectral state filtering layer, and a phase resource bundle readout head. The phase fragment of the task computing power is input into the improved StemGNN model to generate phase resource bundle information. A ternary influence hypergraph is constructed based on task computing power phase segments, phase resource bundle information, node resource data, link status data, and container operation data; The CBBA algorithm, which incorporates a Bid Time-Stamp refresh rule, is used to read the ternary influence hypergraph, perform resource bundle bidding and consensus conflict resolution, and generate candidate computing power dynamic allocation schemes. Perform Bid Time-Stamp refresh processing and local threshold-triggered rebidding on candidate computing power dynamic allocation schemes to generate target computing power dynamic allocation schemes; Based on the target computing power dynamic allocation scheme, computing power scheduling instructions are generated, scheduling feedback data is collected, and written into a standardized computing power operation dataset.
2. The method for dynamic allocation and optimization of computing power performance based on deep learning according to claim 1, characterized in that, The task execution data includes task number, task type, task submission time, task start time, task end time, task log identifier, task queue length, and response latency. The task log identifier includes data loading identifier, computation start identifier, memory swap identifier, network backhaul identifier, and waiting / blocking identifier. The node resource data includes node number, remaining resources of the same type, node running status, and node degradation status data. The link status data includes link number, link endpoint, and link congestion status data. The container execution data includes container number, container-borne task number, node number where the container is located, container migration status, and container congestion status.
3. The method for dynamic allocation and optimization of computing power performance based on deep learning according to claim 1, characterized in that, The preprocessing includes time alignment, field unification, unit unification, missing value completion, outlier removal, normalization, generation of scheduling time window markers, and generation of associated indexes.
4. The method for dynamic allocation and optimization of computing power performance based on deep learning according to claim 1, characterized in that, The generated task computing power phase segment includes: Read the task number, task type, task submission time, task start time, task end time, scheduling time window marker, resource usage curve, task queue length curve, response latency curve, and task log identifier from the standardized computing power operation dataset, and arrange them into a resource variable sequence according to the task number and scheduling time window marker; Resource variable sequences are extracted in units of continuous scheduling time windows. The inverse covariance relationship between resource variables within the same continuous scheduling time window is calculated, and the Toeplitz inverse covariance structure is generated in chronological order. Based on the Toeplitz inverse covariance structure corresponding to adjacent consecutive scheduling time windows, segmentation cost and segment continuity cost are generated. Based on the segmentation cost and segment continuity cost, synchronous segmentation is performed on the resource variable sequence to generate candidate segments for computing power phase. Based on the task type and task log identifier, phase semantic labeling is performed on the candidate segments of computing power phase, and the task number, phase type, phase start and end time, phase duration, resource variable sequence within the phase and Toeplitz inverse covariance structure within the phase are written to generate task computing power phase segments.
5. The method for dynamic allocation and optimization of computing power performance based on deep learning according to claim 1, characterized in that, The generated phase resource bundle information includes: An improved StemGNN model is constructed, which includes a phase semantic spectrogram embedding layer, a degenerate spectral state filtering layer, and a phase resource bundle readout head. The phase semantic spectrum embedding layer reads the phase segment of task computing power, sets the resource variables corresponding to the resource occupancy curve, task queue length curve and response latency curve as resource variable nodes, and writes the task number, phase type, scheduling time window mark and resource type into the resource variable nodes to generate a phase semantic resource graph. The phase semantic spectrum embedding layer performs graph Fourier transform on the phase semantic resource graph and discrete Fourier transform on the sequence of resource variables within the phase to generate a temporal representation of the phase spectrum. The degradation spectrum state filtering layer reads node degradation state data, link congestion state data, and container congestion state to generate node degradation spectrum state. The phase resource bundle readout head reads the node degradation spectrum status, phase semantic resource graph and node resource data, and generates phase node compatibility value, resource elastic boundary, node performance margin and non-migratable flag with associated task number, phase type, scheduling time window mark and node number. The resource elastic boundary includes the compressible lower limit and scalable upper limit of the corresponding resource. The phase node compatibility value, resource elasticity boundary, node performance margin, and non-migration identifier are combined into phase resource bundle information; A training sample set is constructed, which includes task computing power phase segments, phase semantic resource graphs, node degradation state data, ternary influence hypergraphs, and corresponding historical scheduling feedback data. The historical scheduling feedback data includes actual resource usage, actual node performance margin, actual link congestion value, actual migration impairment value, and actual migration result. The training sample set is input into the improved StemGNN model, and the improved StemGNN model is trained using a joint loss function that includes phase resource bundle prediction constraints, node degradation spectrum state reconstruction constraints, phase node compatibility value discrimination constraints, resource elastic boundary constraints, and non-migratable identifier discrimination constraints, resulting in the trained improved StemGNN model.
6. The method for dynamic allocation and optimization of computing power performance based on deep learning according to claim 1, characterized in that, The construction of the ternary influence hypergraph includes: Read task computing power phase segments and phase resource bundle information, generate a candidate computing power node set based on node resource data, generate a candidate transmission link set and link congestion value based on link status data, and generate migration impairment value based on container running data; The intersection of the resource elastic boundary and the remaining resources of the same type in the candidate computing power node set is truncated to generate the resource elastic boundary truncation result, and candidate computing power nodes with empty resource elastic boundary truncation results are deleted. Connect the task computing power phase segment, candidate computing power nodes, and candidate transmission links into a ternary influence hyperedge, and write the task number, phase type, phase start and end time, phase node compatibility value, resource elastic boundary truncation result, node performance margin, migration impairment value, link congestion value, and non-migration flag into the ternary influence hyperedge. The ternary influence hyperedges are grouped according to task number and phase type, and ternary influence hyperedges that do not have a connection between candidate computing power nodes and candidate transmission links are deleted to generate a ternary influence hypergraph.
7. The method for dynamic allocation and optimization of computing power performance based on deep learning according to claim 1, characterized in that, The generated candidate computing power dynamic allocation scheme includes: Read the ternary influence hyperedges in the ternary influence hypergraph, set the candidate computing power nodes as node agents, and generate phase resource bundles according to the combination of task number, phase type, candidate computing power nodes and candidate transmission links in the ternary influence hyperedges. Write the phase resource bundles that have not entered the task bundles of the node agents into the common candidate pool. The node agent reads the phase node compatibility value, node performance margin, migration impairment value and link congestion value in the phase resource bundle, uses the phase node compatibility value and node performance margin as positive bidding items, and uses the migration impairment value and link congestion value as negative deduction items, and generates the resource bundle bidding value based on the positive bidding items minus the negative deduction items. Write the resource bundle bidding value, node agent number, candidate computing power node, task number, phase type, candidate transmission link, phase start and end time, resource elastic boundary truncation result, node performance margin, non-migratable identifier, link congestion value and Bid Time-Stamp field into the bidding record, and arrange the bidding records according to the resource bundle bidding value to generate the node agent task bundle. When task bundles of intelligent agents at different nodes contain the same task number and the same phase type, the bidding record with the higher resource bundle bidding value is retained. When the resource bundle bidding values are the same, the bidding record with the earlier time corresponding to the Bid Time-Stamp field is retained, and a consensus bidding record is generated. Based on the consensus bidding record, extract the task number, phase type, phase start and end time, candidate computing power nodes, candidate transmission links, resource elastic boundary truncation results, node performance margin, non-migratable identifier, link congestion value and bidding timestamp to generate a candidate computing power dynamic allocation scheme.
8. The method for dynamic allocation and optimization of computing power performance based on deep learning according to claim 1, characterized in that, The dynamic allocation scheme for generating target computing power includes: Read the consensus bidding record and Bid Time-Stamp field in the candidate computing power dynamic allocation scheme, perform BidTime-Stamp refresh processing, and record the duration between the start time of the current scheduling time window and the end time of the current scheduling time window as the effective bidding time window; The auction survival time is calculated based on the Bid Time-Stamp field and the current scheduling time window termination time. Consensus auction records with a survival time longer than the effective auction time window are marked as expired auction records, and the phase resource bundles corresponding to the expired auction records are rolled back to the public candidate pool. Resource security thresholds are generated based on the compressible lower bound of resource elasticity boundaries, and link congestion thresholds are generated based on the link congestion values in the candidate transmission link set. When one of the following conditions is met: the node performance margin is less than the resource safety threshold, there are similar resources below the resource safety threshold in the resource elastic boundary truncation results, or the link congestion value is greater than the link congestion threshold, the corresponding candidate computing power node, candidate transmission link, and task computing power phase segment are marked as affected objects. Based on the affected objects, the affected phase resource bundles are extracted from the ternary influence hypergraph. Resource bundle bidding and consensus conflict resolution are re-executed on the affected phase resource bundles to generate updated consensus bidding records. Consensus bidding records that are not marked as expired are merged with the updated consensus bidding records to generate a target computing power dynamic allocation scheme.
9. The method for dynamic allocation and optimization of computing power performance based on deep learning according to claim 1, characterized in that, The generated computing power scheduling instructions include: Read the task number, phase type, candidate computing power nodes, candidate transmission links, resource elastic boundary truncation results, non-migratable flag, node performance margin and bidding timestamp from the target computing power dynamic allocation scheme, and generate scheduling execution records according to task number and phase type; Based on candidate computing power nodes, candidate transmission links, resource elastic boundary truncation results, and node performance margin, task deployment instructions, resource expansion instructions, resource reduction instructions, bandwidth adjustment instructions, and node isolation instructions are generated, and task migration instructions are generated when the non-migration mark is changed to allow migration. Collect the actual resource usage and task completion time after the scheduling is executed, generate the actual link congestion value based on link status data, generate the actual migration damage value based on container operation data, generate the actual node degradation status based on node resource data, and combine them to generate scheduling feedback data. The scheduling feedback data is written into the standardized computing power operation dataset according to the task number, phase type, candidate computing power node, candidate transmission link, and scheduling time window.