Self-adapting elastic expansion method and system for a computer reinforced based on a domestic platform

CN122285266APending Publication Date: 2026-06-26SUZHOU YAOGUO ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU YAOGUO ELECTRONICS CO LTD
Filing Date
2026-03-06
Publication Date
2026-06-26

Smart Images

  • Figure CN122285266A_ABST
    Figure CN122285266A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for adaptive elastic scaling of ruggedized computers based on a domestically developed platform, relating to the field of computer resource management technology. The method includes: collecting computing unit operating state parameters to construct a multi-dimensional state tensor; identifying key driving characteristics through tensor decomposition and causal inference analysis; establishing a mapping relationship between load characteristics and resource requirements; quantifying uncertainty to obtain capacity demand prediction results and confidence intervals; and finally determining the execution strategy for scaling decisions. This invention improves the accuracy of resource allocation and the reliability of scaling decisions, reduces resource waste, and enhances system stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer resource management technology, and in particular to a method and system for adaptive elastic scaling of hardened computers based on a domestic platform. Background Technology

[0002] With the widespread application of information systems and the continuous development of critical businesses, the use of domestically developed ruggedized computer systems in important fields such as national defense, finance, and energy is increasing. This necessitates flexible adjustments to computing resources based on changes in workload, while ensuring security and stability. Ruggedized computer systems possess high security levels and unique physical constraints; resource expansion must consider these system-specific characteristics while simultaneously guaranteeing system reliability and stability.

[0003] Currently, domestic platform hardening and computer resource expansion technologies are mainly implemented through preset rules or simple load monitoring. However, they still suffer from problems such as a lack of analytical capabilities for the complex dependencies between computing units in hardened computer systems, an inability to accurately capture the multidimensional characteristics and evolution of system load, a lack of in-depth understanding of the causal mechanisms behind load changes, difficulty in identifying the key driving factors that truly affect system performance, an inability to quantify the uncertainty of prediction results, a lack of flexible adjustment mechanisms, and difficulty in dynamically adjusting expansion strategies based on prediction confidence. Summary of the Invention

[0004] This invention provides a method and system for adaptive elastic scaling of ruggedized computers based on a domestically developed platform, which can at least solve some of the problems existing in the prior art.

[0005] A first aspect of this invention provides a method for adaptive elastic scaling of ruggedized computers based on a domestically developed platform, comprising: The system collects the operating status parameters of each computing unit in the hardened computer, constructs a multidimensional state tensor based on the operating status parameters, extracts the temporal evolution vector through tensor decomposition, determines the dependency relationship between computing units based on the temporal evolution vector, and obtains the load feature set. Multi-scale temporal decomposition of the load feature set yields a composite load feature vector. The composite load feature vector is analyzed using a causal inference algorithm to obtain causal analysis results. Counterfactual reasoning is then used to identify a subset of features that have a causal impact on resource demand. The causal strength coefficient of each feature in the subset is calculated, and key driving features are determined based on the causal strength coefficient. Based on the key driving characteristics, a correspondence between load characteristics and resource requirements is established through multi-layer nonlinear mapping, and uncertainty quantification is performed to obtain capacity demand prediction results. The confidence interval corresponding to the capacity demand prediction results is calculated through Bayesian inference algorithm. Based on the capacity demand forecast results and the physical constraint parameters corresponding to the ruggedized computer, an expansion decision instruction is determined. The uncertainty metric of the expansion decision instruction is calculated based on the confidence interval, and the execution strategy of the expansion decision instruction is determined based on the uncertainty metric. The expansion decision instruction is executed according to the execution strategy until the expansion is completed.

[0006] In one alternative implementation, The system collects operational status parameters of each computing unit in the hardened computer, constructs a multidimensional state tensor based on these parameters, and extracts a temporal evolution vector through tensor decomposition. Based on this temporal evolution vector, it determines the dependencies between computing units, resulting in a load feature set including: The operating status parameters of each computing unit in the hardened computer are obtained within a preset monitoring period, and the operating status parameters are tensorized and organized according to a preset dimension to obtain a multidimensional state tensor. The multidimensional state tensor is modally decomposed to obtain a time modal factor matrix and the evolution trajectory of each factor component with time dimension is extracted. The evolution trajectory is reconstructed in phase space to obtain the trajectory attractor corresponding to each factor component. The evolution stability index of each factor component is calculated based on the topological invariants of the trajectory attractor. The stable evolution factor component is determined based on the evolution stability index and a preset stability threshold and the temporal evolution vector is reconstructed. The delay embedding of each component in the temporal evolution vector is expanded in the time dimension to obtain a delay embedding matrix. The mutual information of the delay embedding matrix is ​​calculated to obtain the information transmission strength between computing units and a directed association graph is constructed. In the directed association graph, the direct and indirect dependency paths between computing units are identified by transitive closure operation and the key dependency chain is determined by path weight quantization. The path topology metric of each computing unit in the key dependency chain is used as the structural feature of the dependency relationship. The amplitude change rate and phase difference of each component in the temporal evolution vector are used as the dynamic feature of the dependency relationship. The load feature set is obtained by fusing the structural feature and the dynamic feature.

[0007] In one alternative implementation, Multi-scale time-series decomposition of the load feature set yields a composite load feature vector. Causal inference algorithms are then used to analyze this composite load feature vector, yielding causal analysis results including: The evolution data of load features in the load feature set in the time dimension are segmented and trend-fitted according to a preset time window length to obtain trend components. The difference between the evolution data and the trend components is calculated to obtain residual data. Periodic analysis and frequency domain transformation are performed on the residual data to obtain periodic oscillation modes. The time derivative of the trend components is calculated to obtain the long-term evolution direction. The long-term evolution direction and the periodic oscillation mode are spliced ​​together under different time window lengths and superimposed layer by layer according to the time window length to obtain a composite load feature vector. The feature components in the composite load feature vector are used as nodes to construct a conditional independence test matrix, and the independence of each node is statistically tested. Based on the test results, the nodes with conditional dependencies are retained to obtain an undirected association structure. Collision structures are identified in the undirected association structure, and the causal direction of the edges is labeled according to the prior temporal relationship of the incoming edges in the collision structure to obtain a directed causal graph. In the directed causal graph, the node corresponding to each feature component is selected as the intervention node in turn, and the incoming edges of the intervention node are blocked. The numerical distribution distance of the resource demand corresponding to each node before and after the blocking is calculated and used as the causal effect strength. The causal analysis results are obtained by combining each node and its corresponding node identifier and causal effect strength.

[0008] In one alternative implementation, Combining counterfactual reasoning to identify a subset of features that have a causal impact on resource demand, calculating the causal strength coefficient of each feature in the subset, and determining key driving features based on the causal strength coefficient includes: The feature components in the causal analysis results are used as the original feature components, and counterfactual samples are constructed by replacing the values ​​with pre-acquired historical statistical quantile values. The counterfactual samples and the original feature components are weighted and aggregated to obtain counterfactual aggregated feature vectors and original aggregated feature vectors. The counterfactual aggregated feature vectors and the original aggregated feature vectors are queried and matched with the preset resource demand mapping relationship to obtain counterfactual resource demand values ​​and original resource demand values. The absolute difference and numerical change range between the counterfactual resource demand values ​​and the original resource demand values ​​are calculated. An initial feature subset is constructed based on the feature components whose absolute difference exceeds a preset difference threshold, and the causal strength coefficient is calculated by combining the numerical change range. The causal intensity coefficients are sorted in descending order and accumulated one by one to obtain a cumulative sum sequence. The growth rate between adjacent elements in the cumulative sum sequence is subjected to second-order difference to obtain second-order difference values. The position corresponding to the second-order difference value that exceeds the preset mutation threshold is taken as the inflection point position. The features before the inflection point position are taken as the dominant feature set. The features in the dominant feature set are combined in pairs to obtain feature pairs. The difference between the joint probability distribution and the marginal probability distribution of each feature pair is calculated to obtain the mutual information. The key driving features are determined based on the mutual information.

[0009] In one alternative implementation, Based on the aforementioned key driving characteristics, a correspondence between load characteristics and resource requirements is established through multi-layer nonlinear mapping, and uncertainty quantification is performed to obtain capacity demand prediction results. The confidence intervals corresponding to the capacity demand prediction results are calculated using a Bayesian inference algorithm, including: The key driving features are normalized to zero mean to obtain standardized features. Based on the pre-acquired historical load features and historical resource requirements, a mapping weight matrix is ​​constructed using the gradient descent algorithm. Linear transformation features are calculated based on the standardized features and the mapping weight matrix. Nonlinear activation transformation is performed on the linear transformation features to obtain activation features. A high-dimensional abstract feature representation is determined based on the activation features. The mean and variance parameters of the prediction distribution are solved based on the high-dimensional abstract features. The posterior probability distribution of capacity demand is constructed using variational inference methods. Multiple sets of capacity demand prediction samples are generated from the posterior probability distribution using the Monte Carlo sampling method, and statistical analysis is performed to obtain the capacity demand prediction results and prediction stable values. The target confidence level is initialized based on the preset task requirements. Based on the target confidence level, the quantiles of the posterior probability distribution are solved using a Bayesian inference algorithm to obtain the quantile coefficients. The confidence radius is then calculated by combining the predicted stable value. Based on the capacity demand prediction result and the confidence radius, the confidence interval corresponding to the capacity demand prediction result is obtained by using an interval construction method.

[0010] In one alternative implementation, Based on the capacity demand forecast results and the physical constraint parameters corresponding to the ruggedized computer, an expansion decision instruction is determined. The uncertainty metric of the expansion decision instruction is calculated based on the confidence interval, including: Obtain the resource expansion granularity and resource response latency corresponding to the hardened computer, extract the capacity demand point estimate from the capacity demand prediction result and round up according to the resource expansion granularity to obtain the granularity-aligned resource amount, determine whether the granularity-aligned resource amount exceeds the preset maximum scalable resource capacity, if it exceeds, then use the maximum scalable resource capacity as the baseline expansion amount, otherwise use the granularity-aligned resource amount as the baseline expansion amount. Extract the upper and lower bounds of the confidence interval corresponding to the capacity demand forecast result. Calculate the absolute values ​​of the differences between the upper and lower bounds of the confidence interval and the estimated capacity demand point to obtain the upward and downward fluctuation amplitudes. Determine the maximum fluctuation amplitude based on the upward and downward fluctuation amplitudes and round up according to the resource expansion granularity to obtain the uncertainty buffer. Solve for the baseline expansion and the uncertainty buffer to obtain the safety expansion. Generate an expansion decision instruction based on the safety expansion and the resource response delay. Map the uncertainty buffer and the baseline expansion to fuzzy membership features using a fuzzy inference algorithm and perform fuzzy rule inference to obtain a fuzzy risk assessment value. Defuzzify the value to obtain a quantified risk value and solve for the uncertainty metric corresponding to the expansion decision instruction.

[0011] In one alternative implementation, Determining the execution strategy for the extended decision instruction based on the uncertainty measure, and executing the extended decision instruction according to the execution strategy until the extension is completed, includes: A preset risk threshold is obtained, and it is determined whether the uncertainty measure exceeds the risk threshold. If it exceeds the threshold, a preset conservative expansion strategy is adopted. The total amount of expansion resources is determined based on the upper bound of the pre-obtained confidence interval, and the total amount of expansion resources is decomposed into multiple expansion batches. If it does not exceed the threshold, a standard expansion strategy is adopted. The total amount of expansion resources is determined based on the capacity demand prediction result, and resources are configured within a single expansion batch. An execution strategy containing resource allocation and expansion time arrangement is generated based on the conservative expansion strategy or the standard expansion strategy. According to the extension time schedule in the execution strategy, the resource extension operation is performed on the hardened computer in sequence. After each extension batch is completed, the current running status parameters of the hardened computer are collected and the resource utilization rate is calculated. Based on the resource utilization rate and the capacity demand prediction result, it is determined whether the capacity demand has been met. If it is met, the execution of subsequent extension batches is terminated. If the capacity demand is not met, the extension continues until all extension batches in the execution strategy are completed.

[0012] A second aspect of this invention provides a ruggedized computer adaptive elastic expansion system based on a domestically developed platform, comprising: The data acquisition unit is used to collect the operating status parameters of each computing unit in the hardened computer, construct a multidimensional state tensor based on the operating status parameters and extract the temporal evolution vector through tensor decomposition, determine the dependency relationship between computing units based on the temporal evolution vector, and obtain the load feature set. The feature analysis unit is used to perform multi-scale time-series decomposition on the load feature set to obtain a composite load feature vector, analyze the composite load feature vector through a causal inference algorithm to obtain causal analysis results, and identify the feature subset that has a causal impact on resource demand by combining counterfactual reasoning, calculate the causal strength coefficient of each feature in the feature subset, and determine the key driving features based on the causal strength coefficient. The demand forecasting unit is used to establish the correspondence between load characteristics and resource demand through multi-layer nonlinear mapping based on the key driving characteristics and to perform uncertainty quantification assessment to obtain the capacity demand forecasting result. The confidence interval corresponding to the capacity demand forecasting result is calculated through Bayesian inference algorithm. The decision execution unit is configured to determine an expansion decision instruction based on the capacity demand forecast result and the physical constraint parameters corresponding to the ruggedized computer, calculate the uncertainty metric of the expansion decision instruction based on the confidence interval, determine the execution strategy of the expansion decision instruction based on the uncertainty metric, and execute the expansion decision instruction according to the execution strategy until the expansion is completed.

[0013] A third aspect of the present invention provides an electronic device, comprising: A processor and a memory for storing processor-executable instructions, wherein the processor is configured to invoke instructions stored in the memory to perform the aforementioned method.

[0014] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0015] In this invention, by extracting temporal evolution vectors through tensor decomposition and determining the dependencies between computing units, the multidimensional characteristics of system load can be comprehensively captured, providing an accurate foundation for subsequent analysis. By employing a causal inference algorithm combined with counterfactual reasoning, key driving features that substantially affect resource demand can be accurately identified, avoiding misjudgments that may be caused by traditional correlation analysis and improving the accuracy of expansion decisions. Through multi-layer nonlinear mapping and Bayesian inference algorithms, not only is capacity demand predicted, but the uncertainty of the prediction results is also quantified, providing a reliable confidence interval for expansion decisions and enhancing the reliability of the decisions. Based on uncertainty measurement, the execution strategy of expansion decision instructions is determined, which can dynamically adjust resource expansion behavior according to the credibility of the prediction, improving elasticity and resource utilization efficiency. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the adaptive elastic scaling method for ruggedized computers based on a domestic platform, as described in an embodiment of the present invention. Figure 2This is a flowchart illustrating the resource expansion decision and uncertainty calculation process of the adaptive elastic scaling method for hardened computers based on a domestic platform, as described in an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0019] Figure 1 This is a flowchart illustrating the adaptive elastic scaling method for ruggedized computers based on a domestic platform, as described in an embodiment of the present invention. Figure 1 As shown, the method includes: The system collects the operating status parameters of each computing unit in the hardened computer, constructs a multidimensional state tensor based on the operating status parameters, extracts the temporal evolution vector through tensor decomposition, determines the dependency relationship between computing units based on the temporal evolution vector, and obtains the load feature set. Multi-scale temporal decomposition of the load feature set yields a composite load feature vector. The composite load feature vector is analyzed using a causal inference algorithm to obtain causal analysis results. Counterfactual reasoning is then used to identify a subset of features that have a causal impact on resource demand. The causal strength coefficient of each feature in the subset is calculated, and key driving features are determined based on the causal strength coefficient. Based on the key driving characteristics, a correspondence between load characteristics and resource requirements is established through multi-layer nonlinear mapping, and uncertainty quantification is performed to obtain capacity demand prediction results. The confidence interval corresponding to the capacity demand prediction results is calculated through Bayesian inference algorithm. Based on the capacity demand forecast results and the physical constraint parameters corresponding to the ruggedized computer, an expansion decision instruction is determined. The uncertainty metric of the expansion decision instruction is calculated based on the confidence interval, and the execution strategy of the expansion decision instruction is determined based on the uncertainty metric. The expansion decision instruction is executed according to the execution strategy until the expansion is completed.

[0020] In one alternative implementation, The system collects operational status parameters of each computing unit in the hardened computer, constructs a multidimensional state tensor based on these parameters, and extracts a temporal evolution vector through tensor decomposition. Based on this temporal evolution vector, it determines the dependencies between computing units, resulting in a load feature set including: The operating status parameters of each computing unit in the hardened computer are obtained within a preset monitoring period, and the operating status parameters are tensorized and organized according to a preset dimension to obtain a multidimensional state tensor. The multidimensional state tensor is modally decomposed to obtain a time modal factor matrix and the evolution trajectory of each factor component with time dimension is extracted. The evolution trajectory is reconstructed in phase space to obtain the trajectory attractor corresponding to each factor component. The evolution stability index of each factor component is calculated based on the topological invariants of the trajectory attractor. The stable evolution factor component is determined based on the evolution stability index and a preset stability threshold and the temporal evolution vector is reconstructed. The delay embedding of each component in the temporal evolution vector is expanded in the time dimension to obtain a delay embedding matrix. The mutual information of the delay embedding matrix is ​​calculated to obtain the information transmission strength between computing units and a directed association graph is constructed. In the directed association graph, the direct and indirect dependency paths between computing units are identified by transitive closure operation and the key dependency chain is determined by path weight quantization. The path topology metric of each computing unit in the key dependency chain is used as the structural feature of the dependency relationship. The amplitude change rate and phase difference of each component in the temporal evolution vector are used as the dynamic feature of the dependency relationship. The load feature set is obtained by fusing the structural feature and the dynamic feature.

[0021] This technical solution is deployed on a domestically produced ruggedized computer platform built with Phytium CPUs and the Kylin operating system. It acquires the operational status parameters of each computing unit within a preset monitoring period, including but not limited to processor utilization, memory usage, data throughput, and response time. It can collect continuous status data for 72 hours at a 5-minute sampling interval, forming a high-dimensional dataset. The collected status parameters are tensor-organized according to preset dimensions to construct a multi-dimensional state tensor. The preset dimensions include computing unit identifier, parameter type, and timestamp, thus forming a third-order tensor structure. For example, for a scenario with 20 computing units, monitoring 10 status parameters, and sampling 864 time points, a 20×10×864 third-order tensor can be constructed.

[0022] A modal decomposition operation is performed on the constructed multidimensional state tensor, decomposing it into individual modal factor matrices using a tensor decomposition algorithm. The tensor decomposition algorithm is run in a domestically developed hardened computing environment. For the temporal modal factor matrix, the evolution trajectory of each factor component over time is extracted. Taking the factor component corresponding to the CPU utilization of a computing unit as an example, the corresponding temporal modal evolution trajectory can be represented as a one-dimensional vector of length 864. Phase space reconstruction is performed on each evolution trajectory, with an embedding dimension of 5 and a time delay of 3 sampling points, constructing a set of trajectory points in a five-dimensional phase space from the one-dimensional time series. By analyzing the reconstructed phase space trajectories, the topological characteristics of the trajectory attractors are identified. Based on the topological invariants of the trajectory attractors, the evolutionary stability index of each factor component is calculated. The smaller the stability index value, the more stable the system. The calculation method is based on the dispersion and convergence characteristics of the trajectory in the phase space. A stability threshold of 0.35 is set; when the stability index value of a factor component is less than this threshold, it is determined to be a stable evolving factor component; otherwise, it is an unstable component. The identified stable evolutionary factors are recombined to construct a time-series evolution vector for subsequent dependency analysis.

[0023] The reconstructed temporal evolution vectors are subjected to delayed embedding processing, with each component expanded along the time dimension. An embedding dimension of 4 and a delay of 2 sampling points are chosen, transforming each 864-length temporal vector into an 860×4 delayed embedding matrix. The mutual information between the delayed embedding matrices corresponding to different computational units is calculated to determine the information transfer strength between computational units. The mutual information calculation uses a histogram method, dividing the data into 32 equally wide intervals for probability estimation. A directed association graph between computational units is constructed using mutual information as edge weights. Transitive closure is applied to the directed association graph to identify direct and indirect dependency paths between computational units. A mutual information threshold of 0.6 is set; when the mutual information between two computational units exceeds this threshold, a connection is established in the directed graph. Key dependency chains are determined by quantitative analysis of the weights of each path. The path weight is equal to the geometric mean of the weights of each edge on the path; the top 20% of paths with the highest weights are selected as key dependency chains.

[0024] For the identified key dependency chains, structural and dynamic features are extracted. Structural features include path topology metrics for each computational unit in the dependency chain, such as in-degree, out-degree, and centrality. For example, a computational unit with an out-degree of 5 and an in-degree of 2 in the dependency network indicates that this unit outputs data to 5 other units while receiving input data from 2 units. Dynamic features are based on the temporal characteristics of each component in the time-series evolution vector, including the rate of change of amplitude and phase difference. The rate of change of amplitude is calculated as the relative change between consecutive sampling points, and the phase difference is calculated by extracting the phase information of the signal through Hilbert transform and calculating the phase difference between units. For example, if the average phase difference of the time-series signals of two interdependent computational units is 45 degrees, it indicates a significant lead-lag relationship.

[0025] The extracted structural and dynamic features are fused to form a load feature set. The fusion method employs feature vector concatenation and weight normalization, with a weight ratio of 6:4 for structural and dynamic features. The generated load feature set contains the topological structure and dynamic evolution characteristics of the dependencies between computing units, providing a decision-making basis for the adaptive scaling of hardened computers.

[0026] For example, in a domestically produced industrial control system, a ruggedized computer system was analyzed. This system comprises 12 computing units, operates on a 72-hour monitoring cycle, and generates a multidimensional state tensor with dimensions of 12×8×864. The temporal modal factor matrix obtained through modal decomposition contains eight main factor components. Six of these components have stability index values ​​less than the 0.35 threshold and are thus identified as stable evolution components. After delay embedding of the time-series evolution vectors reconstructed from these six stable components, the calculated mutual information matrix between computing units shows that eight pairs of computing units have mutual information values ​​exceeding the 0.6 threshold. In the constructed directed association graph, five key dependency chains were identified through transitive closure operations. One of these key dependency chains contains three computing units, forming a dependency structure of "unit 4 → unit 7 → unit 2". Unit 4 has an out-degree of 3, unit 7 has in-degrees of 2 and 2 respectively, and unit 2 has an in-degree of 4. The average mutual information value of this dependency chain is 0.78, and the average phase difference of the time-series signals is 38 degrees. After integrating structural and dynamic features, the load feature value of this dependency chain is 0.82, ranking first among all dependency chains, indicating that it is the most critical dependency path in the system.

[0027] In this embodiment, by high-dimensional temporal modeling and evolutionary characterization of the operating states of each computing unit in the ruggedized computer, a refined and forward-looking characterization of the load correlation of computing units is achieved, significantly improving the stability, accuracy, and interpretability of load feature characterization. By organizing multi-source operating state parameters into multi-dimensional state tensors and performing mode decomposition, dominant factors with independent evolutionary laws can be separated from high-dimensional, strongly coupled monitoring data, avoiding the feature confusion problem caused by the superposition of different state parameters in traditional methods. By delaying the embedding of the temporal evolution vector and calculating the mutual information, the true information transmission strength between computing units can be characterized. It can not only identify direct dependencies but also reveal indirect dependency paths formed by multi-level transmission, improving the resource management efficiency and operational reliability of the ruggedized computer under complex working conditions.

[0028] In one alternative implementation, Multi-scale time-series decomposition of the load feature set yields a composite load feature vector. Causal inference algorithms are then used to analyze this composite load feature vector, yielding causal analysis results including: The evolution data of load features in the load feature set in the time dimension are segmented and trend-fitted according to a preset time window length to obtain trend components. The difference between the evolution data and the trend components is calculated to obtain residual data. Periodic analysis and frequency domain transformation are performed on the residual data to obtain periodic oscillation modes. The time derivative of the trend components is calculated to obtain the long-term evolution direction. The long-term evolution direction and the periodic oscillation mode are spliced ​​together under different time window lengths and superimposed layer by layer according to the time window length to obtain a composite load feature vector. The feature components in the composite load feature vector are used as nodes to construct a conditional independence test matrix, and the independence of each node is statistically tested. Based on the test results, the nodes with conditional dependencies are retained to obtain an undirected association structure. Collision structures are identified in the undirected association structure, and the causal direction of the edges is labeled according to the prior temporal relationship of the incoming edges in the collision structure to obtain a directed causal graph. In the directed causal graph, the node corresponding to each feature component is selected as the intervention node in turn, and the incoming edges of the intervention node are blocked. The numerical distribution distance of the resource demand corresponding to each node before and after the blocking is calculated and used as the causal effect strength. The causal analysis results are obtained by combining each node and its corresponding node identifier and causal effect strength.

[0029] The evolution data of load features over time in the load feature set is segmented and trend-fitted according to a preset time window length to obtain trend components. The preset time window length is 12 hours, and the monitoring data for 72 consecutive hours can be divided into 6 time windows. A polynomial fitting method is used to extract trends from the load feature evolution data within each time window. The polynomial order is set to 3 to balance fitting accuracy and computational complexity. Taking the processor utilization rate of a computing unit in a ruggedized computer with a domestic processor architecture as an example, the evolution data within 72 hours is divided into 6 time windows, each containing 144 sampling points. A third-order polynomial fitting is performed on the 144 sampling points in the first window to obtain the trend component expression, with fitting parameters of 0.0023, -0.0417, 0.2105, and 42.3768.

[0030] Residual data is obtained by subtracting the evolutionary data from the trend component. For each time window, the point-to-point difference between the original evolutionary data and the corresponding trend component is calculated to form a residual data sequence. The residual data represents the short-term fluctuations of the load characteristics outside the long-term trend. Taking the aforementioned computing unit as an example, the root mean square value of the residual data in the first time window is 3.82, indicating that the actual processor utilization fluctuates around the trend line and is related to the task scheduling strategy under the domestic platform. Periodic oscillation modes are obtained by performing periodic analysis and frequency domain transformation on the residual data. The periodic analysis of the residual data is performed using a combination of autocorrelation function and discrete Fourier transform. The autocorrelation function is used to initially identify potential periodicity, with the maximum lag order being one-third of the length of the residual sequence. The discrete Fourier transform converts the residual data from the time domain to the frequency domain, and the power spectral density is estimated using the piecewise average periodogram method, with a piecewise length of 48 sampling points and an overlap rate of 50%. The first three main frequency components and their corresponding amplitudes and phases are extracted from the power spectral density to form the periodic oscillation modes. Taking the residual data of the first time window as an example, the three main oscillation periods identified under the domestic operating system environment are 12.5 hours, 4.2 hours and 2.1 hours, respectively, with corresponding normalized amplitudes of 0.7, 0.5 and 0.3. These periodic characteristics reflect the inherent laws of task batch processing and scheduling in the domestic system.

[0031] The long-term evolution direction is obtained by calculating the time derivative of the trend component. The central difference method is used to numerically differentiate the trend component, calculating the derivative value at each time point, representing the rate and direction of change of the load characteristics at that moment. To reduce noise, a moving average is applied to the derivative value sequence, with an average window length of 5 sampling points. Taking the trend component of the aforementioned calculation unit as an example, the average value of the derivative value within the first time window is 0.18, indicating that the processor utilization generally shows an upward trend, which is consistent with the characteristic of the domestically hardened platform gradually stabilizing resource usage after startup. The long-term evolution direction and the periodic oscillation mode are concatenated at different time window lengths and superimposed layer by layer according to the time window length to obtain a composite load feature vector. Specifically, the long-term evolution direction (derivative sequence) within each time window is combined with the parameters (frequency, amplitude, phase) of the periodic oscillation mode to form a feature vector. After forming feature vectors for each of the six time windows, weights are set according to the window length, with longer windows having higher weights than shorter windows, and a weighted summation method is used for superposition. The weighting coefficients are set to the inverse square of the window number, i.e., the weight of the first window is 1, the weight of the second window is 1 / 4, and so on. After normalization, the final composite load feature vector is obtained. For hardened computer systems based on domestic processor architecture, the extracted composite load feature vector has 18 dimensions per computing unit.

[0032] The conditional independence test matrix is ​​constructed by treating the eigencomponents in the composite load eigenvector as nodes, and the independence of each node is statistically tested. A conditional independence test matrix is ​​constructed for each number of eigencomponents. Each element in the matrix represents the conditional independence test statistic between eigencomponents given all other eigencomponents. A kernel-based conditional independence test algorithm is adopted, with a Gaussian radial basis function chosen as the kernel function and the bandwidth parameter determined using the median rule. This algorithm can run efficiently on domestic computing platforms with limited memory resources. The significance level is set to 0.05; when the test probability value is greater than the significance level, the two eigencomponents are considered conditionally independent; otherwise, they are considered conditionally dependent. In a test example running under a domestic operating system environment, the test results for a 12-dimensional eigenvector show that there are 28 pairs of eigencomponents with significant conditional dependencies, forming a preliminary dependency structure.

[0033] Based on the test results, node connections with conditional dependencies are retained to obtain an undirected associative structure. The conditional independence test matrix is ​​thresholded, retaining connections between node pairs with test probabilities less than the significance level, forming an undirected graph structure. Collision structures are identified within the undirected associative structure, and causal orientation is assigned to edges based on the prior temporal relationships of the incoming edges in the collision structure, resulting in a directed causal graph. A collision structure refers to a structure resembling a three-node path, where the two endpoints are not directly connected. For each collision structure, the temporal order of the corresponding feature components of the three nodes is analyzed, and the edge direction is determined according to the principle of "causality precedes outcome." When the temporal relationship is unclear, the minimum description length criterion is used for direction determination, which has low computational complexity in domestic computing environments. Of the 15 collision structures identified in the domestically hardened computer system, 12 can be identified by temporal relationships, and the remaining 3 are identified by the minimum description length criterion.

[0034] In a directed causal graph, nodes corresponding to each feature component are sequentially selected as intervention nodes, and their incoming edges are blocked. For each intervention node, the change in system state is simulated when all its incoming edges are blocked (i.e., all causal inputs are cut off). A probabilistic path tracing algorithm is used to calculate the intervention propagation effect, with a tracing depth of 3 layers and a decay factor of 0.7. The numerical distribution distance of resource demand corresponding to each node before and after blocking is calculated and used as the causal effect strength. The numerical distribution distance uses a distance metric, and distribution samples before and after intervention are generated through simulation, with a sample size of 1000. In a causal analysis example running on a domestically developed hardened computing platform, node intervention analysis of 12 key computing units shows that node 5 has the strongest intervention effect, with an average causal effect strength of 0.85, indicating that this node occupies a core position in the resource demand propagation network, consistent with the resource control characteristics of core service components on the domestic platform. The causal analysis results are obtained by combining each node, its corresponding node identifier, and the causal effect strength. For each node, record its identification information (such as computational unit number and feature type), its positional characteristics in the causal network (such as in-degree, out-degree, and centrality), and the strength of the causal effect after intervention, forming a complete set of causal analysis results.

[0035] In this embodiment, by performing trend decomposition and residual analysis on the load characteristic evolution data, the long-term change trend and short-term fluctuation components are effectively separated, so that the slow evolution direction and periodic oscillation mode of resource load can be clearly characterized. Multiple time window lengths are introduced and superimposed layer by layer to construct a composite load feature vector, so that the load features can simultaneously reflect the evolution characteristics under different time scales. This improves the overall ability to characterize changes in resource demand under complex working conditions and cross-scale consistency. The correlation structure between features is constructed through conditional independence test, and causal direction labeling is performed by combining collision structure identification and temporal prior. This overcomes the defect of not being able to distinguish between causal and covariant relationships, effectively reduces the risk of misjudgment and over-configuration, and improves the scientificity and stability of the hardened computer in resource planning and elastic scheduling under dynamic load environment.

[0036] In one alternative implementation, Combining counterfactual reasoning to identify a subset of features that have a causal impact on resource demand, calculating the causal strength coefficient of each feature in the subset, and determining key driving features based on the causal strength coefficient includes: The feature components in the causal analysis results are used as the original feature components, and counterfactual samples are constructed by replacing the values ​​with pre-acquired historical statistical quantile values. The counterfactual samples and the original feature components are weighted and aggregated to obtain counterfactual aggregated feature vectors and original aggregated feature vectors. The counterfactual aggregated feature vectors and the original aggregated feature vectors are queried and matched with the preset resource demand mapping relationship to obtain counterfactual resource demand values ​​and original resource demand values. The absolute difference and numerical change range between the counterfactual resource demand values ​​and the original resource demand values ​​are calculated. An initial feature subset is constructed based on the feature components whose absolute difference exceeds a preset difference threshold, and the causal strength coefficient is calculated by combining the numerical change range. The causal intensity coefficients are sorted in descending order and accumulated one by one to obtain a cumulative sum sequence. The growth rate between adjacent elements in the cumulative sum sequence is subjected to second-order difference to obtain second-order difference values. The position corresponding to the second-order difference value that exceeds the preset mutation threshold is taken as the inflection point position. The features before the inflection point position are taken as the dominant feature set. The features in the dominant feature set are combined in pairs to obtain feature pairs. The difference between the joint probability distribution and the marginal probability distribution of each feature pair is calculated to obtain the mutual information. The key driving features are determined based on the mutual information.

[0037] The feature components from the causal analysis results are used as the original feature components, and counterfactual samples are constructed by replacing the values ​​with pre-acquired historical statistical quantile values. For each feature component, the distribution characteristics of the corresponding historical data are analyzed, and the values ​​of the 25%, 50%, and 75% quantiles are calculated. On a domestically hardened computer platform, the current value of the memory occupancy feature component of a certain computing unit is 68%, and the corresponding 25%, 50%, and 75% quantiles of the historical data are 42%, 53%, and 76%, respectively. When constructing the counterfactual samples, the current value of 68% is replaced with the 50% quantile value of 53%, simulating the scenario of feature reduction. Similar replacement operations are performed on all feature components to form a complete set of counterfactual samples. The counterfactual samples and the original feature components are weighted and aggregated to obtain the counterfactual aggregated feature vector and the original aggregated feature vector, respectively. The weighted aggregation adopts a linear combination method, and the weight coefficients are determined according to the causal centrality of each feature component. Causal centrality is calculated by the connectivity of the feature in the causal network; the higher the connectivity, the wider the influence range of the feature, and the greater the weight is assigned. Taking 12 feature components as an example, the weighted aggregate value of the original feature is 58.7, and the weighted aggregate value of the counterfactual feature is 49.3.

[0038] The counterfactual aggregated feature vector and the original aggregated feature vector are respectively matched with a preset resource demand mapping relationship to obtain the counterfactual resource demand value and the original resource demand value. The preset resource demand mapping relationship is implemented through a lookup table, mapping the aggregated feature vector to specific resource demand quantification indicators. The lookup table covers typical intervals in the feature space, with an interval of 5 units. For values ​​not in the table, the nearest neighbor interpolation method is used to determine the mapping result. For example, on a domestic platform hardened computer, the original aggregated feature vector 58.7 maps to a resource demand value of 3.86 (number of CPU cores), and the counterfactual aggregated feature vector 49.3 maps to a resource demand value of 2.71. The absolute difference and the magnitude of change between the counterfactual resource demand value and the original resource demand value are calculated. The absolute difference is directly calculated as the absolute value of the difference between the two demand values, and the magnitude of change is calculated as the relative rate of change. In the above example, the absolute difference is 1.15, and the magnitude of change is 29.8%.

[0039] An initial feature subset was constructed based on feature components whose absolute differences exceeded a preset difference threshold, and the causal strength coefficient was calculated by combining this with the magnitude of numerical change. The preset difference threshold was 0.5. For each feature component, a counterfactual sample was constructed that replaced only with that feature, and the change in resource demand was calculated. If the absolute difference exceeded the threshold, the feature was included in the initial feature subset. Of the 12 feature components tested in the domestic platform environment, 7 exceeded the threshold and were selected for the initial feature subset. For these 7 feature components, the causal strength coefficient was calculated by a weighted sum of the absolute difference and the magnitude of change, where the absolute difference had a weight of 0.7 and the magnitude of change had a weight of 0.3. The calculated causal strength coefficients for the 7 features were 0.92, 0.83, 0.75, 0.68, 0.54, 0.47, and 0.39, respectively, reflecting the degree of influence of each feature on resource demand.

[0040] The causal strength coefficients are sorted in descending order and summed sequentially to obtain a cumulative sum sequence. The seven causal strength coefficients are arranged in descending order, resulting in the sequence 0.92, 0.83, 0.75, 0.68, 0.54, 0.47, 0.39. These coefficients are then summed sequentially to obtain the cumulative sum sequence 0.92, 1.75, 2.50, 3.18, 3.72, 4.19, 4.58. The second-order difference is performed on the growth rates between adjacent elements in the cumulative sum sequence to obtain the second-order difference values. The growth rate is the ratio of adjacent elements minus 1, resulting in the growth rate sequence 0.902, 0.429, 0.272, 0.170, 0.126, 0.093. The first-order difference is calculated on the growth rate sequence to obtain -0.473, -0.157, -0.102, -0.044, -0.033. The second-order difference values ​​0.316, 0.055, 0.058, and 0.011 are obtained by calculating the difference between the first-order difference sequences. The positions corresponding to the second-order difference values ​​that exceed the preset mutation threshold are taken as inflection points. The preset mutation threshold is 0.05. As can be seen from the second-order difference sequences, the first value of 0.316 is much larger than the threshold, therefore the first position is the inflection point.

[0041] Features preceding the inflection point are selected as the dominant feature set. Based on the inflection point, the dominant feature set contains the two features with the highest causal strength, namely Feature 1 and Feature 2, with causal strength coefficients of 0.92 and 0.83, respectively. Features in the dominant feature set are paired to obtain feature pairs. Each dominant feature set contains two features, forming a feature pair: (Feature 1, Feature 2). The difference between the joint probability distribution and the marginal probability distribution of each feature pair is calculated to obtain the mutual information. The joint probability distribution and marginal probability distribution are calculated using histogram estimation, dividing the feature value range into 10 equal intervals. The frequency of the sample in each interval combination is statistically analyzed to obtain the joint distribution. In historical data samples collected from domestic platforms, the mutual information of feature pairs is 0.37. Key driving features are determined based on mutual information. When the mutual information is low, it indicates that the two features are relatively independent and both are key driving features; when the mutual information is high, it indicates information redundancy between features, and the feature with higher causal strength is selected as the key driving feature. Judging from the current feature mutual information of 0.37, which is lower than the preset threshold of 0.5, it means that the two features are relatively independent and both are key driving features.

[0042] In this embodiment, by constructing counterfactual samples and comparing them with the original features, the true impact of changes in the values ​​of individual feature components on resource demand can be assessed while maintaining consistency with the overall operating context. Combining historical statistical quantiles with numerical replacement makes the counterfactual reasoning results statistically interpretable and realistically attainable, improving the credibility of resource demand change assessments. By accumulating and analyzing causal strength coefficients and introducing a second-order difference inflection point determination mechanism, the critical position where causal contribution changes from significant to marginal can be automatically identified, effectively avoiding feature redundancy or omission of key information caused by subjective threshold selection. By combining dominant features in pairs and using mutual information to measure the difference between their joint and marginal distributions, key feature pairs that have a synergistic amplification or coupled driving effect on resource demand can be identified, improving the accuracy and efficiency of hardened computer resource allocation decisions.

[0043] In one alternative implementation, Based on the aforementioned key driving characteristics, a correspondence between load characteristics and resource requirements is established through multi-layer nonlinear mapping, and uncertainty quantification is performed to obtain capacity demand prediction results. The confidence intervals corresponding to the capacity demand prediction results are calculated using a Bayesian inference algorithm, including: The key driving features are normalized to zero mean to obtain standardized features. Based on the pre-acquired historical load features and historical resource requirements, a mapping weight matrix is ​​constructed using the gradient descent algorithm. Linear transformation features are calculated based on the standardized features and the mapping weight matrix. Nonlinear activation transformation is performed on the linear transformation features to obtain activation features. A high-dimensional abstract feature representation is determined based on the activation features. The mean and variance parameters of the prediction distribution are solved based on the high-dimensional abstract features. The posterior probability distribution of capacity demand is constructed using variational inference methods. Multiple sets of capacity demand prediction samples are generated from the posterior probability distribution using the Monte Carlo sampling method, and statistical analysis is performed to obtain the capacity demand prediction results and prediction stable values. The target confidence level is initialized based on the preset task requirements. Based on the target confidence level, the quantiles of the posterior probability distribution are solved using a Bayesian inference algorithm to obtain the quantile coefficients. The confidence radius is then calculated by combining the predicted stable value. Based on the capacity demand prediction result and the confidence radius, the confidence interval corresponding to the capacity demand prediction result is obtained by using an interval construction method.

[0044] For the identified set of key driving features, the mean and standard deviation of each feature are calculated and transformed using the zero-mean normalization formula. Taking the processor queue length feature in the hardened computer of the domestic platform as an example, the historical data sample has a mean of 5.4, a standard deviation of 2.6, and a current value of 8.5. After zero-mean normalization, the standardized feature value is 1.19. The memory page swap rate feature has a mean of 9.2 times / second, a standard deviation of 3.8 times / second, and a current value of 14.1 times / second. The standardized feature value is 1.29. Based on the pre-acquired historical load features and historical resource requirements, a mapping weight matrix is ​​constructed using the gradient descent algorithm. The operating data of the hardened computer of the domestic platform over the past 30 days is collected, including standardized historical load feature samples and corresponding actual resource requirement value pairs, with a sample size of 720. The loss function is defined as the mean squared error between the predicted value and the actual resource requirement. The mapping weight matrix is ​​initialized with random small values, the learning rate is set to 0.01, the maximum number of iterations is 1000, and the convergence threshold is 0.0001. The gradient descent algorithm was run on a domestic platform. After 768 iterations, the loss function value dropped to 0.00098, which is below the convergence threshold. The final mapping weight matrix has a dimension of 2×4, where the weight vector corresponding to the processor queue length feature is [0.72, -0.35, 0.58, 0.21], and the weight vector corresponding to the memory page swapping rate feature is [0.65, 0.42, -0.28, 0.53].

[0045] Linear transformation features are calculated based on standardized features and mapping weight matrices. The standardized key driving features are multiplied by the mapping weight matrix to obtain the linear transformation features. On a domestically hardened computer platform, the current standardized features are [1.19, 1.29]. Multiplying these with the mapping weight matrix yields the linear transformation features [1.70, 0.12, 0.33, 0.89]. Nonlinear activation transformations are applied to the linear transformation features to obtain activation features. Using the hyperbolic tangent function as the nonlinear activation function, each element of the linear transformation features is transformed, mapping the features to the interval [-1, 1]. The linear transformation feature [1.70, 0.12, 0.33, 0.89] is transformed by nonlinear activation to obtain the activation feature [0.93, 0.12, 0.32, 0.71]. High-dimensional abstract feature representations are determined based on the activation features. High-dimensional abstract feature representations are constructed through feature combination and nonlinear transformations, extracting the interaction information between activation features. Specifically, this involves calculating the pairwise product of activation features and concatenating it with the original activation features to obtain the high-dimensional abstract feature representation. In this embodiment, the 4-dimensional activation feature is expanded into a 10-dimensional high-dimensional abstract feature with values ​​of [0.93, 0.12, 0.32, 0.71, 0.11, 0.30, 0.66, 0.04, 0.09, 0.23].

[0046] The mean and variance parameters of the predicted distribution are obtained by solving for high-dimensional abstract features. These high-dimensional abstract features are input into two independent linear mapping layers: one for calculating the mean parameter and the other for calculating the variance parameter. The weight vector of the mean parameter mapping layer is [0.58, 0.23, -0.15, 0.42, 0.11, -0.08, 0.19, 0.05, -0.03, 0.12], with a bias term of 0.35; the weight vector of the variance parameter mapping layer is [0.12, 0.08, 0.05, 0.10, 0.02, 0.03, 0.07, 0.01, 0.01, 0.04], with a bias term of 0.02. The calculated mean parameter of the predicted distribution is 3.92, and the variance parameter is 0.18. A posterior probability distribution of capacity demand is then constructed using variational inference methods. Based on the obtained mean and variance parameters, a normal distribution was constructed as the posterior probability distribution of capacity demand, i.e., the capacity demand follows a normal distribution with a mean of 3.92 and a standard deviation of 0.42. Multiple sets of capacity demand prediction samples were generated from the posterior probability distribution using the Monte Carlo sampling method. The sampling was set to 1000 times, randomly sampling from the normal distribution with a mean of 3.92 and a standard deviation of 0.42, resulting in 1000 capacity demand prediction samples. The numerical range of the sampling results was [2.89, 5.10], concentrated around the mean. Statistical analysis was performed to obtain the capacity demand prediction results and the prediction stable value. The mean, median, and mode were calculated for the 1000 sampling results, yielding a capacity demand prediction result of 3.92, consistent with the distribution mean. The coefficient of variation (the ratio of the standard deviation to the mean) was calculated to be 0.107, lower than the preset threshold of 0.15, indicating that the prediction result is stable and reliable. The median of 3.90 was taken as the prediction stable value for subsequent interval construction.

[0047] The target confidence level is initialized based on preset task requirements. Considering the criticality of the task and the available resource margins for hardening the domestic platform's computing power, a target confidence level of 95% is set, indicating that the capacity demand prediction result has a 95% probability of falling within the constructed confidence interval. Based on the target confidence level, a Bayesian inference algorithm is used to solve for the quantiles of the posterior probability distribution. For a 95% target confidence level, the 2.5% and 97.5% quantiles of the normal distribution are calculated, yielding the corresponding standard normal distribution quantile coefficients of -1.96 and 1.96. The quantile coefficients are then used in conjunction with the predicted stable value to calculate the confidence radius. The confidence radius is obtained by multiplying the standard normal distribution quantile coefficients by the standard deviation of the posterior distribution. In this embodiment, the standard deviation is 0.42, the quantile coefficient is 1.96, and the calculated confidence radius is 0.82. Based on the capacity demand prediction result and the confidence radius, the confidence interval corresponding to the capacity demand prediction result is obtained using an interval construction method. Adding or subtracting the predicted stable value of 3.90 from the confidence radius of 0.82 yields the upper and lower bounds of the confidence interval, namely [3.08, 4.72], indicating that the capacity demand forecast is between 3.08 and 4.72 at a 95% confidence level.

[0048] In a domestically developed platform-hardened computer system, capacity demand forecasting results are applied to resource elastic scaling decisions. The current system CPU configuration is 4 cores. Based on the forecast result of 3.92 and the confidence interval [3.08, 4.72], the current configuration basically meets the demand, but is close to the upper limit. Considering the risk of business fluctuations, the system reserves 15% resource margin, resulting in a recommended configuration of 3.92 × 1.15 = 4.51 cores, exceeding the current configuration. Further analysis shows the upper limit of the confidence interval is 4.72; increasing the margin by 15% yields 5.43 cores. Based on the aforementioned analysis, the resource scheduling module decides to expand the CPU configuration from 4 cores to 6 cores to meet potential future load growth. In actual deployment, a forecast is executed every 10 minutes, and resource adjustments are triggered based on the consistency of three consecutive forecasts to avoid frequent scaling up and down due to short-term fluctuations. The reserved margin ratio is dynamically adjusted according to the changing trend of the forecast variance parameter; the margin is increased when the variance increases to cope with uncertainty.

[0049] In this embodiment, by standardizing key driving features and constructing mapping weights based on historical load and resource requirements, features participate in modeling at a unified scale, reducing the interference of different dimensions and fluctuation amplitudes on the prediction results and improving the generalization ability to complex load changes. Variational inference is introduced to construct the posterior probability distribution of capacity requirements, and multiple sets of prediction samples are obtained through Monte Carlo sampling, expanding the prediction results from a single numerical value to a distribution form. This avoids the problem of large deviations in capacity estimation caused by occasional load fluctuations. Combined with the target confidence level, Bayesian quantile inference is performed and confidence intervals are constructed, so that the capacity requirement prediction results have clear confidence boundaries. This can quantify the prediction uncertainty and directly map it into a safe interval that can be used for resource allocation, providing a more scientific, robust, and interpretable decision basis for capacity planning and elastic scaling of hardened computers.

[0050] In one alternative implementation, Based on the capacity demand forecast results and the physical constraint parameters corresponding to the ruggedized computer, an expansion decision instruction is determined. The uncertainty metric of the expansion decision instruction is calculated based on the confidence interval, including: Obtain the resource expansion granularity and resource response latency corresponding to the hardened computer, extract the capacity demand point estimate from the capacity demand prediction result and round up according to the resource expansion granularity to obtain the granularity-aligned resource amount, determine whether the granularity-aligned resource amount exceeds the preset maximum scalable resource capacity, if it exceeds, then use the maximum scalable resource capacity as the baseline expansion amount, otherwise use the granularity-aligned resource amount as the baseline expansion amount. Extract the upper and lower bounds of the confidence interval corresponding to the capacity demand forecast result. Calculate the absolute values ​​of the differences between the upper and lower bounds of the confidence interval and the estimated capacity demand point to obtain the upward and downward fluctuation amplitudes. Determine the maximum fluctuation amplitude based on the upward and downward fluctuation amplitudes and round up according to the resource expansion granularity to obtain the uncertainty buffer. Solve for the baseline expansion and the uncertainty buffer to obtain the safety expansion. Generate an expansion decision instruction based on the safety expansion and the resource response delay. Map the uncertainty buffer and the baseline expansion to fuzzy membership features using a fuzzy inference algorithm and perform fuzzy rule inference to obtain a fuzzy risk assessment value. Defuzzify the value to obtain a quantified risk value and solve for the uncertainty metric corresponding to the expansion decision instruction.

[0051] Obtain the resource expansion granularity and resource response latency corresponding to the hardened computer. For hardened computers on domestic platforms, obtain the resource expansion parameters by querying the system configuration parameter table. In this example, the CPU resource expansion granularity is 2 cores, the memory resource expansion granularity is 4GB, and the resource response latency is 90 seconds. Resource expansion granularity represents the smallest unit of resource adjustment, and resource response latency represents the time interval from issuing a resource adjustment command to the actual availability of the resource. Extract the capacity demand point estimate from the capacity demand forecast results and round it up according to the resource expansion granularity to obtain the granularity-aligned resource amount. According to the pre-obtained capacity demand forecast results, the capacity demand point estimate is 3.92 cores. Round up according to the CPU resource expansion granularity of 2 cores. The calculation method is to divide the point estimate by the expansion granularity, take the ceiling value, and then multiply it by the expansion granularity. For 3.92 cores, the calculated granularity-aligned resource amount is 4 cores. Determine whether the granularity-aligned resource amount exceeds the preset maximum expandable resource capacity. The preset maximum expandable resource capacity of the CPU for hardened computers on domestic platforms is 16 cores. Comparing the granularity-aligned resource quantity of 4 cores with the maximum scalable resource capacity of 16 cores, 4 cores is less than 16 cores, so the result is that it does not exceed the maximum scalable resource capacity. Therefore, the granularity-aligned resource quantity of 4 cores is taken as the baseline expansion quantity.

[0052] Extract the upper and lower bounds of the confidence interval corresponding to the capacity demand forecast results. In the previous example, the 95% confidence interval obtained is [3.08, 4.72], with an upper bound of 4.72 and a lower bound of 3.08. Calculate the absolute values ​​of the differences between the upper and lower bounds of the confidence intervals and the estimated capacity demand point to obtain the upward and downward volatility. The estimated capacity demand point is 3.92; the calculated upward volatility is |4.72 - 3.92| = 0.80, and the downward volatility is |3.08 - 3.92| = 0.84. Determine the maximum volatility based on the upward and downward volatility. Compare the upward volatility of 0.80 and the downward volatility of 0.84, and take the larger value of 0.84 as the maximum volatility. Round up according to the resource expansion granularity to obtain the uncertainty buffer. Divide the maximum fluctuation range of 0.84 by the resource expansion granularity of 2 cores, take the ceiling value, and then multiply by the expansion granularity to obtain an uncertainty buffer of 2 cores. Summing the baseline expansion and the uncertainty buffer yields the safe expansion. With a baseline expansion of 4 cores and an uncertainty buffer of 2 cores, the safe expansion is calculated as 4 + 2 = 6 cores. Based on the safe expansion and resource response latency, an expansion decision instruction is generated. Based on the safe expansion of 6 cores and a resource response latency of 90 seconds, an expansion decision instruction is generated, including parameters such as resource type (CPU), target resource quantity (6 cores), response latency (90 seconds), and execution priority (high).

[0053] The uncertainty buffer and baseline expansion are mapped to fuzzy membership features using a fuzzy inference algorithm. First, the uncertainty ratio is calculated, which is the uncertainty buffer divided by the baseline expansion, resulting in 0.5. This uncertainty ratio of 0.5 is mapped to three fuzzy sets: low uncertainty (membership 0.0), medium uncertainty (membership 0.8), and high uncertainty (membership 0.2). Simultaneously, the ratio of the baseline expansion of 4 cores to the maximum scalable resource capacity of 16 cores (0.25) is mapped to three fuzzy sets: low resource usage (membership 0.6), medium resource usage (membership 0.4), and high resource usage (membership 0.0). Fuzzy rule inference is then performed to obtain a fuzzy risk assessment value. The fuzzy rule base contains 9 rules, covering various combinations of uncertainty and resource usage. For the combination of medium uncertainty (0.8), low resource usage (0.6), and medium resource usage (0.4) in this embodiment, two rules are triggered: Rule 1, "If uncertainty is medium and resource usage is low, then the risk is low"; and Rule 2, "If uncertainty is medium and resource usage is medium, then the risk is medium." The activation strength of each rule is determined by the minimum value operator. The activation strength of rule 1 is min(0.8, 0.6) = 0.6, and the activation strength of rule 2 is min(0.8, 0.4) = 0.4. The fuzzy risk set is obtained as follows: low risk (membership degree 0.6), medium risk (membership degree 0.4), and high risk (membership degree 0.0).

[0054] Defuzzification is performed to obtain quantified risk values. The centroid method is used for defuzzification, converting the fuzzy risk set into precise quantified risk values. The centroid value for low risk is 0.2, for medium risk it is 0.5, and for high risk it is 0.8. The weighted average is calculated as: (0.6×0.2+0.4×0.5+0.0×0.8)÷(0.6+0.4+0.0)=0.32. The quantified risk value of 0.32 represents the risk level of the current extended decision. The uncertainty measure corresponding to the extended decision instruction is then calculated. Based on the quantified risk value of 0.32 and the uncertainty ratio of 0.5, the uncertainty measure is calculated. The uncertainty measure is defined as the weighted geometric mean of the quantified risk value and the uncertainty ratio, with weights of 0.7 and 0.3, respectively. The calculated uncertainty measure is 0.32^0.7×0.5^0.3=0.37. The uncertainty measure of 0.37 reflects the degree of uncertainty in the extended decision and is used for subsequent decision evaluation and optimization.

[0055] In this embodiment, by introducing resource expansion granularity and maximum scalable capacity constraints, the estimated capacity demand point is aligned with granularity and its upper limit is verified. This allows the generated baseline expansion quantity to be directly mapped to the actual executable resource expansion operation, improving the feasibility and execution stability of expansion decisions in a hardened computing environment. By combining the confidence interval corresponding to the capacity demand prediction results, the prediction uncertainty is quantitatively analyzed, and the maximum fluctuation amplitude is converted into an uncertainty buffer. This ensures that expansion decisions not only consider average or expected demand but also cover potential uplink load risks. By constructing a safe expansion quantity together with the baseline expansion quantity and the uncertainty buffer, and combining it with resource response latency to generate expansion decision instructions, the expansion behavior can simultaneously meet real-time and security requirements in both time and scale dimensions, improving the foresight and stability in responding to sudden load changes.

[0056] Figure 2 This is a flowchart illustrating the resource expansion decision and uncertainty calculation process of the adaptive elastic scaling method for hardened computers based on a domestic platform, as described in an embodiment of the present invention.

[0057] In one alternative implementation, Determining the execution strategy for the extended decision instruction based on the uncertainty measure, and executing the extended decision instruction according to the execution strategy until the extension is completed, includes: A preset risk threshold is obtained, and it is determined whether the uncertainty measure exceeds the risk threshold. If it exceeds the threshold, a preset conservative expansion strategy is adopted. The total amount of expansion resources is determined based on the upper bound of the pre-obtained confidence interval, and the total amount of expansion resources is decomposed into multiple expansion batches. If it does not exceed the threshold, a standard expansion strategy is adopted. The total amount of expansion resources is determined based on the capacity demand prediction result, and resources are configured within a single expansion batch. An execution strategy containing resource allocation and expansion time arrangement is generated based on the conservative expansion strategy or the standard expansion strategy. According to the extension time schedule in the execution strategy, the resource extension operation is performed on the hardened computer in sequence. After each extension batch is completed, the current running status parameters of the hardened computer are collected and the resource utilization rate is calculated. Based on the resource utilization rate and the capacity demand prediction result, it is determined whether the capacity demand has been met. If it is met, the execution of subsequent extension batches is terminated. If the capacity demand is not met, the extension continues until all extension batches in the execution strategy are completed.

[0058] A preset risk threshold is obtained. Based on the importance of the business and the system stability requirements, the domestic platform hardened computer sets the risk threshold to 0.45. A lower risk threshold indicates a lower tolerance for the risks of expansion decisions and a more conservative expansion strategy; a higher risk threshold indicates that the system allows for a more aggressive expansion strategy. The uncertainty metric is then checked to see if it exceeds the risk threshold. In the previous example, the calculated uncertainty metric was 0.37. Compared to the risk threshold of 0.45, 0.37 is less than 0.45, indicating that the uncertainty metric does not exceed the risk threshold. Since the uncertainty metric does not exceed the risk threshold, a standard expansion strategy is adopted. Under the standard expansion strategy, the total amount of expansion resources is determined based on the capacity demand forecast, and resources are configured within a single expansion batch. The capacity demand forecast indicates a safe expansion of 6 cores. The current system configuration is 4 cores, so the total amount of resources to be expanded is 6-4=2 cores. According to the standard expansion strategy, the 2 core resources are configured within a single expansion batch, that is, the system configuration is expanded from 4 cores to 6 cores in one go.

[0059] Comparative analysis shows that if the uncertainty metric exceeds the risk threshold, a conservative scaling strategy should be adopted. Under this strategy, the total amount of resources to be scaled is determined based on the pre-obtained upper bound of the confidence interval, and this total is then decomposed into multiple scaling batches. The upper bound of the confidence interval is 4.72 cores, which, after rounding up to the nearest integer (6 cores) for resource scaling, results in 6 cores. Since the current system configuration is 4 cores, the total amount of resources to be scaled is 6 - 4 = 2 cores. According to the conservative scaling strategy, these 2 cores are decomposed into multiple scaling batches, for example, 2 batches, with each batch scaling 1 core. An execution strategy, including resource allocation and scaling time scheduling, is generated based on the standard scaling strategy. For the standard scaling strategy, the execution strategy includes: resource type is CPU, scaling quantity is 2 cores, scaling batch is 1 time, scaling time is 30 seconds after the current time point, and the estimated completion time is 120 seconds after the current time point (considering a 90-second resource response latency).

[0060] Resource expansion operations are performed sequentially on the hardened computers according to the expansion schedule in the execution strategy. A resource expansion request is initiated 30 seconds after the current time point, expanding the CPU configuration from 4 cores to 6 cores. After receiving the expansion request, the resource management module requests an additional 2 CPU cores from the underlying resource pool and allocates them to the target hardened computer. After the expansion batch is completed, the current operating status parameters of the hardened computer are collected and resource utilization is calculated. Key operating status parameters such as CPU utilization, memory utilization, and I / O wait time are obtained through the system monitoring interface. Taking CPU resources as an example, the CPU utilization after expansion is 72%, calculated by dividing the actual CPU load by the total number of configured cores, i.e., 4.32 ÷ 6 = 72%. The capacity requirement is determined based on the resource utilization and capacity demand prediction results. The capacity demand prediction result is 3.92 cores, the current system configuration is 6 cores, and the resource utilization is 72%. By comparing two conditions: first, whether the current configuration is greater than or equal to the predicted demand (6 > 3.92, satisfied); second, whether the resource utilization is within a reasonable range (72% < 85%, satisfied), the comprehensive judgment result is that the capacity requirement is satisfied.

[0061] In real-world scenarios, the execution flow differs depending on the conservative scaling strategy. Under this strategy, scaling occurs in two batches, scaling one core at a time. The first batch scales the system from 4 cores to 5 cores, taking 30 seconds after the current time point, with an estimated completion time of 120 seconds. The second batch scales the system from 5 cores to 6 cores, taking 5 minutes after the first batch completes, with an estimated completion time of 5 minutes and 90 seconds. After the first batch, the system is configured with 5 cores. At this point, running status parameters are collected, and the CPU utilization is 86% (calculated as 4.32 ÷ 5 = 86.4%). The system is assessed for capacity requirements: the current configuration of 5 cores exceeds the predicted requirement of 3.92 cores, but the resource utilization of 86% exceeds the reasonable upper limit of 85%. Therefore, the capacity requirement is not met. The second batch of scaling is then executed, scaling the system from 5 cores to 6 cores. After the second batch, running status parameters are collected again, and the CPU utilization has dropped to 72%, indicating that the capacity requirement has been met, and the execution of subsequent scaling batches is terminated.

[0062] For example, when this method is deployed on a domestically-made platform hardened computer, the system dynamically adjusts the risk threshold according to different workload characteristics. For instance, when handling critical tasks, the risk threshold is automatically lowered to 0.35, adopting a more conservative scaling strategy; when handling routine tasks, the risk threshold is raised to 0.55, adopting a more aggressive scaling strategy to improve resource utilization efficiency. Simultaneously, the risk threshold setting is continuously optimized based on historical scaling effects. When five consecutive scaling decisions accurately meet the requirements without resource waste, the risk threshold is increased by 0.05; when resource shortages affect business performance, the risk threshold is lowered by 0.1, gradually forming the optimal risk control strategy. For special scenarios, such as anticipated sudden high loads within a short period, a rapid response resource pool is reserved. By preheating resources 10 minutes in advance, the response latency is reduced from the standard 90 seconds to 30 seconds.

[0063] In a critical business processing scenario, the initial configuration of the hardened computer on the domestic platform was a 4-core CPU, with a predicted capacity requirement of 3.92 cores and a confidence interval of [3.08, 4.72]. The uncertainty metric was 0.37, below the risk threshold of 0.45. A standard scaling strategy was adopted, scaling the CPU configuration from 4 cores to 6 cores in one go. After scaling, the CPU utilization was 72%, within the ideal range (45%~85%). In the following 20 minutes, the business load fluctuated, with the highest CPU utilization reaching 81%, still within a reasonable range, verifying the accuracy of the scaling decision. The entire scaling process took 120 seconds, including 30 seconds for resource request and configuration and 90 seconds for resource response latency. Each scaling decision and its effect were recorded to form a decision knowledge base for rapid decision-making reference in similar subsequent scenarios.

[0064] In this embodiment, by comparing uncertainty metrics with preset risk thresholds, high-risk and normal-risk scenarios can be automatically distinguished before expansion execution. Conservative or standard expansion strategies are selected for different risk levels, ensuring that the total amount of expansion resources and the expansion pace match the current uncertainty level. This effectively reduces the risk of excessive resource investment or amplified system fluctuations when prediction uncertainty is high, and avoids unnecessary expansion delays in low-risk scenarios. By determining the total amount of expansion resources based on the upper bound of the confidence interval and performing multi-batch decomposition, the resource expansion process can be advanced in stages while ensuring the coverage of potential peak demand, reducing the impact of a single large-scale expansion on system stability. By collecting the running status and calculating resource utilization in real time after each expansion batch is completed, the expansion effect is evaluated online, achieving closed-loop control of the expansion process. Subsequent expansions can be terminated in a timely manner after the demand is met, avoiding resource redundancy and unnecessary expansion operations.

[0065] A second aspect of this invention provides a ruggedized computer adaptive elastic expansion system based on a domestically developed platform, comprising: The data acquisition unit is used to collect the operating status parameters of each computing unit in the hardened computer, construct a multidimensional state tensor based on the operating status parameters and extract the temporal evolution vector through tensor decomposition, determine the dependency relationship between computing units based on the temporal evolution vector, and obtain the load feature set. The feature analysis unit is used to perform multi-scale time-series decomposition on the load feature set to obtain a composite load feature vector, analyze the composite load feature vector through a causal inference algorithm to obtain causal analysis results, and identify the feature subset that has a causal impact on resource demand by combining counterfactual reasoning, calculate the causal strength coefficient of each feature in the feature subset, and determine the key driving features based on the causal strength coefficient. The demand forecasting unit is used to establish the correspondence between load characteristics and resource demand through multi-layer nonlinear mapping based on the key driving characteristics and to perform uncertainty quantification assessment to obtain the capacity demand forecasting result. The confidence interval corresponding to the capacity demand forecasting result is calculated through Bayesian inference algorithm. The decision execution unit is configured to determine an expansion decision instruction based on the capacity demand forecast result and the physical constraint parameters corresponding to the ruggedized computer, calculate the uncertainty metric of the expansion decision instruction based on the confidence interval, determine the execution strategy of the expansion decision instruction based on the uncertainty metric, and execute the expansion decision instruction according to the execution strategy until the expansion is completed.

[0066] A third aspect of the present invention provides an electronic device, comprising: A processor and a memory for storing processor-executable instructions, wherein the processor is configured to invoke instructions stored in the memory to perform the aforementioned method.

[0067] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0068] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for adaptive elastic scaling of ruggedized computers based on a domestically developed platform, characterized in that: include: The system collects the operating status parameters of each computing unit in the hardened computer, constructs a multidimensional state tensor based on the operating status parameters, extracts the temporal evolution vector through tensor decomposition, determines the dependency relationship between computing units based on the temporal evolution vector, and obtains the load feature set. Multi-scale temporal decomposition of the load feature set yields a composite load feature vector. The composite load feature vector is analyzed using a causal inference algorithm to obtain causal analysis results. Counterfactual reasoning is then used to identify a subset of features that have a causal impact on resource demand. The causal strength coefficient of each feature in the subset is calculated, and key driving features are determined based on the causal strength coefficient. Based on the key driving characteristics, a correspondence between load characteristics and resource requirements is established through multi-layer nonlinear mapping, and uncertainty quantification is performed to obtain capacity demand prediction results. The confidence interval corresponding to the capacity demand prediction results is calculated through Bayesian inference algorithm. Based on the capacity demand forecast results and the physical constraint parameters corresponding to the ruggedized computer, an expansion decision instruction is determined. The uncertainty metric of the expansion decision instruction is calculated based on the confidence interval, and the execution strategy of the expansion decision instruction is determined based on the uncertainty metric. The expansion decision instruction is executed according to the execution strategy until the expansion is completed.

2. The method according to claim 1, characterized in that, The system collects operational status parameters of each computing unit in the hardened computer, constructs a multidimensional state tensor based on these parameters, and extracts a temporal evolution vector through tensor decomposition. Based on this temporal evolution vector, it determines the dependencies between computing units, resulting in a load feature set including: The operating status parameters of each computing unit in the hardened computer are obtained within a preset monitoring period, and the operating status parameters are tensorized and organized according to a preset dimension to obtain a multidimensional state tensor. The multidimensional state tensor is modally decomposed to obtain a time modal factor matrix and the evolution trajectory of each factor component with time dimension is extracted. The evolution trajectory is reconstructed in phase space to obtain the trajectory attractor corresponding to each factor component. The evolution stability index of each factor component is calculated based on the topological invariants of the trajectory attractor. The stable evolution factor component is determined based on the evolution stability index and a preset stability threshold and the temporal evolution vector is reconstructed. The delay embedding of each component in the temporal evolution vector is expanded in the time dimension to obtain a delay embedding matrix. The mutual information of the delay embedding matrix is ​​calculated to obtain the information transmission strength between computing units and a directed association graph is constructed. In the directed association graph, the direct and indirect dependency paths between computing units are identified by transitive closure operation and the key dependency chain is determined by path weight quantization. The path topology metric of each computing unit in the key dependency chain is used as the structural feature of the dependency relationship. The amplitude change rate and phase difference of each component in the temporal evolution vector are used as the dynamic feature of the dependency relationship. The load feature set is obtained by fusing the structural feature and the dynamic feature.

3. The method according to claim 1, characterized in that, Multi-scale time-series decomposition of the load feature set yields a composite load feature vector. Causal inference algorithms are then used to analyze this composite load feature vector, resulting in the following causal analysis: The evolution data of load features in the load feature set in the time dimension are segmented and trend-fitted according to a preset time window length to obtain trend components. The difference between the evolution data and the trend components is calculated to obtain residual data. Periodic analysis and frequency domain transformation are performed on the residual data to obtain periodic oscillation modes. The time derivative of the trend components is calculated to obtain the long-term evolution direction. The long-term evolution direction and the periodic oscillation mode are spliced ​​together under different time window lengths and superimposed layer by layer according to the time window length to obtain a composite load feature vector. The feature components in the composite load feature vector are used as nodes to construct a conditional independence test matrix, and the independence of each node is statistically tested. Based on the test results, the nodes with conditional dependencies are retained to obtain an undirected association structure. Collision structures are identified in the undirected association structure, and the causal direction of the edges is labeled according to the prior temporal relationship of the incoming edges in the collision structure to obtain a directed causal graph. In the directed causal graph, the node corresponding to each feature component is selected as the intervention node in turn, and the incoming edges of the intervention node are blocked. The numerical distribution distance of the resource demand corresponding to each node before and after the blocking is calculated and used as the causal effect strength. The causal analysis results are obtained by combining each node and its corresponding node identifier and causal effect strength.

4. The method according to claim 1, characterized in that, Combining counterfactual reasoning to identify a subset of features that have a causal impact on resource demand, calculating the causal strength coefficient of each feature in the subset, and determining key driving features based on the causal strength coefficient includes: The feature components in the causal analysis results are used as the original feature components, and counterfactual samples are constructed by replacing the values ​​with pre-acquired historical statistical quantile values. The counterfactual samples and the original feature components are weighted and aggregated to obtain counterfactual aggregated feature vectors and original aggregated feature vectors. The counterfactual aggregated feature vectors and the original aggregated feature vectors are queried and matched with the preset resource demand mapping relationship to obtain counterfactual resource demand values ​​and original resource demand values. The absolute difference and numerical change range between the counterfactual resource demand values ​​and the original resource demand values ​​are calculated. An initial feature subset is constructed based on the feature components whose absolute difference exceeds a preset difference threshold, and the causal strength coefficient is calculated by combining the numerical change range. The causal intensity coefficients are sorted in descending order and accumulated one by one to obtain a cumulative sum sequence. The growth rate between adjacent elements in the cumulative sum sequence is subjected to second-order difference to obtain second-order difference values. The position corresponding to the second-order difference value that exceeds the preset mutation threshold is taken as the inflection point position. The features before the inflection point position are taken as the dominant feature set. The features in the dominant feature set are combined in pairs to obtain feature pairs. The difference between the joint probability distribution and the marginal probability distribution of each feature pair is calculated to obtain the mutual information. The key driving features are determined based on the mutual information.

5. The method according to claim 1, characterized in that, Based on the aforementioned key driving characteristics, a correspondence between load characteristics and resource requirements is established through multi-layer nonlinear mapping, and uncertainty quantification is performed to obtain capacity demand prediction results. The confidence intervals corresponding to the capacity demand prediction results are calculated using a Bayesian inference algorithm, including: The key driving features are normalized to zero mean to obtain standardized features. Based on the pre-acquired historical load features and historical resource requirements, a mapping weight matrix is ​​constructed using the gradient descent algorithm. Linear transformation features are calculated based on the standardized features and the mapping weight matrix. Nonlinear activation transformation is performed on the linear transformation features to obtain activation features. A high-dimensional abstract feature representation is determined based on the activation features. The mean and variance parameters of the prediction distribution are solved based on the high-dimensional abstract features. The posterior probability distribution of capacity demand is constructed using variational inference methods. Multiple sets of capacity demand prediction samples are generated from the posterior probability distribution using the Monte Carlo sampling method, and statistical analysis is performed to obtain the capacity demand prediction results and prediction stable values. The target confidence level is initialized based on the preset task requirements. Based on the target confidence level, the quantiles of the posterior probability distribution are solved using a Bayesian inference algorithm to obtain the quantile coefficients. The confidence radius is then calculated by combining the predicted stable value. Based on the capacity demand prediction result and the confidence radius, the confidence interval corresponding to the capacity demand prediction result is obtained by using an interval construction method.

6. The method according to claim 1, characterized in that, Based on the capacity demand forecast results and the physical constraint parameters corresponding to the ruggedized computer, an expansion decision instruction is determined. The uncertainty metric of the expansion decision instruction is calculated based on the confidence interval, including: Obtain the resource expansion granularity and resource response latency corresponding to the hardened computer, extract the capacity demand point estimate from the capacity demand prediction result and round up according to the resource expansion granularity to obtain the granularity-aligned resource amount, determine whether the granularity-aligned resource amount exceeds the preset maximum scalable resource capacity, if it exceeds, then use the maximum scalable resource capacity as the baseline expansion amount, otherwise use the granularity-aligned resource amount as the baseline expansion amount. Extract the upper and lower bounds of the confidence interval corresponding to the capacity demand forecast result. Calculate the absolute values ​​of the differences between the upper and lower bounds of the confidence interval and the estimated capacity demand point to obtain the upward and downward fluctuation amplitudes. Determine the maximum fluctuation amplitude based on the upward and downward fluctuation amplitudes and round up according to the resource expansion granularity to obtain the uncertainty buffer. Solve for the baseline expansion and the uncertainty buffer to obtain the safety expansion. Generate an expansion decision instruction based on the safety expansion and the resource response delay. Map the uncertainty buffer and the baseline expansion to fuzzy membership features using a fuzzy inference algorithm and perform fuzzy rule inference to obtain a fuzzy risk assessment value. Defuzzify the value to obtain a quantified risk value and solve for the uncertainty metric corresponding to the expansion decision instruction.

7. The method according to claim 1, characterized in that, Determining the execution strategy for the extended decision instruction based on the uncertainty measure, and executing the extended decision instruction according to the execution strategy until the extension is completed, includes: A preset risk threshold is obtained, and it is determined whether the uncertainty measure exceeds the risk threshold. If it exceeds the threshold, a preset conservative expansion strategy is adopted. The total amount of expansion resources is determined based on the upper bound of the pre-obtained confidence interval, and the total amount of expansion resources is decomposed into multiple expansion batches. If it does not exceed the threshold, a standard expansion strategy is adopted. The total amount of expansion resources is determined based on the capacity demand prediction result, and resources are configured within a single expansion batch. An execution strategy containing resource allocation and expansion time arrangement is generated based on the conservative expansion strategy or the standard expansion strategy. According to the extension time schedule in the execution strategy, the resource extension operation is performed on the hardened computer in sequence. After each extension batch is completed, the current running status parameters of the hardened computer are collected and the resource utilization rate is calculated. Based on the resource utilization rate and the capacity demand prediction result, it is determined whether the capacity demand has been met. If it is met, the execution of subsequent extension batches is terminated. If the capacity demand is not met, the extension continues until all extension batches in the execution strategy are completed.

8. A ruggedized computer adaptive elastic expansion system based on a domestically developed platform, used to implement the method of any one of claims 1-7, characterized in that, include: The data acquisition unit is used to collect the operating status parameters of each computing unit in the hardened computer, construct a multidimensional state tensor based on the operating status parameters and extract the temporal evolution vector through tensor decomposition, determine the dependency relationship between computing units based on the temporal evolution vector, and obtain the load feature set. The feature analysis unit is used to perform multi-scale time-series decomposition on the load feature set to obtain a composite load feature vector, analyze the composite load feature vector through a causal inference algorithm to obtain causal analysis results, and identify the feature subset that has a causal impact on resource demand by combining counterfactual reasoning, calculate the causal strength coefficient of each feature in the feature subset, and determine the key driving features based on the causal strength coefficient. The demand forecasting unit is used to establish the correspondence between load characteristics and resource demand through multi-layer nonlinear mapping based on the key driving characteristics and to perform uncertainty quantification assessment to obtain the capacity demand forecasting result. The confidence interval corresponding to the capacity demand forecasting result is calculated through Bayesian inference algorithm. The decision execution unit is configured to determine an expansion decision instruction based on the capacity demand forecast result and the physical constraint parameters corresponding to the ruggedized computer, calculate the uncertainty metric of the expansion decision instruction based on the confidence interval, determine the execution strategy of the expansion decision instruction based on the uncertainty metric, and execute the expansion decision instruction according to the execution strategy until the expansion is completed.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.