IT resource management system based on automatic operation and maintenance

Through deep learning technology, the monitoring data is cleaned and structured mapped, and the IT resource usage status encoding vector time series is generated. Combined with dynamic semantic propagation representation and elastic scaling decision engine, the problem of imbalance between resource utilization and business requirements in the traditional IT resource management model in the cloud-native architecture is solved, and more flexible and intelligent IT resource operation and maintenance management is achieved.

CN120353675AInactive Publication Date: 2025-07-22STATE GRID HENAN INFORMATION & TELECOMM CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510424990.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional IT resource management model is difficult to capture the dynamic coupling relationship between heterogeneous indicators in cloud-native architecture and microservice applications, and cannot effectively explore the time-varying patterns and potential trends of resource usage status, resulting in an imbalance in the dynamic matching of resource utilization rate and business needs, and problems such as expansion delay or resource redundancy.

Method used

The monitoring data is cleaned and structured mapped using deep learning-based data processing technology, and the IT resources are generated using state-coded vector time series, and capacity prediction is performed through dynamic semantic propagation representations of feature independence constraints, and combined with the elastic scaling decision engine to generate scaling decisions.

Benefits of technology

It realizes more flexible adaptation to nonlinear resource requirements in complex load scenarios, reduces expansion delays and resource redundancy, and improves the automation and intelligence level of IT resource operation and maintenance management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353675A_ABST
    Figure CN120353675A_ABST
Patent Text Reader

Abstract

The invention provides an IT resource management system based on automatic operation and maintenance, and relates to the field of intelligent management. The method comprises the following steps of: performing data cleaning and structured mapping on a time sequence data set of monitoring data generated in the running process of an I T infrastructure and an application system by using a data processing technology based on deep learning to obtain an I T resource use state coding vector time sequence; and then decoding based on the dynamic semantic propagation representation of the feature independence constraint of the I T resource use state coding vector time sequence to obtain capacity prediction result data and prediction confidence, and finally inputting the capacity prediction result data and the prediction confidence into an elastic scaling decision engine to obtain a scaling decision. In this way, the generated elastic scaling decision can more flexibly adapt to nonlinear resource requirements in a complex load scene, and then more automatic and intelligent I T resource operation and maintenance management is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent management, and more specifically, to an IT resource management system based on automated operation and maintenance. Background Art

[0002] In the modern IT environment dominated by cloud-native architecture and microservice applications, the infrastructure presents characteristics such as complex distributed topology, transient fluctuations in business traffic, and enhanced resource coupling. The traditional static IT resource management model based on threshold alarms and manual experience decision-making faces severe challenges. Specifically, first, at the level of multi-dimensional resource status perception, discrete single-indicator threshold monitoring is difficult to capture the dynamic coupling relationship between heterogeneous indicators such as CPU, memory, disk IOPS, and network latency, which will lead to blind spots in capacity bottleneck prediction; secondly, traditional statistical analysis methods lack robust feature encoding capabilities for high-noise and non-stationary monitoring time series data, and cannot effectively mine the time-varying patterns and potential trends of resource usage status; in addition, elastic scaling strategies based on fixed rules are difficult to adapt to nonlinear resource requirements under complex load scenarios, and there is a response lag in the iteration of the prediction model with manual intervention. These problems jointly lead to an imbalance in the dynamic matching of resource utilization and business needs, which is specifically manifested in the expansion delay under burst traffic causing service degradation, and the resource redundancy during the business trough period causing cost waste, which seriously restricts the elasticity and operation and maintenance efficiency of distributed systems.

[0003] Therefore, we look forward to an IT resource management system with automated operation and maintenance. Summary of the invention

[0004] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides an IT resource management system based on automated operation and maintenance.

[0005] According to one aspect of the present application, an IT resource management system based on automated operation and maintenance is provided, which includes:

[0006] The monitoring data collection module is used to collect the monitoring data generated during the operation of IT infrastructure and application systems to obtain a time series data set of monitoring data;

[0007] A monitoring data analysis module, used for performing a time series analysis of IT resource usage status on the time series data set of the monitoring data to obtain capacity prediction result data and prediction confidence, wherein the monitoring data analysis module includes: a monitoring data preprocessing unit, used for performing data cleaning and structured mapping on the time series data set of the monitoring data to obtain a time series of IT resource usage status encoding vectors; a monitoring data encoding unit, used for performing dynamic propagation of IT resource usage status characteristics on the time series of the IT resource usage status encoding vectors to obtain the capacity prediction result data and prediction confidence;

[0008] A scaling decision generation module, configured to obtain a scaling decision based on the capacity prediction result data and the prediction confidence.

[0009] Compared with the prior art, the IT resource management system based on automated operation and maintenance provided by this application uses data processing technology based on deep learning to perform data cleaning and structured mapping on the time series dataset of monitoring data generated during the operation of IT infrastructure and application systems to obtain the time series of IT resource usage status encoding vectors. Subsequently, based on the dynamic semantic propagation representation with feature independence constraints of the IT resource usage status encoding vector time series, the capacity prediction result data and the prediction confidence are decoded. Finally, the capacity prediction result data and the prediction confidence are input into the elastic scaling decision engine to obtain a scaling decision. In this way, the generated elastic scaling decision can more flexibly adapt to the non-linear resource requirements in complex load scenarios, thereby realizing more automated and intelligent IT resource operation and maintenance management. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0011] Figure 1 FIG. is a system block diagram of an IT resource management system based on automated operation and maintenance according to an embodiment of this application.

[0012] Figure 2 FIG. is a block diagram of a monitoring data analysis module in an IT resource management system based on automated operation and maintenance according to an embodiment of this application.

[0013] Figure 3 FIG. is a block diagram of a monitoring data preprocessing unit in an IT resource management system based on automated operation and maintenance according to an embodiment of this application.

[0014] Figure 4 FIG. is a block diagram of a monitoring data encoding unit in an IT resource management system based on automated operation and maintenance according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] Next, example embodiments according to this application will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the example embodiments described herein.

[0016] In the current IT field dominated by cloud-native architectures and microservices applications, the infrastructure exhibits characteristics such as a distributed and complex topology, rapidly changing business traffic, and a significant increase in the coupling between resources. The traditional static IT resource management mode based on fixed-threshold alarms and relying on manual experience for decision-making is unable to cope with these changes. Specifically: First, in the process of monitoring the multi-dimensional resource status, the decentralized monitoring method of a single indicator threshold is difficult to accurately reflect the dynamic interaction relationships between different indicators such as CPU usage, memory occupancy, disk IOPS, and network latency, which makes the prediction of capacity bottlenecks insufficient. Second, traditional statistical analysis methods are unable to effectively extract robust features that can describe the evolution pattern and potential trend of resource usage over time when dealing with monitoring time series data with high noise and non-stationary characteristics. Third, the elastic scaling strategy with fixed rules is difficult to handle the non-linear resource demand changes in complex load situations, and due to the need for manual intervention to update the prediction model, the response speed is slow. These problems together result in a mismatch between resource utilization and business requirements, manifested as a decline in service quality caused by slow expansion during burst traffic and cost waste caused by resource idleness during business troughs, greatly limiting the flexibility and operation and maintenance efficiency of distributed systems.

[0017] To address the above technical problems, the technical concept of this application is to first collect the monitoring data generated during the operation of the IT infrastructure and application systems to obtain a time series dataset of the monitoring data, then use deep learning-based data processing techniques to clean and structurally map the time series dataset of the monitoring data to obtain a time series of IT resource usage status encoding vectors, and then decode based on the dynamic semantic propagation representation with feature independence constraints of the IT resource usage status encoding vector time series to obtain capacity prediction result data and prediction confidence. Finally, input the capacity prediction result data and prediction confidence into the elastic scaling decision engine to obtain a scaling decision. In this way, by deeply exploring the potential relationships between different monitoring indicators, it is possible to more comprehensively perceive the operating status of the IT infrastructure and application systems, and then the generated elastic scaling decision can more flexibly adapt to the non-linear resource requirements in complex load scenarios, reduce the expansion delay during burst traffic and resource redundancy during business troughs, so as to provide a more intelligent and automated operation and maintenance IT resource management solution.

[0018] Figure 1 The system block diagram of the IT resource management system based on automated operation and maintenance according to the embodiments of this application. As Figure 1As shown in the figure, in the IT resource management system 100 based on automated operation and maintenance, it includes: a monitoring data collection module 110, which is used to collect monitoring data generated during the operation of IT infrastructure and application systems to obtain a time series dataset of the monitoring data; a monitoring data analysis module 120, which is used to perform time series analysis of the IT resource usage status on the time series dataset of the monitoring data to obtain capacity prediction result data and prediction confidence; a scaling decision generation module 130, which is used to obtain a scaling decision based on the capacity prediction result data and the prediction confidence.

[0019] In the embodiment of the present application, the monitoring data collection module 110 is used to collect monitoring data generated during the operation of IT infrastructure and application systems to obtain a time series dataset of the monitoring data. Specifically, in the embodiment of the present application, the monitoring data includes CPU utilization rate, memory utilization rate, disk IOPS, network bandwidth utilization rate, disk space utilization rate, request response time, TPS / QPS, error rate, connection number, queue length, system load, number of processes, disk queue length, network latency, and CPU context switch. It should be understood that the IT infrastructure and application systems are in a continuously running dynamic process, and various indicators change continuously over time. For example, the CPU utilization rate may fluctuate greatly at different times due to different business processing volumes, and the memory utilization rate will also change in real time as the application programs run and resources are allocated. Collecting the time series dataset can completely record the changes of these indicators in the time series, so as to accurately reflect the dynamic change process of the system operation state. And by collecting the time series dataset of the monitoring data, the correlation relationship between business indicators (such as TPS / QPS, request response time, etc.) and resource indicators (such as CPU utilization rate, memory utilization rate, etc.) in the time dimension can be observed. For example, when TPS / QPS suddenly increases, it may cause the CPU utilization rate and memory utilization rate to rise accordingly. This correlation relationship is crucial for analyzing the resource requirements of the business and optimizing resource allocation. Generally speaking, based on the collected time series dataset of the monitoring data, data analysis algorithms based on deep learning can be used to model and predict the usage of IT resources, and then provide a scientific basis for capacity planning and elastic scaling strategies, so as to effectively avoid service degradation caused by insufficient resources or cost waste caused by excessive resources.

[0020] The following is a detailed description of a specific implementation process of "collecting monitoring data generated during the operation of IT infrastructure and application systems to obtain a time series dataset of the monitoring data":

[0021] First of all, determining the collection metrics and frequency is the starting point and key guidance for the entire collection work. In the current context where cloud-native architectures and microservices-based applications are prevalent, the operating status of IT systems needs to be accurately reflected through a series of key metrics. Metrics such as CPU utilization, memory utilization, disk IOPS, network bandwidth utilization, disk space utilization, request response time, TPS / QPS, error rate, connection count, queue length, system load, process count, disk queue length, network latency, and CPU context switches comprehensively show the operating state of the system from multiple dimensions including computing resources, storage resources, network conditions, and business performance. Taking an e-commerce system as an example, during promotional activities, metrics such as TPS / QPS, CPU utilization, and network bandwidth utilization will change sharply, and these changes are directly related to whether the system can stably handle high-concurrency user access. Therefore, the precise collection and real-time monitoring of these metrics are particularly important. During non-promotional periods, metrics such as disk space utilization and process count help operations and maintenance personnel understand the long-term usage trends of system resources and the overall health of the system.

[0022] Determining the collection frequency cannot be ignored either. Different metrics have different change characteristics. Setting an appropriate collection frequency based on these characteristics and business requirements can avoid wasting system resources caused by over-collection while ensuring sufficient data is obtained to accurately analyze the system state. For metrics like CPU utilization that change relatively frequently, collecting once every 15 seconds can capture its instantaneous fluctuations in a timely manner, providing accurate data for subsequent analysis of the system's real-time load. The change of disk space utilization is relatively slow, and collecting once every 5 minutes is sufficient to meet the monitoring requirements for its usage. In actual operation, flexible adjustment also needs to be made according to the specific situation of the system and the criticality of the business. For example, for core business systems, the collection frequency of key metrics can be appropriately increased during peak business periods to ensure that potential performance issues can be detected and addressed in a timely manner.

[0023] Selecting collection tools and technologies is the core means to achieve efficient data collection. The proxy-based collection method has wide applicability in practical applications. By deploying collection proxy software, such as Prometheus Node Exporter, in the server or container, it can delve into the system internals and collect basic metrics of the server in real time, such as CPU, memory, and disk. Prometheus Node Exporter utilizes the system's underlying interfaces and protocols to accurately obtain these key data and sends the collected data to the Prometheus Server via the HTTP protocol. This method can transmit the system's operation information in a timely and accurate manner. For some devices where proxy software cannot be installed, the agentless collection technology plays an important role. The monitoring data of network devices can be directly obtained by using network protocols (such as SNMP). For example, for a network switch, by reasonably configuring SNMP parameters, information such as port traffic and network latency can be easily collected without installing additional software on the device, reducing the device's burden and maintenance cost.

[0024] Building a data collection architecture is an important guarantee to ensure comprehensive and accurate data collection. In large-scale distributed systems, due to the large number of system nodes and their wide distribution, adopting a distributed collection architecture becomes an inevitable choice. Collection agents are deployed in server clusters in different regions respectively, each responsible for collecting the monitoring data of the servers in its region, and then these data are aggregated into the central data storage system. This architecture mode can not only achieve unified monitoring of the entire distributed system but also effectively reduce the burden on a single collection node, improving the collection efficiency and data reliability. For complex multi-layer architecture applications, the hierarchical collection method has more advantages. Taking a system with a front end, middle layer, and database layer as an example, the front end is responsible for directly interacting with users and collecting metrics such as request response time and connection count, which can intuitively reflect the user experience; the middle layer, as the core layer for business logic processing, collects metrics such as the number of processes and system load, which helps to understand the efficiency of business processing and the overall operation pressure of the system; the database layer stores key data, and collecting metrics such as disk IOPS and disk queue length can timely detect the performance bottleneck of the database. Through hierarchical collection, in-depth analysis of each layer of the system can be carried out to quickly locate the root cause of performance problems.

[0025] Data transmission and storage are key subsequent steps in the data collection process. The collected data needs to be transmitted over the network to a storage system for long-term preservation and subsequent analysis. Message queues (such as Kafka) play an important role in the data transmission process. It acts like a data "buffer", capable of buffering and asynchronously transmitting the collected data. When network fluctuations occur, Kafka can temporarily store the data to avoid data loss and ensure that the data can be stably transmitted to the storage system. In terms of data storage, choosing the right database is crucial. The time series database InfluxDB is specifically optimized for time series data. It can efficiently store and query monitoring data arranged in a time series. This means that when historical data needs to be queried to analyze the system performance trend, InfluxDB can respond quickly and provide accurate data support.

[0026] In the embodiment of the present application, the monitoring data analysis module 120 is used to perform time series analysis of the IT resource usage status on the time series dataset of the monitoring data to obtain capacity prediction result data and prediction confidence. Specifically, Figure 2 It is a block diagram of the monitoring data analysis module in the IT resource management system based on automated operation and maintenance according to the embodiment of the present application. As Figure 2 shown, the monitoring data analysis module 120 includes: a monitoring data preprocessing unit 121, which is used to perform data cleaning and structured mapping on the time series dataset of the monitoring data to obtain a time series of IT resource usage status encoding vectors; a monitoring data encoding unit 122, which is used to perform dynamic propagation of IT resource usage status features on the time series of IT resource usage status encoding vectors to obtain the capacity prediction result data and prediction confidence.

[0027] In the embodiment of the present application, the monitoring data preprocessing unit 121 is used to perform data cleaning and structured mapping on the time series dataset of the monitoring data to obtain a time series of IT resource usage status encoding vectors. Specifically, Figure 3 It is a block diagram of the monitoring data preprocessing unit in the IT resource management system based on automated operation and maintenance according to the embodiment of the present application. As Figure 3 shown, the monitoring data preprocessing unit 121 includes: a monitoring data cleaning subunit 1211, which is used to perform data cleaning on each piece of monitoring data in the time series dataset of the monitoring data to obtain a time series of cleaned monitoring data; a monitoring data structured mapping subunit 1212, which is used to perform structured mapping based on a fully connected layer on the time series of the cleaned monitoring data to obtain the time series of the IT resource usage status encoding vectors.

[0028] In the embodiment of the present application, the monitoring data cleaning subunit 1211 is configured to clean each piece of monitoring data in the time series dataset of the monitoring data to obtain a time series of the cleaned monitoring data. Accordingly, considering that during the acquisition process of the monitoring data, due to the instability of the hardware device, network transmission interference, environmental factors, etc., some noise data may be introduced. These noise data may be manifested as abnormal numerical fluctuations, occasional error data points, etc. In addition, the monitoring system may not be able to completely collect the monitoring data at all time points for various reasons, resulting in missing values in the dataset. And errors may also occur during the acquisition, transmission, or storage of the data, such as data format errors, data values exceeding the reasonable range, etc. That is, there are various quality problems in the original monitoring data. If directly used for analysis, it will lead to model prediction deviation and invalid operation and maintenance decisions. In order to improve the data quality of the monitoring data, in the present application, it is necessary to clean each piece of monitoring data in the time series dataset of the monitoring data to obtain a time series of the cleaned monitoring data. Through data cleaning operations such as removing noise data, handling missing values, and correcting error data, the time series of the processed and cleaned monitoring data can more accurately reflect the real operating state of the IT infrastructure and application systems, which helps to improve the accuracy of subsequent data analysis.

[0029] In the embodiment of the present application, the monitoring data structured mapping sub-unit 1212 is configured to perform structured mapping based on a fully connected layer on the time series of the cleaned monitoring data to obtain the time series of the IT resource usage status encoding vectors. It should be understood that the time series of the cleaned monitoring data contains information in multiple dimensions, such as CPU utilization, memory utilization, etc. The data in each dimension can only reflect the status of a certain aspect of the system when viewed separately. Through the structured mapping based on the fully connected layer, these scattered monitoring data features in different dimensions can be extracted and integrated, and transformed into an encoding vector that can comprehensively represent the IT resource usage status. In this way, the overall usage of IT resources can be described more comprehensively and abstractly, providing more valuable input for subsequent analysis and processing. Specifically, the fully connected layer is a common structure in neural networks, where each neuron is connected to all neurons in the previous layer. In the process of structured mapping based on the fully connected layer, first, the time series of the cleaned and normalized monitoring data is used as the input to ensure that the data scales of different dimensions are similar and to avoid the excessive influence of some dimensions. The fully connected layer contains a weight matrix and a bias term determined according to the dimensions of the input data and the dimensions of the expected output encoding vector. During training, the weights and bias terms are continuously adjusted through the backpropagation algorithm to minimize the loss function and learn the optimal mapping parameters. For the input monitoring data vector at each time step, it is multiplied by the weight matrix through matrix multiplication, added with the bias term, and then passed through an activation function such as ReLU or Sigmoid for non-linear transformation to obtain the IT resource usage status encoding vector corresponding to that time step. This process is repeated for each time step, and finally, a time series of the IT resource usage status encoding vectors is formed. Generally speaking, through the non-linear transformation of the fully connected layer, more complex feature relationships in the data can be learned, and the original monitoring data can be converted into a feature representation form that can better capture the internal patterns and rules of the IT resource usage status. This can provide a more meaningful feature basis for subsequent analysis, thereby improving the accuracy of capacity prediction and the sensitivity to system state changes.

[0030] In the embodiment of the present application, the monitoring data encoding unit 122 is configured to perform dynamic propagation of IT resource usage status features on the time series of the IT resource usage status encoding vectors to obtain the capacity prediction result data and the prediction confidence. Specifically, Figure 4 The block diagram of the monitoring data encoding unit in the IT resource management system based on automated operation and maintenance according to the embodiment of the present application. As Figure 4As shown, the monitoring data encoding unit 122 includes: an IT resource usage status time-series pattern encoding subunit 1221, configured to perform dynamic semantic propagation on the time series of the IT resource usage status encoding vectors based on the independence constraint of IT resource usage status features to obtain IT resource usage status time-series pattern feature vectors; and an IT resource usage status time-series pattern feature decoding subunit 1222, configured to perform time series prediction based on a decoder on the IT resource usage status time-series pattern feature vectors to obtain the capacity prediction result data and prediction confidence.

[0031] In an embodiment of the present application, the IT resource usage status time-series pattern encoding subunit 1221 is configured to perform dynamic semantic propagation on the time series of the IT resource usage status encoding vectors based on the independence constraint of IT resource usage status features to obtain IT resource usage status time-series pattern feature vectors. Specifically, in an embodiment of the present application, the IT resource usage status time-series pattern encoding subunit includes: a positioning reference encoding vector calculation secondary subunit, configured to perform end reference positioning and axial reference positioning on the time series of the IT resource usage status encoding vectors respectively to obtain an IT resource usage status time-series transfer end reference positioning encoding vector and an IT resource usage status time-series transfer axial reference positioning encoding vector; a transfer regulation factor calculation secondary subunit, configured to calculate IT resource usage status time-series transfer regulation factors of each IT resource usage status encoding vector in the time series of the IT resource usage status encoding vectors based on the IT resource usage status time-series transfer end reference positioning encoding vector and the IT resource usage status time-series transfer axial reference positioning encoding vector; and an IT resource usage status time-series pattern feature generation secondary subunit, configured to perform IT resource usage status dynamic constraint transfer encoding on the time series of the IT resource usage status encoding vectors based on the IT resource usage status time-series transfer regulation factors of each IT resource usage status encoding vector to obtain the IT resource usage status time-series pattern feature vectors.

[0032] It should be understood that in the time series of the IT resource usage status encoding vectors, there are long-range dependencies, that is, the information at relatively distant time steps in the sequence may have an important impact on the current state. For example, a resource usage peak some time ago may have a long-term impact on subsequent resource allocation and system performance. In addition, the IT resource usage status contains both global structural information, such as the overall resource usage distribution pattern, and local structural information, such as short-term resource usage fluctuations. In order to comprehensively and accurately grasp the IT resource usage status, it is necessary to balance these two types of information when processing sequence data. Based on this, in this application, a dynamic semantic propagation based on the independence constraint of IT resource usage status features is performed on the time series of the IT resource usage status encoding vectors to obtain the IT resource usage status time series pattern feature vectors. Specifically, by introducing an end reference positioning encoding vector, using the end information of the sequence as a reference point, and combining a dynamic message passing mechanism, the dependencies in the time series of the IT resource usage status encoding vectors can be effectively captured, so as to more accurately describe the changing law of the IT resource usage status over time. At the same time, the global distribution pattern of the sequence (such as the weekday / weekend pattern) is extracted through clustering analysis, and an axial reference positioning encoding vector is generated as a global reference. The transfer regulator dynamically adjusts the weights according to the similarity between the current moment and the global axis and local end anchors, so that the generated IT resource usage status time series pattern feature vectors are more representative and can better reflect the real situation of the IT resource usage status, thereby providing a reliable feature reference basis for capacity prediction and scaling decision-making.

[0033] Specifically, in the embodiment of this application, the positioning reference encoding vector calculation secondary subunit is used to: extract the last IT resource usage status encoding vector from the time series of the IT resource usage status encoding vectors as the IT resource usage status time series transfer end reference positioning encoding vector, and this process can be expressed by the formula:

[0034] O = {v1, v2,..., v i ,..., v t}

[0035] v tai l = v t

[0036] where O is the time series of the IT resource usage status encoding vectors, v1, v2, v i and v t are the 1st, 2nd, i-th, and t-th IT resource usage status encoding vectors in the time series of the IT resource usage status encoding vectors respectively, and v tail is the IT resource usage status time series transfer end reference positioning encoding vector;

[0037] Perform clustering analysis on the time series of the IT resource usage status encoding vectors to obtain the IT resource usage status time series transfer axial reference positioning encoding vectors, and this process can be expressed by the formula:

[0038]

[0039] where Cluster is the clustering analysis operation, max(v i ) and min(v i ) are respectively the maximum and minimum values of v i , η is the adjustment parameter, a i is the i-th IT resource usage status encoding axial reference positioning value in the time series of the IT resource usage status encoding axial reference positioning values, softmax is the normalization function, e i is the i-th IT resource usage status encoding axial reference positioning weight value in the time series of the IT resource usage status encoding axial reference positioning weight values, t is the number of vectors in the O, and v axis is the IT resource usage status time series transfer axial reference positioning encoding vector.

[0040] It should be understood that in the dynamic time series modeling of the IT resource usage status, the end point (i.e., the current state) of the time series of the IT resource usage status encoding vectors usually carries the latest system state information, such as the current resource utilization rate peak or abnormal fluctuation. Taking it as the end reference positioning encoding vector is essentially taking "the current moment" as the boundary condition for dynamic semantic propagation, similar to the terminal state constraint in control theory. In the capacity prediction scenario, the real-time adjustment of resource allocation highly depends on the current state (such as a sudden traffic peak). Anchoring the end information can ensure that the model always takes the latest state as the reference during the message passing process, avoiding semantic deviation caused by historical noise. For example, when the system detects a sudden increase in the current CPU utilization rate, the end reference positioning encoding vector will be used as a key reference point to constrain the direction of subsequent message passing, making the model pay more attention to the historical segments related to the current state (such as a similar peak pattern a few hours ago), thereby improving the prediction sensitivity to sudden loads. That is, by fixing the end reference point, the model can effectively capture the time series patterns strongly correlated with the current state in long-range dependencies (such as the lag effect of periodic peaks), while enhancing the convergence of message passing and preventing the gradient from decaying or exploding in long sequences. This can provide a stable state representation basis for subsequent dynamic adjustment of resource scaling decisions (such as the trigger timing of elastic expansion).

[0041] Accordingly, considering that IT resource usage patterns usually imply global laws, such as the peak load during the day and the low-load period at night on weekdays, or the difference in traffic distribution on weekends. By performing clustering analysis to extract the time-series transfer axial reference positioning coding vector of the IT resource usage status, the essence is to perform "principal component extraction" on the distribution pattern of historical sequences, compressing high-dimensional time-series data into representative global features (such as cluster centers). In capacity prediction, this global axis can be regarded as the "normal mode skeleton" of system operation. For example, after identifying two types of axes, namely the "weekday mode" and the "weekend mode", the model can quickly determine the category to which the current sequence belongs, and then make predictions based on the resource consumption patterns of historical similar modes. For example, if the current sequence is classified as the "weekday mode", the model can preferentially refer to the typical load data at 10 am in the same category, rather than blindly integrating all historical information. That is, the time-series transfer axial reference positioning coding vector of the IT resource usage status introduces global regularization to message passing, forcing the model to follow the main distribution direction of the data (such as the overall load trend) while paying attention to local fluctuations (such as short-term traffic jitters). This avoids the model overfitting to noise. Especially in scenarios with small samples or sparse events (such as e-commerce promotion activities), the axial constraint can maintain the rationality of the prediction results and ensure that the scaling decision (such as the baseline setting of reserved resources) conforms to the long-term operation law of the system.

[0042] Specifically, in the embodiment of the present application, the transfer regulation factor calculation secondary subunit includes: an end regulation factor calculation tertiary subunit, configured to calculate the IT resource usage status time-series transfer end regulation factor of each IT resource usage status coding vector in the time series of the IT resource usage status coding vector relative to the IT resource usage status time-series transfer end reference positioning coding vector; an axial regulation factor calculation tertiary subunit, configured to calculate the IT resource usage status time-series transfer axial regulation factor of each IT resource usage status coding vector in the time series of the IT resource usage status coding vector relative to the IT resource usage status time-series transfer axial reference positioning coding vector; and a message regulation factor calculation tertiary subunit, configured to calculate the IT resource usage status time-series transfer regulation factor of each IT resource usage status coding vector based on the IT resource usage status time-series transfer end regulation factor and the IT resource usage status time-series transfer axial regulation factor of each IT resource usage status coding vector in the time series of the IT resource usage status coding vector.

[0043] More specifically, in the embodiment of the present application, the end regulation factor calculation tertiary subunit is configured to calculate the IT resource usage status time-series transfer end regulation factor of each IT resource usage status coding vector in the time series of the IT resource usage status coding vector relative to the IT resource usage status time-series transfer end reference positioning coding vector, which can be expressed by the following formula:

[0044]

[0045] Among them, f(v i , v tail ) is the terminal regulation factor for calculating between v i and v tail . v ij is the eigenvalue at the j-th position of v i , v tailj is the eigenvalue at the j-th position of v tail . log2 is the logarithmic function value with base 2, n is the number of eigenvalues in v i , exp is the exponential function value with base e (the natural constant), and α i is the IT resource usage status time-series transfer terminal regulation factor corresponding to v i .

[0046] It should be understood that the core of the IT resource usage status time-series transfer terminal regulation factor lies in quantifying the semantic association strength between the historical state and the current state. Collecting the exponential deformation function based on KL divergence instead of the simple time decay function can more accurately depict the non-linear dependence relationship. For example, the similar load peak two weeks ago may have a greater impact on the current state than the irrelevant fluctuations three days ago. This factor realizes an adaptive attention mechanism - assigning higher weights to historical segments highly relevant to the current state. In the capacity planning scenario, this design can automatically focus on the time periods in history similar to the current resource bottleneck (such as the stress test period before an e-commerce promotion), improving the causal relevance of the scaling decision.

[0047] More specifically, in the embodiment of the present application, the axial regulation factor calculation three-level subunit is used to calculate the IT resource usage status time-series transfer axial regulation factor of each IT resource usage status coding vector in the time series of the IT resource usage status coding vector relative to the IT resource usage status time-series transfer axial reference positioning coding vector, which can be expressed by the following formula:

[0048]

[0049] Among them, f(v i , v axis ) is the axial regulation factor for calculating between v i and v axis . ‖·‖2 is used to calculate the Euclidean norm of the vector, arccosh is the inverse hyperbolic cosine function, and β i is the IT resource usage status time-series transfer axial regulation factor corresponding to v i .

[0050] It should be understood that the IT resource usage status time-series transfer axial regulator measures the degree of fit between each time step and the main mode through projection weights, which essentially constructs a structured prior in the feature space. For example, if the low-load period in the early morning of a working day deviates from the typical axial mode (such as a sudden high CPU occupancy), it may imply abnormal task scheduling. That is, the IT resource usage status time-series transfer axial regulator acts as an out-of-distribution detector, suppressing noisy data that deviates from the main mode in message passing and enhancing the model's robustness to trend changes. Especially in scenarios where the load mode changes gradually (such as the natural growth of the user scale), this factor can smooth short-term fluctuations, making the prediction results more in line with the long-term evolution law and avoiding frequent ineffective scaling operations (such as repeated start and stop of virtual machines).

[0051] More specifically, in the embodiment of the present application, the message regulation factor calculation three-level subunit is used to calculate the IT resource usage status time-series transfer regulator of each IT resource usage status coding vector based on the IT resource usage status time-series transfer end regulator and the IT resource usage status time-series transfer axial regulator of each IT resource usage status coding vector in the time series of the IT resource usage status coding vector.

[0052] It should be understood that the fusion (weighted summation) of the IT resource usage status time-series transfer end regulator and the IT resource usage status time-series transfer axial regulator is essentially a game between local and global information. For example, when the system is in a stable state (highly consistent with the axis), the model should tend to global constraints to maintain prediction stability; while when a sudden anomaly is detected (highly related to the end but deviating from the axis), local weights need to be enhanced to respond quickly. In capacity prediction, this dynamic balance can be achieved through learnable parameters: assuming that at a certain moment, the feature vector is highly similar to both the end anchor (current peak) and the axis (low-load mode on weekends), the model may judge it as a "temporary weekend activity", and thus adopt a short-term aggressive strategy (such as on-demand instances) rather than long-term reservation in the capacity expansion decision. That is, the fused IT resource usage status time-series transfer regulator realizes the adaptive fusion of multi-granularity information, enabling the model to capture the impact of short-term events while following the long-term evolution trend.

[0053] More specifically, in the embodiments of the present application, the message regulation factor calculation three-level subunit is configured to: perform lateral convergence covariance on the IT resource usage status time-series transfer end regulation factor and the IT resource usage status time-series transfer axial regulation factor of the IT resource usage status encoding vector using a canonical space constraint matrix to obtain an IT resource usage status time-series transfer end convergence covariance regulation factor and an IT resource usage status time-series transfer axial convergence covariance regulation factor; based on the IT resource usage status time-series transfer end convergence covariance regulation factor and the IT resource usage status time-series transfer axial convergence covariance regulation factor, perform closed-loop constraint optimization based on the spin space on the IT resource usage status time-series transfer end regulation factor and the IT resource usage status time-series transfer axial regulation factor to obtain an optimized IT resource usage status time-series transfer end regulation factor and an optimized IT resource usage status time-series transfer axial regulation factor; perform weighted processing on the optimized IT resource usage status time-series transfer end regulation factor and the optimized IT resource usage status time-series transfer axial regulation factor and then perform activation processing based on sigmoid to obtain the IT resource usage status time-series transfer regulation factor of the IT resource usage status encoding vector. The above process can be expressed by the formula:

[0054]

[0055] y i = Sigmoid(ω1·α i '+ ω2·β i ')

[0056] where δ i and ε i are respectively the IT resource usage status time-series transfer end convergence covariance regulation factor and the IT resource usage status time-series transfer axial convergence covariance regulation factor after lateral convergence covariance of α i and β i , T represents the transpose operation, |·| is the absolute value operation, α i ' and β i ' are respectively the optimized IT resource usage status time-series transfer end regulation factor and the optimized IT resource usage status time-series transfer axial regulation factor corresponding to v i , ω1 and ω2 are respectively weighted hyperparameters, Sigmoid is a weight mapping function, and y i is the IT resource usage status time-series transfer regulation factor corresponding to v i .

[0057] Specifically, in the semantic propagation framework of dynamic time series modeling, the IT resource usage status time series transfer end benchmark positioning coding vector and the IT resource usage status time series transfer axial benchmark positioning coding vector respectively construct a reference benchmark for IT resource usage status time series message passing from the dimensions of local real-time and global regularity. However, there is a potential conflict in the scope of action of the two in the feature space: the lateral propagation field driven by the IT resource usage status time series transfer end regulation factor tends to strengthen the historical correlation features within the current state radiation range (such as the attention focus of sudden load on adjacent time windows), and its radiation attenuation characteristic is vulnerable to short-term noise perturbation, resulting in local semantic overshoot; while the IT resource usage status time series transfer axial regulation factor imposes structural regularization on message passing through the distribution main direction constructed by the global clustering center (such as the trend constraint of the weekday pattern), but when the system state deviates from the historical main pattern, it may cause underfitting risk due to excessive suppression of local mutation signals. To coordinate this contradiction, a lateral convergence covariance mechanism for the message passing field needs to be introduced - first, the feature projections of the two types of regulation factors are jointly optimized based on the canonical space constraint matrix. This matrix extracts the covariant basis on the time series feature manifold through tensor decomposition, maps the radiation attenuation gradient field of the end regulation factor and the global regularization potential field of the axial regulation factor to the same canonical space, and uses the manifold alignment technology to eliminate the heterogeneity in feature scale and propagation direction between the two; subsequently, a spin transformation is performed on the covaried joint representation through the component rotation matrix defined by the reversible space metric. This operation essentially dynamically bends and reparameterizes the feature space while maintaining the information topological structure, enabling the local sensitivity of the end constraint and the global stability of the axial constraint to form positive interaction and complementarity within a finite-dimensional closed subspace, and finally constructing a unified transfer constraint field with adaptive balance ability. For example, when the system encounters a sudden load and deviates from the historical axial pattern, the canonical space constraint matrix will dynamically adjust the radiation weight of the end regulation factor, enhance the representation ratio of the current abnormal state within the closed subspace through component rotation, and at the same time, the axial regulation factor still maintains weak supervision of the long-term operation baseline of the system, enabling the message passing process to quickly capture the local details of the current load increase (such as resource preemption events at the second level) while avoiding completely deviating from the historical law (such as ignoring the baseline load of diurnal periodicity), thus achieving an organic unity of short-term response agility and long-term policy robustness in elastic scaling decisions.

[0058] Specifically, in the embodiment of the present application, the IT resource usage status time series pattern feature generation secondary subunit is used to perform IT resource usage status dynamic constraint transfer coding on the time series of the IT resource usage status coding vector based on the IT resource usage status time series transfer regulation factor of each IT resource usage status coding vector to obtain the IT resource usage status time series pattern feature vector, which can be expressed by the following formula:

[0059]

[0060] Among them, z is the time series pattern feature vector of the IT resource usage status.

[0061] It should be understood that weighted aggregation of the corresponding original IT resource usage status encoding vector based on the time series transfer regulation factor of the IT resource usage status is essentially to convert the original time series into a low-dimensional representation with both local details and global structure. In capacity prediction, this feature vector needs to encode two types of key information simultaneously: 1) recent state fluctuations (such as the variance of CPU usage rate in the past 1 hour); 2) the degree of deviation from the historical pattern (such as whether the current load breaks the quarterly cycle rule). For example, if the feature vector shows "high local fluctuations + significant deviation from the axis", the system may judge it as unpredictable burst traffic, thus triggering a conservative capacity expansion strategy (such as over-provisioning 20% of resources); if it is "low fluctuations + conforming to the axis", the baseline resource allocation is adopted. That is, the finally generated time series pattern feature vector of the IT resource usage status provides an interpretable decision-making basis by explicitly modeling local-global dependencies. It can more accurately predict future resource requirements, and thus can support the elastic scaling decision to achieve Pareto optimality between cost and performance.

[0062] In the embodiment of the present application, the IT resource usage status time series pattern feature decoding subunit 1222 is configured to perform decoder-based time series prediction on the IT resource usage status time series pattern feature vector to obtain the capacity prediction result data and the prediction confidence. It should be understood that the IT resource usage status time series pattern feature vector has integrated core information such as the time-varying pattern and potential trend of resource usage, but it needs to be further transformed into specific capacity prediction results to guide operation and maintenance decisions. The decoder is a key execution unit of the time series prediction model and is good at processing the logical derivation of sequence features. By decoding the IT resource usage status time series pattern feature vector with the decoder, the future resource requirements (such as the number of CPU cores, memory capacity, etc.) can be deduced based on historical features and current inputs. Specifically, the decoder is a core component in the deep learning model. In the time series prediction scenario, it receives the encoded IT resource usage status time series pattern feature vector and analyzes and deduces the features through internal network structures (such as fully connected layers, recurrent units, etc.) to gradually generate the predicted values for future time steps. The generated capacity prediction result data refers to the quantified values of the IT resource requirements in a future period predicted by the decoder, such as the predicted average CPU utilization value for the next hour, the predicted number of virtual machine requirements for the next 24 hours, the disk space prediction value, etc. These data directly represent the specific demand for resources such as CPU, memory, and storage in the future system and are the core basis for elastic scaling decisions. The generated prediction confidence is an evaluation index of the reliability of the capacity prediction result, usually presented in the form of a confidence interval or probability, reflecting the uncertainty of the prediction result. For example, predicting that the memory requirement for a certain period is 8GB with a confidence of 85% means that there is an 85% probability that the prediction result is close to the true value. Through the confidence, the credibility of the prediction can be judged, and the uncertainty can be comprehensively considered in the elastic scaling decision to improve the scientificity of the decision.

[0063] In the embodiment of the present application, the scaling decision generation module 130 is configured to obtain a scaling decision based on the capacity prediction result data and the prediction confidence level. Specifically, in the embodiment of the present application, the scaling decision generation module is configured to: input the capacity prediction result data and the prediction confidence level into an elastic scaling decision engine to obtain the scaling decision. It should be understood that the elastic scaling decision engine is the core decision-making component in the automated operation and maintenance IT resource management system, and undertakes the tasks of judging resource elastic scheduling and generating instructions. Specifically, the elastic scaling decision engine is an intelligent processing system integrating functions such as policy matching, condition judgment, and data processing. It includes a scaling policy management module (for storing and managing predefined scaling policies), an alarm module (receiving signals for triggering scaling by real-time monitoring metrics), a policy matching and condition judgment module (analyzing the matching degree between data and policies), and interfaces with the capacity prediction result data and real-time monitoring data. It takes the capacity prediction result data, prediction confidence level, and real-time monitoring data as inputs, and at the same time associates predefined elastic scaling policies. Through policy matching and condition judgment, it analyzes whether the current IT resource usage status meets the scaling trigger conditions. Finally, it generates a scaling decision on whether to perform elastic scaling operations and what kind of scaling actions to perform (scaling out / down / vertical scaling). That is, the generated scaling decision is a resource adjustment instruction generated by the elastic scaling decision engine, and its core lies in dynamically adapting to system load changes to achieve efficient resource utilization. Specifically, the elastic scaling decision engine parses the scaling policies defined in the scaling policy management module (such as trigger conditions, action types, amplitudes, and cooling times) through a rule engine, combines the capacity prediction results (such as 85% CPU utilization in the next hour) and prediction confidence levels (such as 90% reliability), and comprehensively performs policy matching and condition judgment based on real-time monitoring metrics (current request response time, error rate, etc.). For example, when the predicted CPU utilization exceeds the threshold and the confidence level is met, the engine will match the "scale out 2 servers" policy and check whether the cooling time has ended (such as avoiding repeated scaling out within 10 minutes), and finally generate a specific instruction (such as "add 2 Web server instances"). The scaling decision types specifically include horizontal scaling out / down (increasing or decreasing the number of instances) and vertical scaling (adjusting the instance configuration), and the scaling parameters cover the scaling amplitude (such as scaling out by 10% each time) and the cooling period. In particular, the scaling decision not only focuses on single resource adjustments but also associates with subsequent optimization mechanisms. The elastic scaling decision engine will feedback the scaling effect data (such as the prediction model continuously overestimating or underestimating resource requirements) to the capacity prediction model for adjusting model parameters or optimizing scaling policies (such as adjusting trigger conditions and scaling amplitudes). For example, if it is found that resources are still insufficient after scaling out, the recognition ability of the prediction model for high-load scenarios can be optimized, or the scaling amplitude rule can be adjusted to form an adaptive cycle of "decision execution → effect monitoring → feedback optimization" to achieve dynamic resource adjustment.Generally speaking, by driving decisions through scientific capacity prediction results, the elastic scaling decision engine can avoid over-scaling or under-scaling. This not only prevents business damage caused by insufficient resources but also avoids resource waste, thereby achieving efficient utilization of IT resources and reducing the operation and maintenance costs of enterprises on infrastructure such as servers and storage. At the same time, the elastic scaling decision engine automatically generates scaling decision instructions based on data, without manual experience judgment, and directly triggers the automated orchestration engine to execute scaling operations (such as calling the cloud platform API to create virtual machines and adjusting the number of container replicas). This process significantly shortens the decision-making and execution cycle. Compared with the lagged response of traditional manual intervention, it can respond to transient fluctuations in business traffic in real time, which is beneficial to improving the elastic capacity and operation and maintenance response efficiency of distributed systems.

[0064] In summary, the IT resource management system 100 based on automated operation and maintenance according to the embodiments of the present application is elucidated. It uses data processing technology based on deep learning to perform data cleaning and structured mapping on the time series dataset of monitoring data generated during the operation of IT infrastructure and application systems to obtain the time series of IT resource usage status encoding vectors. Subsequently, based on the dynamic semantic propagation representation with feature independence constraints of the IT resource usage status encoding vector time series, the capacity prediction result data and prediction confidence are decoded. Finally, the capacity prediction result data and prediction confidence are input into the elastic scaling decision engine to obtain scaling decisions. In this way, the generated elastic scaling decisions can more flexibly adapt to the non-linear resource requirements in complex load scenarios, thereby realizing more automated and intelligent IT resource operation and maintenance management.

Claims

1. An IT resource management system based on automated operation and maintenance, characterized in that, Including: A monitoring data acquisition module, which is used to acquire monitoring data generated during the operation of IT infrastructure and application systems to obtain a time series dataset of the monitoring data; A monitoring data analysis module, which is used to perform time series analysis of the IT resource usage status on the time series dataset of the monitoring data to obtain capacity prediction result data and prediction confidence. Among them, the monitoring data analysis module includes: a monitoring data preprocessing unit, which is used to perform data cleaning and structured mapping on the time series dataset of the monitoring data to obtain a time series of IT resource usage status encoding vectors; a monitoring data encoding unit, which is used to perform dynamic propagation of IT resource usage status features on the time series of the IT resource usage status encoding vectors to obtain the capacity prediction result data and prediction confidence; A scaling decision generation module, which is used to obtain a scaling decision based on the capacity prediction result data and prediction confidence.

2. The IT resource management system based on automated operation and maintenance according to claim 1, wherein The monitoring data includes CPU utilization rate, memory utilization rate, disk IOPS, network bandwidth utilization rate, disk space utilization rate, request response time, TPS / QPS, error rate, connection number, queue length, system load, number of processes, disk queue length, network latency, and CPU context switch.

3. The IT resource management system based on automated operation and maintenance according to claim 1, characterized in that, The monitoring data preprocessing unit includes: A monitoring data cleaning subunit, which is used to perform data cleaning on each piece of monitoring data in the time series dataset of the monitoring data to obtain a time series of the cleaned monitoring data; A monitoring data structured mapping subunit, which is used to perform structured mapping based on a fully connected layer on the time series of the cleaned monitoring data to obtain a time series of the IT resource usage status encoding vectors.

4. The IT resource management system based on automated operation and maintenance according to claim 1, wherein The monitoring data encoding unit includes: An IT resource usage status time series pattern encoding subunit, which is used to perform dynamic semantic propagation based on the independence constraint of IT resource usage status features on the time series of the IT resource usage status encoding vectors to obtain IT resource usage status time series pattern feature vectors; An IT resource usage status time series pattern feature decoding subunit, which is used to perform time series prediction based on a decoder on the IT resource usage status time series pattern feature vectors to obtain the capacity prediction result data and prediction confidence.

5. The IT resource management system based on automated operation and maintenance according to claim 4, wherein The IT resource usage status time series pattern encoding subunit includes: A positioning reference encoding vector calculation secondary subunit, which is used to perform end reference positioning and axial reference positioning on the time series of the IT resource usage status encoding vectors respectively to obtain an IT resource usage status time series transfer end reference positioning encoding vector and an IT resource usage status time series transfer axial reference positioning encoding vector; A transfer regulation factor calculation secondary subunit, which is used to calculate the IT resource usage status time series transfer regulation factors of each IT resource usage status encoding vector in the time series of the IT resource usage status encoding vectors based on the IT resource usage status time series transfer end reference positioning encoding vector and the IT resource usage status time series transfer axial reference positioning encoding vector; The IT resource usage status time-series pattern feature generation secondary subunit is used to perform IT resource usage status dynamic constraint transfer encoding on the time series of the IT resource usage status encoding vectors based on the IT resource usage status time-series transfer regulation factors of the respective IT resource usage status encoding vectors to obtain the IT resource usage status time-series pattern feature vectors.

6. The IT resource management system based on automated operation and maintenance according to claim 5, wherein The positioning reference encoding vector calculation secondary subunit is used for: extracting the last IT resource usage status encoding vector from the time series of the IT resource usage status encoding vectors as the IT resource usage status time-series transfer end reference positioning encoding vector; performing cluster analysis on the time series of the IT resource usage status encoding vectors to obtain the IT resource usage status time-series transfer axial reference positioning encoding vector.

7. The IT resource management system based on automated operation and maintenance according to claim 6, wherein The transfer regulation factor calculation secondary subunit includes: The end regulation factor calculation tertiary subunit is used to calculate the IT resource usage status time-series transfer end regulation factors of the respective IT resource usage status encoding vectors in the time series of the IT resource usage status encoding vectors relative to the IT resource usage status time-series transfer end reference positioning encoding vector; The axial regulation factor calculation tertiary subunit is used to calculate the IT resource usage status time-series transfer axial regulation factors of the respective IT resource usage status encoding vectors in the time series of the IT resource usage status encoding vectors relative to the IT resource usage status time-series transfer axial reference positioning encoding vector; The message regulation factor calculation tertiary subunit is used to calculate the IT resource usage status time-series transfer regulation factors of the respective IT resource usage status encoding vectors based on the IT resource usage status time-series transfer end regulation factors and the IT resource usage status time-series transfer axial regulation factors of the respective IT resource usage status encoding vectors in the time series of the IT resource usage status encoding vectors.

8. The IT resource management system based on automated operation and maintenance according to claim 7, characterized in that, The message regulation factor calculation tertiary subunit is used for: using a canonical space constraint matrix to perform lateral convergence covariance on the IT resource usage status time-series transfer end regulation factor and the IT resource usage status time-series transfer axial regulation factor of the IT resource usage status encoding vector to obtain the IT resource usage status time-series transfer end convergence covariance regulation factor and the IT resource usage status time-series transfer axial convergence covariance regulation factor; performing closed constraint optimization based on the spin space on the IT resource usage status time-series transfer end regulation factor and the IT resource usage status time-series transfer axial regulation factor based on the IT resource usage status time-series transfer end convergence covariance regulation factor and the IT resource usage status time-series transfer axial convergence covariance regulation factor to obtain the optimized IT resource usage status time-series transfer end regulation factor and the optimized IT resource usage status time-series transfer axial regulation factor; The IT resource usage status time-series transfer regulatory factor of the IT resource usage status coding vector is obtained by performing weighted processing on the optimized IT resource usage status time-series transfer end regulatory factor and the optimized IT resource usage status time-series transfer axial regulatory factor and then performing activation processing based on sigmoid.

9. The IT resource management system based on automated operation and maintenance according to claim 8, wherein, The scaling decision generation module is configured to: input the capacity prediction result data and the prediction confidence into an elastic scaling decision engine to obtain the scaling decision.