Machine room comprehensive operation and maintenance and information security management method based on cloud platform

By building a prediction model based on long and short-term memory networks on the cloud platform and dynamically adjusting server resource allocation, the problem of difficult to balance tenant service quality and cost control in the cloud computing environment is solved, and efficient resource utilization and operation and maintenance optimization are achieved.

CN119938328AInactive Publication Date: 2025-05-06WUHAN TIANYI DATA TECH DEV CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510027003.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing computer room operation and maintenance management methods are difficult to find a balance between ensuring the quality of tenants' service and controlling costs. Especially in the complex environment of cloud computing, it is difficult to cope with changes in resource demand during business peaks and troughs.

Method used

By collecting server historical resource usage data, building a prediction model based on long-term and short-term memory networks, forming a resource allocation mechanism, and dynamically adjusting server resource allocation during peak periods to ensure smooth services and avoid resource waste.

Benefits of technology

It realizes accurate prediction of server resource usage requirements, assists in intelligent allocation of resources, ensures stable service during peak periods, reduces costs, plans resources in advance, avoids idle waste, and improves operation and maintenance efficiency and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938328A_ABST
    Figure CN119938328A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine room operation and maintenance management, in particular to a machine room comprehensive operation and maintenance and information security management method based on a cloud platform. According to the technical scheme, the method comprises the steps of collecting historical resource usage data of a server, constructing a prediction model based on a long and short-term memory network, forming a resource allocation mechanism based on a real-time cloud resource usage rate and the prediction model, and dynamically adjusting server resource allocation in a peak period. By constructing a prediction model, the use demand of server resources in a future time period can be accurately predicted, intelligent allocation of the resources can be assisted, the resources are reasonably allocated before a peak comes according to a prediction result, smooth service is ensured, overload and jamming are avoided, cost control and operation and maintenance optimization can be assisted, resources are planned in advance, idle waste is avoided, and the method is suitable for popularization and application. Meanwhile, operation and maintenance personnel can deal with potential problems in advance, and the operation and maintenance efficiency and the system stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer room operation and maintenance management, and in particular to a computer room comprehensive operation and maintenance and information security management method based on a cloud platform. Background Art

[0002] With the widespread application of cloud computing technology, cloud platform computer rooms are carrying massive data processing and storage tasks.

[0003] Focusing on the cloud computing environment, the problem becomes more complicated. The cloud platform carries a wide range of businesses from all walks of life. The peaks and troughs of tenants' businesses are irregular, and resource demands are highly fragmented, sudden and changeable. Cloud service providers must ensure that each tenant can enjoy high-quality services that meet the service level agreement (SLA), and they must also carefully control costs to prevent unnecessary waste of resources. However, the existing extensive management methods are simply unable to find a delicate balance between the two, and often fall into a dilemma. Either they cannot fulfill their promises to tenants due to insufficient resource reserves during business peaks, or they bear a heavy cost burden due to over-allocation of resources during business slack periods.

[0004] Even if some companies try to introduce dynamic allocation mechanisms and set rules based on simple real-time monitoring indicators, such as starting a backup link when the network bandwidth usage rate exceeds 80%, it is still difficult to cope with complex and changing business scenarios. On the one hand, such rules are too simple and crude, and do not fully consider the uniqueness of business types. Different businesses have completely different degrees of dependence and sensitivity on network bandwidth, memory and other resources.

[0005] In addition, the traditional static resource management model only allocates resources based on rough estimates in the past, which causes a large number of resources (such as free memory, underutilized disk space, etc.) to be idle in daily operations, increasing costs. Although the dynamic allocation strategy has been improved, it still has obvious shortcomings: first, the rules are simple and do not take into account the differences in business characteristics. Different businesses have different sensitivity to resources and cannot be accurately adapted; second, it only focuses on real-time monitoring at the moment and ignores business development trends. It is difficult to predict traffic peaks in advance, and resources are allocated only when they are in urgent need. This has caused problems such as freezes and delays, affecting user experience, especially in the complex environment of cloud computing, where it is difficult to balance tenant service quality and cost control. Summary of the invention

[0006] The purpose of the present invention is to address the problems existing in the background technology and to propose a method for comprehensive operation and maintenance of a computer room and information security management based on a cloud platform.

[0007] The technical solution of the present invention is a method for comprehensive operation and maintenance of computer rooms and information security management based on a cloud platform, comprising:

[0008] Collect historical server resource usage data and build a prediction model based on long short-term memory network;

[0009] Based on real-time cloud resource usage and prediction models, a resource allocation mechanism is formed to dynamically adjust server resource allocation during peak hours;

[0010] The prediction model building step includes:

[0011] S1: Collect historical server resource usage data to form a data sequence {X t};

[0012] S2: Extract the data sequence {X t}, calculate the first-order difference sequence {ΔX t}, the second-order difference sequence {Δ 2 X t}, calculate the daily average resource utilization rate by daily aggregation, and construct the feature vector F for input into the long short-term memory network t ;

[0013] S3: The feature vector F t Divided into training set and test set, use the training set to train the long short-term memory network;

[0014] S4: Use the test set to evaluate the trained model.

[0015] Preferably, the daily average resource utilization rate is specifically calculated as follows:

[0016]

[0017] Among them, n day is the number of sampling points included on that day, t i is the i-th sampling time of the day, and k is the resource type index;

[0018] The eigenvector F t , the specific expression is:

[0019]

[0020] Preferably, the long short-term memory network is trained using a training set, and the weights are updated using an adaptive moment estimation optimization algorithm. The specific updating steps include:

[0021] S31: Calculate the first-order moment estimate vector m t =β1m t-1 +(1-β1)g t , where m t is the first-order moment estimation vector, β1 is a preset constant, g t is the current gradient vector of the first-order moment estimation vector;

[0022] S32: Calculate the second-order moment estimate vector Among them, v t is the second-order moment estimation vector, β2 is a preset constant, g t 2 is the current gradient vector of the second-order moment estimation vector;

[0023] S33: Calculate update parameters Among them, θ t is the model parameter at time t, α is the learning rate, and ∈ is a positive number;

[0024] S34: Training with mean square error as loss function until the mean square error on the test set converges to a preset threshold. The mean square error is specifically:

[0025]

[0026] Among them, MSE is the mean square error, N is the number of training samples, and Y i is the true value of the i-th sample, is the predicted value of the i-th sample.

[0027] Preferably, the resource allocation mechanism forming step comprises:

[0028] S100: Based on distributed monitoring technology, several monitoring nodes are deployed on the cloud server to obtain real-time cloud resource usage;

[0029] S200: setting a real-time resource usage rate warning threshold based on the cloud server carrying capacity, and activating a prediction model if the detected real-time resource usage rate reaches the warning threshold;

[0030] S300: The prediction model has a built-in intelligent decision-making module, and a resource allocation mechanism is formed through the intelligent decision-making module based on the output results of the prediction model.

[0031] Preferably, the intelligent decision-making module includes encoding through "IF-THEN" logical statements based on historical computer room operation and maintenance data and industry standards to form a rule base.

[0032] Preferably, the intelligent decision-making module also includes building a classification model based on a supervised learning algorithm by collecting multi-dimensional status information of cloud servers, business operation indicators and historical traffic peak data.

[0033] Compared with the existing technology, the beneficial effects of the present invention are: by constructing a prediction model, the present invention can accurately predict the usage demand of server resources in future time periods, and can also assist in the intelligent allocation of resources. According to the prediction results, resources can be reasonably allocated before the peak arrives to ensure smooth service and avoid overload and jamming. It can also assist in cost control and operation and maintenance optimization, plan resources in advance, avoid idle waste, and allow operation and maintenance personnel to deal with potential problems in advance, thereby improving operation and maintenance efficiency and system stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a flow chart of the method of the present invention;

[0035] Figure 2 Flowchart for constructing the prediction model of the present invention. DETAILED DESCRIPTION

[0036] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0037] See attached Figure 1 and Figure 2 , a computer room comprehensive operation and maintenance and information security management method based on a cloud platform, comprising:

[0038] Step 1: Collect historical server resource usage data and build a prediction model based on the long short-term memory network;

[0039] Step 2: Based on the real-time cloud resource usage and prediction model, a resource allocation mechanism is formed to dynamically adjust server resource allocation during peak hours.

[0040] For step 1, the prediction model construction steps include:

[0041] S1: Collect historical server resource usage data to form a data sequence {X t};

[0042] S2: Extract the data sequence {X t}, calculate the first-order difference sequence {ΔX t}, the second-order difference sequence {Δ 2 X t}, calculate the daily average resource utilization rate by daily aggregation, and construct the feature vector F for input into the long short-term memory network t ;

[0043] S3: The feature vector F t Divided into training set and test set, use the training set to train the long short-term memory network;

[0044] S4: Use the test set to evaluate the trained model.

[0045] And for step S1, in this embodiment, historical server resource usage data is collected, covering key indicators such as CPU usage, memory occupancy, network bandwidth utilization, etc., to form a resource usage data sequence {X t This step is the basis for building a prediction model, because only by obtaining enough and representative historical data can we discover the rules and patterns of server resource usage. For example, by collecting data from e-commerce platform servers during multiple past promotional activities and non-promotional activities, we can provide data support for subsequent predictions of resource requirements during promotional activities.

[0046] And for step S2, the daily average resource utilization rate is specifically calculated as follows:

[0047]

[0048] Among them, n day is the number of sampling points included on that day, t i is the i-th sampling time of the day, and k is the resource type index;

[0049] Eigenvector F t , the specific expression is:

[0050]

[0051] In this embodiment, the daily average resource usage rate is a periodic feature, which helps the model capture the long-term patterns of resource usage, such as the patterns of the server's resource usage peak and trough periods every day of the week, providing an important reference for predicting resource usage in the same period in the future.

[0052] And for step S3, the long short-term memory network is trained using the training set, and the weights are updated using the adaptive moment estimation optimization algorithm. The specific updating steps include:

[0053] S31: Calculate the first-order moment estimate vector m t =β1m t-1 +(1-β1)g t , where m t is the first-order moment estimation vector, β1 is a preset constant, g t is the current gradient vector of the first-order moment estimation vector;

[0054] S32: Calculate the second-order moment estimate vector Among them, v t is the second-order moment estimation vector, β2 is a preset constant, g t 2 is the current gradient vector of the second-order moment estimation vector;

[0055] S33: Calculate update parameters Among them, θ t is the model parameter at time t, α is the learning rate, and ∈ is a positive number;

[0056] S34: Training with mean square error as loss function until the mean square error on the test set converges to a preset threshold. The mean square error is specifically:

[0057]

[0058] Among them, MSSE is the mean square error, N is the number of training samples, and Y i is the true value of the i-th sample, is the predicted value of the i-th sample.

[0059] In this embodiment, the first-order difference sequence can capture the rate of change of resource usage, and the second-order difference sequence is used to highlight the acceleration or deceleration of the rate of change of resource usage. These difference sequences can help the model better understand the dynamic change characteristics of resource usage data, such as whether the server load is growing steadily, growing rapidly, or fluctuating.

[0060] By building a prediction model, we can accurately predict the usage demand of server resources in future periods, and assist in the intelligent allocation of resources. Based on the prediction results, we can reasonably allocate resources before the peak arrives to ensure smooth service and avoid overload and lag. It can also help with cost control and operation and maintenance optimization, plan resources in advance, avoid idle waste, and allow operation and maintenance personnel to respond to potential problems in advance, improving operation and maintenance efficiency and system stability.

[0061] For step 2, the steps for forming a resource allocation mechanism include:

[0062] S100: Based on distributed monitoring technology, several monitoring nodes are deployed on the cloud server to obtain real-time cloud resource usage;

[0063] S200: setting a real-time resource usage rate warning threshold based on the cloud server carrying capacity, and activating a prediction model if the detected real-time resource usage rate reaches the warning threshold;

[0064] S300: The prediction model has a built-in intelligent decision-making module, and a resource allocation mechanism is formed through the intelligent decision-making module based on the output results of the prediction model.

[0065] In this embodiment, the intelligent decision-making module includes encoding through "IF-THEN" logical statements based on historical computer room operation and maintenance data and industry standards to form a rule base.

[0066] The rules cover various common business scenarios, and the corresponding resource allocation actions triggered by the output results of the prediction model and the current system status. For example, when it is predicted that the CPU usage rate will exceed the preset threshold in a specific period of time in the future, and the business type is e-commerce promotion, the rules clearly stipulate the number of backup servers to be awakened first and the specific proportion of new resources allocated to key business processes to ensure that the rules can effectively respond to routine operation and maintenance challenges.

[0067] In addition, the intelligent decision-making module also includes building a classification model based on a supervised learning algorithm by collecting multi-dimensional status information of cloud servers, business operation indicators, and historical traffic peak data.

[0068] The classification model has been trained with a large number of historical samples and has learned the underlying rules between resource allocation and business performance in complex environments. When faced with complex decision-making scenarios that the rule base cannot accurately handle, it assists the reasoning engine in making more practical and forward-looking resource allocation decisions, thereby improving the overall intelligence level of the decision-making module.

[0069] In this embodiment, the intelligent decision-making module may further include:

[0070] Step 1: Build a data collection interface: Develop an interface program that is compatible with a variety of data sources, which can collect resource usage data from the server monitoring system in real time and stably, including CPU usage, memory usage, disk I / O rate, network bandwidth consumption, etc.; obtain business indicators from the business platform, such as the number of transactions, user login frequency, and service request response time; and receive output results from the prediction model, such as the predicted demand value of various resources in the future period. These interfaces are highly compatible and can be compatible with monitoring equipment and software systems of different brands and models, ensuring the comprehensiveness and timeliness of data collection, and controlling data transmission delays at the millisecond level;

[0071] Step 2: Data cleaning and verification: Clean the received data and use statistical analysis methods to identify and remove outliers caused by equipment failures and network anomalies. For example, if the CPU usage rate suddenly soars to nearly 100% at a certain moment and lasts for a very short time, it is judged to be abnormal based on the fluctuation range of the data in the previous and subsequent time periods and corrected. At the same time, perform data verification to compare the consistency of related data between different data sources. For example, the network sending traffic recorded by the server should be similar to the outflow traffic of the corresponding port monitored by the network equipment. If differences are found, the causes should be promptly investigated to ensure the accuracy and reliability of the data, providing a solid foundation for subsequent decision-making;

[0072] Step 3: Data normalization and feature extraction: Use a normalization algorithm to convert data of different dimensions to a unified range to facilitate subsequent calculations and comparisons. For example, percentage data such as CPU usage and memory occupancy are normalized to the 0-1 range, and values ​​such as disk read and write rates and network bandwidth are linearly scaled according to the maximum and minimum values ​​of historical data. On this basis, extract key features, such as the load trend characteristics of computing resources (by taking the first-order derivative of the CPU usage over a period of time to determine whether it is an upward, downward or stable trend), and business busy period characteristics (based on historical business indicators to count daily and weekly business peak and trough periods), and use these feature vectors as inputs to subsequent decision-making models to enhance the pertinence of decisions;

[0073] Step 4: Build a rule-based reasoning engine: Organize industry experts and operation and maintenance personnel to jointly sort out resource allocation rules based on historical experience and best practices, and encode them into a rule base in the form of "I F-THEN". For example, "IF it is predicted that the CPU usage rate will exceed 80% in the next half hour AND the business type is e-commerce promotion THEN start 3 backup servers, give priority to order processing business, and adjust the load balancing strategy so that the new server bears 40% of the new load". The rule base has a dynamic update mechanism, which is updated at least once a quarter according to new operation and maintenance cases and business changes to ensure that the rules always fit the actual operation and maintenance scenarios;

[0074] Step 5: Machine learning model training: Collect a large amount of historical operation and maintenance data and the corresponding resource allocation strategy as training samples, and select appropriate machine learning algorithms, such as decision trees, random forests, neural networks, etc. Taking the decision tree as an example, the preprocessed feature vector is used as input, and the resource allocation strategy is used as output for training, so that the model can learn the complex relationship between different system states and optimal resource configuration. During the training process, methods such as cross-validation are used to optimize model parameters, prevent overfitting, and improve the generalization ability of the model. When the accuracy of the model on the validation set reaches more than 80%, it can be used for actual decision-making. It is used to deal with complex and changeable scenarios that the rule base cannot cover, and to assist the inference engine in making more accurate decisions;

[0075] Step 6: Model fusion and arbitration: organically combine the rule-based reasoning engine with the machine learning model and design an arbitration mechanism. When the decision results of the two are consistent, the result is directly adopted; if there is a disagreement, a comprehensive judgment is made based on business priority, risk assessment and other factors. For example, for high-priority businesses such as real-time financial transactions, decisions based on empirical rules are given priority because their stability has been verified by long-term practice; for exploratory new businesses or non-critical businesses, the innovative decisions of the machine learning model can be referred to, giving new business models more opportunities to try. Through this fusion and arbitration, the scientificity and flexibility of decision-making can be improved.

[0076] Its specific effect is to provide a standardized decision-making framework. The procedure library contains a large number of resource allocation rules covering various common business scenarios and resource usage conditions in the form of "IF-THEN" logical statements. These rules are based on industry best practices, past operation and maintenance experience, and in-depth analysis and summary of system performance bottlenecks. For example, "IF the current business type is e-commerce promotion and it is predicted that the network bandwidth utilization rate will exceed 80% in the next half hour, THEN urgently call the third-party CDN service to increase the edge node cache capacity by 50%, and limit the network bandwidth of non-critical businesses (such as automatic update of product recommendation pictures) to 20% of the original." It provides a standardized and followable decision-making framework for the intelligent decision-making module when facing complex and changeable real-time situations, ensuring the timeliness and rationality of decisions and avoiding decision confusion;

[0077] In addition, when the system encounters an emergency, such as hardware failure of some servers, DDoS attacks on the network, etc., the emergency rules in the procedure library will take effect quickly. For example, "IF it is detected that more than 10% of the servers in the server cluster are faulty THEN immediately start the backup data center, migrate key businesses (such as financial core transactions, medical remote diagnosis) to the backup center in order of priority, and complete the switch within 10 minutes, and notify the operation and maintenance personnel to urgently repair the faulty server." By planning the response strategy in advance, the basic functions of the system can be restored in the shortest time, business continuity can be guaranteed, service stability can be maintained, and the negative impact of emergencies on user experience can be reduced.

[0078] In the description of the present invention, it should be understood that the above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0079] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0080] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0081] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0082] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0083] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0084] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0085] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0086] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0087] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.

[0088] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for comprehensive operation and maintenance of computer rooms and information security management based on a cloud platform, characterized in that: include: Collect historical server resource usage data and build a prediction model based on long short-term memory network; Based on real-time cloud resource usage and prediction models, a resource allocation mechanism is formed to dynamically adjust server resource allocation during peak hours; The prediction model building step includes: S1: Collect historical server resource usage data to form a data sequence {X t }; S2: Extract the data sequence {X t }, calculate the first-order difference sequence {ΔX t }, the second-order difference sequence {Δ 2 X t }, calculate the daily average resource utilization rate by daily aggregation, and construct the feature vector F for input into the long short-term memory network t ; S3: The feature vector F t Divided into training set and test set, use the training set to train the long short-term memory network; S4: Use the test set to evaluate the trained model.

2. According to the cloud platform-based computer room comprehensive operation and maintenance and information security management method of claim 1, it is characterized in that: The specific calculation method of the daily average resource utilization rate is: Among them, n day is the number of sampling points included on that day, t i is the i-th sampling time of the day, and k is the resource type index; The eigenvector F t , the specific expression is:

3. According to the cloud platform-based computer room comprehensive operation and maintenance and information security management method of claim 1, it is characterized in that: The long short-term memory network is trained using a training set, and the weights are updated using an adaptive moment estimation optimization algorithm. The specific updating steps include: S31: Calculate the first-order moment estimate vector m t =β1m t-1 +(1-β1)g t , where m t is the first-order moment estimation vector, β1 is a preset constant, g t is the current gradient vector of the first-order moment estimation vector; S32: Calculate the second-order moment estimate vector Among them, v t is the second-order moment estimation vector, β2 is a preset constant, g t 2 is the current gradient vector of the second-order moment estimation vector; S33: Calculate update parameters Among them, θ t is the model parameter at time t, α is the learning rate, and ∈ is a positive number; S34: Training with mean square error as loss function until the mean square error on the test set converges to a preset threshold. The mean square error is specifically: Among them, MSE is the mean square error, N is the number of training samples, and Y i is the true value of the i-th sample, is the predicted value of the i-th sample.

4. According to the cloud platform-based computer room comprehensive operation and maintenance and information security management method of claim 1, it is characterized in that: The resource allocation mechanism forming step comprises: S100: Based on distributed monitoring technology, several monitoring nodes are deployed on the cloud server to obtain real-time cloud resource usage; S200: setting a real-time resource usage rate warning threshold based on the cloud server carrying capacity, and activating a prediction model if the detected real-time resource usage rate reaches the warning threshold; S300: The prediction model has a built-in intelligent decision-making module, and a resource allocation mechanism is formed through the intelligent decision-making module based on the output results of the prediction model.

5. According to claim 4, a cloud platform-based computer room comprehensive operation and maintenance and information security management method is characterized in that: The intelligent decision-making module includes encoding through "IF-THEN" logical statements based on historical computer room operation and maintenance data and industry standards to form a rule base.

6. A cloud platform-based computer room integrated operation and maintenance and information security management method according to claim 5, characterized in that: The intelligent decision-making module also includes building a classification model based on a supervised learning algorithm by collecting multi-dimensional status information of cloud servers, business operation indicators and historical traffic peak data.

Citation Information

Patent Citations

  • Power load prediction method based on improved long-term and short-term memory network

    CN111985719A

  • Cloud service resource dynamic prediction management method

    CN117076882A

  • Edge gateway resource optimization management method based on artificial intelligence

    CN118612065A

  • Task-oriented machine learning and a configurable tool thereof on a computing environment

    US20210397941A1