Operating system updating method based on cloud service
By scanning cloud host system information in the government affairs network, building mirror servers, deploying monitoring agents and building operation and maintenance knowledge bases, the complex patch management problems caused by inconsistent versions of cloud host operating systems are solved, intelligent classification and automatic update of cloud host patches are realized, real-time monitoring and rapid handling of service abnormalities, and the security and stability of cloud platform are improved.
Patent Information
- Application Number
- CN202411992030.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-06-03
AI Technical Summary
The operating system version of cloud host in the government network is not unified, resulting in complex patch management and the inability to directly connect to Microsoft's official update server. It is necessary to build an internal image update server, and operation and maintenance personnel need to deeply understand the technical details of the operating system and services to achieve an automated and intelligent operation and maintenance mechanism.
By scanning the system information of cloud hosts, dividing management packets, and formulating a unified patch strategy and update plan. Build a mirror server, use machine learning clustering algorithm to automatically classify patch files, extract key features and train intelligent classification models. Deploy monitoring agents, collect service operation data in real time, and use machine learning to train anomaly detection models. Build an operation and maintenance knowledge base, summarize operation and maintenance specifications and processes, and make correlation and reasoning through knowledge graph technology to form an intelligent operation and maintenance decision support system.
It realizes intelligent classification and automatic update of cloud host patches, and real-time monitoring and rapid handling of service abnormalities, improving the security and stability of the cloud platform and reducing operation and maintenance costs.
Smart Images

Figure CN120085885A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular, to an operating system update method based on cloud services. Background Art
[0002] Problem Background:
[0003] In the data center of the government network, the operating system versions of cloud hosts are uneven. Some are Windows Server 2008, some are Windows Server 2012, and some are Windows Server 2016. For different operating systems, the required patches are not the same. How to establish a unified patch management mechanism so that different versions of operating systems can obtain the latest security patches in a timely manner is a difficult problem faced by operation and maintenance personnel.
[0004] In addition, due to the relatively closed network environment of the government network, cloud hosts cannot directly connect to Microsoft's official update server. Therefore, it has become an urgent task to build a mirror update server within the government network. However, the update files required for different versions of operating systems vary greatly. How to establish an automatic classification and intelligent recognition mirror source so that different operating systems can find the corresponding patch files requires operation and maintenance personnel to deeply study the internal structure of the operating system and find the commonalities and differences between various versions.
[0005] At the same time, various services running on cloud hosts also need to be monitored in real time by operation and maintenance personnel. Different services have different log formats and different running characteristics. Operation and maintenance personnel need to develop an intelligent log analysis system that can accurately detect abnormal behaviors of various services and give early warnings in a timely manner. And this requires operation and maintenance personnel to have an in-depth understanding of the business logic of various services and know which behaviors are normal and which are abnormal. In short, in the operation and maintenance work of the government network, operation and maintenance personnel need to deeply understand the technical details of the operating system and services, find the commonalities in various heterogeneous environments, and then establish an automated and intelligent operation and maintenance mechanism to ensure the safe and stable operation of the government network. Summary of the Invention
[0006] The present invention provides an operating system update method based on cloud services, mainly including:
[0007] By scanning the system information of cloud hosts, obtaining the operating system versions and patch installation status of different hosts, and dividing the hosts into corresponding management groups according to the operating system type and version number, formulating a unified patch policy and update plan for each group;
[0008] Build a mirror server within the government affairs network, download the security update files required for each operating system version from a trusted patch source, analyze the key attributes such as the applicable systems and version numbers of the patches by extracting the metadata information of the update packages, and use the clustering algorithm of machine learning to automatically classify the patch files into the corresponding directories;
[0009] Extract the key features from the patch files in the mirror server and train an intelligent classification model. When new patch files are added, the model can accurately identify the applicable scope of the patches and automatically place them into the corresponding directories to ensure that matching update files can be found for different operating system versions;
[0010] Deploy a monitoring agent on the cloud host to collect the running metrics and log information of each service process in real time, aggregate the collected data to a centralized log analysis platform, and parse and mine the logs through a batch computing framework to build a behavior model of service operation;
[0011] For the data in the log analysis platform, use machine learning algorithms and train an anomaly detection model through supervised learning. The model can automatically discover various abnormal behaviors during service operation from a large amount of log data and generate alarm information;
[0012] When the monitoring system detects an anomaly in a certain service, according to the pre-configured alarm rules, automatically notify the operation and maintenance personnel of the anomaly information, and at the same time trigger an automated emergency response process. Depending on the severity and impact scope of the anomaly, take different handling measures such as restarting the service process and rolling back the patch version;
[0013] In the operation and maintenance knowledge base, summarize and generalize information such as the configuration parameters, log formats, and common faults of different operating systems and services to form a set of standardized operation and maintenance specifications and processes. Through knowledge graph technology, associate and reason about this knowledge to form an intelligent operation and maintenance decision support system to provide diagnostic and handling suggestions for operation and maintenance personnel.
[0014] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0015] The present invention discloses an intelligent cloud host patch management and anomaly detection method. By scanning the system information of cloud hosts, the hosts are divided into management groups according to the operating system type and version, and a unified patch policy is formulated. An image server is built inside the government network, and machine learning clustering algorithms are used to automatically classify patch files. Monitoring agents are deployed to collect service operation data, and machine learning is used to train an anomaly detection model to discover anomalies and give alarms from a large amount of logs. Combining knowledge graph technology to build an operation and maintenance decision support system to provide diagnostic suggestions for operation and maintenance personnel. The present invention realizes the intelligent classification and automatic update of cloud host patches, as well as the real-time monitoring and rapid disposal of service anomalies, improves the security and stability of the cloud platform, and reduces the operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flowchart of an operating system update method based on cloud services according to the present invention.
[0017] Figure 2 It is a schematic diagram of an operating system update method based on cloud services according to the present invention.
[0018] Figure 3 It is another schematic diagram of an operating system update method based on cloud services according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] As Figures 1-3 , an operating system update method based on cloud services in this embodiment may specifically include:
[0021] S101. By scanning the system information of cloud hosts, obtain the operating system versions and patch installation conditions of different hosts, and divide the hosts into corresponding management groups according to the operating system type and version number, and formulate a unified patch policy and update plan for each group.
[0022] Remotely connect to the cloud host via the API interface or SSH protocol, execute system commands, and obtain system information such as the operating system type, version number, and patch installation list of the cloud host. According to the obtained cloud host system information, use the decision tree algorithm and, in accordance with the preset operating system type and version number rules, determine the management group to which each cloud host belongs. In the configuration management database, update the field of the management group to which it belongs based on the unique identifier of the cloud host, and divide the cloud host into the corresponding group. For each management group, the operation and maintenance personnel formulate a unified system patch policy according to the operating system characteristics of the group, determine the list of patches to be installed and the installation priority. Convert the patch policy into an executable script program, and use automated operation and maintenance tools such as Ansible to batch issue patch installation tasks to all cloud hosts within the group. During the patch installation process, obtain the installation progress and results of each cloud host in real time and record them in the system log for easy tracking and auditing. After the patch installation is completed, scan the system information of the cloud host again to obtain the latest patch installation status, and determine whether it is consistent with the patch policy. If not, perform differential updates to ensure that the cloud hosts within each group meet the requirements of the patch policy.
[0023] Exemplarily, remotely connecting to a cloud host to obtain system information is the basis of cloud platform management. Taking a certain cloud platform as an example, the virtual machine list can be obtained by calling through the RESTful API interface, and then the SSH protocol is used to log in to each virtual machine to execute system commands. Commonly used commands include "uname -a" to obtain the operating system type and version, and "rpm -qa" or "dpkg -l" to list the installed software packages. This information can be used for subsequent grouping and patch management. The decision tree algorithm can efficiently group cloud hosts. For example, the operating system type (Windows / Linux) can be judged first, then the specific distribution (CentOS / Ubuntu, etc.) can be judged, and finally the major version number (such as CentOS 6 / 7) can be judged. This hierarchical judgment can quickly classify cloud hosts into appropriate management groups. The advantage of the decision tree is that the rules are clear and the execution efficiency is high, which is convenient for subsequent maintenance and adjustment of the grouping strategy. The Configuration Management Database (CMDB) is the core of IT asset management. In the CMDB, each cloud host has a unique identifier, such as an asset number or an instance ID. When updating the group to which a cloud host belongs, only need to locate the corresponding record according to this identifier and modify its "management group" field. This centralized management method can ensure the consistency and traceability of asset information. When formulating a patch strategy, multiple factors need to be considered. Taking the Linux system as an example, the patch strategy for CentOS 7 may include: installing all security updates, installing kernel patches of specific versions, excluding some patches that may affect the business, etc. The strategy should also consider the installation order, such as installing dependent packages first and then the main patches. This meticulous strategy formulation can maximize the security and stability of the system. Converting the patch strategy into an executable script is the key to automation. Taking Ansible as an example, a playbook in YAML format can be written to define the sequence of tasks to be executed. For example, the yum module is used to install the specified software package, and the command module is used to execute a specific shell command. The advantage of Ansible lies in its declarative syntax and idempotency, and the same playbook can be executed multiple times without causing system state disorders. Real-time monitoring of the patch installation progress is crucial for large-scale deployments. Callback functions can be added to Ansible tasks to write the execution results of each step into the log system in real time. Commonly used log systems such as the ELK stack (Elasticsearch, Logstash, Kibana) can provide powerful log aggregation and visualization capabilities, facilitating operation and maintenance personnel to grasp the patch deployment situation in real time. Verification after patch installation is a necessary step to ensure that the strategy is implemented in place. The system information collection script can be executed again to compare the software package lists and version numbers before and after installation. If differences are found, a difference report can be generated, and additional patch installation or rollback operations can be decided according to the report content. This closed-loop management can effectively improve the accuracy and integrity of patch management.The automation of the entire process not only improves efficiency but also reduces the risk of human errors. For example, manual grouping may lead to incorrect classification of hosts with certain special configurations, while using algorithms can ensure the consistency of classification. Automated patch deployment can also ensure that all hosts are updated according to the same process and standards, avoiding system configuration inconsistencies caused by human negligence. However, automation also brings new challenges. For example, how to handle unexpected situations that may occur during the patch installation process, such as network interruptions or insufficient disk space. This requires adding more error handling and retry mechanisms to the script. In addition, for critical business systems, it may be necessary to add manual confirmation steps to the automation process to balance efficiency and security. Generally speaking, this method of automated grouping and patch management based on system information can greatly improve the management efficiency and consistency of large-scale cloud environments. By transforming human experience into algorithms and automated scripts, not only can repetitive work be reduced, but also the precise execution of management strategies can be ensured. At the same time, a complete logging and verification mechanism also provides a basis for subsequent auditing and optimization.
[0024] Obtain the operating system version and patch installation status of cloud hosts through scanning, divide management groups according to system types and version numbers, formulate unified patch strategies and update plans for the groups, and determine the patch update solutions for each cloud host.
[0025] Conduct a comprehensive scan of cloud hosts through a scanning tool to obtain the operating system type, version number, and installed patch information of each cloud host. According to the obtained operating system type and version number, use clustering algorithms to automatically group the cloud hosts, and the cloud hosts within each group have similar operating system characteristics. For each group, analyze the patch installation status of the cloud hosts within the group, identify existing patch deficiencies and vulnerability risks, and determine the patch update priorities according to the risk levels. Combine the cloud host grouping information and patch priorities to formulate unified patch strategies and update plans for each group, clarify the time nodes and specific operation steps of patch updates. When formulating patch strategies, use association rule mining algorithms to analyze the dependencies and compatibilities between different patches to ensure the rationality and security of patch updates. According to the unified patch strategies and update plans, generate personalized patch update solutions for each cloud host, clarify the list of patches to be installed and the specific operation procedures. Use automated operation and maintenance tools to batch update and repair patches for cloud hosts according to the patch update solutions, and ensure the smoothness and controllability of the patch update process through monitoring and rollback mechanisms.
[0026] Exemplarily, cloud host scanning is the basis of patch management. Through professional scanning tools such as Nessus or OpenVAS, comprehensive information about cloud hosts can be obtained. For example, the scan results may show that a certain host is running CentOS 7.6, has installed the kernel-3.10.0-957.el7 patch, but lacks the latest security updates. These detailed information provide a basis for subsequent grouping and policy formulation. Clustering algorithms can automatically group similar hosts. For instance, using the K-means algorithm, with the operating system type and version as features, hosts can be divided into groups such as Windows Server 2016 group, Ubuntu 18.04 group, etc. This grouping method makes subsequent patch management more targeted and efficient. Patch analysis is a key step in identifying risks. By comparing the installed patches with the latest released patches, potential vulnerabilities can be discovered. For example, a certain Windows Server group lacks the MS17-010 patch and has the high-risk EternalBlue vulnerability, and should be given the highest update priority. This risk-based priority ranking ensures that the most critical security issues are resolved in a timely manner. Formulating a patch policy requires comprehensive consideration. In addition to security, business continuity also needs to be considered. For example, for critical business systems, updates may need to be carried out during off-peak hours and rollback time should be reserved. Specifically, the database server may be scheduled for patch updates at 2-4 am on weekends and a 2-hour observation period is set aside. This meticulous planning can minimize the impact on the business. Analyzing the dependency relationships between patches is crucial. Through association rule mining, rules such as the.NET Framework update must be installed prior to certain application patches can be discovered. This analysis can avoid system instability caused by patch conflicts and improve the update success rate. Personalized patch solutions take into account the characteristics of each host. For example, for web servers, in addition to regular system patches, special attention needs to be paid to the security updates of Apache or Nginx. For database servers, on the other hand, the patches of Oracle or MySQL need to be focused on. This customized solution ensures that each host receives the most suitable updates. Automation tools such as Ansible greatly improve the efficiency of patch updates. By writing playbooks, batch and orderly patch installations can be achieved. For example, non-critical servers can be set to be updated first, and after observing for a period of time, the core business servers can be updated. At the same time, by setting checkpoints and rollback mechanisms, if a certain patch causes system instability, it can be quickly restored to the state before the update, ensuring the controllability of the whole process. This series of steps forms a complete patch management closed-loop. From scanning, analysis, policy formulation to execution and monitoring, each link is carefully designed and supports each other. This systematic method not only improves the security of cloud hosts, but also optimizes the efficiency and reliability of the entire patch management process. By continuously improving this process, the overall security level and operation and maintenance quality of the cloud environment can be continuously enhanced.
[0027] S102. Build a mirror server within the government affairs network, download the security update files required for each operating system version from a trusted patch source, analyze the key attributes such as the applicable systems and version numbers of the patches by extracting the metadata information of the update packages, and automatically classify the patch files into the corresponding directories using the clustering algorithm of machine learning.
[0028] Obtain the available server resources within the government affairs network, and determine the hardware configuration and network environment for building the mirror server. Obtain the security update files for each operating system version from a trusted patch source and download them to the specified storage location of the mirror server. For the downloaded patch files, extract their metadata information to obtain the key attributes such as the applicable systems and version numbers of the patches. According to the extracted key attributes of the patches, use the K-means clustering algorithm to automatically classify the patch files according to the applicable systems and version numbers. Based on the classification results obtained by the clustering algorithm, determine the operating system and version to which each patch file belongs and move it to the corresponding directory on the mirror server. Configure a Web service such as Nginx on the mirror server to provide the download service for the patch files, ensuring that each system within the government affairs network can access and download the required security updates. Continuously monitor the updates of the trusted patch source, automatically download the newly released patch files through scheduled tasks, and repeat the above classification and release processes to keep the patch library on the mirror server synchronized with the trusted source.
[0029] Exemplarily, building a mirror server within the e-government network is a crucial step to ensure timely, reliable and secure system updates. First, it is necessary to evaluate the network environment and hardware resources and select an appropriate server configuration. For example, a server with a dual-way Intel Xeon processor, 128GB of memory, and 10TB of storage space, equipped with a gigabit network interface, can be selected to meet the storage requirements of a large number of patch files and high-concurrency download needs. Obtaining security update files from trusted patch sources is the basis for ensuring patch reliability. Patch files can be synchronized from official sources through tools such as rsync, such as Windows Update, Red Hat Satellite, etc. The downloaded patch files need to be initially classified according to the operating system type and version and stored in the specified directory of the mirror server. Extracting patch metadata information is the key to achieving automated classification. Scripts can be used to parse the description information of patch files and extract attributes such as the applicable operating system, version number, release date, etc. For example, for Windows patches, the XML description of the.msu file can be parsed; for Linux patches, the metadata of RPM or DEB packages can be parsed. Using the K-means clustering algorithm to automatically classify patches is an effective method to improve efficiency. First, convert the key attributes of the patches into numerical vectors. For example, the operating system type can be mapped to different numerical values. Then, set the number of clustering centers, such as the total number of versions of Windows and Linux. Through iterative calculations, similar patches are grouped into the same category. This method can effectively process a large number of patch files and automatically identify new operating system versions. According to the clustering results, the patch files can be moved to the corresponding directory structure. For example, directories such as Windows / Server2016 and Linux / CentOS7 can be created, and the clustered patch files can be stored according to categories for subsequent management and download. Configuring a Web service is the key to providing a patch download service for the internal systems of the e-government network. Nginx can be used as a Web server to configure virtual hosts and access controls, allowing only internal IPs of the e-government network to access. At the same time, caching and compression functions can be enabled to improve download efficiency. To ensure security, HTTPS can be configured and a self-signed certificate can be used to encrypt the transmission. Continuously monitoring the patch source and automatically updating is an important means to maintain the timeliness of the mirror server. A scheduled task can be written to synchronize new patch files from the official source every early morning. After the synchronization is completed, the patch classification and release process are automatically triggered to ensure that new patches can be provided to the internal systems of the e-government network in a timely manner. This automated mechanism can greatly reduce manual intervention and improve the efficiency and accuracy of patch management. Through the above steps, an efficient and reliable patch mirror service can be built within the e-government network to provide timely security updates for each system and effectively enhance the overall network security. This local mirror service can not only accelerate the patch distribution speed but also effectively control the patch source, reduce the dependence on the external network, and is an important safeguard measure for the security operation and maintenance of the e-government network.
[0030] S103. Extract the key features of the patch files in the mirror server and train an intelligent classification model. When new patch files are added, the model can accurately identify the applicable scope of the patches and automatically place them in the corresponding directories to ensure that matching update files can be found for different operating system versions.
[0031] Obtain all the patch file metadata in the mirror server, parse the file metadata to obtain the file name, size, hash value, and the associated information list, and get the preliminary classification set for the patch files of different operating systems. Determine the operating system field and version field according to the associated information list of each patch file in the preliminary classification set, extract the applicable operating system version information of the patch files, and label the patch file set according to the version information to obtain the sample data. Parse the sample data to generate a binary data stream, conduct numerical statistics based on the binary data stream. For example, for the byte frequency, if the occurrence frequency of a specific byte conforms to the normal distribution, then calculate the file difference feature vector through the statistical feature value construction method. Adopt the support vector machine algorithm, combine with the extracted file difference feature vector, and train to obtain a supervised learning classifier. According to the labeled data, perform iterative calculations repeatedly. If the error value between the predicted value and the marked value continues to decrease, then finally converge to determine the accuracy of the training result. Obtain the newly added patch files, generate a binary data stream, conduct numerical statistics based on the binary data stream, and construct the file difference feature vector of the new patch files according to the input feature value requirements defined by the classifier. Input the file difference feature vector of the new patch files into the supervised learning classifier to obtain the classification of the applicable scope of the files, get the classification result, and determine the specific operating system version directory where the patch should be placed. According to the classification result, migrate the newly added patch files to the specific classification directory in the server, update the patch file index, and if the content of the download list of the patch files associated with the operating system is completely refreshed, then all online users of the server are automatically synchronized.
[0032] Exemplarily, to obtain all patch file metadata in the mirror server, it is first necessary to traverse the server storage directory and use the file system API to read the metadata information of each patch file. For example, use the `stat` command in the Linux system or the `GetFileAttributesEx` function in the Windows system to obtain basic information such as file name, size, and hash value. The hash value can be calculated through the SHA-256 algorithm to ensure the integrity and uniqueness of the file. The associated information list includes the operating system type, version number, etc. applicable to the patch. These information are usually stored in the metadata of the patch file or the attached description file. After parsing the file metadata, the patch files are initially classified according to the operating system field and version field. For example, the patch files for the Windows system may contain fields such as "Windows10" and "WindowsServer2016", while the patch files for the Linux system may contain fields such as "Ubuntu20.04" and "CentOS7". These key information are extracted through regular expression matching or string parsing methods to form a preliminary classification set. Based on the preliminary classification set, the operating system version information applicable to the patch file is further extracted. For example, if a patch file metadata contains "Applicable to Windows10 version 1809", then "Windows10" and "1809" are extracted as key information. According to these information, the patch file set is labeled to form sample data. The sample data can be represented in the form of {(file name, size, hash value, operating system, version number)}. Parse the sample data to generate a binary data stream, and analyze the characteristics of the file by counting the byte frequencies in the binary data stream. For example, count the frequency of each byte value in the file. If the frequency distribution of a specific byte conforms to the normal distribution, then these statistical characteristic values can be used to construct a file difference feature vector. The feature vector can contain multiple dimensions, such as the number of occurrences of high-frequency bytes, the variance of byte distribution, etc. Adopt the support vector machine (SVM) algorithm, combined with the extracted file difference feature vector, to train a supervised learning classifier. During the training process, through repeated iterative calculations, the parameters of the classifier are continuously adjusted to make the error value between the predicted value and the marked value continuously decrease. When the error value converges below a certain threshold, it is considered that the training result reaches a high accuracy rate. For example, after 1000 iterations, the error value drops from 0.5 to 0.01, indicating that the classifier has learned the characteristics of the sample data well. After obtaining the newly added patch file, similarly generate a binary data stream and perform byte frequency statistics. According to the input feature value requirements defined by the classifier, construct the file difference feature vector of the new patch file. For example, if the byte frequency statistics result of the new patch file shows that the frequency of a specific byte is highly similar to a certain category in the training sample, then this feature vector is input into the trained classifier.Based on the results output by the classifier, determine the operating system and version applicable to the new patch file. For example, if the classifier output is "Windows 10 version 1903", then migrate the patch file to the corresponding "Windows 10 / 1903" directory on the server. At the same time, update the patch file index to ensure that each system within the e-government network can obtain the latest patch information in a timely manner. Finally, determine whether the content of the patch file download list associated with the operating system has been fully refreshed. For example, by checking the cache update status of the patch management tools of each system, ensure that all online users can automatically synchronize the latest patch files. This not only improves the automation level of patch management but also ensures the security and stability of the internal systems of the e-government network. Through the above steps, using machine learning and statistical methods, the automatic classification and efficient management of patch files are achieved, enhancing the security protection ability of the e-government network. Each step is closely linked, forming a strict logical relationship and thinking chain, ensuring the accuracy and timeliness of patch management.
[0033] S104. Deploy a monitoring agent on the cloud host to collect the running metrics and log information of each service process in real time, aggregate the collected data into a centralized log analysis platform, and parse and perform data mining on the logs through a batch computing framework to construct a behavior model of service operation.
[0034] Deploy an agent on the cloud host to collect the running metric values and log streams of each process in each service process of the cloud host and upload the data. According to the preset log stream data format specification, the log analysis platform receives the data uploaded by the agent and performs data classification operations on the log streams generated by different service processes, obtaining different identifiers for different classifications. Through the data content of each field in the log stream, the parser module in the analysis platform performs information extraction operations to obtain the key log information, and determines the identifiers of different service processes for the logs of different service processes. Use a pre-trained machine learning model to perform anomaly detection operations on different processes, input the key log information and identifiers into a binary classification model to obtain the classification results of whether each process is normal or abnormal. Obtain the running metric values of each process, construct a set of running metric values for each process, input the set of running metric values into a pre-trained regression model, determine the output of the regression model, and obtain the performance prediction results of each process. According to the output results of the regression model and the running metric values, perform statistical analysis operations on the load status of each service process in the service process set, and obtain a behavior sequence by comparing the load conditions of processes with similar behaviors in historical data. Construct a behavior library based on the behavior sequence and the current service process behavior, represent the behavior sequence using a Markov model, and determine the next behavior of the service process.
[0035] Exemplarily, the deployment agent collects the running metric values and log streams of each service process on the cloud host and uploads them to the log analysis platform. Suppose there is a cloud host running three service processes: a web server, a database server, and a cache server. The agent will monitor the running metrics such as CPU usage, memory occupancy, and network traffic of these processes in real time and collect the log information generated by the processes. The log stream data format specification is preset to the JSON format, which includes fields such as timestamp, process ID, log level, and log content. For example, the logs of the web server may contain information such as the URL of the HTTP request and the response time; the logs of the database server may contain SQL query statements and execution times; the logs of the cache server may contain cache hit rates and expiration times. After the log analysis platform receives the data uploaded by the agent, it first classifies the log streams generated by different service processes. By parsing the process ID field in the log, the log streams are divided into three categories: web server logs, database server logs, and cache server logs, and different identifiers are assigned. Next, the parser module performs information extraction operations on each type of log. For example, the URL and response time are extracted from the web server logs, the SQL statement and execution time are extracted from the database server logs, and the cache hit rate and expiration time are extracted from the cache server logs. These key information will be used for subsequent anomaly detection and performance prediction. Anomaly detection is performed using a pre-trained machine learning model. Suppose a binary classification model based on decision trees is used. The key log information and identifiers are input, and the normal or abnormal classification results of each process are output. For example, if the response time of the web server suddenly increases, the model may mark it as abnormal; if the SQL query execution time of the database server is abnormally long, the model will also mark it as abnormal. The running metric values of each process are obtained, and a set of running metric values for each process is constructed. For example, the set of running metric values of the web server may include CPU usage, memory occupancy, network traffic, etc. These metric values are input into a pre-trained regression model to predict the performance of each process. Suppose a regression model based on random forest is used, and the model will predict the performance metrics for a period of time in the future based on historical data. According to the output results of the regression model and the current running metric values, the load status of each service process is calculated. For example, by comparing the current CPU usage of the web server with the load conditions of processes with similar behaviors in historical data, it can be determined whether the current load status of the web server is high, medium, or low. By comparing the load conditions of processes with similar behaviors in historical data, a behavior sequence is obtained. For example, historical data shows that whenever the CPU usage of the web server exceeds 80%, the query execution time of the database server will increase significantly. This behavior sequence can be represented by a Markov model to predict the next behavior of the service process. A behavior library is constructed based on the behavior sequence and the current behavior of the service process.For example, if the CPU usage of the current web server has reached 85%, according to the Markov model in the behavior library, it can be predicted that the query execution time of the database server may increase, and thus optimization measures can be taken in advance, such as increasing the cache or optimizing SQL queries. Through these steps, not only can the log stream be monitored and classified in real time, but also the performance and abnormal behaviors of processes can be predicted, and intervention can be carried out in advance to ensure the stability and efficiency of the system. Such technical effects include improving system availability, reducing fault response time, optimizing resource allocation, etc. For example, in practical applications, when it is detected that the response time of the web server increases abnormally, the system can automatically trigger an expansion operation to increase server instances, thereby alleviating the load pressure; when the SQL query execution time of the database server is abnormally long, the system can automatically optimize the query statement or add indexes to improve query efficiency. Through this multi-dimensional monitoring and prediction mechanism, the service processes on the cloud host can run in an efficient and stable environment, greatly improving the overall performance of the system and the user experience.
[0036] S105. For the data in the log analysis platform, use machine learning algorithms and train an anomaly detection model through supervised learning. This model can automatically discover various abnormal behaviors during the service operation from a large amount of log data and generate alarm information.
[0037] Construct a log dataset based on the log content. This log dataset should cover various log entries in historical data records, where normal and abnormal log entries are marked with data labels. Perform data preprocessing on the log dataset with data labels, extract features, obtain a numerical dataset, and divide it into a training set and a test set. The sample proportions of different service states in the training set should be balanced. Train a model of the support vector machine algorithm using the training set to obtain an initial detection model. Test the test set using the initial detection model and calculate various metrics for the detection results of the initial detection model on the test set. Obtain the log content generated in the real environment and perform data preprocessing, and obtain the previously trained model of the time series as a benchmark detection model. Use the benchmark detection model to process these log data to generate preliminary prediction labels, and then perform label determination. Perform label determination based on the preliminary prediction labels. For a preliminary prediction label, if a log record has a preliminary prediction label of abnormal, and at the same time the time and log content of this log record match the log data of an abnormal record in the existing marked database, then determine the final prediction label as abnormal. Perform abnormal determination on the final prediction labels. If a log has a final prediction label of abnormal, then judge the alarm level corresponding to this log and use different-level alarm schemes for alarming. Record the final marking results in the marked database, and then continue the determination. If it is determined that all final prediction labels are normal, then it is determined that there is a defect in the detection model. Judge the false alarm quantity ratio in the marked database. If it exceeds the set ratio, then take out all abnormal label samples in the marked database for retraining to obtain a detection model of the new time series, and then redeploy the model.
[0038] Exemplarily, to construct a log dataset based on log content, it is first necessary to extract various log entries from historical data, including normal and abnormal records. For example, in the order processing system of an e-commerce platform, normal logs may include entries such as "Order created successfully" and "Payment completed", while abnormal logs may include "Payment failed" and "Insufficient inventory". Each log record needs to be labeled with a data tag, such as "normal" or "abnormal". In the data preprocessing stage, the log data is cleaned and feature extracted. Suppose a log record is "2023-10-01 10:00:00 User 123 payment failed", the features extracted after preprocessing may include timestamp, user ID, operation type, and result status, etc. These features are converted into numerical data, such as the timestamp is converted into a Unix timestamp, and the operation type and result status are one-hot encoded. The preprocessed data is divided into a training set and a test set to ensure that the sample ratio of different service states in the training set is balanced. For example, the ratio of normal logs to abnormal logs in the training set is 1:1 to avoid the model being biased towards a certain type of data. The support vector machine (SVM) algorithm is used for model training. Suppose the training set contains 10,000 log records, the SVM model learns the features of these records to establish a classification boundary. After training, the test set is used to verify the model, and metrics such as accuracy, recall, and F1 score are calculated. For example, the test set contains 2,000 records, the model accuracy is 90%, the recall is 85%, and the F1 score is 0.875. In a real environment, new log data is continuously collected and preprocessed. Suppose 1,000 new logs are collected one day, and the previously trained baseline detection model is used for processing to generate preliminary prediction labels. If a certain log is preliminarily predicted as abnormal and its time and content match an abnormal record in the labeled database, such as "2023-10-02 11:00:00 User 456 payment failed", then the final prediction label is confirmed as abnormal. Abnormal determination is performed for the final prediction label. If a certain log is finally predicted as abnormal, the alarm level is judged according to its severity. For example, payment failure may trigger a medium-level alarm, while system crash triggers a high-level alarm. The alarm information is notified to relevant personnel via email, SMS, etc., and the final labeled results are recorded in the labeled database. If the proportion of false alarm quantities in the labeled database exceeds the set threshold, such as the false alarm rate reaches 20%, then the model needs to be retrained. All abnormal label samples are extracted from the database, combined with new data to retrain the SVM model, generate a new time series detection model, and redeploy it. When constructing a behavior library, the Markov model is used to represent the behavior sequence. Suppose the historical behavior sequence of a service process is "Start → Run → Stop", and the next behavior is predicted through the Markov model. If the current behavior is "Run", the model predicts that the probability of the next behavior being "Stop" is relatively high. The advantage of this method is that through continuous learning and optimization, the model can more accurately identify abnormalities and reduce false alarms and missed detections.Meanwhile, by combining historical data and real-time data, it is possible to comprehensively grasp the running status of the service process, improving the stability and reliability of the system. In practical applications, the construction of the log dataset and model training are dynamic processes that require continuous iteration and optimization. In this way, not only can anomalies be detected and processed in a timely manner, but also strong support can be provided for system optimization and upgrading. For example, by analyzing the abnormal logs and finding that a certain interface reports errors frequently, targeted optimization can be carried out to improve system performance. To sum up, through steps such as constructing the log dataset, data preprocessing, model training and optimization, anomaly determination and warning, it is possible to effectively improve the monitoring and management level of the service process on the cloud host and ensure the stable operation of the system.
[0039] Abnormal logs are obtained by analyzing attributes such as log level and event time, the type of anomaly is judged based on attributes such as request path and response status, and a random forest algorithm is used to train an anomaly detection model to obtain the anomaly detection result.
[0040] Obtain the log data generated by the service operation, and parse each log data to include log level and event time attribute information. If the log level meets the pre-determined abnormal level, extract this log as an abnormal log to obtain multiple abnormal log samples. Obtain each abnormal log, parse the request path and response status attribute information of each abnormal log, match the request path with the pre-constructed path rule library to obtain matching information, match the response status with the pre-determined abnormal status table to obtain abnormal status information. If the matching information meets a certain type and the abnormal status meets a certain type, comprehensively judge the abnormal type based on the two to obtain the abnormal type of each abnormal log. Extract abnormal log features, where feature extraction includes abnormal log text numerical features and abnormal log time series features. Divide the abnormal logs into a training set and a test set according to the quantity. Half of the abnormal logs in the training set are used to construct the model, and the other half of the abnormal logs are used for model training to obtain all the training set data. Determine the abnormal log feature set of the training set according to all the training set data generated during the data preprocessing operation. Construct the feature matrix of each abnormal log in the training set. For the feature matrix constructed for the abnormal log feature set of all the training sets, calculate the information gain rate of each training set abnormal log feature according to the statistical values of each training set abnormal log feature in different categories, sort the training set abnormal log features according to the information gain rate, and sequentially input the sorted feature matrix into the constructed random forest model to train the random forest model. Iterate in a loop until the convergence condition is reached, and then obtain the trained random forest model. Obtain each abnormal log in the test set during the data preprocessing operation, determine the abnormal log feature set of each test set, and obtain the information gain rate of each test set abnormal log feature by calculating the statistical values of each test set abnormal log feature in different categories. Sort the test set abnormal log features according to the information gain rate. Obtain the sorted test set feature matrix. Input the sorted test set abnormal log feature matrix according to the trained random forest model for forward operation to obtain the predicted output value, compare it with the test set label value, and obtain the prediction accuracy value of each test set abnormal log. Calculate the average value to obtain the model accuracy. Judge whether to perform the next iteration according to the accuracy. After the iteration ends, obtain the optimal detection model. Receive online logs through the interface, parse the log level and time information included in the logs. If the log level is the pre-set abnormal level, obtain the log sample. Parse the request path and response status information of the log, perform path rule matching according to the request path to obtain matching information, perform status matching according to the response status to obtain status information. If the matching information and the status information match the pre-set abnormal type, classify this log as abnormal, and perform detection operation on the abnormal log according to the optimal detection model to obtain the detection result of whether this log is abnormal.
[0041] Exemplarily, the log analysis platform implements anomaly detection through machine learning algorithms. First, it is necessary to obtain the log data generated during the service operation. Taking a Web server as an example, the log may contain information such as access time, request path, response status code, etc. When parsing the log data, key attributes such as log level and event time are focused on. For example, logs at the ERROR level are regarded as anomaly samples, which can quickly filter out potential problem logs. For further analysis of anomaly logs, it is necessary to parse the request path and response status. Suppose there is a log showing that the " / api / user" path returns a 500 status code. By matching with predefined path rules and anomaly status tables, it can be determined that this is an error in the user interface server. This method can quickly locate the problem and is beneficial for timely handling of faults. Feature extraction is a key step in model training. For text-based logs, methods such as TF-IDF can be used to extract numerical features; for time series data, sliding window techniques can be considered to extract time-related features. For example, counting the occurrence frequency of specific errors within 30 minutes helps to discover periodic or sudden anomalies. In the model training stage, the random forest algorithm is widely used due to its excellent performance and interpretability. By calculating the information gain rate to select the most discriminative features, the efficiency and accuracy of the model can be improved. For example, if it is found that the "response time" feature has the highest information gain rate, it indicates that it is the most critical for anomaly judgment and should be considered first. Model evaluation uses precision as an indicator, which helps to measure the model's ability to identify anomalies. Suppose the precision reaches 95% on the test set, which means the model can accurately identify 95% of the anomaly cases, but there are still 5% misjudgments and further optimization is needed. Finally, the trained model is applied to real-time log analysis. When new logs are received, the system quickly determines whether it is an anomaly. For example, if it is detected that a certain API frequently returns errors in a short period of time, the model may mark it as an anomaly and trigger an alarm mechanism, enabling the operation and maintenance personnel to intervene and handle it in a timely manner. This machine learning-based log analysis method has stronger adaptability and accuracy compared to traditional rule matching. It can automatically learn complex anomaly patterns, adapt to the changing system environment, and effectively improve the stability and reliability of the service. At the same time, through continuous model updates and optimizations, the system's anomaly detection ability will be continuously improved, bringing long-term benefits to the enterprise's IT operation and maintenance.
[0042] S106. When the monitoring system detects an anomaly in a certain service, according to the pre-configured alarm rules, it automatically notifies the operation and maintenance personnel of the anomaly information, and at the same time triggers an automated emergency response process. Depending on the severity and impact scope of the anomaly, different handling measures are taken, such as restarting the service process, rolling back the patch version, etc.
[0043] According to the log information of the collection server cluster, through log structured processing, a log information flow is formed. Using natural language processing algorithms to identify unstructured abnormal texts in the log information flow, comparing the vectors corresponding to the abnormal texts with the vectors in the pre-established feature library, calculating the matching degree. If the matching degree exceeds the preset threshold of 85%, the alarm program is triggered. Using a time series prediction model to analyze the sequence of matching degrees, predicting the time points of subsequent error logs. If the time point sequence is dense, a conclusion of service process failure is obtained, and it is judged that the system reaches a severe level. The historical state information is obtained as a feature parameter and input into the pre-constructed multi-modal fault prediction model to obtain the output result of the need to restart the service. Training a recurrent neural network model for the fault information sequence, generating a set of fault troubleshooting solutions by predicting the next character of the information sequence, scoring the confidence of different steps in the set of fault troubleshooting solutions by comparing the actual service log content, obtaining the best fault troubleshooting strategy, and judging that the first step of this strategy needs to obtain the modification records of the code management library. Extracting the recent code modification record files and historical version information files from the code management library, analyzing the differences between each file to obtain an associated file group, calculating the difference ratio before and after modification for each file in the associated file group. If the difference ratio exceeds the limit, the modification record file is marked as a suspicious file. Transmitting the suspicious file information to a vulnerability scanning tool for file security assessment, creating an abstract syntax tree according to the code syntax structure, obtaining the number of vulnerabilities by traversing and comparing the nodes of the abstract syntax tree. When it is judged that the number of vulnerabilities exceeds the security threshold, determining the correlation degree between the marked file and the vulnerabilities. If the number of vulnerabilities exceeds the preset threshold of 8, obtaining the set of patch information deployed recently, performing vectorization processing on the set of patch information, representing the current version state using vector space model technology, obtaining the set of difference points between the current version state and the historical version state through a version comparison tool, and obtaining the set of patch information files to be rolled back. Constructing a reverse operation instruction set according to the set of patch information files to be rolled back and the current version information, performing clustering analysis operations on the new log information vectors generated after the execution of the instructions using a clustering analysis algorithm, judging the risk degree of the instruction set by evaluating the tightness within the clusters of the clustering results, and determining whether to execute the reverse operation instruction set.
[0044] Exemplarily, collecting the log information of the server cluster is the starting point of anomaly detection. Suppose the server cluster of an e-commerce platform generates several gigabytes of log data every day, and these logs record detailed information such as user access, transaction processing, and database operations. Through log structured processing, unstructured text is converted into a structured data format, such as JSON or CSV, for subsequent analysis. Natural Language Processing (NLP) algorithms are used to identify unstructured abnormal text. For example, the TF-IDF algorithm is used to extract keywords from the logs, and combined with word embedding techniques (such as Word2Vec), the text is converted into a vector representation. Suppose a log record is "Database connection failed". After NLP processing, the vector corresponding to this text is compared with the "Database anomaly" vector in the feature library, and the cosine similarity is calculated. If the similarity exceeds 85%, the alarm program is triggered. Time series prediction models such as ARIMA or LSTM networks are used to analyze the matching degree sequence. Suppose that in the past week, the matching degree of the log of database connection failure has frequently exceeded the threshold, and the model predicts that more similar errors may occur within the next 24 hours. If the predicted time point sequence is dense, it is judged that there may be a fault in the service process, and the system reaches a serious level. Historical state information, such as CPU usage, memory occupancy, network latency, etc., is obtained and passed as feature parameters into a pre-built multi-modal fault prediction model. This model may contain a hybrid structure of Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN), and the output result recommends restarting the service to alleviate the fault. A recurrent neural network model is trained for the fault information sequence, and by predicting the next character of the information sequence, a set of fault troubleshooting solutions is generated. For example, the model predicts "Check the database connection configuration", and combined with the actual service log content, a confidence score is given to this step, and finally the best fault troubleshooting strategy is determined as "The first step is to obtain the code modification records in the code management library". The recent code modification record files and historical version information files are extracted from the code management library, and a file comparison tool (such as diff) is used to analyze the differences to obtain an associated file group. Suppose the difference ratio of a file before and after modification exceeds 30%, then it is marked as a suspicious file. The information of this suspicious file is transmitted to a vulnerability scanning tool (such as SonarQube), and by creating an Abstract Syntax Tree (AST) to traverse each node, the number of vulnerabilities is found by comparison operations. If the number of vulnerabilities exceeds 8, it is determined that the marked file is highly associated with the vulnerabilities. The set of recently deployed patch information is obtained, vectorized, and the current version state is represented using Vector Space Model (VSM) technology. The set of difference points is obtained through a version comparison tool (such as GitDiff), and the set of patch information files to be rolled back is obtained. According to the set of patch information files to be rolled back and the current version information, a set of reverse operation instructions is constructed. The clustering analysis algorithm (such as K-means) is used to perform clustering analysis on the new log information vectors generated after the execution of the instructions, and the tightness within the clusters of the clustering results is evaluated to judge the risk degree of the instruction set.If the degree of tightness within the cluster is high, it indicates that the risk of reverse operation is low, and the instruction set can be executed. Through the above steps, not only is it possible to automatically detect abnormal behaviors from log data and generate warning messages, but also through multi-level model analysis and operation instruction generation, the stability and security of the system are ensured. This comprehensive abnormal detection and troubleshooting mechanism can effectively improve the operation and maintenance efficiency and response speed of the system, and reduce the business interruption time caused by failures.
[0045] S107. In the operation and maintenance knowledge base, summarize and generalize information such as configuration parameters, log formats, and common faults of different operating systems and services to form a set of standardized operation and maintenance specifications and processes. Through knowledge graph technology, associate and reason about this knowledge to form an intelligent operation and maintenance decision support system, providing diagnostic and disposal suggestions for operation and maintenance personnel.
[0046] According to the operating system and service configuration information in the operation and maintenance knowledge base, through knowledge graph technology for association and reasoning, obtain a standardized configuration parameter template to form a configuration specification library. Adopt natural language processing technology to parse and extract the log formats in the operation and maintenance knowledge base, obtain key fields and event types, construct a log parsing model, and realize the automated analysis of logs. For the common fault information in the operation and maintenance knowledge base, through knowledge graph technology for correlation analysis, obtain the causes, impact scopes, and solutions of faults to form a fault diagnosis knowledge base. According to the configuration specification library, log parsing model, and fault diagnosis knowledge base, use a rule-based inference engine to realize the automated decision-making of operation and maintenance processes and generate standardized operation and maintenance operation steps. Through machine learning algorithms, train historical operation and maintenance data to obtain classification models and prediction models for operation and maintenance events, and realize the early warning and early disposal of potential faults. Integrate the operation and maintenance decision support system with the monitoring platform to obtain the running state data of the system and services in real time. Through knowledge graph reasoning and machine learning models, judge whether there are abnormalities and give diagnostic suggestions. According to the diagnostic suggestions and standardized operation and maintenance operation steps, generate automated disposal scripts and execute them through the operation and maintenance automation platform to achieve rapid repair of faults and system recovery, improving operation and maintenance efficiency and system availability.
[0047] Exemplarily, the operation and maintenance knowledge base is the core of IT operation and maintenance management, containing a large amount of configuration information, log formats, and fault handling experiences. Through knowledge graph technology, this information can be structured and analyzed for associations to form a standardized configuration specification library. For example, for the security configuration of Linux servers, key parameters such as the maximum number of failed logins and password complexity requirements can be extracted, and the associations between them can be established. This can not only quickly check whether the configuration complies with the specifications but also infer potential security risks. Log parsing is an important part of operation and maintenance automation. Through natural language processing technology, key information can be extracted from unstructured log texts. For instance, for a log of "user login failed", fields such as the username, IP address, and reason for failure can be identified. These structured log data provide the basis for subsequent analysis and decision-making. The construction of the fault diagnosis knowledge base utilizes the reasoning ability of the knowledge graph. By analyzing historical fault cases, associations can be established between fault symptoms, causes, and solutions. For example, when it is found that a certain Web service responds slowly, the system can infer possible causes such as database connection pool exhaustion and network congestion based on the knowledge graph and give corresponding troubleshooting steps. The rule-based inference engine is the key to realizing automated decision-making. It combines configuration specifications, log analysis results, and fault diagnosis knowledge to generate standardized operation and maintenance operation processes. For example, when it is detected that the CPU usage rate of a server continuously exceeds 90%, the system can automatically generate a series of troubleshooting steps, including checking the process status and analyzing the load source. Machine learning algorithms play an important role in operation and maintenance warning. By training on historical data, a prediction model for system anomalies can be established. For example, by analyzing past disk usage trends, the system can predict when the disk space may be exhausted and issue a warning in advance. Integrating the operation and maintenance decision support system with the monitoring platform can achieve real-time anomaly detection and diagnosis. For example, when it is monitored that the query latency of a certain database service suddenly increases, the system can immediately start the diagnostic process, analyze whether there are slow queries, index inefficiencies, etc., and give optimization suggestions. Finally, the generation and execution of automated disposal scripts are the key to improving operation and maintenance efficiency. For example, when it is diagnosed that a certain service anomaly is caused by a memory leak, the system can automatically generate a script to restart the service and execute it through the operation and maintenance platform to achieve rapid recovery. This not only reduces manual intervention but also greatly shortens the fault repair time. Through this series of technologies and processes, an intelligent operation and maintenance system can be constructed, which can not only quickly respond to and solve problems but also predict and prevent potential faults, thereby significantly improving the availability and stability of the system.
[0048] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A method for updating an operating system based on cloud services, characterized in that: The method comprises: By scanning the system information of the cloud host, the operating system version and patch installation status of different hosts are obtained. According to the operating system type and version number, the hosts are divided into corresponding management groups, and a unified patch strategy and update plan are formulated for each group; Build a mirror server inside the government network to download security update files required for each operating system version from a trusted patch source. By extracting metadata information from the update package and analyzing key attributes such as the applicable system and version number of the patch, the machine learning clustering algorithm is used to automatically classify the patch files into corresponding directories. For the patch files in the image server, key features are extracted and an intelligent classification model is trained. When a new patch file is added, the model can accurately identify the applicable scope of the patch and automatically put it into the corresponding directory to ensure that different operating system versions can find matching update files. Deploy monitoring agents on cloud hosts to collect the operating indicators and log information of each service process in real time, aggregate the collected data to a centralized log analysis platform, parse and mine the logs through a batch computing framework, and build a behavioral model for service operation. Based on the data in the log analysis platform, we use machine learning algorithms to train an anomaly detection model through supervised learning. This model can automatically discover various abnormal behaviors during service operation from massive log data and generate alarm information. When the monitoring system finds that a service is abnormal, it automatically notifies the operation and maintenance personnel of the abnormal information according to the pre-configured alarm rules, and triggers the automated emergency response process. Different processing measures are taken according to the severity and impact scope of the abnormality, such as restarting the service process, rolling back the patch version, etc. In the operation and maintenance knowledge base, the configuration parameters, log formats, common faults and other information of different operating systems and services are summarized and organized to form a set of standardized operation and maintenance specifications and processes. Through knowledge graph technology, this knowledge is associated and inferred to form an intelligent operation and maintenance decision support system to provide diagnosis and disposal suggestions for operation and maintenance personnel.
2. The method according to claim 1, characterized in that The system information of the cloud host is scanned to obtain the operating system version and patch installation status of different hosts. The hosts are divided into corresponding management groups according to the operating system type and version number, and a unified patch strategy and update plan are formulated for each group, including: Remotely connect to the cloud host through the API interface or SSH protocol, execute system commands, and obtain system information such as the cloud host's operating system type, version number, and patch installation list; Based on the obtained cloud host system information, a decision tree algorithm is used to determine the management group to which each cloud host belongs according to the preset operating system type and version number rules; In the configuration management database, according to the unique identifier of the cloud host, the field of the management group to which it belongs is updated, and the cloud host is divided into the corresponding group; For each management group, the operation and maintenance personnel formulate a unified system patch strategy based on the characteristics of the operating system of the group, determine the list of patches to be installed and the installation priority; Convert patch policies into executable scripts and use automated operation and maintenance tools such as Ansible to batch distribute patch installation tasks to all cloud hosts in the group. During the patch installation process, the installation progress and results of each cloud host are obtained in real time and recorded in the system log for easy tracking and auditing; After the patch installation is completed, scan the system information of the cloud host again to obtain the latest patch installation status and determine whether it is consistent with the patch policy. If not, perform differential updates to ensure that the cloud hosts in each group meet the requirements of the patch policy. It also includes: obtaining the cloud host operating system version and patch installation status through scanning, dividing management groups according to system type and version number, formulating a unified patch strategy and update plan for the groups, and determining the patch update plan for each cloud host.
3. The method according to claim 2, characterized in that The method of obtaining the cloud host operating system version and patch installation status through scanning, dividing management groups according to system type and version number, formulating a unified patch strategy and update plan for each group, and determining the patch update plan for each cloud host includes: Use scanning tools to perform a comprehensive scan of the cloud host to obtain the operating system type, version number, and installed patch information of each cloud host; Based on the obtained operating system type and version number, a clustering algorithm is used to automatically group the cloud hosts. The cloud hosts in each group have similar operating system characteristics. For each group, analyze the patch installation status of the cloud hosts in the group, identify the existing patch missing and vulnerability risks, and determine the priority of patch updates based on the risk level; Based on the cloud host grouping information and patch priority, formulate a unified patch strategy and update plan for each group, and clarify the time nodes and specific operation steps for patch updates; When formulating patch strategies, we use association rule mining algorithms to analyze the dependencies and compatibility between different patches to ensure the rationality and security of patch updates. Generate a personalized patch update plan for each cloud host based on a unified patch strategy and update plan, clearly specifying the patch list that needs to be installed and the specific operation process; Adopt automated operation and maintenance tools to batch update and repair cloud hosts according to the patch update plan, and ensure a smooth and controllable patch update process through monitoring and rollback mechanisms.
4. The method according to claim 1, characterized in that The system builds a mirror server inside the government network, downloads the security update files required by each operating system version from a trusted patch source, extracts metadata information of the update package, analyzes key attributes such as the applicable system and version number of the patch, and uses a machine learning clustering algorithm to automatically classify the patch files into corresponding directories, including: Obtain the server resources available within the government network and determine the hardware configuration and network environment for building the mirror server; Obtain security update files for each operating system version from a trusted patch source and download them to the designated storage location of the mirror server; Extract metadata information from downloaded patch files to obtain key attributes such as applicable systems and version numbers of the patches. Based on the extracted patch key attributes, the K-means clustering algorithm is used to automatically classify patch files according to applicable systems and version numbers; The classification results obtained by the clustering algorithm are used to determine the operating system and version to which each patch file belongs, and then the patch file is moved to the corresponding directory on the mirror server. Configure Nginx and other Web services on the mirror server to provide patch file download services to ensure that all systems within the government network can access and download the required security updates; Continuously monitor the updates of trusted patch sources, automatically download newly released patch files through scheduled tasks, and repeat the above classification and release process to keep the patch library of the mirror server synchronized with the trusted source.
5. The method according to claim 1, characterized in that The method extracts key features from the patch files in the mirror server and trains an intelligent classification model. When a new patch file is added, the model can accurately identify the applicable scope of the patch and automatically put it into the corresponding directory to ensure that different operating system versions can find matching update files, including: Obtain metadata of all patch files in the image server, parse the file metadata to obtain the file name, size, hash value and related information list, and obtain a preliminary classification collection of patch files of different operating systems; Determine the operating system field and the version field according to the list of associated information of each patch file in the preliminary classification set, extract the operating system version information applicable to the patch file, and annotate the patch file set according to the version information to obtain sample data; Parse sample data, generate binary data stream, and perform numerical statistics based on the binary data stream.
6. The method according to claim 1, characterized in that The monitoring agent is deployed on the cloud host to collect the operation indicators and log information of each service process in real time, and the collected data is aggregated to a centralized log analysis platform. The logs are parsed and data mined through a batch computing framework to build a behavior model for service operation, including: Deploy the agent to collect the operating indicator values and log streams of each process in the cloud host service process set and upload the data; According to the preset log stream data format specification, the log analysis platform receives the data uploaded by the agent, performs data classification operations on the log streams generated by different service processes, and different classifications are marked differently; Through the data content of each field in the log stream, the Taichung parser module performs information extraction operations to obtain key log information, logs of different service processes, and determine the identifiers of different service processes; Use pre-trained machine learning models to perform anomaly detection operations on different processes, input log key information and identifiers into the binary classification model, and obtain the normal or abnormal classification results of each process; Obtain the running index value of each process, build a running index value set for each process, input the running index value set into the pre-trained regression model, determine the regression model output, and obtain the performance prediction result of each process; According to the output results of the regression model and the operating index values, the calculation block performs statistical analysis on the load status of each service process in the service process set, and obtains a behavior sequence by comparing the load status of processes with similar behaviors in historical data; A behavior library is constructed based on the behavior sequence and the current service process behavior. The behavior sequence is represented by a Markov model to determine the next behavior of the service process.
7. The method according to claim 1, characterized in that The data in the log analysis platform is trained using a machine learning algorithm through supervised learning to train an anomaly detection model. The model can automatically discover various abnormal behaviors in the service operation process from massive log data and generate alarm information, including: Construct a log dataset based on the log content. The log dataset should cover various log entries of historical data records, where normal and abnormal log entries are marked with data labels. Perform data preprocessing on the labeled log data set, extract features, obtain a numerical data set, and divide it into a training set and a test set. The proportion of samples with different service states in the training set should be balanced. The support vector machine algorithm is trained through the training set to obtain an initial detection model, the test set is tested through the initial detection model, and various indicators of the initial detection model for the test set detection results are calculated; Obtain the log content generated in the real environment and perform data preprocessing, and obtain the previously trained model of the time series as the benchmark detection model. Use the benchmark detection model to process these log data, generate preliminary prediction labels, and then perform label determination; The label is determined based on the preliminary predicted label. For the preliminary predicted label, if the preliminary predicted label of a log record is abnormal, and the time and log content of the log record meet the matching conditions with the log data of an abnormal record in the existing label database, the final predicted label is determined to be abnormal. Perform anomaly determination on the final predicted labels. If the final predicted label of a log is abnormal, determine the alarm level corresponding to the log, use different levels of alarm schemes to issue an alarm, record the final labeling results in the labeled database, and then continue to perform the determination. If it is determined that the final predicted labels are all normal, it is determined that the detection model has defects. Determine the ratio of false positives in the labeled database. If it exceeds the set ratio, take out all abnormal label samples in the labeled database for retraining to obtain a new time series detection model, and then redeploy the model; It also includes: obtaining abnormal logs by analyzing attributes such as log level and event time, judging the abnormal type according to attributes such as request path and response status, and using random forest algorithm to train the abnormal detection model to obtain the abnormal detection results.
8. The method according to claim 7, characterized in that The abnormal log is obtained by analyzing the attributes such as log level and event time, the abnormal type is determined according to the attributes such as request path and response status, and the abnormal detection model is trained by using the random forest algorithm to obtain the abnormal detection results, including: Obtain the log data generated by the service operation, parse each log data to include the log level and event time attribute information, and if the log level meets the predetermined abnormal level, extract the log as an abnormal log to obtain multiple abnormal log samples; Obtain each exception log, parse the request path and response status attribute information of each exception log, match the request path with the pre-built path rule library to obtain matching information, match the response status with the pre-determined exception status table to obtain exception status information, and if the matching information matches a certain type and the exception status matches a certain type, then combine the two to determine the exception type and obtain the exception type of each exception log; Extract abnormal log features. Feature extraction includes abnormal log text numerical features and abnormal log time series features. The abnormal logs are divided into training set and test set according to the number. Half of the abnormal logs in the training set are used to build the model, and the other half are used for model training to obtain all the training set data. Determine the abnormal log feature set of the training set according to all training set data generated during the data preprocessing operation; Construct the feature matrix of each abnormal log in the training set; A feature matrix is constructed for the abnormal log feature set of all training sets. The information gain rate of each abnormal log feature of the training set is calculated according to the statistical value of each abnormal log feature of the training set in different categories. The abnormal log features of the training set are sorted according to the information gain rate. The sorted feature matrix is input into the constructed random forest model in sequence. The random forest model is trained and iterated until the convergence condition is reached to obtain a trained random forest model. Obtain each abnormal log of the test set during the data preprocessing operation, determine the feature set of each test set abnormal log, obtain the statistical value of each test set abnormal log in different categories, and calculate the information gain rate of each test set abnormal log feature; Sort the abnormal log features of the test set by information gain rate; Get the sorted test set feature matrix; According to the trained random forest model, the sorted test set abnormal log feature matrix is input, and forward operation is performed to obtain the predicted output value, which is compared with the test set label value to obtain the accurate prediction value of each test set abnormal log; Calculate the average value to get the model accuracy; Determine whether to proceed to the next round of iteration based on accuracy; After the iteration, the optimal detection model is obtained; Receive online logs through the interface, parse the log level and time information contained in the log, and obtain log samples if the log level is the preset abnormal level; Parse the request path and response status information of the log, match the path rules according to the request path to obtain the matching information, and match the status according to the response status to obtain the status information. If the matching information and the status information match the preset anomaly type, the log is classified as abnormal. Perform detection operations on the abnormal log according to the optimal detection model to obtain the detection result of whether the log is abnormal.
9. The method according to claim 1, characterized in that: When the monitoring system finds that a service is abnormal, it automatically notifies the operation and maintenance personnel of the abnormal information according to the pre-configured alarm rules, and triggers the automated emergency response process. Different processing measures are taken according to the severity and impact scope of the abnormality, such as restarting the service process, rolling back the patch version, etc., including: According to the collected server cluster log information, log information stream is formed through log structured processing, and the unstructured abnormal text in the log information stream is identified by natural language processing algorithm. The vector corresponding to the abnormal text is compared with the vector in the pre-established feature library, and the matching degree is calculated. If the matching degree exceeds the preset threshold of 85%, the alarm program is triggered; Use the time series prediction model to analyze the matching degree sequence and predict the time point of subsequent error log occurrence. If the time point sequence is dense, the conclusion of service process failure is obtained, and the system is judged to have reached a serious level. The historical status information is obtained as a feature parameter and passed into the pre-built multi-modal fault prediction model to obtain the output result that the service needs to be restarted; A recurrent neural network model is trained for the fault information sequence, and a set of troubleshooting solutions is generated by predicting the next character in the information sequence. By comparing the actual service log content, the confidence scores of different steps in the set of troubleshooting solutions are performed to obtain the best troubleshooting strategy. It is determined that the first step of this strategy requires obtaining the modification record of the code management library. Extract recent code modification record files and historical version information files from the code management library, analyze the differences between each file to obtain the associated file group, and calculate the difference ratio before and after modification for each file in the associated file group. If the difference ratio exceeds the limit, mark the modification record file as a suspicious file; The suspicious file information is transferred to the vulnerability scanning tool to perform file security assessment. An abstract syntax tree is created based on the code syntax structure. The number of vulnerabilities is obtained by traversing each node of the abstract syntax tree and performing comparison operations. When it is determined that the security threshold is exceeded, the correlation between the marked file and the vulnerability is determined. If the number of vulnerabilities exceeds the preset threshold of 8, obtain the patch information set deployed recently, vectorize the patch information set, use vector space model technology to represent the current version status, use version comparison tools to obtain the difference point set between the current version status and the historical version status, and obtain the patch information file set that needs to be rolled back; A reverse operation instruction set is constructed based on the patch information file set that needs to be rolled back and the current version information. A clustering analysis algorithm is used to perform clustering analysis operations on the new log information vector generated after the instruction execution. By evaluating the tightness within the clustering result cluster, the risk level of the instruction set is judged to determine whether to execute the reverse operation instruction set.
10. The method according to claim 1, characterized in that In the operation and maintenance knowledge base, the configuration parameters, log formats, common faults and other information of different operating systems and services are summarized and summarized to form a set of standardized operation and maintenance specifications and processes. Through the knowledge graph technology, these knowledge are associated and reasoned to form an intelligent operation and maintenance decision support system to provide diagnosis and disposal suggestions for operation and maintenance personnel, including: Based on the operating system and service configuration information in the operation and maintenance knowledge base, the knowledge graph technology is used to associate and infer information, obtain standardized configuration parameter templates, and form a configuration specification library; Use natural language processing technology to parse and extract the log format in the operation and maintenance knowledge base, obtain key fields and event types, build a log parsing model, and realize automatic analysis of logs; For common fault information in the operation and maintenance knowledge base, we use knowledge graph technology to perform correlation analysis to obtain the cause, impact range and solution of the fault, and form a fault diagnosis knowledge base; Based on the configuration specification library, log parsing model and fault diagnosis knowledge base, a rule-based reasoning engine is used to realize automated decision-making of the operation and maintenance process and generate standardized operation and maintenance operation steps; Through machine learning algorithms, historical operation and maintenance data is trained to obtain classification models and prediction models for operation and maintenance events, thus achieving early warning and early disposal of potential failures. Integrate the operation and maintenance decision support system with the monitoring platform to obtain real-time operating status data of systems and services, determine whether there are abnormalities through knowledge graph reasoning and machine learning models, and provide diagnostic suggestions; Generate automated handling scripts based on diagnostic recommendations and standardized operation and maintenance procedures, and execute them through the operation and maintenance automation platform to achieve rapid fault repair and system recovery, thereby improving operation and maintenance efficiency and system availability.
Citation Information
Cited By
Server firmware version management method and program product
CN120315752A
Server firmware version management method and program product
CN120315752B
Running state monitoring method and system applied to server cluster
CN121093316A
Log analysis method and electronic equipment
CN122451298A
Configuring drive heterogeneous monitoring data mapping and cmdb enhancement method and system
CN122547887A