Cross-cluster infrastructure automated monitoring method, device, equipment and storage medium

By deploying data acquisition components in each cluster infrastructure and using the central monitoring platform for unified management and processing, the problem of cross-cluster monitoring is solved, efficient and accurate infrastructure monitoring and fault handling is achieved, and the stability and operation and maintenance efficiency of the system are improved.

CN119377039BActive Publication Date: 2025-08-26BEIJING BIG DATA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411388678.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-08-26
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

The existing monitoring solutions are difficult to effectively cover multiple clusters, lack a unified monitoring view, cannot process and analyze massive data in real time, have inconsistent alarm strategies, complex configurations, cannot detect and handle faults in a timely manner, cannot meet the needs of unified operation and maintenance, ignore non-core but important infrastructure components, and the monitoring indicators are not detailed enough to reflect the actual operating status of the system. The monitoring of infrastructure in cross-cluster environments is independent, resulting in dispersed operation and maintenance operations and inefficient efficiency.

Method used

Deploy preset data acquisition components in each infrastructure of each cluster, store and process data centrally through the central monitoring platform, select and configure data acquisition components using automated deployment tools to realize automatic discovery mechanisms, perform standardized processing and display, build a unified visual interface, and provide cross-cluster monitoring and troubleshooting capabilities.

Benefits of technology

It realizes unified monitoring and management across cluster infrastructure, improves the stability, reliability and efficiency of system operation, can promptly detect and deal with faults, simplify operation and maintenance operations, reduces manual participation, improves the accuracy and scope of monitoring, and ensures efficient utilization of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119377039B_ABST
    Figure CN119377039B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method, apparatus, device and storage medium for automated monitoring of cross-cluster infrastructure, the method comprising: deploying a preset data acquisition component in each infrastructure of each cluster; collecting operating data of each infrastructure of each cluster through the preset data acquisition component, and filling the operating data of the infrastructure into a preset monitoring template; standardizing the operating data in each preset monitoring template to obtain standardized data; receiving an input operating data access request, and displaying corresponding standardized data based on the display rules in the operating data access request. By monitoring the operating data of each infrastructure of each cluster under the cross-cluster, the stability, reliability and efficiency of the system operation in a multi-cluster environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data analysis technology, and in particular, to a method, apparatus, computer device, and storage medium for automated monitoring of cross-cluster infrastructure. Background Art

[0002] With the rapid development of cloud computing and big data technologies, more and more companies are building multi-cluster environments to meet high availability and scalability requirements. However, cross-cluster infrastructure management and monitoring has become a challenge.

[0003] The monitoring solutions in related technologies are often designed for a single cluster, making it difficult to effectively cover multiple clusters and achieve unified management and monitoring. For example, there is a lack of a unified monitoring view, difficulty in real-time processing and analysis of massive data, inconsistent alarm strategies, or complex configurations.

[0004] Monitoring solutions in related technologies may suffer from infrequent data collection or processing delays, leading to untimely updates of monitoring information, delayed fault detection and resolution (results), and an inability to quickly respond to system anomalies, impacting service stability and user experience. Current monitoring may primarily focus on core systems or critical services, while neglecting other, less-core but equally important, infrastructure components. Lack of comprehensive coverage of multi-cluster, cross-regional infrastructure can prevent timely detection and resolution of issues in certain regions or clusters. Existing monitoring may focus solely on basic performance metrics, such as CPU usage and memory usage, while overlooking key business or user experience indicators. Monitoring metrics may be insufficiently detailed and fail to accurately reflect the actual system operation or potential issues. Independent monitoring of infrastructure in a multi-cluster environment makes centralized monitoring and management difficult, resulting in fragmented and inefficient operations and maintenance. This makes efficient global monitoring and troubleshooting difficult, and fails to meet unified operations requirements. Cross-cluster infrastructure may involve multiple operating systems, middleware, databases, and other components. The complexity and diversity of these components pose significant challenges to automated data collection and monitoring. Summary of the Invention

[0005] Embodiments of the present application provide a method, apparatus, computer device, and storage medium for automated monitoring of cross-cluster infrastructure.

[0006] A first aspect of an embodiment of the present application provides a method for automated monitoring of cross-cluster infrastructure, including:

[0007] Deploy preset data collection components in each infrastructure of each cluster;

[0008] The preset data collection component collects the operating data of each infrastructure in each cluster, and fills the operating data of the infrastructure into the preset monitoring template;

[0009] Standardize the operating data in each preset monitoring template to obtain standardized data;

[0010] An input operation data access request is received, and corresponding standardized data is displayed based on a display rule in the operation data access request.

[0011] In an optional embodiment of the present application, deploying a preset data collection component in each infrastructure of each cluster includes:

[0012] Building a central monitoring platform, where the central monitoring platform is used to centrally store, process, and display data collected from each infrastructure in different clusters;

[0013] Select the corresponding preset data collection component based on the type of each infrastructure in each cluster, and ensure that the preset data collection component supports the automatic discovery mechanism.

[0014] In an optional embodiment of the present application, selecting a corresponding preset data collection component according to the type of each infrastructure of each cluster to ensure that the preset data collection component supports an automatic discovery mechanism includes:

[0015] Based on the type of infrastructure in each cluster, use the preset automated deployment tool to obtain the data collection components corresponding to the current infrastructure from the component database of the central monitoring platform;

[0016] Deploy the acquired data collection components to the corresponding infrastructure, and configure the infrastructure according to the configuration template corresponding to the infrastructure;

[0017] The data acquisition component deployed to the infrastructure is configured so that the data acquisition component is connected to the central monitoring platform, and corresponding data acquisition sources and data acquisition targets are set on the data acquisition component.

[0018] In an optional embodiment of the present application, the method further includes:

[0019] Configure the auto-discovery mechanism for each cluster;

[0020] Based on the automatic discovery mechanism, the infrastructure in the cluster is automatically discovered using a preset cluster management tool, and the discovered infrastructure information is stored in the configuration management database of the central monitoring platform.

[0021] In an optional embodiment of the present application, after automatically discovering the infrastructure in the cluster using a preset cluster management tool, the method further includes:

[0022] Determining whether a data collection component corresponding to the infrastructure has been deployed on the discovered infrastructure;

[0023] For infrastructures where data collection components have been deployed, the current infrastructure information is stored in the configuration management database of the central monitoring platform;

[0024] For infrastructure that has not deployed data collection components, first deploy the preset data collection components in the current infrastructure, and then store the infrastructure information after deployment in the configuration management database of the central monitoring platform.

[0025] In an optional embodiment of the present application, when there are more than one infrastructures of the same type in the same cluster, the step of deploying a preset data collection component in each cluster includes:

[0026] For the same type of infrastructure in the same cluster, if the physical locations of all infrastructures are distributed within a preset range, or the network conditions of the infrastructure do not meet the preset network requirements, a data acquisition server is set between the central monitoring platform and the infrastructure. The data acquisition server is used to collect the operating data of all infrastructures and perform preliminary processing. All infrastructures are connected to the data acquisition server, and the data acquisition server is connected to the central monitoring platform. Data acquisition components are deployed on each infrastructure and data acquisition server.

[0027] If the distribution of the physical locations of all infrastructures exceeds the preset range, or the network conditions of the infrastructures meet the preset network requirements, a data acquisition server and data acquisition components are deployed on each infrastructure. The data acquisition server is used to perform preliminary processing on the local data collected by the data acquisition components and send the processed data to the central monitoring platform.

[0028] In an optional embodiment of the present application, the operating data of all infrastructures are collected and preliminarily processed, including:

[0029] Obtain the operating data of all infrastructure and extract the feature vectors of the pre-processed operating data;

[0030] Build an operation data sensitivity scoring model to evaluate the sensitivity of operation data and classify it, and encrypt sensitive data based on the classification results;

[0031] Decrypt the exchanged data to store the running data, and build a visual interface to display the running data in real time.

[0032] A second aspect of an embodiment of the present application provides a cross-cluster infrastructure automated monitoring device, including:

[0033] Deployment module, used to deploy preset data collection components in each infrastructure of each cluster;

[0034] The filling module is used to collect the operating data of each infrastructure of each cluster through the preset data collection component, and fill the operating data of the infrastructure into the preset monitoring template;

[0035] The standardization module is used to standardize the operating data in each preset monitoring template to obtain standardized data;

[0036] The display module is configured to receive an input operation data access request and display corresponding standardized data based on a display rule in the operation data access request.

[0037] According to a third aspect of an embodiment of the present application, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any of the above methods for automated monitoring of cross-cluster infrastructure are implemented.

[0038] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program implements the steps of any of the above methods for automated monitoring of cross-cluster infrastructure.

[0039] The above technical solutions provided by the embodiments of the present application have at least some or all of the following advantages compared to the prior art:

[0040] The cross-cluster infrastructure automated monitoring method described in the embodiment of the present application deploys a preset data acquisition component in each infrastructure of each cluster; collects the operating data of each infrastructure of each cluster through the preset data acquisition component, and fills the operating data of the infrastructure into a preset monitoring template; standardizes the operating data in each preset monitoring template to obtain standardized data; receives an input operating data access request, and displays the corresponding standardized data based on the display rules in the operating data access request. By monitoring the operating data of each infrastructure of each cluster under the cross-cluster, the stability, reliability and efficiency of the system operation in a multi-cluster environment can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0042] Figure 1 A flowchart of a method for automated monitoring of cross-cluster infrastructure provided by one embodiment of the present application;

[0043] Figure 2 A flowchart for collecting operating data of each infrastructure of each cluster provided in one embodiment of the present application;

[0044] Figure 3 A flowchart of an agent configuration provided for one embodiment of the present application;

[0045] Figure 4 A schematic diagram of the structure of a cross-cluster infrastructure automated monitoring device provided in one embodiment of the present application;

[0046] Figure 5 A schematic diagram of the computer device structure provided for one embodiment of the present application. DETAILED DESCRIPTION

[0047] See Figure 1 The cross-cluster infrastructure automated monitoring method provided in the embodiment of the present application includes the following steps 100 to 400:

[0048] Step 100: Deploy a preset data collection component in each infrastructure of each cluster;

[0049] Step 200: Collect the operating data of each infrastructure of each cluster through the preset data collection component, and fill the operating data of the infrastructure into the preset monitoring template;

[0050] Step 300: Standardize the operating data in each preset monitoring template to obtain standardized data;

[0051] Step 400: Receive an input operation data access request, and display corresponding standardized data based on a display rule in the operation data access request.

[0052] In an optional embodiment of the present application, in step 100, deploying a preset data collection component in each infrastructure of each cluster includes:

[0053] Building a central monitoring platform, where the central monitoring platform is used to centrally store, process, and display data collected from each infrastructure in different clusters;

[0054] Select the corresponding preset data collection component based on the type of each infrastructure in each cluster, and ensure that the preset data collection component supports the automatic discovery mechanism.

[0055] In an optional embodiment of the present application, selecting a corresponding preset data collection component according to the type of each infrastructure of each cluster to ensure that the preset data collection component supports an automatic discovery mechanism includes:

[0056] Based on the type of infrastructure in each cluster, use the preset automated deployment tool to obtain the data collection components corresponding to the current infrastructure from the component database of the central monitoring platform;

[0057] Deploy the acquired data collection components to the corresponding infrastructure, and configure the infrastructure according to the configuration template corresponding to the infrastructure;

[0058] The data acquisition component deployed to the infrastructure is configured so that the data acquisition component is connected to the central monitoring platform, and corresponding data acquisition sources and data acquisition targets are set on the data acquisition component.

[0059] In an optional embodiment of the present application, the method further includes:

[0060] Configure the auto-discovery mechanism for each cluster;

[0061] Based on the automatic discovery mechanism, the infrastructure in the cluster is automatically discovered using a preset cluster management tool, and the discovered infrastructure information is stored in the configuration management database of the central monitoring platform.

[0062] In an optional embodiment of the present application, after automatically discovering the infrastructure in the cluster using a preset cluster management tool, the method further includes:

[0063] Determining whether a data collection component corresponding to the infrastructure has been deployed on the discovered infrastructure;

[0064] For infrastructures where data collection components have been deployed, the current infrastructure information is stored in the configuration management database of the central monitoring platform;

[0065] For infrastructure that has not deployed data collection components, first deploy the preset data collection components in the current infrastructure, and then store the infrastructure information after deployment in the configuration management database of the central monitoring platform.

[0066] In an optional embodiment of the present application, when there are more than one infrastructures of the same type in the same cluster, the step of deploying a preset data collection component in each cluster includes:

[0067] For the same type of infrastructure in the same cluster, if the physical locations of all infrastructures are distributed within a preset range, or the network conditions of the infrastructure do not meet the preset network requirements, a data acquisition server is set between the central monitoring platform and the infrastructure. The data acquisition server is used to collect the operating data of all infrastructures and perform preliminary processing. All infrastructures are connected to the data acquisition server, and the data acquisition server is connected to the central monitoring platform. Data acquisition components are deployed on each infrastructure and data acquisition server.

[0068] If the distribution of the physical locations of all infrastructures exceeds the preset range, or the network conditions of the infrastructures meet the preset network requirements, a data acquisition server and data acquisition components are deployed on each infrastructure. The data acquisition server is used to perform preliminary processing on the local data collected by the data acquisition components and send the processed data to the central monitoring platform.

[0069] In an optional embodiment of the present application, how to deploy data collection components when there are many infrastructures of the same type in the same cluster?

[0070] First, choose between centralized agent deployment and distributed deployment. Centralized agent deployment: If there are many infrastructures in the cluster, but the physical locations are relatively concentrated, or the network conditions are poor, you can consider using a centralized data acquisition server. This server is responsible for collecting data from all infrastructures and performing preliminary processing. This method may require the deployment of collection components separately on each infrastructure, and also on the agent server; Distributed deployment: If the cluster is large, the infrastructure is widely distributed, and the network conditions are good, it is recommended to adopt a distributed deployment method. Deploy data acquisition agents (Agents) on each infrastructure or a group of infrastructures. These agents are responsible for the collection and preliminary processing of local data, and then send the data to the central server for aggregation and analysis. This method can improve the flexibility and scalability of data collection, while reducing dependence on network bandwidth.

[0071] Second, use standardized configurations to ensure that all infrastructures of the same type use the same configuration template. This can be achieved through configuration management. Standardized configurations help reduce problems caused by inconsistent configurations and simplify the update and maintenance process.

[0072] Third, automated deployment, using automation tools or scripts to automatically deploy data acquisition components. For example, in a K8s environment, you can use Helm charts (Helm charts is a packaging format for K8s resources that simplifies application deployment and management) to deploy data acquisition components; in other environments, you can use Ansible playbook (an automation script written in YAML for configuring, deploying, and orchestrating applications) or other similar automation tools.

[0073] Fourth, dynamic configuration. If the number of infrastructure in the cluster changes frequently (for example, in a cloud-native environment), you should consider using dynamic configuration methods to automatically discover newly added nodes and automatically install and configure data collection components on these nodes. This can be achieved through service discovery mechanisms (such as Consul, etcd, etc.).

[0074] Fifth, centralized management. For large-scale clusters, a centralized configuration management and service registration system is used to centrally control the configuration and status of data collection components, making it easier to monitor and troubleshoot.

[0075] Sixth, resource isolation ensures that data acquisition components do not occupy too many resources and affect the performance of the host or application. You can assign fixed resource limits to data acquisition components, or use lighter data acquisition tools in resource-intensive applications.

[0076] In an optional embodiment of the present application, how to deploy data collection components on different infrastructures in the same cluster,

[0077] First, determine the infrastructure type: clarify what types of infrastructure are included in the cluster, such as physical servers, virtual machines, database servers, network devices (switches, routers), storage devices, etc.

[0078] Second, analyze monitoring requirements: For each infrastructure, analyze its key performance indicators (KPIs) and monitoring requirements. For example, physical servers may need to monitor CPU usage, memory usage, disk I / O, etc.; database servers may need to monitor database query performance, number of connections, etc.; network devices may need to monitor network traffic, packet loss rate, etc.

[0079] Third, distributed deployment: Deploy dedicated data collection agents or lightweight monitoring services on each infrastructure. These agents or services will be responsible for collecting data from their respective infrastructures and sending it to the central monitoring server for centralized processing.

[0080] Fourth, centralized processing and analysis: Establish a central monitoring server to receive data from various infrastructures and perform unified processing, analysis, and storage. The central monitoring server should have strong data processing capabilities and high availability to ensure the real-time and accuracy of data.

[0081] Fifth, security and rights management: Implement strict security measures to ensure data transmission and storage security during the data collection process. Set reasonable rights management policies to limit access to data collection components and monitoring data.

[0082] In an optional embodiment of the present application, for infrastructure in different clusters, data collection components are deployed through the following steps:

[0083] First, build a central monitoring platform (proxy gateway): Select or build a central monitoring platform to centrally store, process, and display collected data. Ensure that the monitoring platform supports multi-tenancy, permission management, and scalability to manage data from multiple clusters.

[0084] Second, data collection component deployment: select appropriate data collection components based on the infrastructure type and ensure that the data collection components support automatic discovery and configuration.

[0085] a. Based on the infrastructure information, use automated deployment tools to pull the image or software package of the data collection component from the central warehouse.

[0086] b. Deploy the data collection components to the corresponding infrastructure and configure them according to the configuration template.

[0087] c. Configure the data collection component to connect to the central monitoring platform (agent gateway) and set the correct data collection source and target.

[0088] Third, automated collection implementation: Use the cloud service provider's cluster management tools to automatically discover the infrastructure in the cluster, and store the discovered infrastructure information (such as IP address, port, type, etc.) in a configuration management database or key-value storage.

[0089] In an optional embodiment of the present application, the operation data of all infrastructures are collected and preliminarily processed, including:

[0090] Obtain the operating data of all infrastructure and extract the feature vectors of the pre-processed operating data;

[0091] Build an operation data sensitivity scoring model to evaluate the sensitivity of operation data and classify it, and encrypt sensitive data based on the classification results;

[0092] Decrypt the exchanged data to store the running data, and build a visual interface to display the running data in real time.

[0093] In an optional embodiment of the present application, a feature vector of the pre-processed running data is extracted by presetting an autoencoder network model.

[0094] In an optional embodiment of the present application, the constructing of the application system information data sensitivity scoring model to evaluate the application system information data sensitivity and classify the application system information data includes collecting historical application system information data and extracting historical feature vectors, and calculating the mean of the historical feature vectors and setting them as the benchmark data vector;

[0095] Collect application system information data in real time and extract real-time feature vectors. Use the K-means clustering analysis algorithm to cluster the real-time feature vectors, and select the center point of each cluster as the reference data vector.

[0096] Combining the RBF kernel function with the integral, the cumulative similarity A(x) between the benchmark data vector and the feature vector of the application system information data is calculated. The formula is:

[0097]

[0098] Where x is the application system information data feature vector, x0 is the benchmark data vector, and x i is the historical feature vector of the i-th application system information data;

[0099] Perform logarithmic transformation B(x) on the accumulated similarity A, the formula is:

[0100] B(x)=log(1+A(x));

[0101] The hyperbolic tangent function is introduced to smooth the cumulative similarity of the reference feature vector to obtain the smoothed cumulative similarity C(x), which is expressed as follows:

[0102]

[0103] Where M is the number of reference eigenvectors, x j is the jth reference eigenvector;

[0104] A sensitivity scoring model is constructed to evaluate the sensitivity score S(x) of the feature vector of application system information data. The formula is:

[0105]

[0106] Collect the sensitivity scores of historical application system information data and set an evaluation threshold. Compare the sensitivity scores of the application system information data feature vectors with the evaluation threshold. If the sensitivity score of the application system information data feature vector is greater than and equal to the evaluation threshold, it is determined to be sensitive data. If the sensitivity score of the application system information data feature vector is less than the evaluation threshold, it is determined to be ordinary data.

[0107] In an optional embodiment of the present application, encrypting the sensitive data of the application system information data based on the classification result refers to randomly generating a key K using a random number generator;

[0108] Use exponential function and sine function to perform nonlinear transformation on sensitive data to obtain the nonlinear transformation result H(x′), which is as follows:

[0109]

[0110] Where x′ is sensitive data;

[0111] The sensitive data is nonlinearly transformed using the logarithmic function and the smooth inverse tangent function to obtain the multi-level nonlinear transformation result E(x′), which is as follows:

[0112] E(x')=tan -1 (x' 2 k+log(x'+k));

[0113] Construct the encryption formula:

[0114] S(x′)=H(x′)+E(x′);

[0115] Bring sensitive data into the encryption formula for encryption.

[0116] In an optional embodiment of the present application, decrypting the exchanged data and storing the application system information data refers to receiving the transmitted encrypted data and the corresponding key K;

[0117] Use the decryption formula to decrypt the received data and obtain the decrypted sensitive data O(S). The formula is:

[0118]

[0119] The collected application system information and sensitive data generated by analysis are stored in the database, the application system information and sensitive data generated by analysis in the database are backed up in the cloud, and the backup data is regularly checked for integrity.

[0120] In an optional embodiment of the present application, see Figure 2In step 200, the operation data of each infrastructure of each cluster is collected by the preset data collection component, including:

[0121] Automatically collect the operating data of each infrastructure in each cluster through preset data collection components;

[0122] According to the correspondence between the operation data type and the collection type in the preset collection management, the collection type of the collected operation data is determined, wherein the collection type includes real-time data collection, scheduled data collection and event-triggered data collection,

[0123] Real-time data collection is one of the platform's main collection methods. It uses efficient data transmission protocols to ensure data accuracy and real-time performance. It collects infrastructure operation data in real time through agent nodes deployed in each cluster. This data includes but is not limited to key indicators such as server hardware, network equipment, storage area network storage devices, cloud services, power supply, etc., server load, network bandwidth utilization, storage device health status, etc.

[0124] Scheduled data collection is an important means for the platform to periodically monitor infrastructure. According to the preset time interval (such as hourly, daily, weekly, etc.), the agent node will automatically trigger the data collection task to collect the operating data within a specific time period. This method is suitable for data items that do not require real-time monitoring but require periodic attention.

[0125] When a specific event occurs in the infrastructure (such as equipment failure, network interruption, etc.), the platform will trigger the event collection mechanism. The agent node will immediately collect data related to the event and send the data to the central server for further processing and analysis. This method can ensure that relevant data can be quickly obtained when a critical event occurs, providing support for rapid response and troubleshooting.

[0126] Collection management allows you to configure collection objects. Using dedicated agent plug-ins, Simple Network Management Protocol (SNMP), Java Management Extensions (JME), and intelligent platform management interfaces, it collects local client data, sends it to the server, and stores the data in a database. Users can view monitoring data directly on the front-end interface. Collection management allows you to directly use monitoring templates with pre-defined indicator data. When adding collected data, you can also specify whether to enable asset logging. This provides foundational data for asset management, integrating business and monitoring, and avoiding mismatches between business data and monitored object data.

[0127] In an optional embodiment of the present application, in step 300, the standardization of the operating data in each preset monitoring template to obtain standardized data includes:

[0128] The operating data in each preset monitoring template (such as CPU usage, memory usage, network traffic, disk I / O, etc.) is integrated into a unified format through standardization processing, which is used as standardized data.

[0129] In an optional embodiment of the present application, before displaying the corresponding standardized data based on the display rule in the operation data access request, the method further includes:

[0130] Build a unified visual monitoring interface that allows users to view and manage the status of all clusters through a single portal, including resource utilization, anomaly detection, historical trend analysis, etc., to quickly grasp the overall operation status and identify potential problem areas;

[0131] Customizable visual instrument views use data visualization technology to display data to users in the form of curve graphs, bar graphs, pie charts, etc., allowing users to intuitively understand the distribution and changing trends of data. Users can also customize monitoring views according to their needs, including selecting which indicators to display and how to layout charts, to meet their personalized monitoring needs.

[0132] Monitoring data query allows you to intuitively and simply view the indicator data of the monitored objects, and update the latest data in real time. It also provides charts and historical data viewing, traces the causes of alarms and analyzes them. Through data, you can understand the utilization of existing resources, discover resource bottlenecks and waste, and thus optimize resource allocation and improve resource utilization efficiency.

[0133] In an optional embodiment of the present application, the method further includes:

[0134] Update the status information of each cluster in real time to ensure the consistency and accuracy of monitored operation data;

[0135] Automatically adjust resource allocation based on business needs and cluster resource conditions to improve resource utilization;

[0136] When a failure occurs, quickly locate the cause of the problem and initiate appropriate recovery measures.

[0137] In an optional embodiment of the present application, the method further includes:

[0138] Draw a topology diagram for the visual network structure of each infrastructure in each cluster, graphically displaying each node in the network structure and the connection relationship between them, making the network structure and connection method clear at a glance;

[0139] Analyze and evaluate network performance through topology diagrams to help identify possible bottlenecks or failure points;

[0140] Plan network expansion with topology diagrams to ensure new devices and connections can be effectively integrated into the existing network structure.

[0141] When a network problem occurs, the topology map can be used to quickly identify the root cause of the problem, such as a hardware or software failure, thereby speeding up problem resolution.

[0142] Simplify the configuration process through topology diagrams, which make it easier to understand and remember the configuration of network components through intuitive graphical forms.

[0143] In an optional embodiment of the present application, the method further includes:

[0144] Provides alarm rule setting services, such as triggering an alarm when a certain indicator exceeds or falls below a certain threshold;

[0145] Once an abnormal situation is detected, the alarm system will quickly issue an alarm and notify the operation and maintenance personnel through various means such as email reminders and text messages. The problem can be quickly located and diagnosed through detailed alarm information.

[0146] Save historical alarm information for operation and maintenance personnel to query and analyze. By analyzing historical alarm data, potential problems in the system can be discovered and prevented and optimized in advance.

[0147] In an optional embodiment of the present application, see Figure 3 , the method further comprises:

[0148] Through agent configuration management, a unified configuration interface is defined to receive configuration requests from users. These interfaces include but are not limited to: basic information configuration of agent nodes (such as IP address, port number, operating system type, etc.), data collection configuration (such as collection items, collection frequency, data format, etc.), alarm policy configuration (such as alarm threshold, alarm method, etc.), etc. By defining a unified interface, unified configuration of agent nodes in different clusters, different operating systems, and different network environments can be achieved.

[0149] Automated deployment of agent nodes is achieved through the agent configuration management's automated deployment tool. Upon receiving a user's configuration request, the module generates a deployment script based on the parameters defined in the configuration interface and pushes the script to the target agent node for execution through the automated deployment tool. During the deployment process, the module monitors the deployment progress and results in real time to ensure successful deployment and startup of the agent node.

[0150] The following is an application example of the cross-cluster infrastructure automated monitoring method of this application, including the following steps:

[0151] The first step is to deploy data collection components: Deploy data collection components in each cluster and configure the corresponding collection rules and parameters. Based on the cluster type, environment, and monitoring requirements, select appropriate data collection components and develop a deployment strategy to ensure efficient and stable data collection. Data collection components generally refer to software tools or service modules designed to automatically collect monitoring data from cross-cluster infrastructure (such as physical servers, virtual machines, containers, databases, network devices, etc.). These components serve as the data source for the monitoring platform, responsible for collecting, extracting, and forwarding various monitoring metrics and log information, converting it into a format suitable for analysis, alerting, and visualization. These components are typically lightweight agents that can run on different computing nodes (such as servers, virtual machines, or containers) and communicate with the monitoring system's backend. Non-agent components can connect to servers through protocols to collect infrastructure data. These protocols eliminate the need to install additional management software on servers or other network devices, simplifying management processes, improving efficiency, and reducing maintenance costs. Furthermore, these protocols provide interoperability between devices from different vendors, allowing administrators to use unified tools to manage the entire network environment.

[0152] Step 2: Configure the automatic discovery mechanism: To ensure that the platform can discover the infrastructure in the cluster in real time, you need to configure the automatic discovery mechanism. Set the scanning period and scanning range of the automatic discovery mechanism to ensure that the platform can discover the infrastructure in the cluster in real time.

[0153] Step 3: Start Data Collection: After configuring the data collection components and the auto-discovery mechanism, you can start the data collection process. Start the data collection component to collect monitoring data from each cluster, ensure that the component is functioning properly, start collecting monitoring data, and send the data to the central storage system.

[0154] Step 4: Configure the data analysis and display components: Configure the display rules and styles for the data analysis and display components based on user needs. Clarify the display requirements and monitoring objectives, including the monitoring indicators to be displayed, chart types, and the time range for data display. Also configure the display rules for the data analysis and display components.

[0155] Step 5: User Access and Monitoring: Users can view the operating status and performance indicators of each cluster through a unified monitoring view. The platform also provides an alarm function that automatically triggers an alarm to notify the user when monitoring data exceeds the preset threshold.

[0156] The cross-cluster infrastructure automated monitoring method of the present application establishes a unified monitoring architecture for cluster infrastructures across different regions, different cloud environments, or different technology stacks, solving the problem that traditional monitoring systems find it difficult to monitor multi-cluster resources from a global perspective. It enables automated collection of cross-regional and cross-cluster hybrid infrastructures. Regardless of where the infrastructure is located or whether it belongs to different clusters or platforms, the platform can perform unified and automated data collection, greatly improving the efficiency and scope of monitoring. It automatically adjusts the data collection frequency and resource allocation according to the cluster load, ensuring that the monitoring system itself does not become a burden on the cluster, while optimizing the overall monitoring efficiency.

[0157] The cross-cluster infrastructure automated monitoring method of the present application automates the monitoring of cross-regional and cross-cluster hybrid infrastructure, covering multiple types of monitoring data indicators such as server hardware, network equipment, storage area network storage equipment, cloud services, power supply, etc., to achieve comprehensive, efficient and automated monitoring management, understand the utilization of existing resources, discover resource bottlenecks and waste, thereby optimizing resource allocation, improving resource utilization efficiency, and greatly improving the monitoring scope and efficiency. The comprehensive monitoring capabilities ensure that no matter where the infrastructure is distributed, it can be effectively included in the monitoring scope, greatly improving the efficiency and accuracy of monitoring. Through automated data collection and monitoring, the manual participation links are significantly reduced, and the workload of operation and maintenance personnel is reduced. The unified visual interface enables the operation and maintenance team to quickly identify problems and understand the health status of the cluster, so as to respond faster and improve operation and maintenance efficiency. It can immediately discover performance bottlenecks and abnormal behaviors, automatically trigger early warnings, and help the operation and maintenance team take action before problems affect the business, reduce the impact of failures, and improve service quality.

[0158] It should be understood that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0159] See Figure 4 One embodiment of the present application provides a cross-cluster infrastructure automated monitoring device 400, including:

[0160] A deployment module 410 is used to deploy a preset data collection component in each infrastructure of each cluster;

[0161] Filling module 420, for collecting the operating data of each infrastructure of each cluster through a preset data collection component, and filling the operating data of the infrastructure into a preset monitoring template;

[0162] The standardization module 430 is used to standardize the operation data in each preset monitoring template to obtain standardized data;

[0163] The display module 440 is configured to receive an input operation data access request and display corresponding standardized data based on a display rule in the operation data access request.

[0164] For the specific limitations of the apparatus 400, please refer to the limitations of the pseudo-random sequence generation method described above and will not be repeated here. Each module in the apparatus 400 may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in hardware form, or may be stored in a memory in a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0165] The cross-cluster infrastructure automated monitoring device of this application solves the problem of infrastructure management and monitoring in a multi-cluster environment. Through automated collection mechanisms, rich monitoring indicators, data processing and analysis, and cross-cluster collaborative management functions, it achieves comprehensive monitoring of infrastructure such as servers, networks, and storage in the cluster, and provides accurate early warnings and fault handling suggestions through intelligent analysis algorithms, greatly improving the operation and maintenance efficiency and stability of the cluster, providing users with efficient and reliable monitoring services, and ensuring the stable operation of the business.

[0166] The cross-cluster infrastructure automated monitoring device of the present application designs corresponding monitoring strategies and collection schemes for different types of components, comprehensively covers all objects that need to be monitored, realizes real-time collection and efficient transmission of cross-cluster data, and ensures the timeliness and accuracy of monitoring data; at the same time, by expanding the monitoring scope, covering all key infrastructure components, and eliminating monitoring blind spots; in the setting of monitoring indicators, comprehensive consideration will be given to performance indicators, business indicators and user experience indicators to ensure the comprehensiveness and richness of monitoring data, through unified management and intelligent analysis functions, simplifying the configuration and management process of the monitoring platform, improving the intelligence level of monitoring, and solving the complexity and diversity problems of cross-cluster infrastructure by optimizing system architecture and data processing algorithms, ensuring the stability and scalability of the monitoring platform.

[0167] The cross-cluster infrastructure automated monitoring device of this application improves data collection efficiency: it adopts automated collection technology to achieve fast and accurate data capture of multiple cluster infrastructures. Through intelligent algorithms and efficient data processing mechanisms, the platform can significantly reduce manual intervention, reduce data collection costs, and improve data accuracy and real-time performance; achieve comprehensive monitoring: it can not only monitor traditional infrastructure indicators such as CPU, memory, disk usage, etc., but also comprehensively monitor network status, application performance, security events, etc. This cross-cluster and cross-level monitoring capability enables a comprehensive grasp of the operating status of the entire system, and timely discovery and handling of potential problems; enhance cross-cluster data consistency and integration capabilities: unified data models and protocols ensure that data from different clusters can be collected, integrated and analyzed under unified standards, providing a global perspective monitoring view, which facilitates operation and maintenance personnel to fully understand and manage the status of the entire distributed system; real-time early warning and fault location: it has powerful data analysis capabilities, can monitor the operating status of the infrastructure in real time, and trigger early warnings based on preset thresholds or rules. Once a failure or abnormal situation occurs, the problem can be quickly located, and detailed fault information and handling suggestions can be provided to users, thereby shortening the fault recovery time and reducing the risk of system downtime; Visual management and operation: Provides a rich visual interface and tools, allowing users to intuitively understand the operating status of the cluster infrastructure. At the same time, it also supports customized monitoring items, alarm rules, etc. to meet the personalized needs of different users. In addition, through visual operations, administrators can configure, manage and optimize the system more conveniently; Flexible expansion and integration: Taking into account the actual needs and scenario differences of different users, the platform has designed a flexible expansion mechanism that can easily integrate third-party monitoring tools or systems. In addition, it also supports multi-tenant mode, which can meet the resource sharing and isolation needs between different organizations or departments.

[0168] In one embodiment, a computer device is provided, wherein the internal structure diagram of the computer device can be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a pseudo-random sequence generation method as described above is implemented. It includes: a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, any step in the pseudo-random sequence generation method as described above is implemented.

[0169] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any step in the above pseudo-random sequence generation method can be implemented.

[0170] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0171] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0172] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0173] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0174] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0175] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A cross-cluster infrastructure automated monitoring method, characterized in that: include: Deploy preset data collection components in each infrastructure of each cluster; The preset data collection component collects the operating data of each infrastructure in each cluster, and fills the operating data of the infrastructure into the preset monitoring template; Standardize the operating data in each preset monitoring template to obtain standardized data; Receive an input operation data access request, and display corresponding standardized data based on the display rules in the operation data access request, The deployment of a preset data collection component in each infrastructure of each cluster includes: Building a central monitoring platform, where the central monitoring platform is used to centrally store, process, and display data collected from each infrastructure in different clusters; Select the corresponding preset data collection component according to the type of each infrastructure in each cluster, and ensure that the preset data collection component supports the automatic discovery mechanism. When there are more than one infrastructures of the same type in the same cluster, the deployment of the preset data collection components in each cluster includes: If the physical location distribution of all infrastructures exceeds the preset range, or the network conditions of the infrastructures meet the preset network requirements, a data acquisition server and a data acquisition component are deployed on each infrastructure. The data acquisition server is used to perform preliminary processing on the local data collected by the data acquisition component and send the processed data to the central monitoring platform. For the same type of infrastructure in the same cluster, if the physical locations of all infrastructures are distributed within a preset range, or the network conditions of the infrastructure do not meet the preset network requirements, a data acquisition server is set between the central monitoring platform and the infrastructure. The data acquisition server is used to collect the operating data of all infrastructures and perform preliminary processing. All infrastructures are connected to the data acquisition server, and the data acquisition server is connected to the central monitoring platform. Data acquisition components are deployed on each infrastructure and data acquisition server. The collection of all infrastructure operation data and preliminary processing include: Obtain the operating data of all infrastructure and extract the feature vectors of the pre-processed operating data; Build an operation data sensitivity scoring model to evaluate the sensitivity of operation data and classify it, and encrypt sensitive data based on the classification results; Decrypt the exchanged data and store the running data, build a visual interface to display the running data in real time, A sensitivity scoring model is constructed to evaluate the sensitivity score S(x) of the feature vector of application system information data. The formula is: Collect the sensitivity scores of historical application system information data and set an evaluation threshold. Compare the sensitivity scores of the application system information data feature vectors with the evaluation threshold. If the sensitivity score of the application system information data feature vector is greater than or equal to the evaluation threshold, it is determined to be sensitive data. If the sensitivity score of the application system information data feature vector is less than the evaluation threshold, it is determined to be ordinary data. Combining the RBF kernel function with the integral, the cumulative similarity A(x) between the benchmark data vector and the feature vector of the application system information data is calculated. The formula is: Where x is the application system information data feature vector, x0 is the benchmark data vector, and x i is the historical feature vector of the i-th application system information data; Perform logarithmic transformation B(x) on the accumulated similarity A, the formula is: B(x)=log(1+A(x)); The hyperbolic tangent function is introduced to smooth the cumulative similarity of the reference feature vector to obtain the smoothed cumulative similarity C(x), which is expressed as follows: Where M is the number of reference eigenvectors, x j is the jth reference eigenvector.

2. The method according to claim 1, characterized in that The method of selecting a corresponding preset data collection component based on the type of each infrastructure in each cluster to ensure that the preset data collection component supports the automatic discovery mechanism includes: Based on the type of infrastructure in each cluster, use the preset automated deployment tool to obtain the data collection components corresponding to the current infrastructure from the component database of the central monitoring platform; Deploy the acquired data collection components to the corresponding infrastructure, and configure the infrastructure according to the configuration template corresponding to the infrastructure; The data acquisition component deployed to the infrastructure is configured so that the data acquisition component is connected to the central monitoring platform, and corresponding data acquisition sources and data acquisition targets are set on the data acquisition component.

3. The method according to claim 1, characterized in that The method further comprises: Configure the auto-discovery mechanism for each cluster; Based on the automatic discovery mechanism, the infrastructure in the cluster is automatically discovered using a preset cluster management tool, and the discovered infrastructure information is stored in the configuration management database of the central monitoring platform.

4. The method according to claim 3, characterized in that After automatically discovering the infrastructure in the cluster using the preset cluster management tool, the method further includes: Determining whether a data collection component corresponding to the infrastructure has been deployed on the discovered infrastructure; For infrastructures where data collection components have been deployed, the current infrastructure information is stored in the configuration management database of the central monitoring platform; For infrastructure that has not deployed data collection components, first deploy the preset data collection components in the current infrastructure, and then store the infrastructure information after deployment in the configuration management database of the central monitoring platform.

5. A cross-cluster infrastructure automated monitoring device, characterized in that: include: Deployment module, used to deploy preset data collection components in each infrastructure of each cluster; The filling module is used to collect the operating data of each infrastructure of each cluster through the preset data collection component, and fill the operating data of the infrastructure into the preset monitoring template; The standardization module is used to standardize the operating data in each preset monitoring template to obtain standardized data; A display module is configured to receive an input operation data access request and display corresponding standardized data based on the display rules in the operation data access request. The deployment of a preset data collection component in each infrastructure of each cluster includes: Building a central monitoring platform, where the central monitoring platform is used to centrally store, process, and display data collected from each infrastructure in different clusters; Select the corresponding preset data collection component according to the type of each infrastructure in each cluster, and ensure that the preset data collection component supports the automatic discovery mechanism. When there are more than one infrastructures of the same type in the same cluster, the deployment of the preset data collection components in each cluster includes: If the physical location distribution of all infrastructures exceeds the preset range, or the network conditions of the infrastructures meet the preset network requirements, a data acquisition server and a data acquisition component are deployed on each infrastructure. The data acquisition server is used to perform preliminary processing on the local data collected by the data acquisition component and send the processed data to the central monitoring platform. For the same type of infrastructure in the same cluster, if the physical locations of all infrastructures are distributed within a preset range, or the network conditions of the infrastructure do not meet the preset network requirements, a data acquisition server is set between the central monitoring platform and the infrastructure. The data acquisition server is used to collect the operating data of all infrastructures and perform preliminary processing. All infrastructures are connected to the data acquisition server, and the data acquisition server is connected to the central monitoring platform. Data acquisition components are deployed on each infrastructure and data acquisition server. The collection of all infrastructure operation data and preliminary processing include: Obtain the operating data of all infrastructure and extract the feature vectors of the pre-processed operating data; Build an operation data sensitivity scoring model to evaluate the sensitivity of operation data and classify it, and encrypt sensitive data based on the classification results; Decrypt the exchanged data and store the running data, build a visual interface to display the running data in real time, A sensitivity scoring model is constructed to evaluate the sensitivity score S(x) of the feature vector of application system information data. The formula is: Collect the sensitivity scores of historical application system information data and set an evaluation threshold. Compare the sensitivity scores of the application system information data feature vectors with the evaluation threshold. If the sensitivity score of the application system information data feature vector is greater than or equal to the evaluation threshold, it is determined to be sensitive data. If the sensitivity score of the application system information data feature vector is less than the evaluation threshold, it is determined to be ordinary data. Combining the RBF kernel function with the integral, the cumulative similarity A(x) between the benchmark data vector and the feature vector of the application system information data is calculated. The formula is: Where x is the application system information data feature vector, x0 is the benchmark data vector, and x i is the historical feature vector of the i-th application system information data; Perform logarithmic transformation B(x) on the accumulated similarity A, the formula is: B(x)=log(1+A(x)); The hyperbolic tangent function is introduced to smooth the cumulative similarity of the reference feature vector to obtain the smoothed cumulative similarity C(x), which is expressed as follows: Where M is the number of reference eigenvectors, x j is the jth reference eigenvector.

6. A computer device comprising: The method comprises a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of the cross-cluster infrastructure automatic monitoring method according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the cross-cluster infrastructure automatic monitoring method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Automatic monitoring operation and maintenance platform supporting cross-region and cross-cluster hybrid infrastructure

    CN116300564A

  • Alarm monitoring method, system and device in cloud native environment and medium

    CN117614853A