A Multivariate Real-time Measurement and Status Testing Method for Data Centers
By integrating high-precision sensors and building index structures in the data center, and combining machine learning models for intelligent analysis and dynamic scheduling, the problems of multivariate correlation analysis and dynamic scheduling in the data center are solved, and resource utilization efficiency and stability are improved.
Patent Information
- Application Number
- CN202510007759.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The prior art is difficult to realize correlation analysis and dynamic scheduling between multivariables in data centers, resulting in low resource utilization efficiency and slow failure response speed, which cannot meet the requirements of high efficiency and high stability.
By integrating high-precision temperature and humidity sensors and devices in the data center, the original data is collected for filtering, cleaning and formatting, the distributed file system is used to store and build an inverted index and B-tree index structure, and intelligent analysis is combined with machine learning models to dynamically schedule resource configuration.
Multivariable real-time monitoring and status testing of data centers are realized, resource utilization efficiency and operation stability are improved, and fault risk and energy consumption burden are reduced.
Smart Images

Figure CN119782187B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer technology and data testing, and particularly to a method for multi-variable real-time measurement and status testing for a data center. Background Art
[0002] With the expansion of the scale of data centers and the diversification of application requirements, the resource management and operation and maintenance efficiency of data centers have become important research directions. Currently, various sensors and devices in data centers provide rich environmental data and performance data, including key indicators such as temperature and humidity, power consumption, CPU utilization rate, and memory occupancy. However, traditional data collection and monitoring methods are mostly limited to the monitoring of single indicators, and it is difficult to achieve correlation analysis and dynamic scheduling between multiple variables, resulting in low resource utilization efficiency and slow fault response speed, and unable to effectively meet the requirements of modern data centers for high performance and high stability.
[0003] Therefore, intelligent monitoring and dynamic management technologies based on multi-source sensor data have emerged. By real-time monitoring and intelligent analysis of multiple variables in the data center, combined with machine learning algorithms to predict system state changes, and performing intelligent scheduling to achieve optimal resource allocation. This technology can not only improve the resource utilization rate of data centers, but also enhance the operation stability of the system, ensuring the reliability and continuity of data centers under complex environments and high load conditions. However, there is currently a lack of a complete method to achieve efficient data collection, analysis, and scheduling to adapt to the growing management needs of data centers. Summary of the Invention
[0004] The present invention provides a method for multi-variable real-time measurement and status testing for a data center to solve the problem of how to achieve multi-variable real-time monitoring, intelligent analysis, and dynamic scheduling based on multi-source sensor data and key performance indicators in the data center, and improve the resource utilization efficiency and operation stability of the data center.
[0005] A method for multi-variable real-time measurement and status testing for a data center includes:
[0006] Collecting raw data from high-precision temperature and humidity sensors, devices, and applications in the data center, and performing preliminary filtering, cleaning, and formatting processing on the raw data on edge computing devices to generate preprocessed data in a standardized JSON format carrying device IDs and timestamps;
[0007] Storing the preprocessed data in a distributed file system, and constructing an inverted index and a B-tree index structure based on the device IDs and timestamps;
[0008] By introducing a machine learning model, performing intelligent analysis on the stored data, predicting device status, and calculating the probability of failure;
[0009] The expression for predicting the device state is as follows:
[0010]
[0011] Among them, J[u] is a functional with respect to the function u; u is the prediction function of the device state; u t is the time derivative; is the spatial gradient; α, β, γ are model parameters; f(x, t) is the external excitation;
[0012] According to the intelligent analysis result, perform dynamic scheduling and optimal allocation operations of resources; the expression of the dynamic adjustment includes:
[0013]
[0014] Among them, J[g] represents an objective functional with respect to the function g; represents the joint integral over the time interval [t1, t2] and the spatial region S; X(t, s) is the state input vector of the system at time t and spatial position s; g(X(t, s)) is the decision function; Y(t) is the operation result or state change response of the system at time t; L(·) is the loss function.
[0015] Furthermore, the high-precision temperature and humidity sensor further includes:
[0016] An integrated high-precision temperature and humidity sensor continuously monitors the environmental conditions in each area of the data center, including the temperature and humidity at the inlet and outlet of the server room, cold and hot channels, and air conditioning system.
[0017] Furthermore, the content of introducing a machine learning model to perform intelligent analysis on the stored data further includes:
[0018] The data collected from the DCIM system includes temperature T, humidity H, current I, voltage V, fan speed F, CPU utilization U cpu , memory utilization U mem , disk I / O speed D, and network bandwidth usage N;
[0019] Normalize the data and decompose it into different frequency band signals using wavelet transform (WT).
[0020] Furthermore, the content of constructing an inverted index and a B-tree index structure based on the device ID and timestamp further includes:
[0021] Index construction includes identifying key fields and frequently queried attributes, such as device ID, timestamp, and exception code;
[0022] Create an index based on the key fields of the data using an inverted index, B-tree, or other efficient index structures;
[0023] Inverted indexes are suitable for text searches, and B-trees are suitable for range queries and sorting operations.
[0024] The key innovation points include:
[0025] (1) Multivariate data processing and storage: By filtering, cleaning, and standardizing the data, data consistency is ensured, and a distributed file system is used to achieve redundant backup and efficient indexing of large-scale data, providing support for complex data analysis.
[0026] (2) Intelligent analysis and prediction: Based on intelligent analysis using machine learning models, anomaly detection and prediction of key performance indicators can be performed to detect problems that may affect system stability in advance.
[0027] (3) Dynamic resource scheduling: According to the analysis results, the allocation of server, storage, and network resources is adjusted in real time to ensure the stable operation of the data center in a high-load environment, improving resource utilization efficiency and reducing the need for manual intervention.
[0028] The beneficial effects of the present invention at least include:
[0029] The present invention adopts a multivariate real-time measurement and status testing method for a data center. The raw data collected by various sensors in the data center is stored in a distributed file system after preliminary filtering, cleaning, and formatting. The use of inverted index and B-tree index structures realizes efficient retrieval and redundant backup of data. Compared with traditional methods, the present invention greatly improves data retrieval efficiency by constructing an optimized index structure and supports fast access to multi-dimensional data. In addition, the introduced machine learning model and automated operation and maintenance strategy can identify the abnormal status of key performance indicators in the data center and perform intelligent prediction on resource usage, thus realizing dynamic scheduling and optimized allocation of resources, effectively reducing the failure risk and energy consumption burden of the data center. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a schematic flowchart of a multivariate real-time measurement and status testing method for a data center provided by an embodiment of the present invention.
[0031] Figure 2 is a structural block diagram of a multivariate real-time measurement and status testing method for a data center provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0033] Referring to Figure 1 , which is a schematic flowchart of a multi-variable real-time measurement and status testing method for a data center according to an embodiment of the present invention. This process may at least include steps S100-S400:
[0034] S100. Collect raw data from sensors, devices, and applications in the data center, and perform preliminary filtering, cleaning, and formatting on the raw data to generate preprocessed data.
[0035] S200. Store the preprocessed data in a distributed file system, construct an inverted index and a B-tree index structure for efficient data retrieval and redundant backup.
[0036] S300. Through the introduction of machine learning models and automated operation and maintenance strategies, perform intelligent analysis on the stored data to identify abnormal states of key performance indicators in the data center and predict resource usage.
[0037] S400. According to the intelligent analysis results, perform dynamic scheduling and optimized allocation operations on resources.
[0038] Step S100 at least includes steps S110-S140:
[0039] S110. Automatically collect real-time and historical data from various devices and applications in the data center by integrating multiple sensor interfaces and APIs. Specifically:
[0040] ① Integrate high-precision temperature and humidity sensors to continuously monitor the environmental conditions in each area of the data center, including the server room, hot and cold aisles, and the inlet and outlet temperatures and humidities of the air conditioning system, to ensure that the equipment is in the best working environment; Power Monitoring Unit (PMU) and Uninterruptible Power Supply (UPS) interfaces: By directly connecting to the communication ports of the PMU and UPS, obtain real-time current, voltage, power consumption, and battery status information for monitoring the energy usage efficiency and stability of the data center; Use standard interfaces such as RESTful API and SNMP (Simple Network Management Protocol) to regularly or on-demand extract key performance indicators such as CPU utilization, memory occupancy, disk I / O, and network throughput from network devices such as servers, switches, and routers.
[0041] ②By setting a reasonable sampling frequency, the module can capture the real-time data streams of each device and sensor, including the peak CPU load of the server, instantaneous network traffic spikes, or sudden changes in ambient temperature; for long-term trend analysis and historical comparison, the module stores data records over a past period, which can be daily, weekly, or monthly averages, or snapshots triggered by specific events, such as device restarts or system failures.
[0042] ③Before the data is transmitted to the central processing unit, the edge device performs preliminary normalization on the collected raw data, such as unifying data units, converting data formats (e.g., from binary to JSON), and removing invalid or incorrect readings, to ensure data consistency and quality.
[0043] ④All data is encrypted during transmission to ensure data security and integrity, preventing interception or tampering during transit. Additionally, sensitive information such as personal identity data is anonymized at the collection stage to comply with data protection regulations.
[0044] ⑤The edge device has preliminary anomaly detection capabilities and can identify abnormal data points based on preset thresholds or historical behavior patterns, such as the server CPU utilization suddenly soaring above 95% or the ambient temperature exceeding the safe range, thus triggering an alarm or taking corrective measures in a timely manner.
[0045] S120. Using edge computing technology, preliminary processing of the collected data is performed near the data source, including data cleaning, anomaly detection, format conversion, and preliminary analysis. Specifically:
[0046] ①When data is collected from sensors, devices, and applications in the data center, the edge computing device immediately initiates the data cleaning process. This includes checking data integrity, verifying the existence of data fields, and excluding or correcting incorrect entries, such as handling missing temperature readings or abnormal CPU utilization values. Through data cleaning, the quality and consistency of the data are ensured, preparing for subsequent analysis and storage.
[0047] ②Using preset thresholds and dynamic baselines, the edge computing device monitors abnormal patterns in the data stream in real time. For example, when it detects that the CPU utilization of a certain server suddenly soars above 90% while historical data indicates that its normal range should be around 60%, the system will mark this event as abnormal. Anomaly detection helps quickly identify potential problems, such as device overload or system failure, so as to take measures in a timely manner to avoid interruptions in the operation of the data center.
[0048] ③ The edge computing device is responsible for converting the original data in different formats and encodings into a unified format for subsequent analysis and storage. For example, converting the temperature data received from different sensors from Celsius to Fahrenheit uniformly used within the system, or converting binary device status information into JSON format for easy reading and processing by the central server.
[0049] ④ After data cleaning, anomaly detection, and format conversion, the edge computing device performs preliminary data analysis. This includes calculating basic statistical metrics such as the mean, standard deviation, identifying data trends and periodic patterns, and conducting simple predictive analysis. For example, predicting the power consumption trend for the next hour based on historical data, or assessing the risk of overheating of the device in the next few hours. These preliminary analysis results not only contribute to immediate decision support but also reduce the amount of data that needs to be transmitted to the central server.
[0050] ⑤ To further reduce the burden of network transmission, the edge computing device compresses the processed data and generates a data summary. For example, compressing consecutive temperature readings into the average value per minute, or summarizing a series of device status updates into a one-time status report. Through data compression and summarization, it is ensured that only the processed and filtered useful data will be sent to the central server, significantly reducing the data transmission latency and the required bandwidth.
[0051] ⑥ Before the data leaves the edge device, to protect the security and privacy of the data, all transmitted data will be encrypted. Secure transmission protocols such as TLS / SSL are used to ensure that the data will not be intercepted or tampered with during the transmission from the edge device to the central server.
[0052] S130. According to predefined rules and algorithms, intelligently filter out useless and duplicate data, aggregate similar data to reduce storage requirements while maintaining data integrity. Specifically:
[0053] ① The data center operation and maintenance management system collects a large amount of data, including but not limited to device status, performance metrics (such as CPU utilization, memory usage, disk I / O), environmental monitoring data (such as temperature, humidity), network traffic, log records, etc. The intelligent filtering mechanism first screens out the data points that are of practical significance for operation and maintenance management and analysis according to predefined rules and algorithms, removing redundant or insignificant information. For example, if the consecutive CPU utilization readings do not change much in a short period of time, the intelligent filtering will only retain one representative value and ignore the remaining duplicate or similar data points, thus reducing the burden of data transmission and storage. In addition, for outliers or obviously incorrect data, the intelligent filtering will also eliminate them to avoid interfering with data analysis.
[0054] ②The aggregation function, based on intelligent filtering, further integrates similar or related data to reduce the storage space requirements. For example, the readings of multiple sensors of the same type within the same time period can be averaged or summed to generate a single aggregated value; for log data, the aggregation operation involves combining similar events into a single event, recording the occurrence times and timestamps of the events, which not only maintains the integrity of the data but also saves storage resources; during the aggregation process, the system also considers the time series nature of the data to ensure that time-related data is correctly combined and the time continuity of the data is not disrupted.
[0055] ③Predefined rules and algorithms are the basis for intelligent filtering and aggregation. These rules include but are not limited to the importance level of the data, data type, data timeliness, and specific business logics. For example, for the operating status data of critical devices, even if the data volume is large, complete records must be retained; while for non-critical environmental data, a more relaxed aggregation strategy can be used to reduce storage. In terms of algorithms, machine learning models are used to automatically identify data patterns and predict which data will become important in the future, thereby determining whether to retain or discard certain data points.
[0056] ④Even during the filtering and aggregation processes, the system must ensure that the integrity of the data is not damaged. This means that when the data is simplified or combined, the key attributes and relationships of the original data must be retained. For example, the aggregation operation should not affect the statistical characteristics of the data, and the filtering operation should not result in the loss of important information. To ensure this, the system will retain sufficient metadata to record the original state and processing process of the data so that it can be traced and restored when needed.
[0057] S140. Monitor changes in key metrics, and immediately trigger the warning mechanism when an anomaly is detected, notify the operation and maintenance personnel, and automatically start the corresponding emergency procedures. Specifically:
[0058] ①Collect data in real time from various sensors, devices, and applications in the data center, including but not limited to device status, performance metrics (such as CPU utilization, memory usage, disk I / O), environmental monitoring data (such as temperature, humidity), network traffic, log records, etc. Before reaching the real-time monitoring and warning module, this data has been preliminarily filtered and aggregated by the intelligent edge collection and preprocessing module to reduce noise and improve processing efficiency.
[0059] ②Use real-time analysis algorithms to continuously monitor the input data. These algorithms can identify patterns and trends in the data, especially in terms of device performance, network conditions, and environmental conditions. Anomaly detection algorithms judge whether the current data deviates from the normal range based on historical data and predefined thresholds. For example, if the CPU utilization suddenly soars to 95%, this is a warning signal indicating that the server is overloaded.
[0060] ③When the anomaly detection algorithm identifies an abnormal change in a key metric, the real-time monitoring and warning module responds immediately and triggers the warning mechanism. The warning mechanism includes but is not limited to sending alerts to the operations and maintenance team, recording abnormal events, automatically adjusting system configurations, or starting pre-set emergency procedures. Alerts can take various forms, such as emails, text messages, phone calls, or push notifications through an integrated communication platform, ensuring that operations and maintenance personnel can respond promptly.
[0061] ④In some cases, the warning mechanism is not limited to notifications only. It also automatically executes pre-set emergency procedures to mitigate the problem. For example, if the temperature in a certain area is detected to be too high, the system will automatically increase the power of the cooling system or reroute network traffic to relieve the load on the affected servers. Automated responses can significantly shorten the problem-solving time and prevent small problems from escalating into larger disasters.
[0062] ⑤After a warning occurs, the system records detailed information about the abnormal event, including the time, type, severity, and response measures of the anomaly. This information is used for subsequent analysis to help the operations and maintenance team understand the root cause of the problem and optimize the warning algorithm and emergency procedures. Based on the data of historical warning events, the system can also continuously learn and adjust to improve the accuracy of anomaly detection and the efficiency of response.
[0063] Step S200 includes at least steps S210 - S250:
[0064] S210. Use distributed file system and database technologies to achieve efficient storage and redundant backup of large-scale data. Specifically:
[0065] ①Receive the data stream that has been preliminarily cleaned, filtered, and aggregated from the intelligent edge acquisition and preprocessing module. This data includes but is not limited to device status, performance metrics, environmental monitoring data, network traffic, and log records, etc. Before the data enters the distributed storage system, additional processing is required, such as data format conversion, metadata addition, and consistency check, to ensure the integrity and readability of the data.
[0066] ②Use a distributed file system such as Hadoop HDFS or a distributed database such as Cassandra to split the data into multiple smaller data blocks. Each data block contains a part of the data and necessary metadata. These data blocks are then distributed and stored on multiple nodes in the network. Each node stores one or more data blocks, depending on the system configuration and the size of the data blocks.
[0067] ③To ensure the high availability and durability of data, the system creates multiple copies for each data block, usually storing each copy on different nodes. In this way, even if a certain node fails, the data can still be restored from the copies on other nodes, thus avoiding data loss. The data persistence strategy ensures that data is permanently saved and can be restored even after node restart or failure. This is usually achieved by regularly writing copies of data blocks to stable storage media (such as hard disks).
[0068] ④The distributed data storage system provides efficient indexing and query mechanisms, enabling operation and maintenance personnel and analysis modules to quickly retrieve the required data. This includes various retrieval methods such as keyword search, time range query, and data type filtering. The system also supports real-time access to data and backtracking of historical data, enabling the operation and maintenance team to monitor the status of the data center in real time and also conduct in-depth analysis of past events.
[0069] ⑤In a distributed environment, maintaining data consistency is a major challenge. The system adopts mechanisms such as consistent hashing, distributed locks, and distributed transactions to ensure data consistency among multiple nodes. The transaction processing ability enables the system to perform atomic operations among multiple data blocks, ensuring data integrity and transaction isolation.
[0070] ⑥The distributed data storage system itself is also equipped with a monitoring mechanism that can monitor the health status of nodes, the distribution of data blocks, and the data access patterns in real time. When a node failure or data block corruption is detected, the system can automatically trigger a fault recovery process, such as replicating the lost data blocks again, migrating the data blocks to healthy nodes, or adjusting the data distribution strategy to ensure the continuous operation of the system.
[0071] S220. Build an inverted index and a B-tree efficient index structure to optimize the data retrieval speed. Specifically:
[0072] ①Before the input data enters the intelligent data indexing module, it has been preliminarily cleaned, filtered, and aggregated by the intelligent edge collection and preprocessing module. This data includes, but is not limited to, device status, performance metrics, environmental monitoring data, network traffic, and log records, etc. The index construction process begins with an in-depth analysis of the data to identify key fields and frequently queried attributes, such as device ID, timestamp, exception code, etc. These fields will be used as the basis for index construction. The system uses an inverted index, B-tree, or other efficient index structures to create indexes based on the key fields of the data. The inverted index is particularly suitable for text search, which maps keywords to a list of all documents containing the word, while the B-tree is suitable for range queries and sorting operations.
[0073] ②As new data continuously flows in, the index must be updated in real time to reflect the state of the latest data. The system adopts an incremental update strategy, only modifying or adding index entries corresponding to the new data, rather than rebuilding the entire index, thus maintaining the real-time nature and efficiency of the index. Index maintenance also includes periodic reconstruction and optimization to prevent index bloat and performance degradation. For example, when the data volume increases, the system automatically adjusts the index partitioning strategy to maintain the compactness of the index structure and query efficiency.
[0074] ③The intelligent data index module is built with a query optimizer that can analyze query requests, select the optimal index path and query strategy to minimize the data scan range and improve query speed. The query optimizer also utilizes a caching mechanism to temporarily store frequently queried results, reducing the overhead of repeated queries and further enhancing the response speed.
[0075] ④Operation and maintenance personnel can submit query requests through a graphical interface or command-line tools. These requests involve device status within a specific time range, tracking of abnormal events, trend analysis of performance metrics, etc. The system provides flexible query syntax and filtering options, supporting exact matching, fuzzy search, range query, and aggregation operations, enabling operation and maintenance personnel to customize query conditions according to specific requirements.
[0076] ⑤The query results are quickly returned to the user in the form of data records, statistical summaries, charts, and visualization reports, etc., facilitating understanding and analysis by operation and maintenance personnel. The system also supports the export and sharing of results, facilitating collaboration and communication among team members.
[0077] S230. Dynamically adjust storage and computing resources according to the data volume and processing requirements, and support horizontal expansion.
[0078] Specifically:
[0079] ①The system continuously monitors the current resource usage situation, including storage capacity, CPU utilization, memory usage, and network bandwidth, etc., while collecting and analyzing the growth trend of the data volume and changes in processing requirements. Based on historical data and current trends, the system can predict future resource requirements, providing a basis for dynamic resource adjustment.
[0080] ②When it is detected that the resource usage is approaching the threshold or it is predicted that the resource demand will exceed the existing capacity, the system automatically triggers the resource expansion mechanism. According to the results of demand assessment, the system can dynamically increase storage nodes or computing nodes to expand storage capacity and computing power. For example, when the data volume surges, the number of DataNodes in Hadoop HDFS or the number of nodes in a Cassandra cluster can be automatically increased to increase storage space. For computing requirements, the system adds Worker nodes of Apache Spark or MapReduce, or automatically scales virtual machine instances in a cloud environment to improve the efficiency of data processing and analysis.
[0081] ③As resources are expanded, the system automatically performs load balancing to ensure that the newly added nodes can evenly share data storage and processing tasks, avoiding the emergence of hotspots and bottlenecks. The data redistribution strategy ensures that data is evenly distributed among nodes to improve the efficiency of data access and the overall performance of the system. For example, data blocks are redistributed through a hashing algorithm to ensure that the amount of data on each node is roughly equal.
[0082] ④When the data processing demand decreases or the data volume stabilizes, the system can identify the situation of resource surplus and automatically recycle unnecessary resources to avoid resource waste and reduce operating costs. The resource recycling process includes reducing queued computing nodes, decreasing the number of storage nodes, or releasing unused storage space, while ensuring that the integrity of the data and the stability of the system are not affected.
[0083] ⑤The elastic expansion capability module adopts an adaptive adjustment strategy and automatically increases or decreases resources according to the real-time monitored system status and pre-set policies. The system can learn and optimize resource allocation policies, and over time, continuously improve the accuracy and efficiency of resource scheduling through machine learning algorithms.
[0084] S240. Integrate a big data analysis framework to support batch and real-time data analysis, and use statistical, machine learning, and deep learning technologies to mine potential patterns and trends in the data. Specifically:
[0085] ①Load the data to be analyzed from the distributed storage system, which includes device status, performance metrics, environmental monitoring data, network traffic, and log records, etc. Before entering the analysis process, the data needs to be preprocessed, such as filling missing values, detecting outliers, and converting data types, to ensure the quality and consistency of the data.
[0086] ②Use big data analysis frameworks such as Apache Spark or Flink to execute batch data analysis tasks. These tasks include data cleaning, data aggregation, statistical analysis, and pattern recognition, etc. The RDD (Resilient Distributed Dataset) and DataFrame API of Apache Spark provide efficient data processing capabilities and can handle data volumes in the petabyte range, while Flink is good at stream processing and can process continuously updated data streams in real time. Batch data analysis is usually used for in-depth mining of historical data, such as trend analysis, periodic pattern recognition, and anomaly detection, to assist in operation and maintenance decision-making.
[0087] ③For scenarios that require immediate response, such as monitoring system health status or detecting real-time anomalies, real-time data analysis is crucial. Flink's stream processing engine can process data streams in real time and provide low-latency analysis results, such as real-time monitoring of CPU utilization, network traffic, or temperature changes, etc. Real-time data analysis supports rule-based anomaly detection and immediate alerts to ensure that operation and maintenance personnel can quickly respond to potential problems.
[0088] ④Combine machine learning and deep learning algorithms to perform pattern recognition and predictive analysis on data. For example, use regression analysis to predict the remaining useful life of equipment, or use clustering algorithms to identify energy consumption patterns within a data center. Deep learning techniques, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), can be used to identify complex patterns in image or time series data, such as abnormal vibration patterns before equipment failure or signs of network attacks.
[0089] ⑤The results of data analysis and mining are presented in an intuitive form, such as charts, dashboards, or reports, to help operation and maintenance personnel understand the meaning behind the data. The analysis results can be used to optimize resource allocation, predict equipment maintenance requirements, improve energy efficiency, or enhance the security protection of the data center. The decision support system provides automated suggestions or triggers predefined operation and maintenance actions based on the data analysis results, such as automatically restarting services, adjusting the cooling system, or optimizing network traffic.
[0090] S250. Provide intuitive data visualization tools and automatically generated reports to help operation and maintenance personnel understand the meaning behind the data and support decision-making. Specifically:
[0091] ①First, extract the processed data from the elastic distributed storage system. This data includes, but is not limited to, equipment status information, performance metrics, environmental monitoring data, network traffic statistics, and various application logs. The data comes from multiple sources, including sensors, servers, network devices, and applications, so data integration is required to ensure data consistency and integrity.
[0092] ②Before data visualization, the data needs to be further cleaned and transformed to meet the requirements of specific visualization tools and report templates. This stage involves operations such as adjusting data formats, scaling numerical ranges, smoothing time series, and filtering outliers, ensuring the quality of the data presented finally.
[0093] ③Use data visualization tools (such as Tableau, Grafana, or Kibana) to design and generate charts, including line charts, bar charts, pie charts, heatmaps, scatter plots, and geographic distribution charts, etc. The specific types depend on the data characteristics and analysis objectives. Chart design needs to consider color coding, labels, and annotations to enhance readability and information conveyance, enabling operation and maintenance personnel to quickly identify key trends and anomalies.
[0094] ④Reports can be generated automatically at regular intervals or dynamically according to user requests, and include content such as key performance indicators (KPIs), summaries of abnormal events, historical trend comparisons, and predictive analytics. Custom reports support flexible data filtering and grouping, and operation and maintenance personnel can view data according to specific time periods, device types, or geographical locations, so as to deeply analyze the operation conditions in specific fields.
[0095] ⑤Provide interactive exploration functions, allowing operation and maintenance personnel to deeply explore data details through operations such as drilling down, rolling up, and filtering, and discover hidden associations and patterns. The data visualization interface and reports support multi-platform access, facilitating information sharing and collaboration among operation and maintenance team members, ensuring that all relevant parties can make decisions based on the latest data.
[0096] ⑥Data visualization and reports are not only tools for information display, but also can be part of a decision support system. By setting thresholds and warning mechanisms, they can automatically trigger operation and maintenance strategies or notifications, such as equipment maintenance reminders, resource optimization suggestions, or security event responses.
[0097] Step S300 includes at least steps S310 - S350:
[0098] S310. Use a machine learning model to predict the failure probability of equipment, plan maintenance strategies in advance, and reduce unplanned downtime. Specifically:
[0099] ①First, the data collected from the DCIM system includes temperature T, humidity H, current I, voltage V, fan speed F, CPU utilization U cpu 、memory utilization U mem 、disk I / O speed D, network bandwidth usage N, etc. Standardize this data and use wavelet transform (WT) to decompose it into different frequency band signals to identify patterns at different time scales.
[0100] ②For each signal x(t), apply the wavelet transform W x (a, b), where a is the scale parameter and b is the translation parameter. The wavelet coefficients are given by:
[0101]
[0102] where ψ * is the complex conjugate of the wavelet function. This step will help extract features at different scales, thus capturing the changing patterns of the device state.
[0103] ③Using the wavelet coefficients and the original data as inputs, construct a dynamic model based on partial differential equations (PDEs) to predict the state of the device. Define a functional J, which represents the change of the device state over time:
[0104]
[0105] where u is the predicted function of the device state, u t is the time derivative, is the spatial gradient, α, β, γ are model parameters, and f(x, t) is the external excitation, i.e., the input function composed of the wavelet coefficients and the original data. Solve for u by minimizing the functional J.
[0106] ④To solve the above functional minimization problem, use the variational method and the finite element method (FEM). This involves solving the variational derivative of the functional J with respect to u and setting it to zero to obtain a system of partial differential equations. Discretize this system of equations using the new method and find the optimal solution u through an iterative algorithm.
[0107] ⑤With the predicted device state u, the failure probability P f can be calculated by comparing the deviation of the predicted state from the normal operating range. If the deviation exceeds a predetermined threshold θ, a maintenance alert is triggered.
[0108]
[0109] where u0 is the normal state of the device.
[0110] ⑥Establish a closed-loop feedback mechanism to adjust the model parameters α, β, γ according to the results of maintenance actions to improve the prediction accuracy. This involves using the Bayesian update rule to update the prior distribution of the model parameters based on the actual effects of the maintenance actions.
[0111] S320. Analyze the resource usage of the data center and propose optimization suggestions, such as adjusting server configurations, optimizing network traffic, or improving energy usage efficiency. Specifically:
[0112] (I) Algorithm design:
[0113] ① Let \(X = \{X_1, X_2, \ldots, X\ n \}\) be a matrix of real-time and historical data collected from the data center, where each \(X\ i \) represents metrics such as CPU utilization, memory usage, network traffic, energy consumption, etc. Let \(n\) be the number of features and \(m\) be the number of observation samples, then \(X\) is an \(m\times n\) matrix.
[0114] ② Use a non-linear dynamics model to describe the variation of resource usage in the data center over time. Let \(y(t)\) represent the resource usage level at time \(t\), and \(u(t)\) be the control vector, representing adjustable parameters such as server configuration, network bandwidth allocation, etc.
[0115]
[0116] Here \(f\) is a non-linear function that describes the relationship between the resource usage state and time and the control vector.
[0117] ③ Define the objective function \(J\), aiming to minimize the resource usage cost \(C\) while ensuring the maximization of the quality of service \(QoS\) and energy efficiency \(EE\):
[0118] \(J(u)=w_1C(u)-w_2QoS(u)-w_3EE(u)\)
[0119] where \(w_1\), \(w_2\), \(w_3\) are the weight coefficients of cost, quality of service, and energy efficiency respectively.
[0120] ④ Due to limited resources, constraint conditions need to be set to ensure that resources are not over-allocated or below the threshold:
[0121] \(g\) i (u)\leq0, i = 1, 2, \ldots, k
[0122] Here \(g\) i represents the \(i\)-th constraint condition, such as the maximum server configuration limit, network bandwidth upper limit, etc.
[0123] ⑤ Use the Lagrange multiplier method to transform the constrained optimization problem into an unconstrained optimization problem:
[0124]
[0125] Solve for the optimal value of \(u\) through gradient descent or Newton-Raphson method, that is:
[0126]
[0127] where represents the gradient with respect to \(u\).
[0128] ⑥ The obtained optimal control vector \(u\) *Implement it in the resource management of the data center, and at the same time monitor the actual effect and adjust u and the weight w to form a closed-loop control system.
[0129] (2) Operation process
[0130] ① Data collection and feature extraction: Obtain the real-time and historical data in X from the data center, and perform preprocessing and feature extraction.
[0131] ② Dynamic system modeling: Use a non-linear dynamics model to describe the resource usage situation.
[0132] ③ Construct the objective function: Define J(u) to quantify the economy and efficiency of resource usage.
[0133] ④ Set the constraint conditions: Ensure that g i (u) ≤ 0 to prevent overuse of resources.
[0134] ⑤ Solve the optimal control vector: Use the Lagrange multiplier method and optimization algorithm to solve u * .
[0135] ⑥ Implementation and feedback: Implement u * and monitor the effect, and adjust the strategy if necessary.
[0136] S330. Based on the preset rules and thresholds, the system automatically executes operation and maintenance operations, such as restarting services, adjusting resource allocation, or sending alerts. Specifically:
[0137] Data input source and metrics: X t represents the vector of all sensor data obtained from the data center at time t, including CPU utilization, memory usage, disk I / O rate, network traffic, and environmental temperature, etc. Y t represents the operation result executed by the system or the change in the data center state, such as resource allocation, cooling system adjustment, or energy consumption level.
[0138] ① In order to accurately simulate the dynamic changes inside the data center, a new system of multivariate differential equations is established to describe the state changes of the data center, which takes into account various factors in time t and space s, including the dynamic changes of computing resources, storage resources, network resources, and power resources. The system of equations is as follows:
[0139]
[0140] where f i is a non-linear function describing the change rate of the i-th type of data, which includes the physical characteristics of the data center, the influence of operation instructions, and external environmental changes; X i(t, s) represents the state of the i-th type of resource; X(t, s) is a vector containing the states of all resources; Y(t) is the operation result executed by the system; t is the time; s is the spatial location of the data center.
[0141] ② To find the optimal decision function g(X(t, s)), functional analysis is used to find the function g such that a certain objective functional J[g] reaches a minimum value, that is:
[0142]
[0143] where L is the loss function, which measures the quality of the decision-making effect, S is the spatial domain of the data center; g(X(t, s)) is the decision function; L is the loss function, which is used to measure the quality of the decision-making effect; J[g] is the objective functional, representing the overall effect of the decision function; represents the joint integral over the time interval [t1, t2] and the spatial region S.
[0144] ③ To solve the above differential equations and functional extremum problems, high-order partial derivatives are used to analyze the sensitivity of the system state to time and space, as well as its dependence on the decision function. Numerical analysis methods, such as the finite difference method or the spectral method, are used to approximately solve these complex equations.
[0145]
[0146] where h i is a function describing the second-order rate of change.
[0147] ④ The reinforcement learning framework is combined with the decision tree algorithm to construct a new intelligent decision-making model. Reinforcement learning learns a policy function π(a|X t ) by interacting with the environment, where a is the action and X t is the state, with the aim of maximizing the long-term reward R. The decision tree helps the system understand the optimal action sequence in different states.
[0148]
[0149] where, r t is the reward at time t, T is the time step; R is the long-term reward.
[0150] ⑤ Dynamic programming is used to find the optimal future decision sequence given the current state X t and the historical decision sequence. At the same time, state estimation techniques are used to process noisy data and predict the future state X t+1 .
[0151] X t+1 = F(X t , u t) + w t
[0152] where F is the state transition function, u t is the control input, w t is the process noise.
[0153] ⑥ Dynamically adjust the algorithm parameters, such as the learning rate α and the exploration rate ∈, according to the decision-making effect and environmental feedback to ensure that the model can adapt to the dynamic changes of the data center.
[0154] α = α0 / (1 + ηt)
[0155] where α0 is the initial learning rate, η is the decay factor, t is the time step, and α is the learning rate.
[0156] ⑦ To enhance the robustness and computational efficiency of the algorithm, a new topology is also introduced to describe the network structure of the data center, and the connectivity and distance metrics in graph theory are used to optimize resource scheduling. In addition, group theory in abstract algebra is used to analyze the symmetry and reversibility of data center operations to help the system understand which operations can cancel or simplify each other.
[0157] G(V, E)
[0158] where G is the representation of the graph, V is the set of vertices, and E is the set of edges.
[0159] ⑧ Finally, the result Y of the executed operation t is fed back to the system and compared with the predicted result to evaluate the effectiveness of the decision. If there is a deviation between the actual result and the expectation, the model parameters are updated through the backpropagation algorithm to achieve self-optimization.
[0160]
[0161] where θ is the model parameter, Δθ is the parameter update amount, and η is the learning rate.
[0162] S340. Combine historical data and real-time analysis results to provide data-based decision-making suggestions to help the operation and maintenance team make more accurate decisions. Specifically:
[0163] (I) Algorithm model formula
[0164] Let D be the set of log data of all sensors, devices, and applications in the data center, where each data point d i contains multiple attributes such as CPU utilization x i,1 , memory usage x i,2 , disk I / O rate x i,3 , network traffic x i,4 , and environmental temperature x i,5 , etc.
[0165] ① First, preprocess the data, including standardization and denoising, to ensure data quality. The preprocessed data is represented as
[0166] ② Map the preprocessed data to a functional space F, where each data point d i is represented as a function f i (t), which reflects the resource usage over time t. The functional space allows the use of functional analysis tools to handle the time series characteristics of the data.
[0167] ③ Construct a multiple regression model R(·) to predict the relationship between resource usage and environmental variables:
[0168] y = R(x1, x2, x3, x4, x5) + ∈
[0169] where y represents the resource usage level and ∈ is the error term. The model can be expressed as:
[0170] y = β0 + β1x1 + β2x2 + β3x3 + β4x4 + β5x5 + ∈
[0171] ④ Use differential equations to simulate the trend of resource usage over time. Assume that the resource usage follows a certain dynamic law, which can be described by the following non - linear differential equation:
[0172]
[0173] where g(·) is a non - linear function that depends on the current resource usage level y and environmental variables x i .
[0174] (2) Algorithm model operation process
[0175] ① Collect log data from sensors, devices, and applications in the data center and preprocess it to ensure data accuracy and availability.
[0176] ② Map the data to the functional space F and use functional analysis methods to extract time series features.
[0177] ③ Use historical data to train the multiple regression model R(·) to determine the coefficients β i to accurately predict the relationship between resource usage and environmental variables.
[0178] ④ Solve the differential equation to obtain the predicted values of resource usage over time.
[0179] ⑤ Based on the prediction results, generate decision suggestions, such as adjusting resource allocation, optimizing workloads, or taking preventive measures.
[0180] ⑥ Regularly evaluate the decision-making effect, and update the model parameters according to new data to continuously optimize the prediction accuracy.
[0181] S350. Through continuous learning and self-optimization, the system can gradually improve the accuracy and efficiency of decision-making and adapt to the dynamic changes of the data center. Specifically:
[0182] ① The system continuously ingests real-time data from sensors, devices, and applications in the data center, including but not limited to key metrics such as CPU utilization, memory usage, disk I / O rate, network traffic, and environmental temperature. This data serves as the input for the algorithm model and forms the basis for self-learning and adaptation.
[0183] ② By implementing anomaly detection algorithms, the system can identify metric values that deviate from the normal range, which indicates potential problems or impending failures. At the same time, using pattern recognition technology, the system can discover hidden patterns in the data to provide a basis for prediction and decision-making.
[0184] ③ Based on functional analysis and multiple regression analysis, the model automatically adjusts its parameters according to the latest data. For example, in the differential equation for resource usage prediction, the non-linear function g(·) will be adjusted over time to reflect the actual changes in the data center environment. This dynamic adjustment ensures the continuous relevance and prediction accuracy of the model.
[0185] ④ The system adopts an online learning mechanism and can instantaneously update the model as new data flows in. This means that even when the operating conditions of the data center change, the system can quickly adapt and adjust its decision-making strategy without manual intervention.
[0186] ⑤ Decision trees and reinforcement learning algorithms are used to formulate specific action plans. Decision trees are based on a series of tests and conditions to guide the system's responses; while reinforcement learning enables the system to learn which behaviors are most beneficial for optimizing goals, such as maximizing resource utilization efficiency or minimizing energy consumption, through a reward mechanism.
[0187] ⑥ Through a feedback loop mechanism, the system continuously evaluates the effect of decisions, collects new data after implementing decisions, and uses it to further optimize the model. This includes monitoring whether resource usage reaches the expected goals and evaluating the system's response speed and decision-making accuracy.
[0188] ⑦ In addition to short-term online learning, the system also has a long-term memory function, which can accumulate historical data and lessons learned to form a knowledge base. This helps the system make faster and more accurate decisions when facing similar situations.
[0189] ⑧ Once a decision is made, the system can automatically execute corresponding operations, such as adjusting the cooling system settings, optimizing the workload distribution, or starting backup resources, to respond to changes within the data center.
[0190] Step S400 includes at least steps S410 - S450:
[0191] S410. Continuously monitor the physical and virtual resources of the data center, including computing, storage, network, and power, to ensure reasonable allocation and utilization of resources. Specifically:
[0192] ① The system automatically collects real - time data from all levels of the data center, including but not limited to: computing resources: CPU utilization, number of processes, thread load, etc.; storage resources: disk usage, I / O read - write rate, cache hit rate, etc.; network resources: bandwidth usage, latency, packet loss rate, etc.; power resources: power consumption, UPS status, generator readiness, etc.; This data is captured by sensors and monitoring devices and transmitted to the DCIM system for processing.
[0193] ② The collected data is efficiently stored and classified, and the DCIM system uses advanced data - processing techniques to analyze it. The system utilizes statistical analysis, trend prediction, and anomaly - detection algorithms to ensure that any unusual patterns or potential bottlenecks in resource utilization can be detected in a timely manner.
[0194] ③ The system evaluates the resource utilization status based on the data - analysis results and generates resource - utilization reports. These reports include the real - time status, historical trends, and prediction models of resources, helping the operations and maintenance personnel understand the resource - utilization efficiency and potential optimization opportunities.
[0195] ④ To adapt to the dynamic changes in the data center, the system dynamically adjusts the resource - utilization thresholds based on historical data and current load conditions. When the resource utilization approaches the threshold, the system triggers an early - warning mechanism to notify the operations and maintenance team to take corresponding measures.
[0196] ⑤ Based on the monitoring results, the system can propose resource - optimization suggestions, such as dynamically adjusting virtual - machine configurations, migrating workloads, optimizing storage strategies, or adjusting network architectures, to ensure efficient utilization of resources.
[0197] ⑥ For some common resource - allocation problems, the system can automatically execute resource - allocation strategies, such as automatically expanding computing nodes, dynamically adjusting the size of the storage pool, or intelligently allocating network bandwidth, to reduce the need for human intervention.
[0198] ⑦ In addition to real - time monitoring, the system also supports long - term resource planning. By analyzing historical data, the system can predict future resource requirements, helping the data - center managers plan for hardware upgrades or capacity expansions in advance.
[0199] ⑧ Resource monitoring not only focuses on resource - utilization efficiency but also regularly checks security policies and compliance requirements to ensure that all resources are managed in accordance with the established security standards and industry regulations.
[0200] S420 Dynamically adjust resource allocation according to the current load and predicted demand, such as dynamically scaling in or out cloud service instances, to balance resource utilization and cost. Specifically:
[0201] ① The intelligent scheduling module continuously collects real-time data from multiple dimensions of the data center, including but not limited to: computing resource metrics such as CPU utilization, memory usage, disk I / O rate; network performance metrics such as network traffic, latency, packet loss rate; storage resource metrics such as storage space utilization rate, read / write speed; physical environment metrics such as power consumption, ambient temperature; This data is collected through sensors and monitoring devices and then efficiently processed by the DCIM system to ensure rapid analysis and response of the data.
[0202] ② The system uses machine learning algorithms to deeply analyze the collected data, identify the current load patterns and historical trends, and then predict future resource requirements. The prediction model is trained based on historical data sets and can take into account the impacts of seasonal, periodic, and sudden business activities.
[0203] ③ Based on the analysis results of the current load and predicted demand, the intelligent scheduling module formulates resource allocation strategies. Including: Dynamic scaling out: When it is predicted that high load is about to occur, the system automatically increases the number of cloud service instances, such as starting additional virtual machines or containers, to cope with the upcoming increased workload; Resource scaling in: When it is detected that resources are idle or predicted that demand will decrease, the system automatically reduces cloud service instances to avoid unnecessary cost expenditures.
[0204] ④ After determining the resource adjustment strategy, the intelligent scheduling module directly interacts with the underlying infrastructure to perform dynamic scaling in or out of resources. This usually involves communicating with the API interface of the cloud platform to adjust virtual machine configurations, container orchestration, or storage allocation strategies.
[0205] ⑤ After resource adjustment, the system continues to monitor the adjustment effect to verify whether the expected performance goals and cost savings are achieved. If not, the system will re-adjust the resource allocation strategy based on new data, forming a continuously optimized feedback loop.
[0206] ⑥ The intelligent scheduling module also regularly generates cost analysis reports to show how resource adjustment affects the overall operating cost, helping data center managers understand the economic benefits brought by intelligent scheduling.
[0207] S430 Design and execute a fault recovery strategy to ensure that in case of a failure, it can quickly switch to standby resources and maintain the normal operation of the data center. Specifically:
[0208] ①The system continuously collects various key indicator data through sensors and monitoring devices throughout the data center, including but not limited to: Hardware health status: The operating status of servers, storage devices, and network devices; System performance indicators: CPU usage, memory occupancy, disk I / O, network throughput; Environmental conditions: Temperature, humidity, power supply status; This data is analyzed in real time to identify potential abnormal patterns or performance bottlenecks and trigger the warning mechanism.
[0209] ②Using big data analysis and machine learning technologies, the system can predict the nature of equipment failures and evaluate the impact of potential failures on the overall operation of the data center. This includes: Failure prediction: By analyzing historical data, identifying early signs leading to failures; Risk quantification: Evaluating the impact degree and recovery time under different failure scenarios and prioritizing the handling of high-risk events.
[0210] ③When a failure is detected, the system immediately starts an automated process to locate the failure source and isolate it to prevent the problem from spreading. This involves: Automatic isolation: Shutting down or disconnecting the faulty component to reduce the impact on other systems; Alarm notification: Sending an immediate notification to the operation and maintenance team, providing details of the failure and preliminary diagnosis.
[0211] ④The system is designed with a redundant architecture to ensure seamless switching to backup resources when the primary resources fail. This includes: Resource switching: Automatically transferring the workload to a backup server or storage array; Network redundancy: Activating backup network paths to ensure the continuity of data transmission.
[0212] ⑤For major failures or catastrophic events (such as natural disasters), the system executes pre-planned disaster recovery strategies, including: Off-site backup and recovery: Recovering data and applications from a remote data center; Business continuity plan: Enabling pre-defined business recovery processes to ensure the continuous operation of critical business functions.
[0213] ⑥After the failure is recovered, the system conducts a review and analysis of the entire event, identifies the root cause of the failure, and updates the failure recovery strategy and disaster recovery plan to improve the ability to handle similar events in the future.
[0214] S440. Analyze system performance bottlenecks and provide optimization suggestions, such as adjusting the cache strategy, optimizing database queries, or improving network configuration, to improve system response speed and user experience. Specifically:
[0215] ①The performance tuning module collects data from multiple sources in the data center, including but not limited to: computing resources: CPU utilization, memory usage, disk I / O read and write rates; network resources: network bandwidth usage, latency, packet loss rate; storage resources: disk space usage, storage I / O response time; applications: service response time, error rate, transaction processing speed; databases: query response time, index usage efficiency, cache hit rate; This data is first collected and preprocessed to remove noise, fill in missing values, and standardize the format for subsequent analysis.
[0216] ②The system uses big data analysis techniques and machine learning algorithms to deeply analyze the collected data and identify performance bottleneck points. This includes: trend analysis: identifying trends in resource usage over time and predicting future requirements; anomaly detection: finding patterns that deviate from normal behavior, which may be due to improper configuration or hardware failures; correlation analysis: exploring the mutual influence between different resources and applications to understand the overall system behavior.
[0217] ③Based on the analysis results, the system can locate the specific causes of performance degradation, such as: insufficient computing resources: CPU or memory overload; network congestion: high latency or low bandwidth; inefficient storage: frequent I / O operations or disk fragmentation; database performance issues: inefficient queries or missing indexes.
[0218] ④According to the diagnostic results, the system generates specific tuning suggestions, including: adjusting the cache policy: increasing the cache size, optimizing the cache replacement algorithm; optimizing database queries: creating or modifying indexes, rewriting query statements; improving network configuration: adjusting routing policies, optimizing network protocols; optimizing resource allocation: dynamically adjusting the allocation of computing resources such as CPU and memory; adjusting software configuration: modifying application configuration parameters such as thread pool size or connection timeout.
[0219] ⑤The operation and maintenance personnel implement the adjustments according to the tuning suggestions provided by the system. The system then monitors the effects of these changes to verify whether the performance bottleneck has been truly resolved and the system response speed and user experience have been improved.
[0220] ⑥Performance tuning is an ongoing process. The system continuously monitors the optimized performance metrics and repeats the above steps regularly to adapt to the dynamic changes and growing demands of the data center.
[0221] S450. Identify cost-saving opportunities through resource usage analysis, such as idle resource recovery, optimized energy usage, and procurement cost control. Specifically:
[0222] ① The cost optimization module collects relevant data from multiple perspectives of the data center, including but not limited to: Equipment status: The operating status and health metrics of servers, storage units, and network devices; Resource utilization: The real-time and historical usage of CPU, memory, storage, and network; Energy consumption data: Power consumption, energy consumption statistics of the cooling system; Environmental conditions: Temperature, humidity, air flow distribution, etc.; Maintenance records: Historical records of equipment maintenance and predictive maintenance requirements; Cost information: Equipment depreciation, electricity bills, maintenance contracts, and service fees; These data are integrated and preprocessed to form a unified data view for subsequent analysis.
[0223] ② Using big data analysis and machine learning techniques, the module deeply analyzes the collected data and establishes a cost model. The key points of the analysis are: Resource efficiency: Identifying which resources (such as servers, storage) are over-provisioned or under-utilized; Energy consumption optimization: Analyzing energy consumption patterns and identifying energy-saving opportunities, such as adjusting equipment operation time or improving thermal management; Cost trend: Predicting future costs based on historical data and market trends.
[0224] ③ Based on the results of data analysis, the module identifies opportunities to reduce costs: Recycling of idle resources: Automatically identifying resources that have not been used for a long time and proposing suggestions for reallocation or recycling to reduce waste; Optimization of energy use: Intelligently adjusting power supply and cooling strategies according to equipment load and environmental conditions to reduce energy consumption; Control of procurement costs: Optimizing the procurement timing of spare parts and new equipment through predictive analysis to avoid additional costs brought by inventory backlogs or emergency purchases.
[0225] ④ The operation and maintenance team takes actions according to the suggestions of the cost optimization module, such as: Adjusting resource allocation: Reallocating or shutting down idle servers and optimizing storage resources; Improving facility management: Implementing more efficient cooling solutions and adjusting the data center layout; Adjusting procurement strategies: Formulating procurement plans based on predictive analysis and negotiating more favorable contract terms; After implementation, the system continuously monitors changes in costs and resource usage to evaluate the effectiveness of optimization measures.
[0226] ⑤ Cost optimization is a continuous process. The system will regularly execute the above steps and make dynamic adjustments according to the operating conditions of the data center and external market conditions to ensure that the cost control strategy is always effective.
[0227] Key innovation points include:
[0228] (1) Multivariate data processing and storage: Through data filtering, cleaning, and standardization, data consistency is ensured, and a distributed file system is used to achieve redundant backup and efficient indexing of large-scale data, providing support for complex data analysis.
[0229] (2) Intelligent analysis and prediction: Based on machine learning models, intelligent analysis can perform anomaly detection and prediction on key performance indicators, and discover problems that may affect system stability in advance.
[0230] (3) Dynamic resource scheduling: According to the analysis results, it adjusts the allocation of server, storage, and network resources in real time, ensuring the stable operation of the data center in a high-load environment, improving resource utilization efficiency, and reducing the need for manual intervention.
[0231] The beneficial effects of the present invention at least include:
[0232] The present invention adopts a multi-variable real-time measurement and status testing method for a data center. The raw data collected by various sensors in the data center is stored in a distributed file system after preliminary filtering, cleaning, and formatting. The inverted index and B-tree index structures are used to achieve efficient data retrieval and redundant backup. Compared with traditional methods, the present invention greatly improves data retrieval efficiency by constructing an optimized index structure, supporting fast access to multi-dimensional data. In addition, the introduced machine learning model and automated operation and maintenance strategy can identify the abnormal status of key performance indicators in the data center and perform intelligent prediction on resource usage, thereby realizing dynamic scheduling and optimized allocation of resources, effectively reducing the failure risk and energy consumption burden of the data center.
[0233] Figure 2 It is a structural block diagram of a multi-variable real-time measurement and status testing method for a data center provided by an embodiment of the present invention. The system may include the following modules:
[0234] Data acquisition module 10. The data acquisition module is used to obtain real-time status information from various data sources in the data center, including temperature and humidity sensors, power monitoring units, device performance data (such as CPU utilization, memory occupancy, disk I / O), and network traffic. Through standardized interfaces (such as RESTful API, SNMP), it ensures that the system can automatically collect environmental and performance parameters from various devices and applications, providing diverse and timely data sources for subsequent data processing.
[0235] Edge computing and preprocessing module 20. The edge computing and preprocessing module is used to perform preliminary processing on the collected data near the data source, including operations such as data cleaning, anomaly detection, and format conversion. By using intelligent algorithms to screen out invalid or abnormal data, it effectively reduces the data transmission volume while ensuring data quality, ensuring the consistency and accuracy of the data entering the storage and analysis stages.
[0236] Data Storage and Indexing Module 30. The data storage and indexing module uses a distributed storage system to efficiently store the processed data, and constructs an inverted index and a B-tree structure to optimize the data's fast retrieval ability. The system performs distributed storage and redundant backup of the data to ensure high availability within the data center, and supports multi-dimensional queries based on time, device type, and data type.
[0237] Machine Learning Prediction and Anomaly Detection Module 40. The machine learning prediction and anomaly detection module is responsible for trend prediction of the key performance indicators (such as energy consumption, device utilization, etc.) of the data center through a trained machine learning model, and performs real-time anomaly detection based on historical data and preset thresholds. When an abnormal change is detected, the system will automatically trigger an alarm to prompt the operation and maintenance personnel to handle it in a timely manner, ensuring the safe operation of the data center.
[0238] Intelligent Scheduling and Resource Optimization Module 50. The intelligent scheduling and resource optimization module is used to intelligently adjust the resource configuration of the data center according to the real-time monitored data and prediction results. The system can dynamically expand or reduce cloud service instances, adjust the load distribution, and optimize resource allocation through an adaptive strategy to improve the resource utilization rate of the data center, reduce energy consumption, and reduce operating costs.
[0239] Data Visualization and Report Generation Module 60. The data visualization and report generation module provides an intuitive user interface and visualization tools to display multivariate data and analysis results to users in the form of charts, dashboards, etc. The system supports generating real-time and historical reports to help operation and maintenance personnel and managers comprehensively grasp the resource usage, performance bottlenecks, and optimization suggestions of the data center.
[0240] System Monitoring and Adaptive Module 70. The system monitoring and adaptive module continuously monitors the operating status of each module and uses adaptive learning technology to optimize the system parameters to ensure the efficient and stable operation of the system in a changing environment. This module can adjust the system settings according to the feedback to improve the accuracy and efficiency of decision-making and adapt to the dynamically changing requirements of the data center.
[0241] Through the collaborative work of multiple modules, the invention realizes multi-variable real-time measurement and status testing of the data center. The system not only improves the efficiency of data collection and processing, optimizes resource management and scheduling, reduces operation and maintenance costs, but also enhances the stability and security of the data center through intelligent monitoring and anomaly detection. The data visualization function provides a clear and intuitive management interface for users, helping managers quickly understand and apply the data analysis results, thus significantly improving the operation and maintenance efficiency and resource utilization rate of the data center.
[0242] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions thereof, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A multi-variable real-time measurement and status testing method for a data center, characterized in that It includes the following steps: Collect raw data from high-precision temperature and humidity sensors, devices, and applications within the data center, and perform preliminary filtering, cleaning, and formatting on the raw data on edge computing devices to generate preprocessed data in a standardized JSON format carrying device IDs and timestamps; Store the preprocessed data in a distributed file system, and construct an inverted index and a B-tree index structure based on the device IDs and timestamps; By introducing a machine learning model, perform intelligent analysis on the stored data, predict device status, calculate failure probabilities, and automatically execute operation and maintenance operations; The expression for predicting the device status is: Among them, is a functional with respect to the function ; is a prediction function of the device state; is the time derivative; is the spatial gradient; , , are model parameters; is the external excitation; The process of automatically executing operation and maintenance operations includes: Among them, is a non-linear function describing the change rate of the th class of data; represents the state of the th class of resources; is a vector containing all resource states; is the operation result executed by the system; is time; is the spatial location of the data center; Use functional analysis to find the optimal decision function, and the expression for the functional analysis is: Among them, represents an objective functional regarding the function ; represents the joint calculus over the time interval and the spatial region ; is the state input vector of the system at the current time , spatial position ; is the decision function; is the operation result or state change response executed by the system at the time ; is the loss function; According to the intelligent analysis results, perform dynamic scheduling and optimal allocation operations of resources.
2. The multivariate real-time measurement and status testing method for a data center according to claim 1, wherein The high-precision temperature and humidity sensor further includes: Integrate high-precision temperature and humidity sensors to continuously monitor the environmental conditions in various areas within the data center, including the server room, hot and cold aisles, and the inlet and outlet temperatures and humidity of the air conditioning system.
3. The multivariate real-time measurement and status testing method for a data center according to claim 1, wherein The content of introducing a machine learning model and performing intelligent analysis on the stored data further includes: The data collected from the DCIM system includes temperature T, humidity H, current I, voltage V, fan speed F, CPU utilization , memory usage , disk I / O speed D, and network bandwidth usage N; Standardize the data and decompose it into different frequency band signals using wavelet transform.
4. The method for multivariate real-time measurement and status testing for a data center according to claim 1, characterized in that The content of constructing an inverted index and a B-tree index structure based on device IDs and timestamps further includes: Index construction includes identifying key fields and frequently queried attributes, including device IDs, timestamps, and exception codes; Use an inverted index, B-tree, or other efficient index structures to create indexes based on the key fields of the data; The inverted index is suitable for text search, and the B-tree is suitable for range queries and sorting operations.
Citation Information
Patent Citations
Data management system for intelligent operation and maintenance service
CN117194399A
Computing power data management system and method based on distributed computing
CN119025283A