Server full life cycle monitoring and optimizing system
By configuring ultrasonic transceivers and phase jitter detection units at both ends of the server's airflow channel, the limitations of perception dimensions and response lag in existing technologies are solved, enabling low-cost, low-resource-consumption early fault warnings, which are suitable for monitoring and optimization throughout the server's entire lifecycle.
Patent Information
- Application Number
- CN202511079405.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-03
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies for server monitoring have limitations in perception dimensions, making it difficult to effectively detect early system-level and global physical anomalies. Response mechanisms are lagging, and advanced analysis methods are difficult to deploy universally at low cost, resulting in the difficulty in early detection of unknown types of physical faults.
An ultrasonic transmitter and receiver are configured at both ends of the server's airflow channel. Combined with a phase jitter detection unit and a baseline management module, the stability of the medium flow field in the airflow channel is continuously scanned. The XOR logic gate and counter are used to achieve early fault warning with low resource consumption, and it has the ability to self-diagnose and self-compensate.
It achieves overall stability awareness of the server's internal operating environment at low cost and low resource consumption, captures early physical faults of unknown types, avoids the risk of missed detections caused by sensor aging, and is applicable to various types of servers, including resource-constrained edge computing devices.
Smart Images

Figure CN120950334A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a server lifecycle monitoring and optimization system, belonging to the field of server fault prediction and health management technology. Background Technology
[0002] The currently prevalent technical approach is based on monitoring discrete physical quantities of key components within the server. For example, sensors deployed on processors or hard drives acquire parameters such as temperature, voltage, or vibration. Changes in these parameters are then analyzed to determine the system's health. This point-based sensing method, which monitors specific points, is intuitive and easy to implement, and has become standard practice in server health management. However, with increasing demands for server reliability and operational efficiency, the inherent limitations of this point-based sensing method are becoming increasingly apparent. As a highly coupled and complex thermodynamic system, the operating status of a server's internal components is monitored through airflow, a crucial factor. Because the media are closely interconnected, many early, physical potential faults, such as minor dust accumulation on heat sinks, slight displacement of internal cables due to long-term vibration, or slight deformation of chassis structural components, do not immediately cause drastic changes in the temperature or vibration parameters of a single component beyond the normal threshold in their initial stages. These early anomalies first disrupt the overall stability of the airflow field inside the server, that is, they introduce continuous, weak turbulence into the originally orderly laminar airflow field. Existing point-based sensing methods, due to the limitations of their monitoring focus, lack the ability to directly perceive this system-level, diffuse degradation of airflow stability, thus creating an inherent monitoring blind spot and response delay in the initial stage of fault evolution.
[0003] To improve the accuracy of predictions, the industry has also tried to introduce more complex machine learning or artificial intelligence algorithms to build predictive models by performing correlation analysis on historical data from multiple sensors. However, this approach also faces its own challenges. On the one hand, it is highly dependent on massive training datasets with accurate fault labels, while low-probability atypical physical fault samples such as cable displacement are almost impossible to obtain in reality. On the other hand, complex models consume huge amounts of computing resources on the edge of the server, making it difficult to achieve low-cost, universal deployment on all devices.
[0004] Specifically, existing technologies suffer from the following shortcomings: 1. Limitations in perception dimensions: They lack effective means to detect system-level, global early physical anomalies that do not directly cause drastic changes in monitoring parameters of specific components; 2. Lag in response mechanisms: Responses are typically only possible after a fault has developed to a certain extent and produced clear symptoms at local measurement points, missing effective early intervention opportunities; 3. Limitations in the universality of advanced analysis methods: Complex data models, due to their dependence on specific fault samples and high computational resource requirements, are difficult to deploy widely and economically across all servers. Therefore, the technical problem this invention aims to solve is to avoid the limitations of traditional point-based perception methods and achieve direct perception of the overall stability of the server's internal operating environment in a low-cost, low-resource-consumption manner that does not rely on specific fault samples, thereby capturing unknown types of nascent physical anomalies. Summary of the Invention
[0005] This invention provides a server full lifecycle monitoring and optimization system. Its main purpose is to solve the problem of how to avoid the limitations of traditional point-based sensing methods and achieve direct sensing of the overall stability of the server's internal operating environment in a low-cost and low-resource-consumption manner, thereby capturing early physical faults of unknown types.
[0006] To achieve the above objectives, the present invention provides a server lifecycle monitoring and optimization system, comprising:
[0007] At least one ultrasonic transmitter and at least one ultrasonic receiver are respectively disposed at the inlet and outlet of an airflow channel inside the server.
[0008] A reference signal generation unit is connected to an ultrasonic transmitter and is configured to generate a reference signal with a predetermined stable delay based on the transmitted signal of the ultrasonic transmitter.
[0009] The phase jitter detection unit is connected to the ultrasonic receiver and the reference signal generation unit, and is configured to: perform phase comparison between the received signal received by the ultrasonic receiver and the reference signal in real time; generate a pulse signal characterizing the instantaneous change in the phase difference between the received signal and the reference signal based on the phase comparison result; and count the pulse signal to obtain a count value characterizing the stability of the medium flow field in the air flow channel.
[0010] Preferably, the system is further configured to: periodically drive a physical actuator disposed within the server to generate an air disturbance with a standard waveform in the airflow channel; and instruct a phase jitter detection unit to capture and determine a transient response quantization feature of the count value caused by the air disturbance with the standard waveform; and determine a health status parameter of the ultrasonic receiver based on the transient response quantization feature.
[0011] Preferably, the phase jitter detection unit includes an XOR logic gate, which is configured to perform a bitwise XOR operation on the received signal and the reference signal to achieve phase comparison and generate a pulse signal.
[0012] Preferably, the system further includes a baseline management module, which is configured to: determine a statistical baseline for the count value when the server is in a pre-defined healthy operating state; and generate an early warning signal when the count value exceeds the statistical baseline for a predetermined threshold duration.
[0013] Preferably, the physical actuator is at least one cooling fan of the server; the system is configured to generate air disturbance with a standard waveform by instantaneously changing the rotation speed of at least one cooling fan at a predetermined rate.
[0014] Preferably, the system is further configured to: in its initial deployment state, record the transient response quantization characteristics as an initial baseline entropy value. In subsequent periodic drives, the current instantaneous response quantization characteristic is determined as the current response entropy change value. And calculate the attenuation coefficient of the ultrasonic receiver according to the following formula. , ,in, The attenuation coefficient is... The current response entropy change value, The initial reference entropy value; and the calculated attenuation coefficient. The statistical baseline value is dynamically adjusted.
[0015] Preferably, the phase jitter detection unit is integrated as a hardware logic module into a baseboard management controller of the server; and the operating frequency of the ultrasonic transmitter and ultrasonic receiver is set to 40 kHz.
[0016] Preferably, the reference signal generation unit includes a digital delay phase-locked loop to ensure that the reference signal has a predetermined stable delay relative to the transmitted signal, unaffected by temperature and voltage fluctuations in the server's operating environment; the pulse signal is an instantaneous level flip generated by phase comparison with a duration on the order of nanoseconds.
[0017] Preferably, the baseline management module is further configured to: collect basic count values for calculating the statistical baseline during the initial running cycle when the CPU load of the server is lower than a pre-stored low load threshold; and, in subsequent runs, perform a smooth adjustment of the statistical baseline using a moving average algorithm combined with historical count values at a pre-stored update cycle, while excluding the count values corresponding to the time period in which the warning signal has been triggered when executing the moving average algorithm.
[0018] Preferably, the system further includes a task scheduling module, which is configured to: monitor the load status of a central processing unit of the server in real time; and only initiate the periodic drive operation of the physical actuator when the load status of the central processing unit is continuously lower than a pre-stored low load threshold for a pre-stored trigger duration, so as to ensure that the process of determining the health status parameters of the ultrasonic receiver does not compete for computing resources with the high load of the server.
[0019] Compared with the prior art, the beneficial effects of the present invention are:
[0020] 1. This invention provides a way to identify unknown physical faults. Existing technologies rely on monitoring point-like physical quantities such as temperature and rotation speed of specific components, which is based on the premise that the fault mode has been predicted and a correlation has been established. This invention, on the other hand, uses ultrasonic transceivers configured at both ends of the server air duct to continuously perform linear scanning of the medium uniformity of the entire airflow channel. Its core phase jitter detection unit does not analyze specific physical parameters, but is sensitive to systemic entropy increase events that can cause wind turbulence and thus lead to irregular jitter of ultrasonic wave propagation phase. This allows the system to capture early physical anomalies in the blind spots of existing technologies, such as accidental detachment of internal cables and dust accumulation on heat sinks. The core of its monitoring is whether the overall stable state of the system has been disrupted, rather than whether a specific component has known failure symptoms.
[0021] 2. This invention establishes a monitoring system capable of maintaining its initial detection sensitivity throughout its entire lifecycle. The physical attenuation of sensors during long-term operation is a common problem restricting the long-term reliability of all monitoring systems. This invention generates a standardized, weak air disturbance by periodically driving the existing cooling fan in the server and reusing its core phase jitter detection unit to capture the system's response characteristics to this standard signal. This allows the system to online infer the health status of the ultrasonic sensor itself, i.e., the attenuation coefficient, without additional hardware costs. Furthermore, the system dynamically adjusts its internal early warning judgment baseline based on this attenuation coefficient. This self-diagnostic and self-compensating closed-loop logic avoids the risk of missed detections due to sensor aging and ensures the effectiveness of monitoring throughout the entire process of server deployment, operation, and retirement.
[0022] 3. This invention frees effective fault early warning capabilities from heavy reliance on computing resources and expertise, making it feasible for low-cost, large-scale deployment. Its core phase jitter detection is implemented at the hardware level of the baseboard management controller through a simplified XOR logic gate and counter. This process does not involve complex signal processing algorithms or machine learning models and does not occupy the server's main computing resources. At the same time, its subsequent sensor self-calibration mechanism is also completed by reusing instructions from existing fans and performing simple proportional calculations. This design, which transforms the perception of complex physical phenomena into low-resource-consumption logical judgments, enables highly sensitive early fault warning functions to become a basic configuration of every server, including resource-constrained edge computing devices, at a lower cost and power consumption, rather than an exclusive feature of a few high-performance computing clusters. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the system architecture of a server lifecycle monitoring and optimization system according to the present invention;
[0024] Figure 2 This is a comparison chart of the phase jitter count value of the present invention and the response characteristics of traditional temperature measuring points to duct blockage faults;
[0025] Figure 3 This is a schematic diagram of the logic for the entire lifecycle operation state transition of this invention.
[0026] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0028] This application discloses a server lifecycle monitoring and optimization system. Its system architecture includes an ultrasonic transceiver, a reference signal generation unit, a phase jitter detection unit, and a baseline management module, all configured within the server's internal airflow channel. The ultrasonic transceiver generates and receives acoustic signals passing through the airflow channel. The phase jitter detection unit compares the received signal with a reference signal generated by the reference signal generation unit in real time to obtain a count value characterizing the macroscopic stability of the medium flow field within the airflow channel. The baseline management module determines the server's health status based on the statistical characteristics of this count value and maintains its lifecycle monitoring reliability through an online self-calibration procedure. In a specific deployment scenario, such as a high-density server requiring long-term uninterrupted operation, the compact layout of its internal components allows for easy monitoring of even the smallest physical structures. Changes can affect overall heat dissipation performance, and existing monitoring methods based on discrete temperature measurement points have an inherent lag in perceiving such system-level risks. To address this challenge, the ultrasonic transmitter and receiver of this invention are respectively configured at the inlet and outlet of the main airflow channel inside the server, for example, installed on both sides of the front air vent and the rear exhaust fan module of the chassis, forming a through-type and non-contact scanning path. The operating frequency of the ultrasonic transmitter is set to 40 kHz. This frequency was selected after taking into account both signal diffraction capability and resistance to environmental electromagnetic noise interference. In this way, by continuously scanning the medium uniformity of the entire airflow channel, the system can capture global wind field turbulence caused by early physical anomalies such as trace dust accumulation on the heat sink or slight displacement of internal cables due to vibration, thus providing a way to identify unknown physical faults.
[0029] To convert the weak turbulent disturbances in the airflow field into quantifiable digital signals, while avoiding the dependence on high-precision clocks and computational complexity of traditional time-of-flight measurement methods, the system employs a phase comparison-based detection procedure. Therefore, the reference signal generation unit is connected to the ultrasonic transmitter and configured to generate a reference signal with a predetermined stable delay based on the transmitter's transmitted signal. To ensure that this delay is unaffected by temperature and voltage fluctuations in the server's operating environment, the reference signal generation unit is specifically implemented using a digital time-delay phase-locked loop (PDL). Correspondingly, the core logic of the phase jitter detection unit is implemented through an XOR logic gate, which performs a bitwise XOR operation on the received signal and the reference signal. When the airflow channel... When the internal medium is uniform and stable, the phase difference between the received signal and the reference signal is constant, and the output of the XOR logic gate is a stable level. However, when turbulence occurs in the flow field and causes irregular jitter in the phase of the received signal, the output of the XOR logic gate will generate an instantaneous level flip with a duration on the order of nanoseconds, i.e., a pulse signal. The phase jitter detection unit then counts the pulse signal to obtain a count value that characterizes the stability of the medium flow field in the air flow channel per unit time. This procedure transforms the analysis of flow field stability into the counting of digital pulse signals, enabling a highly sensitive early fault warning function to be integrated into the server's baseboard management controller without occupying the server's central processing unit computing resources.
[0030] In the phase jitter detection unit, the two inputs of the XOR logic gate are each connected to the output of a Schmitt trigger. The received signal from the ultrasonic receiver and the reference signal from the reference signal generation unit are converted into square wave digital signals after entering the Schmitt trigger. The hysteresis voltage window of the Schmitt trigger is set to 1.2 times the peak-to-peak value of the background noise signal collected by the server fan in a silent state. Simultaneously, when the physical actuator performs periodic drive operations, its speed is controlled by a preset trapezoidal wave command sequence. This sequence includes a 500ms linear rising edge, a 1000ms plateau period, and a 500ms linear falling edge. The board management controller performs closed-loop adjustment of the pulse width modulation duty cycle based on feedback from the fan tachometer until the integral error between the actual speed curve and the target trapezoidal wave command sequence is less than 5%. To ensure that the obtained count values have practical early warning indication significance, a judgment benchmark needs to be established to distinguish between normal operating conditions and potential abnormal conditions. Therefore, the system also includes a baseline management module, which is configured to operate in a pre-calibrated healthy state. During operation, a statistical baseline for the count values is determined. The specific procedure for determining this statistical baseline is as follows: During the initial operating cycle when the CPU load of the server is lower than a pre-stored low load threshold, the baseline management module continuously collects basic count values for calculating the statistical baseline. This is to eliminate the interference of normal hot air disturbances under high load conditions on the baseline calibration. After collecting sufficient samples, the system calculates its statistical mean and standard deviation, and stores, for example, the mean plus three times the standard deviation, as the initial statistical baseline. In subsequent operation, the baseline management module uses a moving average algorithm combined with historical count values at a pre-stored update cycle to perform smooth adjustments to the statistical baseline to adapt to slow environmental changes. At the same time, when executing the moving average algorithm, the system will actively exclude the count values corresponding to the time period that has triggered the warning signal to prevent the count values caused by real faults from continuously being too high and contaminating the accuracy of the statistical baseline. When the real-time acquired count value exceeds the time of the dynamically adjusted statistical baseline for a continuously determined threshold duration, such as 30 consecutive seconds, the baseline management module generates a warning signal.
[0031] Considering that any physical sensor may experience physical attenuation during long-term operation, leading to decreased sensitivity and the risk of missed detections, thus compromising the long-term reliability of the monitoring system, this invention establishes an online self-calibration mechanism. This mechanism periodically drives a physical actuator installed within the server to generate air disturbances with a standard waveform within the airflow channel. In a specific deterministic procedure, the physical actuator is at least one cooling fan of the server. The system monitors the CPU load status of the server in real time through a task scheduling module, and only initiates the periodic drive operation when the load status remains below a pre-stored low load threshold for a pre-stored trigger duration. This operation specifically... To generate air disturbances with a standard waveform by instantaneously changing the speed of at least one cooling fan at a predetermined rate, the task scheduling module, in addition to monitoring the CPU load status in real time, also has a maximum calibration interval timer with a period set to 72 hours. If the low load threshold condition for triggering periodic drive operations is not met within this period, the timer will time out and initiate a micro-disturbance calibration procedure. This procedure instructs the physical actuator to increase its speed from the current operating reference by a fixed increment, which is calibrated to 50% of the minimum response threshold of the server's thermal management strategy. This small increment speed is maintained for 60 seconds, during which the phase jitter detection unit continuously integrates its count value, and the integration result is used to calculate the attenuation coefficient. While generating an air disturbance with a standard waveform, the phase jitter detection unit is instructed to capture and determine an instantaneous response quantization characteristic of the count value caused by the disturbance. To quantify and compensate for sensor attenuation, the system performs this calibration procedure once in its initial deployment state and records the obtained instantaneous response quantization characteristic as an initial reference entropy value. In subsequent periodic drives, the system determines the current instantaneous response quantization characteristic as the current response entropy change value. And calculate the attenuation coefficient of the ultrasonic receiver according to the following formula. , in, The attenuation coefficient is... The current response entropy change value, The calculated attenuation coefficient is based on the initial reference entropy value. This is then used to dynamically adjust the value of the statistical baseline; for example, the adjusted statistical baseline can be set to the initial statistical baseline and... The product of these factors, through this closed-loop calibration procedure system, can effectively offset the decrease in sensitivity caused by sensor aging, maintaining the effectiveness of its monitoring throughout the server's entire lifecycle.
[0032] Example 1: In a continuously high-density data center environment, a server began experiencing intermittent and irregular service interruptions in the sixth month after deployment. The server's baseboard management controller logs did not record any explicit hardware error codes, and the operations team's continuous monitoring of the server's routine metrics, including CPU temperature, fan speeds, and core voltage, did not detect any abnormal fluctuations outside the normal range. After deploying a monitoring system on the server and completing an initial low-load operation cycle, the baseline management module automatically established a statistical baseline for the count values of its internal airflow channels. In subsequent continuous monitoring, the phase jitter detection unit recorded a stable but consistently higher count value than this statistical baseline. This continuous deviation of a single count value directly pointed to a system-level, physical state variation from an untraceable, multivariate performance problem. This shifted the diagnostic focus from speculating on the parameters of multiple independent components to confirming a macroscopic physical state.
[0033] Based on the warning signal issued by the system, engineers conducted an unpacking and physical inspection of the server. They discovered a non-critical data cable used to connect to the backplane, one end of which had detached from the cable management channel and entered the main cooling airflow, creating a continuous, weak air vortex at the cable end. After re-secured the cable, the system's counter value returned to its statistical baseline level within a few minutes, and the intermittent service interruptions of the server disappeared. Following this incident, the server continued to operate for more than two years. During this period, its built-in task scheduling module triggered the cooling fan to perform instantaneous speed adjustment during low-load windows in each preset 72-hour cycle. The system then captured the response entropy change value caused by this standard air disturbance. And based on the relation Continuously update the attenuation coefficient of the ultrasonic receiver This allows for dynamic adjustment of the statistical baseline. The synergistic operation of this sensing and calibration mechanism solidifies the one-time fault detection capability into a stable status monitoring capability that spans the entire lifecycle of the equipment. Its sensitivity to the stability of the medium flow field is independent of its deployment duration. In the subsequent operating cycle, the server did not experience any similar logless service interruption events, and the count value output by the phase jitter detection unit remained stable within the dynamically adjusted statistical baseline range.
[0034] Example 2: To objectively quantify the detection sensitivity and response characteristics of this technical solution for early degradation of the system operating environment, a controlled experiment was conducted. The experiment used a standard 2U rack-mount server as the test platform, which was placed in a constant temperature and humidity shielded environment to eliminate external interference. Simultaneously, the server's central processing unit was subjected to a constant load, and the speed of all cooling fans was fixed. This was intended to create a stable and repeatable internal benchmark operating condition. The core parameters monitored in the experiment included the count values output by the system of this invention, and the temperatures measured by high-precision thermocouples deployed on the CPU core and memory modules as a reference comparison. The sampling period for the count values was set to 1 second, which was set to ensure the capture of millisecond-level phase jitter events. The experiment simulated the gradual blockage of the heat dissipation airflow caused by long-term dust accumulation by setting a precisely adjustable grille at the server's main air intake. The steps were as follows: First, the server was run continuously for 60 minutes at a baseline state with an air intake blockage rate of 0% to achieve internal thermal equilibrium, and the baseline management module determined the statistical baseline of the count value. Subsequently, the air intake blockage rate was set to 5%, 10%, 15%, and 20% in sequence, and kept constant for 30 minutes at each blockage rate level to ensure that the system reached a new stable state. During this period, various monitoring parameters were continuously recorded. Table 1 shows the data collected after each blockage rate level reached a stable state during the experiment.
[0035] Table 1: Comparison of system state parameters under different blocking rates.
[0036]
[0037] Referring to Table 1, when the intake blockage rate slightly increases from 0% to 5%, the system count increases from 12 times per minute to 85 times per minute, while the average temperature of the CPU core and the average temperature of the memory module increase by no more than 0.5% during the same period. As the blockage rate further increases, the count shows an exponential growth trend, while traditional temperature indicators show significant response lag; before the blockage rate reaches 15%, the magnitude of change remains within the normal measurement fluctuation range. The mechanism of this phenomenon lies in the fact that a small intake blockage introduces continuous and low-energy turbulence into the airflow channel. Although this turbulence... While insufficient to immediately cause significant accumulation of macroscopic heat, its disruption of the uniformity of the air medium instantly alters the propagation phase of the ultrasonic signal, which is then sensitively captured by the phase jitter detection unit in the form of pulse counting. In the very initial stage of a progressive physical fault, before any identifiable changes occur in traditional monitoring indicators, this technical solution provides a significant and quantifiable early warning signal. This detection capability shifts the fault prediction window from relying on the consequences of the fault, namely the occurrence of heat accumulation, to the direct perception of the cause of the fault, namely the disruption of the stability of the physical environment.
[0038] Example 3: This example combines Figures 1 to 3 This describes a server lifecycle monitoring and optimization system, such as... Figure 1 As shown, within an internal airflow channel of a server, an ultrasonic transmitter generating a 40kHz acoustic signal and an ultrasonic receiver capturing sensor signals are deployed, forming an ultrasonic propagation path. The ultrasonic transmitter is connected to a digital time-delayed phase-locked loop (PDL) serving as a reference signal generation unit. This unit generates a reference signal with a stable delay based on the transmitted signal. Both the ultrasonic receiver and the reference signal generation unit are connected to a core phase jitter detection unit. This unit quantizes the phase difference between the received signal and the reference signal into a count value using XOR logic gates and pulse counting. This count value is sent to a baseline management module, which dynamically adjusts and compares the statistical baseline to generate an early warning signal when an early physical fault occurs. Simultaneously, the system includes a task scheduling module to issue calibration trigger commands during low-load windows. These commands drive a cooling fan, acting as a physical actuator, to momentarily adjust its speed to generate a standardized, weak air disturbance. The system calculates the attenuation coefficient characterizing the sensor state by capturing the response characteristics induced by this standardized disturbance. The coefficient is then fed back to the baseline management module to dynamically adjust its statistical baseline, thus forming a closed-loop self-calibration and fault early warning system.
[0039] like Figure 2 As shown in the figure, the horizontal axis represents the inlet blockage rate (%), the left vertical axis represents the phase jitter count (times / minute) on a logarithmic scale, and the right vertical axis represents the temperature. The figure clearly shows that as the air intake blockage rate gradually increases from 0% to 20%, the phase jitter count curve, which characterizes the monitoring results of this invention, exhibits an exponential and rapid increase. In contrast, the CPU core temperature curve and the memory module temperature curve only show a small and lagging linear increase. This phenomenon proves that this system can provide a more sensitive early warning signal at the nascent stage of physical anomalies, that is, before significant heat accumulation occurs.
[0040] like Figure 3 As shown, after system startup, it first enters the system initialization state, completing operations such as configuring the ultrasonic transceiver, establishing the initial reference entropy value, and setting the statistical baseline. Then it enters the core normal monitoring state, where phase jitter detection is performed in real time, and the count value is continuously collected and compared with the statistical baseline. This state has three main transition paths: First, when the low load window reaches the calibration cycle, the system enters the self-calibration state, updating the statistical baseline by driving the cooling fan, capturing the response entropy change value, and calculating the attenuation coefficient. After completion, it automatically returns to the normal monitoring state. Second, when the count value exceeds the baseline and the duration exceeds the threshold, the system enters the early warning state to generate an early warning signal, record abnormal data, and notify the maintenance system. Third, when a chassis open detection signal is received, the system enters the maintenance suppression state, suspending the early warning logic and waiting for the chassis to close to avoid false alarms caused by maintenance operations. After maintenance is completed, the system is guided to the subsequent fault handling state or resume normal monitoring.
[0041] Example 4: During the firmware integration and calibration phase of a specific server model, the following parameter calibration procedure is executed to set a set of operating parameters and judgment thresholds related to the physical characteristics of the server model for the system's self-calibration mechanism and early warning logic. To determine the operating parameters of the physical actuator used to generate the standard waveform air disturbance, the server is placed under a stable low-load condition, and its phase jitter detection unit continuously outputs a background count value with a period of 1 second. The system drives the cooling fan with a preset small increment and records the resulting response entropy change value. If the If the value does not reach more than five standard deviations above the mean of the background count values, the instantaneous increase in fan speed will be increased in increments of 0.5%, and the driving and recording process will be repeated until a certain response entropy change value is reached. When the statistical significance requirement is met for the first time, and the instantaneous temperature fluctuation caused by the driving operation at the CPU core is less than 0.1 degrees Celsius, the system will solidify the current fan speed increase and maintenance duration as standardized operating parameters for subsequent periodic self-calibration of this server model.
[0042] To establish an initial health status model, the server needs to run continuously for 30 minutes under a predefined low-load condition, after confirming that its internal physical structure and airflow channels are in a preset state that meets factory standards. This means that the CPU load remains below 10% for 15 minutes. During this period, the baseline management module collects all count values and calculates the statistical average based on the dataset. with standard deviation The initial statistical baseline was then set to Next, the system calls the standardized operating parameters calibrated in the previous stage to drive the cooling fan to generate a standard waveform air disturbance, and records the captured response entropy change value as the server's unique initial baseline entropy value. To calibrate the threshold duration parameter in the early warning logic, the testers applied a set of standardized, brief physical impacts to the server chassis to simulate non-faulty physical disturbances that may occur during operation and maintenance. The system recorded the duration of the instantaneous pulse of the count value caused by each impact, forming a dataset of instantaneous disturbance durations. Ultimately, the system's threshold duration was set to the 99th percentile value in this dataset. This was intended to filter out short-duration noise signals introduced by the external environment and only respond to count value deviations that exceed the threshold and indicate a permanent change in the internal state. After executing this offline calibration procedure, the core thresholds and operating parameters within the system were set to a set of values associated with the physical characteristics of the specific server model, thereby establishing an initial benchmark for health status assessment at the beginning of lifecycle monitoring.
[0043] Example 5: In a scenario where routine hardware maintenance is being performed in a data center where this monitoring system has been deployed, when the server chassis is opened, the electrical signal generated by the chassis intrusion detection switch synchronously triggers the monitoring system to enter a preset maintenance suppression state. In this state, the phase jitter detection unit and the counter continue to record changes in the count value, but the logic judgment function of the baseline management module used to generate early warning signals is temporarily suspended for a preset duration or until a signal indicating that the chassis is closed and reset is received. After the suppression state is lifted, the baseline management module automatically performs a baseline compliance check. If the count value converges and stabilizes within the statistical baseline range in the following few cycles, no historical alarm is generated. If the count value continues to be significantly higher than the statistical baseline, the system generates a specific type of early warning indicating a potential physical state abnormality after maintenance.
[0044] Furthermore, the warning signal generated by the system is not an isolated Boolean flag, but is constructed into a data packet containing contextual information. In addition to the warning event itself, this data packet also encapsulates the unique identifier of the server that triggered the warning, the real-time count value at the time the warning occurred, the corresponding dynamic statistical baseline value at that time, and the attenuation coefficient, which characterizes the current health status of the sensor and is calculated by the self-calibration mechanism. This structured early warning information allows the upper-level central operations and maintenance management platform to determine the extent to which the count value deviates from the baseline, combined with the server's business criticality level and known attenuation coefficient. It automatically classifies the risk level of early warning events from different servers and dynamically sorts the handling priorities.
[0045] Example 6: On a server that has had this monitoring system deployed and running for an extended period, after one of the cooling fans was replaced during routine maintenance, the system executed an online cross-validation and recalibration procedure for the physical actuators. This procedure drove each cooling fan in the server, including the newly replaced fan, individually and independently, performing an instantaneous speed change according to the initially calibrated standardized operating parameters to generate airflow disturbance within the airflow channel. The system recorded the response entropy change value generated by each individual fan, denoted as . , … and compared with the initial baseline entropy value stored in the system. The system compares the response entropy change value generated by the new fan with that of other fans. If the deviation between the response entropy change value generated by the new fan and that of other fans exceeds a preset actuator consistency threshold, the system determines that the aerodynamic characteristics of the new fan have changed and automatically re-executes the iterative optimization process for determining the optimal physical actuator operating parameters for that fan to update its standardized operating parameters. If the response entropy change values generated by all fans are consistent with each other, but are proportionally lower than the initial reference entropy value, the system will determine the change. This further verifies that the deviation originates from the physical attenuation of the ultrasonic receiver, thus confirming that the attenuation coefficient is the cause. The accuracy of the sensor health status characterized.
[0046] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0047] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A server lifecycle monitoring and optimization system, characterized in that, include: At least one ultrasonic transmitter and at least one ultrasonic receiver are respectively disposed at the inlet and outlet of an airflow channel inside the server. A reference signal generation unit is connected to an ultrasonic transmitter and is configured to generate a reference signal with a predetermined stable delay based on the transmitted signal of the ultrasonic transmitter. The phase jitter detection unit is connected to the ultrasonic receiver and the reference signal generation unit, and is configured to: perform phase comparison between the received signal received by the ultrasonic receiver and the reference signal in real time; and generate a pulse signal characterizing the instantaneous change in the phase difference between the received signal and the reference signal based on the phase comparison result. And perform counting on the pulse signal to obtain a count value characterizing the stability of the medium flow field within the airflow channel.
2. The server full lifecycle monitoring and optimization system according to claim 1, characterized in that, The system is also configured to: periodically drive a physical actuator located within the server to generate an air disturbance with a standard waveform in the airflow channel; and instruct the phase jitter detection unit to capture and determine an instantaneous response quantization characteristic of the count value caused by the air disturbance with the standard waveform; And a health status parameter of the ultrasonic receiver is determined based on the instantaneous response quantization characteristics.
3. The server full lifecycle monitoring and optimization system according to claim 1, characterized in that, The phase jitter detection unit includes an XOR logic gate, which is configured to perform a bitwise XOR operation on the received signal and the reference signal to achieve phase comparison and generate a pulse signal.
4. The server full lifecycle monitoring and optimization system according to claim 1, characterized in that, The system also includes a baseline management module, which is configured to: determine a statistical baseline for the count value when the server is in a pre-defined healthy operating state; and generate an early warning signal when the count value exceeds the statistical baseline for a predetermined threshold duration.
5. A server full lifecycle monitoring and optimization system according to claim 2, characterized in that, The physical actuator is at least one cooling fan of the server; the system is configured to generate air disturbance with a standard waveform by instantaneously changing the speed of at least one cooling fan at a predetermined rate.
6. The server full lifecycle monitoring and optimization system according to claim 2, characterized in that, The system is also configured to record the transient response quantization characteristics as an initial baseline entropy value in its initial deployment state. In subsequent periodic drives, the current instantaneous response quantization characteristic is determined as the current response entropy change value. And calculate the attenuation coefficient of the ultrasonic receiver according to the following formula. , ,in, The attenuation coefficient is... The current response entropy change value, The initial reference entropy value; and the calculated attenuation coefficient. The statistical baseline value is dynamically adjusted.
7. The server full lifecycle monitoring and optimization system according to claim 1, characterized in that, The phase jitter detection unit, as a hardware logic module, is integrated into the server's baseboard management controller; and the operating frequency of the ultrasonic transmitter and ultrasonic receiver is set to 40 kHz.
8. A server full lifecycle monitoring and optimization system according to claim 1, characterized in that, The reference signal generation unit includes a digital delay phase-locked loop to ensure a predetermined stable delay of the reference signal relative to the transmitted signal, unaffected by temperature and voltage fluctuations in the server's operating environment; the pulse signal is an instantaneous level flip generated by phase comparison with a duration on the order of nanoseconds.
9. A server full lifecycle monitoring and optimization system according to claim 4, characterized in that, The baseline management module is also configured to: collect basic count values for calculating the statistical baseline during the initial running cycle when the CPU load of the server is lower than a pre-stored low load threshold; and, in subsequent runs, perform smooth adjustments to the statistical baseline using a moving average algorithm combined with historical count values at a pre-stored update cycle, while excluding count values corresponding to the time period in which the warning signal has been triggered when executing the moving average algorithm.
10. A server full lifecycle monitoring and optimization system according to claim 2, characterized in that, The system also includes a task scheduling module, which is configured to: monitor the load status of a central processing unit of the server in real time; and only initiate the periodic drive operation of the physical actuator when the load status of the central processing unit is continuously lower than a pre-stored low load threshold for a pre-stored trigger duration.