Computer maintenance decision dispatching method based on environmental factor data

By constructing a computer maintenance decision-making and scheduling method driven by environmental factor data, the task inflow rate and cooling fan control are adjusted in real time, solving the problem of frequent oscillations of computer clusters under environmental factor fluctuations, and realizing continuous control of hardware health status and stability of computing power output.

CN122632999APending Publication Date: 2026-08-25DIBES (CHENGDU) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610591221.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing computer clusters suffer from frequent oscillations in maintenance decisions under environmental fluctuations, leading to hardware thermal stress damage and fluctuations in computing resource utilization. The lack of dynamic coupling modeling of environmental stress and load flow makes it difficult to achieve stable computing power output in high-reliability scenarios.

Method used

By constructing a computer maintenance decision-making and scheduling method based on environmental factor data, utilizing feedback loops and heat dissipation path efficiency evaluation models, the task inflow rate is monitored and adjusted in real time, pulse load is injected to control the cooling fan to enter variable speed resonance mode, and a virtual tension model is established to smooth task migration, thereby achieving continuous control of hardware health status.

Benefits of technology

It eliminates frequent maintenance decision oscillations under critical environmental conditions, avoids thermal stress shocks, maintains the stability of computing resources and the operational consistency of the distributed system, and improves the system's adaptability to extreme conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122632999A_ABST
    Figure CN122632999A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer maintenance, and discloses a computer maintenance decision scheduling method based on environmental factor data, which comprises the following steps: obtaining load data and environmental parameter data, and inputting the load data into a feedback loop to adjust scheduling gain; extracting the lag time of temperature change of a computer node relative to load change; using a heat dissipation path efficiency evaluation model to associate and map the lag time with the environmental parameter data, and calculating a heat dissipation efficiency conduction index; when the environmental parameter fluctuation is stable and the change rate of the conduction index exceeds a threshold value, generating a maintenance scheduling instruction; injecting a pulse load to induce resonance of a heat dissipation system, verifying the maintenance effect according to the second derivative of the response time, and canceling the maintenance scheduling instruction; the present application uses thermal inertia residual error generated by feedback regulation to realize logical compensation sensing of heat dissipation efficiency, eliminates environmental interference, realizes closed-loop verification of maintenance effect, and guarantees the operation stability of a computing cluster under complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a computer-based maintenance decision-making and scheduling method based on environmental factor data, belonging to the field of computer maintenance technology. Background Technology

[0002] Current computer clusters are constantly affected by environmental factors such as ambient temperature, humidity, physical vibration, and voltage fluctuations during operation. Existing technologies typically use environmental sensors to collect environmental data and trigger maintenance actions such as task migration or forced frequency reduction based on preset discrete thresholds. From the perspective of semiconductor physics, the failure process of electronic components is closely related to the thermal stress cycle they endure. Fluctuations in environmental factors are transformed into changes in the internal temperature gradient of the hardware through adjustments to the computing load. Existing discrete triggering mechanisms have step characteristics at the decision boundary, causing the system to frequently start and stop tasks or jump in frequency near the critical environmental conditions. This oscillation of maintenance decisions induces high-frequency thermal stress pulses within the hardware, and the resulting mechanical fatigue losses exceed the steady-state effects of the environmental factors themselves.

[0003] While the industry has attempted to introduce multi-level buffer thresholds or extend the sampling period to smooth out oscillations, most improvements are passive responses. While hardware structural limitations become apparent, flawed control methods also restrict system health. For example, the utility model patent CN213634316U discloses a heat dissipation chassis for installing a computer monitoring system, utilizing a multi-fan layout and a side-panel removable structure to improve internal ventilation. However, this approach remains at the stage of passive physical heat dissipation. Identification of hidden faults such as dust accumulation within the heat dissipation path relies on discrete sensor feedback or manual inspection, lacking dynamic coupling modeling of environmental stress and load flow. In industrial edge computing or high-reliability scenarios, hardware structural improvements cannot perceive the evolution of heat conduction path resistance in real time, making it difficult to adaptively compensate for attenuation based on the physical characteristics of the equipment under complex operating conditions. The industry has previously adopted... Introducing multi-level buffer thresholds or extending the sampling period to mitigate the aforementioned oscillations is a passive response, leading to system lag under extreme and sudden operating conditions. Existing technologies mainly suffer from the following drawbacks: a lack of continuous evolutionary logic between maintenance actions and environmental evolution, causing maintenance decision oscillations in the system at critical states; a lack of dynamic coupling modeling of environmental stress and load flow, making it difficult to achieve smooth flow of computing power output; and a lack of adaptive compensation capability for hardware performance degradation in the scheduling logic, making it difficult to guarantee scheduling accuracy throughout the entire lifecycle. Constructing a continuous feedback mechanism that can characterize environmental stress features in real time and fine-tune task scheduling weights accordingly to eliminate systemic maintenance damage caused by discrete decisions is a current technical requirement in related fields. Existing static maintenance methods are insufficient to meet the requirements of stable computing power output in high-reliability computing scenarios.

[0004] Therefore, the technical problem to be solved by this invention is how to construct a computer maintenance decision scheduling method based on environmental factor data, which can offset environmental stress fluctuations by adjusting the task inflow rate in real time and achieve continuous control of the system health status. Summary of the Invention

[0005] To address the problems mentioned in the background art, the technical solution of the present invention is as follows: A computer-based maintenance decision scheduling method based on environmental factor data, the method comprising the following steps: Step 101: Obtain the load data and environmental parameter data of the computer node, and input the load data into the feedback loop to perform resource scheduling. The feedback loop adjusts the scheduling correction gain based on the environmental parameter data. Step 102: Monitor the adjustment process of the feedback loop in real time and extract the lag time of the temperature change of the computer node relative to the load data change. Step 103: Construct a heat dissipation path performance evaluation model. The heat dissipation path performance evaluation model includes an evolution function that characterizes the heat conduction relationship. The lag time is correlated and mapped with environmental parameter data through the heat dissipation path performance evaluation model in order to calculate the heat dissipation performance conduction index of the heat dissipation path. Step 104: Determine whether the fluctuation range of the environmental parameter data is within the preset stable range. If it is within the preset stable range and the rate of change of the heat dissipation efficiency conduction index exceeds the preset threshold, then determine that the heat dissipation path has a blockage fault and generate a maintenance scheduling instruction. Step 105: Execute the maintenance scheduling command by injecting a pulse load into the computer node to control the cooling fan to enter the variable speed resonance mode, and extract the second derivative of the response time of the computer node during the pulse load execution to verify the maintenance effect. The maintenance scheduling command is released when the second derivative recovers to the preset range.

[0006] Preferably, the process of adjusting the scheduling correction gain in step 101 specifically includes: calculating the ratio of the number of instructions issued to the number of real-time responses of computer nodes to determine the scheduling sensitivity; and weighting the scheduling correction gain according to the scheduling sensitivity to offset the thermal response lag caused by the performance degradation of computer nodes.

[0007] Preferably, the operation of extracting the lag time in step 102 specifically includes: obtaining the temperature change curve fed back by the temperature sensor after the scheduling correction gain action; and calculating the lag time of the temperature change curve relative to the starting point of the jump in the load data. ; Calculate the lag time Deviation from the hysteresis constant under the preset calibration conditions ;in, , The deviation is used as a criterion. For the lag time, This is the preset calibration phase lag.

[0008] Preferably, the evolution function in the heat dissipation path effectiveness evaluation model is configured to: map the lag time to a heat dissipation resistance characteristic value; calculate the rate of change of the heat dissipation resistance characteristic value within a preset monitoring period; and determine the physical attenuation characteristics of the heat dissipation path based on the rate of change.

[0009] Preferably, the criteria for determining that the heat dissipation path is blocked in step 104 specifically include: extracting the trend component of the heat dissipation performance conduction index; when the absolute value of the slope of the trend component exceeds a preset attenuation threshold, and the variance of the load data is less than a preset variance threshold, it is determined that the heat dissipation path is blocked by dust accumulation.

[0010] Preferably, the operation of controlling the cooling fan to enter the variable speed resonance mode in step 105 specifically includes: controlling the injection frequency of the pulse load to match the natural frequency of the cooling fan, and inducing the cooling fan to generate resonance displacement through the periodic change of the load data.

[0011] Preferably, the operation of verifying the maintenance effect and releasing the maintenance scheduling command when the second derivative recovers to the preset range specifically includes: calculating the second derivative of the response time with respect to the pulse load execution time; if the second derivative is within the preset range, determining that the air resistance of the heat dissipation path has recovered to the normal state, and terminating the injection of the pulse load.

[0012] Preferably, the method further includes: generating an environmental data correction operator when the evolution trend of environmental parameter data does not match the degradation characteristics of heat dissipation efficiency conduction index; and using the environmental data correction operator to compensate for the environmental parameter data in order to correct the sensor sampling error.

[0013] Preferably, the method further includes: by adjusting the task migration step size, limiting the impact of maintenance actions on global throughput within a preset elastic range, so as to maintain the consistency of the operating frequency of the distributed cluster during maintenance.

[0014] Preferably, the method further includes: real-time monitoring of the task migration dispersion of the computer node during the execution of pulse load; when the task migration dispersion exceeds a preset threshold, reducing the scheduling correction gain to achieve overheat protection of the computer node by reducing the load injection intensity.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. In computer maintenance decision-making and scheduling, the discrete threshold triggering in traditional maintenance decision-making is transformed into continuous state adjustment based on environmental fluctuations through the closed-loop synergistic effect of environmental stress characteristic values, load carrying capacity characteristic values, and scheduling correction gains. During the operation of computer nodes, the evolution of environmental factors usually has a continuous cumulative characteristic. This invention uses a combination of feedforward compensation and closed-loop feedback to make the task scheduling priority smoothly fine-tuned with the fluctuation of environmental stress characteristic values. This eliminates the scheduling oscillations caused by the frequent triggering of protective frequency reduction or task migration under critical environmental conditions, avoids secondary thermal stress shocks caused by frequent load jumps, and maintains the stability of computing resource output.

[0016] 2. By establishing a virtual tension model between communicating computing nodes, the system dynamically offsets the maintenance actions of a single node with the overall load flow of the cluster. When the target computer node increases its maintenance depth due to environmental degradation, the system uses the maintenance depth compensation to adjust the virtual tension gain between nodes and synchronously controls the migration speed between the task source node and the task receiving node. This cascade scheduling mechanism ensures that the maintenance intervention of local nodes no longer causes a sudden increase in the load of neighboring nodes or network bandwidth blockage, limiting the impact of local maintenance actions on global throughput to a preset elastic range and ensuring the consistency of the operating rhythm of the distributed system under large-scale concurrent pressure.

[0017] 3. By extracting the task execution time deviation and constructing the causal relationship between environmental factors and virtual forward sliding parameters, the authenticity of environmental monitoring data is logically verified. In industrial edge environments, physical drift generated by sensors often leads to incorrect scheduling decisions. This invention uses system operation feedback as an auxiliary reference system to generate a correction operator when the fluctuation trend of environmental factors does not match the virtual forward sliding parameters. It performs real-time drift compensation on the environmental pressure vector input to the adjustment logic and uses logical calculation methods to compensate for the monitoring defects of physical hardware. Without adding redundant sensing hardware, it improves the system's adaptability to extreme and complex working conditions. Attached Figure Description

[0018] Figure 1 This is a flowchart of the computer maintenance decision scheduling process for environmental feedback adjustment and resonance verification of the present invention. Figure 2 This is a schematic diagram of the hierarchical architecture of the system that integrates thermal inertia analysis and active excitation functions according to the present invention. Detailed Implementation

[0019] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. The following embodiments are intended to explain the present invention and are not intended to limit the scope of protection of the present invention.

[0020] This invention provides a computer maintenance decision-making and scheduling method based on environmental factor data, primarily applied to industrial edge computing environments or large-scale data center nodes characterized by high-frequency vibration and significant temperature and humidity fluctuations. It addresses scheduling oscillations and hardware thermal stress damage caused by discrete threshold-based maintenance mechanisms near environmental critical points. The method transforms the continuous evolution of environmental factor data into real-time gain adjustment for resource scheduling. It utilizes the thermal inertia residual characteristics during hardware operation to achieve non-intrusive perception of the health status of heat dissipation paths. When a decrease in heat dissipation efficiency is detected, pulse excitation generated by the computing load induces resonance in the cooling fans to perform self-cleaning maintenance. Industrial environmental factor fluctuations have continuous and cumulative effects. If a discrete threshold triggering mechanism is used, the system will frequently perform frequency reduction or task migration when environmental parameters are at the judgment boundary, leading to significant fluctuations in computing resource utilization and inducing secondary thermal stress shocks. To address this challenge, this invention acquires load data and environmental parameter data of the computer node during execution. Environmental parameter data includes real-time collected temperature, humidity, vibration frequency, and voltage fluctuation amplitude. The system inputs the load data into a feedback loop to execute resource scheduling, and the feedback loop adjusts the scheduling correction gain based on the environmental parameter data. The system integrates multidimensional environmental parameters into a single environmental stress characteristic value through a pre-defined attenuation correlation model. The characteristic value of the load-bearing capacity that the current node can withstand. The system utilizes feedback control logic to determine the load-bearing margin characteristic value. With target value The deviation is taken as the input. The specific calculation process is as follows: The real-time collected temperature, humidity, and vibration frequency are subtracted from their rated reference values ​​and divided by the maximum safe range to obtain normalization coefficients between 0 and 1. These coefficients are then accumulated according to weight ratios of 0.6, 0.25, and 0.15 to obtain the environmental stress characteristic value. A base gain of 1.0 minus 0.5 times this characteristic value is used as the reference. Furthermore, based on the load deviation, the gain is accumulated in increments of 0.02 every 10ms, thus combining the environmental stress characteristic value... Feedforward compensation is used to calculate the scheduling correction gain. ,in, To adjust the gain for scheduling, This is a load-bearing margin characteristic value, representing the current load margin that a node can withstand. This represents the target value for load-bearing capacity margin. The kernel scheduler adjusts the gain based on the environmental stress characteristic value. Adjust task distribution weights in real time.

[0021] By smoothly altering the task inflow rate, the impact of environmental stress on hardware lifespan is mitigated, thus ensuring sufficient load-bearing capacity. Stabilizing within the preset target range, when the environmental stress characteristic value When the value increases from 1.0 to 1.2, the scheduling correction gain is adjusted. Corresponding adjustments were made to smoothly reduce the task inflow rate from 100% to 85%, thereby reducing maintenance oscillations under critical operating conditions and maintaining the stability of computing power output. In industrial environments, external data collected by environmental sensors sometimes cannot accurately reflect the dust accumulation or physical wear state inside the heat dissipation path. Blockages in heat dissipation ducts are often concealed, and effective non-invasive detection methods are lacking. This embodiment of the invention monitors the adjustment process of the feedback loop in real time, extracting the lag time of computer node temperature changes relative to load data changes. The system obtains the scheduling correction gain. After the action, the temperature change curve fed back by the temperature sensor is used to calculate the lag time of the curve relative to the starting point of the load data jump. The system calculates the lag time. The hysteresis constant under the preset calibration conditions Criterion deviation ,in, The deviation is used as a criterion. The lag time for real-time extraction is expressed in milliseconds (ms). The preset calibration phase lag; the criterion deviation. To objectively characterize the changes in the physical properties of the hardware heat conduction path, a heat dissipation path performance evaluation model is constructed. This model utilizes an evolution function representing the heat conduction relationship to determine the lag time. The mapping is represented by a heat dissipation resistance characteristic value. The specific mapping logic is as follows: the heat dissipation resistance reference table stored in the system memory is retrieved. When the lag time increases by 10ms compared to the calibration value, the heat dissipation resistance characteristic value increases by 5% on the 100% baseline. This reference table is pre-written by performing 10 sets of different heat dissipation gradient experiments on the equipment before it leaves the factory, and is associated with environmental parameter data to calculate the heat dissipation efficiency conduction index of the heat dissipation path. The system determines whether the fluctuation range of the environmental parameter data is within the preset stable range. When the fluctuation is stable and the change rate of the heat dissipation efficiency conduction index exceeds the preset threshold, a maintenance scheduling command is generated. When the external environment is stable but the heat dissipation efficiency index drops rapidly, it is determined that the change is caused by dust accumulation and blockage in the heat dissipation path, thus realizing the perception of the health status of the heat dissipation path.

[0022] Troubleshooting physical blockages typically relies on manual intervention. In unmanned factories or edge cabinet scenarios, manual maintenance is costly. To achieve automated closed-loop maintenance, this invention executes maintenance scheduling commands by injecting pulse loads into computer nodes to control the cooling fans into a variable-speed resonance mode. The system determines the natural frequency of the cooling fans based on their physical specifications and controls the injection frequency of the pulse loads to match the natural frequency of the cooling fans. Fluctuations in computing load induce resonant displacement in the cooling fans, and the mechanical energy generated by resonance loosens accumulated dust. The second derivative of the computer node's response time during pulse load execution is extracted as a characteristic parameter representing the implicit stress within the hardware. If this second derivative recovers to a preset range, the air resistance of the heat dissipation path is determined to have returned to normal, and the maintenance scheduling command is released. After long-term service, computer hardware experiences electromigration or thermal fatigue, leading to time-varying attenuation in its response to the same scheduling pulses. The system needs calibration capabilities to ensure scheduling accuracy throughout its lifecycle. Therefore, this invention calculates the ratio of the command issuance quantity to the real-time response quantity of the computer node to determine the scheduling sensitivity and defines this ratio as the system's equivalent stiffness characteristic value. The scheduling correction gain is adjusted based on this system's equivalent stiffness characteristic value. Weighted corrections are applied, and when the system's equivalent stiffness characteristic value falls below a preset threshold (i.e., the hardware response slows down), the scheduling correction gain is automatically increased. The proportional coefficient, through algorithmic intervention to compensate for physical performance degradation, ensures that the maintenance intervention strength can penetrate the damping of the physical layer. In response to the physical drift that may occur in the sensor, the system generates an environmental data correction operator when the evolution trend of environmental parameter data does not match the deterioration characteristics of the heat dissipation efficiency conduction index. The specific judgment logic is as follows: when the monitored ambient temperature is kept at a constant 25℃ for more than 300s, but the heat dissipation efficiency conduction index shows a continuous downward trend of more than 1% per second, the system determines that the sensor has a temperature drift deviation, and generates a correction coefficient with the absolute value of the downward slope as the environmental data correction operator, which is used to compensate for the environmental parameter data.

[0023] In distributed computing cluster scenarios, maintenance actions on a single node can easily trigger sudden changes in load flow and overload neighboring nodes. By adjusting the task migration step size, the impact of maintenance actions on global throughput is limited to a preset elastic range. The system establishes a virtual tension model among communicating computing nodes, exchanging real-time task processing queue lengths of each node via synchronous heartbeat messages at a frequency of 50ms. The difference in queue lengths between adjacent nodes is calculated as the tension input. When this difference increases by 10 pending tasks, the task source node increases the migration step size from the default 5% to 8%, thereby achieving smooth control of load flow and thus... The virtual tension model is used to characterize the coupling strength of computational task flows between nodes. When the output resistance of the target node increases due to maintenance depth compensation, the system controls the task migration speed between the task source node and the task receiving node by adjusting the virtual tension gain between nodes. The load migration process and the receiving process are synchronized on the time axis, thereby eliminating load surges caused by local maintenance and maintaining the operational stability at the cluster level. In addition, the system monitors the task migration dispersion of computer nodes in real time during pulse load execution. When the task migration dispersion exceeds a preset threshold of 0.15, the scheduling correction gain is automatically reduced. The calibration value of 0.15 was obtained by controlling each computing node to execute 5 sets of synchronous migration instructions within the 10% to 90% load range during the system initialization phase, calculating the standard deviation of the load rate of each node and taking twice the maximum standard deviation of the 5 sets of experiments. Overheat protection of computer nodes is achieved by reducing the load injection intensity. The maintenance depth compensation amount is determined by calculating the difference in heat dissipation efficiency conduction index of computer nodes before and after executing pulse load. The virtual tension model calculates the flow regulation damping between the task source node and the task receiving node based on the maintenance depth compensation amount. When the maintenance depth compensation amount increases, the system synchronously increases the task migration rate of the task source node and the task receiving redundancy of the task receiving node, thereby using the phase compensation of the task flow on the time axis to offset the computing power fluctuations caused by local maintenance. This process transforms the surface maintenance action of a single node into a smooth load evolution at the cluster level.

[0024] Task migration dispersion measures the deviation level of load distribution among nodes in a cluster. Its calculation process includes obtaining the real-time load rate of communicating computing nodes, calculating the sample variance of the real-time load rate of each node and extracting the square root of this sample variance. The system then adjusts the scheduling gain based on the task migration dispersion. The system employs nonlinear feedback; when the task migration dispersion exceeds a preset threshold of 0.15, it reduces the scheduling correction gain. The task distribution priority of the target node is reduced to alleviate the sudden pressure on the cluster interconnect bandwidth by utilizing the local load absorption mechanism. During system initialization, a standard operation sequence is selected as the test stimulus, and the ambient temperature is maintained at 20°C to 25°C. The temperature feedback curve of the computer node is recorded as it increases from standby load to 80% full load. By identifying the time coordinate corresponding to the maximum slope of the temperature feedback curve, the difference between the coordinate and the starting time of the load jump signal is calculated to determine the preset calibration phase lag. ,in, Characterizes the intrinsic heat conduction delay under clean conditions of the heat dissipation path, in milliseconds; executes maintenance scheduling commands and selects the natural frequency of the cooling fan. As a baseline, by adjusting the interrupt triggering cycle of the kernel task scheduler, the power spectrum of the injected pulse load is made to exhibit a certain characteristic. To determine the center frequency energy distribution, a computer node pulse width modulation controller is used to generate speed oscillations in response to temperature fluctuations. This causes the cooling fan to undergo radial displacement at the mechanical structure's resonant point, peeling away the deposits on the heatsink fins. Given the natural frequency of the cooling fan (in Hz), calculate the second derivative of the response time by collecting task response delay data for 100 consecutive sampling periods. The rate of change of delay fluctuation is characterized by calculating the first-order difference between adjacent sampling periods' delay differences. The second derivative is obtained by performing a difference operation on the first-order difference. ,when The numerical envelope returns to the preset interval [0, 0.05] ms / s Once the air resistance of the heat dissipation path is determined to have returned to normal, the maintenance command is lifted.

[0025] Example 1: In an edge computing node scenario deployed along a rail transit base station, the external environment is characterized by high-frequency mechanical vibrations from passing trains and temperature and humidity fluctuations accompanying changes in the outdoor environment. When the edge computing node undertakes real-time video stream analysis tasks, the processor load rate remains above 85%. Traditional protection mechanisms based on discrete temperature thresholds trigger thermal frequency modulation of the processor near the critical point where the ambient temperature rises to 55°C, causing task processing latency to fluctuate frequently between 20ms and 150ms. This scheduling instability causes logical queuing blockage in data processing. Frequent switching of the processor core voltage causes thermomechanical fatigue damage at the packaging level. This computer maintenance decision scheduling method uses real-time collected environmental parameter data such as temperature, humidity, and vibration frequency, and utilizes a preset attenuation correlation model to fuse multi-dimensional parameters to generate environmental stress characteristic values. Meanwhile, the kernel monitoring module calculates the current load capacity characteristic value. The feedback loop uses the load carrying capacity characteristic value With target value The deviation is taken as the input, and the environmental stress characteristic value is introduced. The disturbance compensation term is used to calculate the real-time scheduling correction gain. ,in, To adjust the gain for scheduling, The characteristic value of environmental stress, This is the characteristic value of load-bearing capacity margin. The scheduler adjusts its gain based on the target value of the load capacity margin. Adjusting the task distribution weight limits the task inflow rate from 100% to 82.5%. Since this adjustment process does not involve discrete switching threshold logic, the processor core temperature is maintained within a constant range of 78°C, eliminating scheduling fluctuations caused by thermal frequency modulation and secondary thermal stress shocks caused by voltage jumps.

[0026] The heat dissipation path efficiency evaluation model extracts the lag time of temperature change relative to load data change by real-time monitoring of the temperature response hysteresis of the feedback loop after load adjustment. Calculate the lag time Compared to the hysteresis constant under the calibration condition Criterion deviation ,in, The deviation is used as a criterion. The lag time for real-time extraction is expressed in milliseconds. The preset calibration phase lag, when the lag duration is... The time increased from 180ms in the initial operating state to 315ms, while the environmental stress characteristic value There was no upward fluctuation, and the criterion deviation was not observed. The system characterizes the increase in physical resistance of the heat dissipation airflow path, identifies the dust accumulation on the heat sink surface, and generates a maintenance scheduling command. The system executes the maintenance scheduling command, injecting a 60Hz pulsed load into the processor. This causes the thermal pulses generated by the computing load to match the inherent physical frequency of the cooling fan, inducing the cooling fan to produce variable-speed resonant displacement to peel off the dust adhering to the fin surface. The lag time is recorded 300 seconds after the maintenance action is executed. The response time dropped from 315ms to 195ms and remained stable under load disturbances. The second derivative of the response time recovered to the preset range near zero. The system determined that the physical conductivity of the heat dissipation path had been restored and released the maintenance scheduling command to restore the computer node to the preset computing power output state.

[0027] Example 2: In an industrial-grade computing node containing a processor, cooling fan, and multi-dimensional environmental sensors, the operational stability of the maintenance decision-making scheduling method is verified by acquiring real-time environmental parameter data. Gaussian white noise with a signal-to-noise ratio of 20dB and power frequency interference harmonics at a frequency of 50Hz are superimposed on the data acquisition link to simulate the electromagnetic interference environment of an industrial site. The sampling period is set to balance the real-time performance of data sensing with the computational load of the processor. When the fluctuation bandwidth of the monitored environmental parameters is in the range of 10Hz to 60Hz, the sampling period is determined to be 50ms. An experimental control group and the sample group of this invention are set up. The control group adopts a resource scheduling mode based on discrete temperature thresholds, while the sample group of this invention adopts a mode based on environmental stress characteristic values. The continuous scheduling mode of the drive simulates environmental stresses of different intensities by changing the mechanical vibration acceleration and the external ambient temperature. The test points cover the gradient working conditions where the vibration acceleration increases from 0.5g to 5.0g, as shown in Table 1.

[0028] Table 1: Comparison of System Operating Parameters under Different Vibration Stress Intensities

[0029] in, The lag time of temperature change relative to load data change is expressed in milliseconds (ms); the oscillation frequency characterizes the fluctuation frequency of resource scheduling weights, expressed in Hz; the heat dissipation efficiency conduction index is the quantified percentage of heat conduction efficiency of the heat dissipation path; the second derivative of response time characterizes the rate of change of implicit stress within the hardware, expressed in ms / s²; the results recorded in Table 1 show that when the vibration acceleration is in the range of 0.5g to 2.0g, the scheduling oscillation frequency of the control group increases with the increase of environmental stress. The sample group of this invention corrects the gain through scheduling. Smooth adjustment keeps the scheduling oscillation frequency below 0.1Hz, utilizing environmental stress characteristic values. Adjusting the task inflow rate to avoid frequent flipping of the discrete threshold at the decision boundary, when the vibration acceleration reaches 4.0g, the hysteresis time of the sample group of this invention... The latency increased from 188.5ms to 325.8ms, and the thermal efficiency conductivity index dropped to 78.6%. The system detected that the environmental parameter fluctuations were within a stable range, determined that dust accumulation had blocked the heat dissipation path, and generated a maintenance scheduling command. This command induced resonance in the cooling fan by injecting a pulse load into the processor. The maintenance action was executed with a delay of [duration missing]. The response time dropped back to 194.2ms, and the second derivative of the response time recovered from 48.4ms / s² to 1.8ms / s². When the vibration acceleration exceeded 4.5g, the degradation rate of the heat dissipation performance conductivity index exceeded the repair rate of the maintenance action. Because the task migration dispersion exceeded the preset threshold of 0.15, the system reduced the scheduling correction gain. To implement frequency reduction protection, this phenomenon confirms that the parameter range is determined based on the physical thermal limitations of the hardware. The thermal inertia residual generated by feedback adjustment is used to identify the resistance changes of the heat dissipation path, and the closed-loop recovery of heat dissipation efficiency is achieved through load excitation, so as to ensure the stability of computing power output of computer nodes under complex working conditions.

[0030] Example 3: This example combines Figures 1 to 2 The computer-based maintenance decision-making and scheduling method based on environmental factor data is explained, such as... Figure 1 As shown, step 101 acquires the load data and environmental parameter data of the computer node and inputs the load data into the feedback loop to perform resource scheduling. The feedback loop adjusts the scheduling correction gain based on the environmental parameter data. Step 102 then monitors the adjustment process of the feedback loop in real time and extracts the lag time of the computer node's temperature change relative to the load data change. Step 103 then constructs a heat dissipation path efficiency evaluation model containing the heat conduction relationship evolution function. The model maps the lag time to the environmental parameter data to calculate the heat dissipation efficiency conduction index of the heat dissipation path. Step 104 then determines whether the fluctuation range of the environmental parameter data is within a preset stable range. If it is within a stable range and the rate of change of the heat dissipation efficiency conduction index exceeds a threshold, a heat dissipation path blockage fault is determined and a maintenance scheduling command is generated. Finally, step 105 executes the maintenance scheduling command, injects pulse load to control the cooling fan to enter variable speed resonance mode, extracts the second derivative of the response time to verify the maintenance effect, and releases the maintenance scheduling command when the second derivative recovers to the preset range.

[0031] like Figure 2 As shown, the system architecture uses a full-duplex data interaction bus as the core communication hub, connecting the upper-layer logic operation unit and the lower-layer physical sensing execution unit. The upper layer of the system deploys a thermal inertial analysis engine, a feedback decision core, and an active excitation module. The thermal inertial analysis engine contains a lag time extraction unit and a heat dissipation performance evaluation model. The feedback decision core integrates a scheduling gain regulator and a fault and maintenance judgment unit. The active excitation module is equipped with a pulse load generation unit and a second-order derivative verification unit. The lower layer of the system connects an environmental perception layer, a computing node execution layer, and a distributed cooperative domain. The environmental perception layer contains temperature and humidity sensors, vibration sensors, and signal conditioning circuits. The computing node execution layer is equipped with a central processing unit (CPU) and a variable-speed cooling fan that supports resonant response. The distributed cooperative domain contains neighboring cooperative nodes and virtual tension control logic.

[0032] Example 4: In scenarios involving the deployment of large-scale data center nodes, physical differences exist in the heat sink fin gaps and airflow damping between different batches of servers. The system faces initial heat dissipation baseline drift and environmental noise interference. To determine the calibration phase lag of each node... In addition to maintaining the trigger threshold, the system executes the initialization calibration procedure, defines the initial state of environmental parameter data, acquires the original sampled values ​​of temperature, humidity, and vibration frequency, and uses a linear normalization function to map each factor to an interval. Based on the Arrhenius model, the weighting coefficients for temperature (0.6), humidity (0.25), and vibration stress (0.15) were determined, and the initial environmental stress characteristic values ​​were calculated by weighted summation. .

[0033] The system adjusts the load to generate a step response curve, monitors temperature feedback, and uses a heat dissipation path efficiency evaluation model to map the observed lag time into a heat dissipation efficiency conduction index. The specific mapping formula is as follows: ,in, For heat dissipation performance conductivity indicators; It is the thermal inertia constant, determined by the product of the thermal conductivity and specific heat volume of the radiator material; The lag time of temperature change relative to load data change, in milliseconds; the system calculates the heat dissipation efficiency conduction index during a 24-hour no-load operation cycle. The sample mean and sample standard deviation, using The principle is to determine the trigger threshold for maintenance scheduling commands, setting this threshold as three times the sum of the sample standard deviation and the sample mean. The variance of environmental parameter data is collected and analyzed in real time. When the variance value is lower than the preset signal-to-noise ratio threshold, the environment is considered to be within a stable range. At this point, the real-time calculated heat dissipation efficiency conductivity index... If the deviation from the sample mean exceeds the trigger threshold, a maintenance scheduling instruction is generated.

[0034] Example 5: In the node deployment scenario of a cross-regional distributed computing center, due to differences in power supply ripple and hardware aging levels at different geographical locations, the system executes an initial response calibration process to determine the baseline distribution of the system's equivalent stiffness characteristic value. The system distributes standard floating-point operation sequences with preset instruction issuance amounts to the nodes, collects real-time response data from the nodes in real time, calculates the ratio of instruction issuance amount to real-time response amount, and determines the system's equivalent stiffness characteristic value. Defined as ,in, For the amount of instructions issued, To determine the real-time response, a moving average is applied to the ratio over 500 consecutive calibration periods to filter out fluctuations caused by instantaneous task scheduling, thus establishing the initial sensitivity benchmark for the node, which serves as the scheduling correction gain. The basis for adjusting the proportional coefficient.

[0035] When the system faces a cluster operation with multi-node collaborative scheduling, the system executes an optimization procedure for the virtual tension model gain among the communicating computing nodes to balance the load migration pressure generated by maintenance actions. The system monitors the task migration dispersion between the task source node and the task receiving node to obtain the initial coupling strength of the task flow. The iteration step size of the virtual tension gain is set to 0.05. The feedback loop is used to monitor the changing trend of the global throughput. When the target node performs maintenance actions, causing the deviation between the load migration speed and the receiving speed to exceed 15%, the system triggers the gain compensation mechanism to adjust the virtual tension gain between nodes so that the global throughput is maintained within the preset elastic range. Based on the evolution curve of the second derivative of the response time under different maintenance depths, the system determines the maintenance intervention intensity that satisfies the recovery of heat dissipation performance without causing node overload.

[0036] Example 6: In the pre-deployment and debugging scenario of a multi-node heterogeneous computing power cluster, the system executes a response latency baseline establishment program to eliminate the interference of underlying hardware physical differences on scheduling sensitivity. By acquiring instruction issuance and real-time response data under idle conditions, the length of the sliding time window is set to 2000 sampling periods, and the equivalent stiffness characteristic value of the system in each period is calculated and stored. The discrete distribution of, where, The equivalent stiffness eigenvalue of the system. For the amount of instructions issued, The equivalent stiffness eigenvalue of the system is the real-time response quantity. The calculation formula is Then, an anomaly filtering procedure based on the standard deviation ratio elimination algorithm is executed to remove response noise points that deviate from the mean by more than 2 standard deviations, and the initial scheduling correction benchmark value of the node is determined by the mathematical expectation value of the remaining valid samples.

[0037] When the cluster enters the dynamic task scheduling stress test mode, the system executes a virtual tension model parameter tuning program for load migration dispersion to lock the scheduling efficiency of the system without triggering overheat protection. By controlling the task source node to increase the task migration speed, the incremental step size of the virtual tension gain is set to 0.02, and the second derivative characteristic data of the processor's response time is obtained in each iteration cycle. When the rate of change of the second derivative of the response time exceeds 10% and the task migration dispersion reaches the critical point of 0.12, the system stops the upward optimization action of the virtual tension gain and performs forced gain clamping to maintain the current parameters as the upper limit of the maintenance scheduling gain for steady-state operation.

[0038] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0039] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A computer-based maintenance decision-making and scheduling method based on environmental factor data, characterized in that, The method includes the following steps: Step 101: Obtain the load data and environmental parameter data of the computer node, and input the load data into the feedback loop to perform resource scheduling. The feedback loop adjusts the scheduling correction gain based on the environmental parameter data. Step 102: Monitor the adjustment process of the feedback loop in real time and extract the lag time of the temperature change of the computer node relative to the load data change. Step 103: Construct a heat dissipation path performance evaluation model. The heat dissipation path performance evaluation model includes an evolution function that characterizes the heat conduction relationship. The lag time is correlated and mapped with environmental parameter data through the heat dissipation path performance evaluation model in order to calculate the heat dissipation performance conduction index of the heat dissipation path. Step 104: Determine whether the fluctuation range of the environmental parameter data is within the preset stable range. If it is within the preset stable range and the rate of change of the heat dissipation efficiency conduction index exceeds the preset threshold, then determine that the heat dissipation path has a blockage fault and generate a maintenance scheduling instruction. Step 105: Execute the maintenance scheduling command by injecting a pulse load into the computer node to control the cooling fan to enter the variable speed resonance mode, and extract the second derivative of the response time of the computer node during the pulse load execution to verify the maintenance effect. The maintenance scheduling command is released when the second derivative recovers to the preset range.

2. The computer-based maintenance decision-making and scheduling method based on environmental factor data according to claim 1, characterized in that, The process of adjusting the scheduling correction gain in step 101 specifically includes: calculating the ratio of the number of instructions issued to the number of real-time responses of computer nodes to determine the scheduling sensitivity; and weighting the scheduling correction gain based on the scheduling sensitivity.

3. The computer-based maintenance decision-making and scheduling method based on environmental factor data according to claim 1, characterized in that, Step 102, which involves extracting the lag time, specifically includes: obtaining the temperature change curve fed back by the temperature sensor after the scheduling correction gain action; and calculating the lag time of the temperature change curve relative to the starting point of the jump in the load data. ; Calculate the lag time Deviation from the hysteresis constant under the preset calibration conditions ;in, , The deviation is used as a criterion. For the lag time, This is the preset calibration phase lag.

4. The computer-based maintenance decision-making and scheduling method based on environmental factor data according to claim 1, characterized in that, The evolution function in the heat dissipation path effectiveness evaluation model is configured as follows: mapping the lag time to the heat dissipation resistance characteristic value; calculating the rate of change of the heat dissipation resistance characteristic value within the preset monitoring period; and determining the physical attenuation characteristics of the heat dissipation path based on the rate of change.

5. The computer-based maintenance decision-making and scheduling method based on environmental factor data according to claim 1, characterized in that, The criteria for determining that the heat dissipation path is blocked in step 104 include: extracting the trend component of the heat dissipation performance conduction index; when the absolute value of the slope of the trend component exceeds the preset attenuation threshold, and the variance of the load data is less than the preset variance threshold, it is determined that the heat dissipation path is blocked by dust accumulation.

6. The computer-based maintenance decision-making and scheduling method based on environmental factor data according to claim 1, characterized in that, The operation of controlling the cooling fan to enter the variable speed resonance mode in step 105 specifically includes: controlling the injection frequency of the pulse load to match the natural frequency of the cooling fan, and inducing the cooling fan to generate resonance displacement through the periodic changes of the load data.

7. The computer-based maintenance decision-making and scheduling method based on environmental factor data according to claim 1, characterized in that, The operation of verifying the maintenance effect and releasing the maintenance scheduling command when the second derivative recovers to the preset range specifically includes: calculating the second derivative of the response time relative to the pulse load execution time; if the second derivative is within the preset range, determining that the air resistance of the heat dissipation path has returned to normal, and terminating the injection of the pulse load.

8. The computer-based maintenance decision-making and scheduling method based on environmental factor data according to claim 1, characterized in that, The method also includes: generating an environmental data correction operator when the evolution trend of environmental parameter data does not match the degradation characteristics of heat dissipation efficiency conduction index; and using the environmental data correction operator to compensate for the environmental parameter data in order to correct the sensor sampling error.

9. The computer-based maintenance decision-making and scheduling method based on environmental factor data according to claim 1, characterized in that, The method also includes: by adjusting the task migration step size, the impact of maintenance actions on global throughput is limited to a preset elastic range, so as to maintain the consistency of the operating frequency of the distributed cluster during maintenance.

10. The computer-based maintenance decision-making and scheduling method based on environmental factor data according to claim 1, characterized in that, The method also includes: real-time monitoring of task migration dispersion of computer nodes during pulsed load execution; when the task migration dispersion exceeds a preset threshold, reducing the scheduling correction gain to achieve overheat protection of computer nodes by reducing the load injection intensity.

Citation Information

Patent Citations

  • Heat dissipation case for installing computer monitoring system

    CN213634316U