Temperature monitoring method of server and electronic device
By meticulously dividing the heat dissipation area in the server and dynamically adjusting the sampling frequency, the problem of insufficient area division and anomaly location capabilities in traditional server temperature control technology is solved. This enables accurate identification and rapid response of temperature monitoring in high-density servers, improving monitoring accuracy and resource utilization efficiency.
Patent Information
- Application Number
- CN202511190995.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-25
AI Technical Summary
Traditional server temperature control technology suffers from weak regional segmentation and anomaly localization capabilities, as well as crude dynamic sampling and threshold management in high-density scenarios, making it difficult to balance monitoring accuracy, resource consumption, and fault response speed.
By meticulously dividing the server plane into multiple heat dissipation areas, deploying temperature monitoring points according to hardware layout and heat dissipation characteristics, dynamically adjusting the sampling frequency, and combining the temperature change rate and mean standard deviation for anomaly detection.
It enables accurate identification and rapid response of temperature monitoring in high-density servers, improving monitoring accuracy, reducing resource consumption, and shortening troubleshooting time.
Smart Images

Figure CN120687328B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of servers, and in particular to a temperature monitoring method of a server and an electronic device. BACKGROUND
[0002] With the improvement of hardware integration, the temperature abnormality risk of the flat heat dissipation area of storage and computing servers has significantly increased, and the hardware thermal coupling problem is prominent in multi-disk storage servers or multi-processor computing servers. Abnormal single-point heat dissipation may cause a chain failure.
[0003] The temperature control scheme of the traditional server has the following limitations when dealing with such scenarios:
[0004] (1) Weak area division and abnormal positioning capability
[0005] The flat area division is extensive, and the monitoring points are sparse and difficult to identify single-point abnormalities. Some technologies increase the monitoring points but lack temperature gradient time analysis, and the abnormality recognition is delayed.
[0006] (2) Extensive dynamic sampling and threshold management
[0007] The sampling strategy does not consider the thermal sensitivity of the area, resulting in delayed monitoring of critical areas or waste of resources. The temperature threshold setting does not consider real-time variables, and complex conditions are prone to misjudgment.
[0008] The above limitations make it difficult to balance the monitoring accuracy, resource consumption and fault response speed in high-density server scenarios. SUMMARY
[0009] The present application provides a temperature monitoring method and an electronic device for a server to at least solve the problems of weak area division and abnormal positioning capability, and extensive dynamic sampling and threshold management in the current flat temperature control technology of the server.
[0010] The present application provides a temperature monitoring method of a server, comprising the following steps: obtaining a first temperature change rate of at least one temperature monitoring point in the server; if the first temperature change rate of any of the temperature monitoring points is greater than or equal to a first preset change rate, determining a target heat dissipation area where any of the temperature monitoring points is located, and adjusting the sampling frequency of the temperature monitoring points in the target heat dissipation area to a target sampling frequency; based on the target sampling frequency, detecting the second temperature change rate of the temperature monitoring points in the target heat dissipation area within a preset time length, if there is a target temperature monitoring point in the target heat dissipation area whose second temperature change rate within the preset time length is greater than a second preset change rate, determining that the target temperature monitoring point is temperature abnormal, otherwise, obtaining a temperature monitoring result of the target heat dissipation area based on the standard deviation of the temperature mean value of the target heat dissipation area within the preset time length.
[0011] The application further provides a temperature monitoring system of a server, comprising: an acquisition module, configured to acquire a first temperature change rate of at least one temperature monitoring point in the server; an adjustment module, configured to, if the first temperature change rate of any of the temperature monitoring points is greater than or equal to a first preset change rate, determine a target heat dissipation area where any of the temperature monitoring points is located, and adjust a sampling frequency of the temperature monitoring points in the target heat dissipation area to a target sampling frequency; and a monitoring module, configured to detect a second temperature change rate of the temperature monitoring points in the target heat dissipation area within a preset time length based on the target sampling frequency, and if the second temperature change rate of a target temperature monitoring point in the target heat dissipation area within the preset time length is greater than a second preset change rate, determine that the target temperature monitoring point is abnormal in temperature, otherwise, obtain a temperature monitoring result of the target heat dissipation area based on a standard deviation of a temperature mean value of the target heat dissipation area within the preset time length.
[0012] The application further provides an electronic device, comprising: a memory, configured to store a computer program; and a processor, configured to execute the computer program to implement the steps of the temperature monitoring method of the server.
[0013] The application further provides a non-volatile computer readable storage medium, wherein the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the temperature monitoring method of the server.
[0014] The application further provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the temperature monitoring method of the server.
[0015] According to the application, if the first temperature change rate of any of the temperature monitoring points in the server is greater than or equal to a first change rate, the target heat dissipation area of any of the temperature monitoring points is determined, the sampling frequency of the temperature monitoring points in the target heat dissipation area is adjusted to a target sampling frequency, the second temperature change rate of the temperature monitoring points in the target heat dissipation area within a preset time length is detected based on the target sampling frequency, if the second temperature change rate of a target temperature monitoring point in the target heat dissipation area within the preset time length is greater than a second change rate, it is determined that the target temperature monitoring point is abnormal in temperature, otherwise, the temperature monitoring result of the target heat dissipation area is obtained based on the standard deviation of the temperature mean value of the target heat dissipation area within the preset time length. Therefore, the problems of weak region division and abnormal positioning ability, extensive dynamic sampling and threshold management of the current server plane temperature control technology are solved, the plane region division is refined, the sampling frequency is automatically adjusted according to the temperature change amplitude, the accurate identification of single-point abnormality in the region is realized, and the monitoring accuracy is improved, and in the high-density server scene, the monitoring accuracy, resource consumption and response speed can be balanced. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort based on these drawings.
[0017] Figure 1 The flow chart of the temperature monitoring method of the server according to an embodiment of the present application;
[0018] Figure 2 The schematic diagram of the area division of the server according to an embodiment of the present application;
[0019] Figure 3 The flow chart of the temperature monitoring method of the server according to an embodiment of the present application;
[0020] Figure 4 The schematic diagram of the temperature monitoring system of the server according to an embodiment of the present application;
[0021] Figure 5 The schematic diagram of the electronic device structure according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort fall within the protection scope of the present application.
[0023] It should be noted that, in the description of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover the non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0024] With the improvement of hardware integration, the temperature abnormality risk of the flat heat dissipation area of storage and computing servers is significantly increased. In multi-disk storage servers or multi-processor computing servers, the hardware thermal coupling problem in the same flat heat dissipation area is prominent, and a single point of abnormal heat dissipation may trigger a chain failure. The traditional temperature control scheme has obvious limitations in dealing with such scenarios: the static sampling strategy is difficult to capture the sudden temperature rise, the fixed temperature threshold cannot adapt to the hardware thermal characteristics under different working conditions (such as the temperature difference between the training task and the idle state of the computing server), and the flat area division is rough (such as the entire hard disk backplane as a single monitoring unit), which makes it difficult to accurately locate the fault source. Although the BMC (Baseboard Management Controller) as the core of the server out-of-band management has basic temperature collection capability, the existing technology is mostly limited to single-point data recording and fixed strategy alarm level, lacking dynamic analysis of temperature gradient and adaptive response mechanism within the flat area.
[0025] In the flat temperature control technology, the existing scheme mainly has the following problems:
[0026] (1) Weak area division and abnormality positioning capability: Most schemes simply divide the server plane into several large areas without fine segmentation according to the hardware layout, resulting in sparse temperature monitoring points within the area and difficulty in identifying single-point abnormalities. Some technologies increase the number of monitoring points, but lack temperature gradient analysis in the time dimension, and abnormality recognition has obvious delay.
[0027] (2) Dynamic sampling and threshold management are rough: The sampling strategy mostly uses fixed frequency or simple threshold switching, without considering the differences in regional thermal sensitivity, leading to delayed monitoring of critical areas or resource waste. Temperature thresholds are mostly set based on hardware nominal values without considering real-time load, historical temperature, etc., which may lead to misjudgment under complex working conditions.
[0028] The core bottleneck of the current flat temperature control technology for storage and computing servers is that the low abnormality positioning accuracy caused by insufficient area division granularity and the static strategy that cannot adapt to the dynamic thermal characteristics of hardware. These limitations make it difficult for traditional schemes to balance monitoring accuracy, resource consumption, and fault response speed in high-density server scenarios.
[0029] To solve the above problems, the embodiment of the present application provides a temperature monitoring method of a server, as shown in Figure 1 .
[0030] As shown in Figure 1 , the temperature monitoring method of the server comprises the following steps:
[0031] Step S101, obtaining a first temperature change rate of at least one temperature monitoring point in the server.
[0032] Optionally, in some embodiments, before acquiring the first temperature change rate of the at least one temperature monitoring point in the server, the server is divided into a plurality of heat dissipation regions based on the heat dissipation characteristics of the hardware regions of the server, wherein the plurality of heat dissipation regions are respectively a backplane heat dissipation region, a power supply heat dissipation region, a central processing unit heat dissipation region, and an interface card heat dissipation region; the number of temperature monitoring points in the backplane heat dissipation region is determined according to the area and heat dissipation characteristics of the backplane heat dissipation region, the number of temperature monitoring points in the power supply heat dissipation region is determined according to the area and heat dissipation characteristics of the power supply heat dissipation region, the number of temperature monitoring points in the central processing unit heat dissipation region is determined according to the area and heat dissipation characteristics of the central processing unit heat dissipation region, and the number of temperature monitoring points in the interface card heat dissipation region is determined according to the area and heat dissipation characteristics of the interface card heat dissipation region.
[0033] The server comprises a plurality of heat dissipation regions divided by a hardware layout, and each heat dissipation region is provided with at least one temperature monitoring point, which can be one or multiple.
[0034] Specifically, as shown in Figure 2 The system first divides the server plane into a plurality of independent heat dissipation regions according to the hardware layout. Taking a storage server as an example, in the storage server, the hard disk backplane is divided into a front-mounted backplane hard disk heat dissipation region, an in-built backplane hard disk heat dissipation region, and a rear-mounted backplane hard disk heat dissipation region due to the differences in heat dissipation conditions caused by the different positions of the front and rear. The front-mounted backplane hard disk heat dissipation region is close to the air inlet of the server, and the air flow is relatively smooth, but the heat dissipation may be affected by the accumulation of dust from the external environment; the in-built backplane hard disk region is inside the server, and the heat dissipation mainly depends on the internal air duct; the rear-mounted backplane hard disk region is close to the air outlet, and the heat is quickly discharged, but the high temperature of the air outlet may also affect it.
[0035] In addition, the power supply unit (PSU, Power Supply Unit) heat dissipation region is separately divided due to its large amount of heat generated by itself, and its stability is related to the overall operation of the server; the central processing unit (CPU, Central Processing Unit) heat dissipation region is also separately divided as the operation core of the server, which generates concentrated heat and is extremely sensitive to temperature; the interface card heat dissipation region contains various network and storage interface cards, which are densely arranged and generate heat when working, and is also divided as an independent region.
[0036] In each heat dissipation region, the system reasonably deploys a plurality of temperature monitoring points in each heat dissipation region according to the area, hardware heat intensity, and other factors. For example, in the front-mounted backplane hard disk heat dissipation region, temperature sensors are arranged at equal intervals along the hard disk arrangement direction to ensure that the temperature changes of the entire region can be covered, thereby forming a high-density monitoring network to achieve comprehensive and accurate capture of the temperature changes of each part of the server.
[0037] Through the technical solution, the number of monitoring points is dynamically determined according to the area and heat dissipation characteristics (such as air duct design, equipment heat density, heat conduction efficiency, etc.) of each heat dissipation area, efficient allocation of monitoring resources is realized, and the temperature monitoring accuracy is improved by real-time response to heat dissipation abnormalities in different areas through the partition independent monitoring system.
[0038] Specifically, after the system is started, based on the division of each heat dissipation area of the server and the deployment of the monitoring points, an initial sampling frequency is set for each temperature monitoring point, all temperature monitoring points in all heat dissipation areas perform temperature sampling at the initial sampling frequency, and the first temperature change rate of all temperature monitoring points is monitored.
[0039] In the embodiment of the application, the initial sampling frequency is a sampling frequency of once every 10 seconds.
[0040] Step S102, if the first temperature change rate of any temperature monitoring point is greater than or equal to the first preset change rate, the target heat dissipation area where the temperature monitoring point is located is determined, and the sampling frequency of all temperature monitoring points in the target heat dissipation area is adjusted to the target sampling frequency.
[0041] The first preset change rate is 0.5℃ / min, and the second preset change rate is 5℃ / min.
[0042] When the server as a whole is in a temperature stable state, that is, the first temperature change rate of all temperature monitoring points in all heat dissipation areas is within the range of ±0.5℃ / min, the system uniformly sets the temperature monitoring points to an initial sampling frequency of once every 10 seconds for temperature data collection in each area. This uniform low-frequency collection strategy can not only meet the basic monitoring needs of the temperature of the server in the normal running state, but also effectively reduce the resource consumption of the baseboard management controller, and ensure the stable operation of the system.
[0043] If the first temperature change rate of any temperature monitoring point in all temperature monitoring points is greater than or equal to 0.5℃ / min within a preset time (for example, 5 minutes), the target heat dissipation area where the temperature monitoring point is located is determined, and the sampling frequency of all temperature monitoring points in the target heat dissipation area is adjusted to the target sampling frequency.
[0044] Optionally, in some embodiments, adjusting the sampling frequency of all temperature monitoring points in the target heat dissipation area to the target sampling frequency includes: determining the target temperature change interval of the first temperature change rate of any temperature monitoring point, and determining the target sampling frequency according to the target temperature change interval.
[0045] Optionally, in some embodiments, the target temperature change interval of the first temperature change rate of any temperature monitoring point is determined, and the target sampling frequency is determined according to the target temperature change interval, including: judging whether the first temperature change rate of any temperature monitoring point is greater than or equal to a first preset change rate and less than or equal to a second preset change rate; if the first temperature change rate of any temperature monitoring point is greater than or equal to the first preset change rate and less than or equal to the second preset change rate, it is determined that the target temperature change interval is a first interval, and the first sampling frequency is taken as the target sampling frequency of the temperature monitoring point in the target heat dissipation area; if the first temperature change rate of any temperature monitoring point is greater than the second preset change rate, it is determined that the target temperature change interval is a second interval, and the second sampling frequency is taken as the target sampling frequency of the temperature monitoring point in the target heat dissipation area.
[0046] In some embodiments, the first preset change rate is 0.5℃ / min, the second preset change rate is 5℃ / min, the first interval is greater than or equal to 0.5℃ / min and less than or equal to 5℃ / min, the second interval is greater than 5℃ / min, the first sampling frequency is 1 / s, and the second sampling frequency is 10 / s.
[0047] In some embodiments, the first interval is greater than or equal to 0.5℃ / min and less than or equal to 5℃ / min, and the second interval is greater than 5℃ / min.
[0048] If the first temperature change rate of any temperature monitoring point is greater than or equal to 0.5℃ / min and less than or equal to 5℃ / min, it is determined that the target temperature change interval is the first interval, and if the first temperature change rate of any temperature monitoring point is greater than 5℃ / min, it is determined that the target temperature change interval is the second interval, and the target sampling frequency is determined according to the target temperature change interval.
[0049] In some embodiments, the first sampling frequency is 1 / s, which is a medium frequency sampling in the application, and the second sampling frequency is 0.1 / s, which is a high frequency sampling in the application.
[0050] Specifically, the adjustment mode of the sampling frequency is as follows:
[0051] Let f be the temperature sampling frequency (unit: times / s), ΔTi be the temperature change rate of the ith temperature monitoring point in the region (unit: ℃ / min), ΔTmax=max(ΔT1, ΔT2, …, ΔTn) be the maximum temperature change rate of all temperature monitoring points in the region, and the sampling frequency rule can be expressed as:
[0052] (1) Initial sampling frequency:
[0053] If and only if the temperature change rate of all temperature monitoring points in the heat dissipation area satisfies ∣ΔTi∣≤0.5℃ / min (for all i):
[0054] The initial sampling frequency of the temperature monitoring point is maintained: that is, f=0.1 (i.e., sampling once every 10 seconds);
[0055] (2) Medium frequency sampling
[0056] When the temperature change rate of at least one temperature monitoring point in the heat dissipation area satisfies 0.5℃ / min<ΔTi≤5℃ / min:
[0057] Adjust the sampling frequency of all temperature monitoring points in the heat dissipation area to the target sampling frequency: f=1 (i.e., sampling once every second);
[0058] (3) High frequency sampling trigger
[0059] When the temperature change rate of at least one monitoring point in the area satisfies ΔTi>5℃ / min:
[0060] Adjust the sampling frequency of all temperature monitoring points in the heat dissipation area to the target sampling frequency: that is, f=10 (i.e., sampling once every 0.1 seconds).
[0061] For example, in a certain heat dissipation area, when the first temperature change rate of a temperature monitoring point rises to the interval of 0.5℃ / min-5℃ / min, the system will preferentially increase the sampling frequency of all temperature monitoring points in the heat dissipation area where the temperature monitoring point is located to the first sampling frequency, i.e., sampling once every second. Figure 2 For example, if the temperature change rate of hard disk 1 in the front backplane hard disk area reaches 0.8℃ / min, the system will immediately increase the sampling frequency of all temperature monitoring points in the front backplane hard disk area to the first sampling frequency, i.e., the target sampling frequency of all temperature monitoring points is sampling once every second, so as to more timely and accurately obtain the temperature change data in the area.
[0062] If the first temperature change rate of a temperature monitoring point in a certain heat dissipation area exceeds 5℃ / min, the system determines that the area is facing an emergency overheating risk and immediately starts the high frequency sampling mode of 0.1 times / second for all temperature monitoring points in the area. For example, if the temperature change rate of a temperature monitoring point in the CPU area reaches 7℃ / min, all temperature monitoring points in the CPU area will enter the high frequency sampling state, i.e., sampling once every 0.1 seconds, to capture temperature data at a very high frequency and ensure accurate grasp of the rapidly changing temperature in the area.
[0063] Through the above technical solution, when the temperature change rate of at least one temperature monitoring point in a certain heat dissipation area is in the interval of 0.5℃ / min-5℃ / min, the system will increase the sampling frequency of the monitoring points in this area to 1 time per second. This adjustment can obtain temperature data at an appropriate frequency, avoiding data redundancy caused by over-sampling, and timely capturing the subtle changes in temperature, providing accurate and time-sensitive data support for subsequent analysis, which helps to accurately judge the temperature change trend and potential problems. If the temperature change rate of a temperature monitoring point in the area exceeds 5℃ / min, the system immediately starts a high-frequency sampling mode of 1 time per 0.1 second. In the face of an emergency overheating risk, this ultra-high frequency sampling can capture the rapid fluctuations of temperature in real time and accurately, not missing any moment of temperature change, ensuring that the most detailed and accurate temperature data is obtained, providing a key basis for dealing with emergency situations.
[0064] Optionally, in some embodiments, after obtaining the first temperature change rate of the monitoring points of each heat dissipation area, it includes: in the absence of any temperature monitoring point having a first temperature change rate greater than or equal to the first preset change rate within the preset time period, controlling the temperature monitoring points in the heat dissipation area to maintain the current sampling frequency.
[0065] The preset time period can be a threshold value set by the user in advance, can be a threshold value obtained through a limited number of experiments, or can be a threshold value obtained through a limited number of computer simulations, without specific limitation.
[0066] It can be understood that after obtaining the first temperature change rate of all temperature monitoring points of each heat dissipation area, if the first temperature change rate of all temperature monitoring points within the preset time period is less than the first preset change rate, then all temperature monitoring points in all heat dissipation areas are controlled to maintain the current sampling frequency, with sampling once every 10 seconds.
[0067] The first preset change rate set by the system is 0.5℃ / min, the preset time period is 5 minutes, and the initial sampling frequency of all temperature monitoring points is once every 10 seconds.
[0068] After the system starts running, each temperature monitoring point is sampled once every 10 seconds, and the temperature is recorded. Every 5 minutes (i.e., the preset time period), the system calculates the temperature change rate of each temperature monitoring point within the 5 minutes. For example, for temperature monitoring point A, the temperature at the start time is 25℃, and the temperature after 5 minutes is 26℃, so its temperature change rate within 5 minutes is:
[0069] (26-25)÷5=0.2℃ / min;
[0070] The temperature of temperature monitoring point B at the start time is 23℃, and the temperature after 5 minutes is 24℃, and the temperature change rate is:
[0071] (24-23)÷5=0.2℃ / minute;
[0072] By analogy, the rate of temperature change at all monitoring points in each heat dissipation area is calculated.
[0073] The system evaluates the calculated temperature change rate for all temperature monitoring points and finds that none of them have a temperature change rate greater than or equal to 0.5℃ / minute (the first preset change rate). In this case, all temperature monitoring points continue to maintain their current sampling frequency of once every 10 seconds, without increasing or decreasing the sampling frequency.
[0074] The above technical solution determines whether to adjust the sampling frequency based on the rate of temperature change at all monitoring points in each heat dissipation area within a preset time period. When the rate of temperature change at all monitoring points is less than a first preset rate (e.g., 0.5℃ / minute), the current sampling frequency is maintained. This dynamic adjustment mechanism ensures that the sampling frequency matches the actual temperature changes, avoiding unnecessary resource consumption such as data storage space, transmission bandwidth, and processor computing resources caused by using high-frequency sampling when temperature changes are gradual, thus achieving efficient resource utilization.
[0075] Step S103: Based on the target sampling frequency, detect the second temperature change rate of the temperature monitoring point in the target heat dissipation area within a preset time. If the second temperature change rate of the target temperature monitoring point in the target heat dissipation area within the preset time is greater than the second preset change rate, then determine that the temperature of the target temperature monitoring point is abnormal. Otherwise, obtain the temperature monitoring result of the target heat dissipation area based on the standard deviation of the mean temperature of the target heat dissipation area within the preset time.
[0076] It should be understood that the temperature monitoring points in the target heat dissipation area sample the temperature based on the target sampling frequency. During the continuous collection of temperature data, the second temperature change rate within a preset time period is calculated in real time. The server's temperature anomaly monitoring results are obtained based on the second temperature change rate of the temperature monitoring points in the target heat dissipation area within the preset time period.
[0077] Specifically, during the continuous acquisition of temperature data, the system calculates the second temperature change rate at each temperature monitoring point in real time. This second temperature change rate is obtained by calculating the magnitude of temperature increase per unit time. For example, if the temperature at a monitoring point rises from 25℃ to 35℃ within 1 minute, its second temperature change rate is (35℃ - 25℃) ÷ 1 minute = 10℃ / minute. Calculating the second temperature change rate provides crucial information for determining whether temperature changes are abnormal.
[0078] Specifically, the system continuously compares the second temperature change rates of all temperature monitoring points, detects whether there is a target temperature monitoring point with a second temperature change rate greater than the second preset change rate, and if there is a target temperature monitoring point, it means that the temperature rising speed of the target temperature monitoring point is faster than that of other temperature monitoring points in the neighborhood, and the target temperature monitoring point is marked as an abnormal temperature point.
[0079] Optionally, in some embodiments, after determining that the temperature of the target temperature monitoring point is abnormal, the method further includes: performing temperature abnormality warning based on the target temperature monitoring point, and sending alarm information of the temperature abnormality of the target temperature monitoring point to a preset terminal.
[0080] The temperature abnormality warning is performed based on the abnormal temperature point, and alarm information of the temperature abnormality of the target temperature monitoring point in the server is sent to a preset terminal.
[0081] For example, when the temperature rising speed of a temperature monitoring point of a certain hard disk is significantly faster than that of other points in the neighborhood, and the temperature rising amplitude difference per minute exceeds 5°C, the point will be marked as an abnormal temperature point.
[0082] If the temperature of a certain temperature monitoring point rises from 25°C to 35°C in 1 minute, its second temperature change rate is (35°C-25°C) ÷ 1 minute = 10°C / minute, which is greater than 5°C / minute. At this time, it is determined that the temperature monitoring point is an abnormal temperature point. If the second temperature change rate of all temperature monitoring points is less than 5°C / minute in 1 minute, it is determined that all temperature monitoring points are normal temperature points.
[0083] Through the above technical solution, once the second temperature change rate of a temperature monitoring point exceeds the second preset change rate, it can be quickly marked as an abnormal temperature point. This real-time monitoring mechanism can timely detect abnormal temperature changes, which saves valuable time for subsequent warning and processing, and prevents the abnormal situation from further deteriorating. When an abnormal temperature point is detected, the system will immediately perform temperature abnormality warning based on the abnormal point and send alarm information to a preset terminal. This enables relevant personnel to learn about the temperature abnormality in the first time, whether it is an on-site operation and maintenance personnel or a remote monitoring personnel, who can take timely measures to reduce the probability of failure.
[0084] Optionally, in some embodiments, the temperature monitoring result of the target heat dissipation area is obtained based on the temperature mean standard deviation of the target heat dissipation area within a preset time length, and the method further includes: calculating the temperature mean standard deviation of the temperature monitoring points in the target heat dissipation area within the preset time length; if the temperature mean standard deviation is greater than a preset threshold value corresponding to the target heat dissipation area, it is determined that the temperature of the target heat dissipation area is abnormal, and temperature abnormality warning is performed, and alarm information of the temperature abnormality of the target heat dissipation area is sent to a preset terminal.
[0085] It should be understood that if there is no target temperature monitoring point with a second temperature change rate greater than the second preset change rate, it means that all temperature monitoring points are normal temperature points, and at this time, it is necessary to determine whether the entire target heat dissipation region is a temperature abnormal region according to the temperature mean standard deviation of the temperature monitoring points.
[0086] Specifically, the system evaluates the temperature distribution of the target heat dissipation region by calculating the temperature mean and the temperature mean standard deviation of all temperature monitoring points in the target heat dissipation region. If the temperature mean standard deviation of the target heat dissipation region is greater than a preset threshold, it means that the target heat dissipation region is temperature abnormal, temperature abnormality warning needs to be performed, and alarm information of the temperature abnormality of the target heat dissipation region in the server is sent to a preset terminal.
[0087] It should be noted that different heat dissipation regions have different normal fluctuation ranges of temperature distribution due to different hardware compositions and working characteristics. Therefore, the system sets exclusive standard deviation thresholds according to the hardware characteristics and historical operation data of each region. For example, the preset threshold is set to 3°C for the PSU region because the heat is stable and the temperature fluctuation is small when the PSU region is normally working. The preset threshold is set to 4°C for the interface card region because the working state of the device in the interface card region is variable and the temperature fluctuation is large. When the calculated standard deviation of a certain region exceeds the preset threshold, the system immediately triggers a regional temperature abnormality warning.
[0088] For example, if the target heat dissipation region is the PSU region, the temperature mean standard deviation of the PSU region in a preset time length is calculated. If the temperature mean standard deviation of the PSU region in the preset time length is 2°C, which is less than 3°C, it means that the temperature of the PSU region is normal. If the temperature mean standard deviation of the PSU region in the preset time length is 5°C, which is greater than 3°C, it means that the temperature of the PSU region is abnormal, at which time temperature abnormality warning is performed, and alarm information of the temperature abnormality of the PSU region is sent to a preset terminal.
[0089] Through the above technical solution, the data of all monitoring points are comprehensively considered, which can comprehensively understand the overall distribution characteristics of the temperature in the target heat dissipation region, help to find potential temperature abnormal regions and temperature change trends, and provide comprehensive and accurate information for subsequent temperature management and fault prevention. When the temperature mean standard deviation of the target heat dissipation region is greater than the preset threshold, the system immediately determines that the region is temperature abnormal and triggers the warning mechanism, ensuring that relevant operation and maintenance personnel can know the abnormal situation in the first time. Different heat dissipation regions have different normal fluctuation ranges of temperature distribution due to different hardware compositions and working characteristics. The system sets exclusive standard deviation thresholds according to the hardware characteristics of each region, fully considers the heating characteristics and heat dissipation requirements of different hardware devices, and better adapts to the characteristics of regional temperature changes.
[0090] Optionally, in some embodiments, after sending the alarm information of the temperature abnormality of the target temperature monitoring point to the preset terminal, the method further includes: generating a fault diagnosis report and a processing suggestion according to the alarm information of the temperature abnormality of the target temperature monitoring point.
[0091] During the operation of the temperature monitoring system, when the system detects an abnormal temperature at the target temperature monitoring point, an alarm mechanism is triggered immediately, and alarm information containing the key information of the abnormal temperature at the target temperature monitoring point is sent to the pre-set terminal device. This pre-set terminal can be the mobile phone or computer of the relevant operation and maintenance personnel, ensuring that the operation and maintenance personnel can be informed of the abnormal situation in a timely manner.
[0092] After successfully sending the alarm information, the system uses the built-in fault diagnosis algorithm and knowledge base to conduct a comprehensive and in-depth analysis and judgment of the abnormal situation based on the obtained alarm information of the abnormal temperature at the target temperature monitoring point. Through analysis of the degree of temperature abnormality, change trend, historical data comparison, and other factors, a detailed fault diagnosis report is generated. This report clearly indicates the possible causes of the temperature abnormality, such as equipment heat dissipation failure, high environmental temperature, sensor failure, etc.
[0093] At the same time, the system also provides targeted treatment recommendations for operation and maintenance personnel based on the fault diagnosis results, combined with pre-set treatment strategies and experience data. The treatment recommendations may include specific operation steps, such as checking whether the equipment heat dissipation fan is operating normally, adjusting the environmental temperature control parameters, replacing the faulty sensor, etc., as well as the priority of the treatment recommendations and the estimated processing time, etc., helping the operation and maintenance personnel to quickly and effectively solve the temperature abnormality problem.
[0094] Through the above technical solutions, by generating a fault diagnosis report and treatment recommendations immediately after sending the alarm information, the operation and maintenance personnel do not need to spend a lot of time to investigate the fault cause and develop a treatment plan, and can quickly carry out maintenance work according to the report and recommendations provided by the system, greatly shortening the fault handling time and improving the operation efficiency and stability of the entire system.
[0095] Optionally, in some embodiments, after detecting the second temperature change rate of the temperature monitoring point in the target heat dissipation area within the preset time period based on the target sampling frequency, it includes: judging whether the second temperature change rate of the temperature monitoring point in the target heat dissipation area is less than the first preset change rate; in the case that the second temperature change rate of the temperature monitoring point in the target heat dissipation area is less than the first preset change rate, the target sampling frequency of the temperature monitoring point in the target heat dissipation area is reduced to the initial sampling frequency.
[0096] It can be understood that after all temperature monitoring points in the target heat dissipation area detect the second temperature change rate within the preset time period based on the target sampling frequency, if the second temperature change rate of all temperature monitoring points in the target heat dissipation area is less than the first preset change rate, the target sampling frequency of all temperature monitoring points is reduced to the initial sampling frequency.
[0097] Specifically, the system continuously monitors the temperature change trend of all temperature monitoring points in the target heat dissipation region, and when the second temperature change rate of all temperature monitoring points in the target heat dissipation region falls back to the stable interval (i.e., the temperature change rate is less than the first preset change rate), the sampling frequency of all temperature monitoring points is reduced to the initial sampling frequency, i.e., the frequency of once every 10 seconds.
[0098] For example, if the PSU region previously enters high-frequency sampling due to excessive temperature change rate caused by load increase, when the temperature change rate of all temperature monitoring points in the PSU region stabilizes after the load decreases, the sampling frequency will gradually return to the initial sampling frequency, avoiding unnecessary high-frequency sampling and reducing system resource waste.
[0099] Through the above technical solutions, high-frequency sampling and medium-frequency sampling will cause the system to be in a high-load running state, which is easy to cause system overload and problems such as lag and crash. When the temperature change trend of the temperature monitoring point tends to be stable, reducing the sampling frequency can reduce the working pressure of the system, avoid unnecessary high-frequency sampling, and reduce system resource waste.
[0100] So that those skilled in the art can further understand the temperature monitoring method of the server of the embodiments of the present application, the following will be described in detail in conjunction with specific embodiments, such as Figure 3 As shown in the drawings.
[0101] In step S301, the heat dissipation region layout and the temperature monitoring point deployment. The system first divides the CPU, hard disk backplane, PSU and other independent heat dissipation regions according to the server hardware layout, and deploys monitoring points in each heat dissipation region.
[0102] In step S302, the initial sampling frequency of the temperature monitoring point is set. When initially running, the temperature is collected at a frequency of once every 10 seconds.
[0103] In step S303, the sampling is started.
[0104] In step S304, it is judged whether the first temperature change rate of any temperature monitoring point is greater than or equal to 0.5℃ / min. If it is greater, step S305 is executed, and if the first temperature change rate of all temperature monitoring points is less than 0.5℃ / min, step S306 is executed.
[0105] In step S305, it is judged whether the first temperature change rate of the temperature monitoring point is greater than 5℃ / min. If it is greater, step S307 is executed, and if it is greater than or equal to 0.5℃ / min and less than or equal to 5℃ / min, step S308 is executed.
[0106] In step S306, the initial sampling frequency is maintained.
[0107] Step S307: adjust the sampling frequency of all temperature monitoring points in the target heat dissipation region where the temperature monitoring point is located to a high frequency sampling frequency, and execute step S309.
[0108] Step S308: adjust the sampling frequency of all temperature monitoring points in the target heat dissipation region where the temperature monitoring point is located to a medium frequency sampling frequency, and execute step S309.
[0109] Step S309: calculate whether there is a single point high temperature in all temperature monitoring points of the target heat dissipation region. If there is a single point high temperature, execute step S310, and if there is no single point high temperature, execute step S311.
[0110] Step S310: mark the single point as an abnormal temperature point, and trigger a single point temperature abnormality warning.
[0111] Step S311: calculate the temperature mean standard deviation of all temperature monitoring points in the target heat dissipation region, and execute S312.
[0112] Step S312: determine whether the temperature mean standard deviation exceeds a preset threshold. If it exceeds, execute step S313, and if it does not exceed, execute step S314.
[0113] Step S313: trigger a target heat dissipation region temperature abnormality warning.
[0114] Step S314: continue next monitoring.
[0115] In summary, the technical effects brought by the embodiments of the present application are as follows.
[0116] (1) Significantly improve the server temperature monitoring efficiency:
[0117] Based on the hardware region heat dissipation characteristics, the temperature change rate is collected and calculated at a second level, which can quickly capture subtle temperature changes. For example, in the hard disk backplane region, local temperature changes caused by a single hard disk failure can be detected in time. Compared with traditional monitoring methods, the monitoring accuracy is greatly improved.
[0118] (2) In terms of fault diagnosis, single point and regional abnormality judgment are used together. The neighborhood comparison gradient of the monitoring point is dynamically determined to accurately locate the single point abnormality. Through the exclusive standard deviation threshold, the regional abnormality is judged. In the interface card region, the overall overheating caused by device density or failure can be detected in advance, which greatly shortens the fault troubleshooting time and effectively reduces the risk of server downtime.
[0119] (3) In terms of energy saving and consumption reduction, the system adopts a dynamic sampling strategy. Low-frequency sampling is performed when the temperature is stable to reduce BMC resource consumption; the frequency is increased when the temperature fluctuates, avoiding invalid data collection, significantly reducing system operating energy consumption, and extending the service life of server hardware. In addition, the fault diagnosis reports and handling suggestions output by the system effectively improve operation and maintenance efficiency and have high practical value and economic benefits in scenarios such as data centers.
[0120] According to the server temperature monitoring method proposed in this embodiment, if the first temperature change rate of any temperature monitoring point within the server is greater than or equal to the first change rate, a target heat dissipation area for that temperature monitoring point is determined, and the sampling frequency of the temperature monitoring points within the target heat dissipation area is adjusted to the target sampling frequency. Based on the target sampling frequency, a second temperature change rate of the temperature monitoring points within the target heat dissipation area is detected within a preset time period. If the second temperature change rate of the target temperature monitoring points within the target heat dissipation area is greater than the second change rate within the preset time period, the temperature of the target temperature monitoring point is determined to be abnormal; otherwise, the temperature monitoring result of the target heat dissipation area is obtained based on the standard deviation of the mean temperature of the target heat dissipation area within the preset time period. This solves the problems of weak area division and anomaly location capabilities, and coarse dynamic sampling and threshold management in current server planar temperature control technology. It refines the planar area division, automatically adjusts the sampling frequency according to the temperature change amplitude, and achieves accurate identification of single-point anomalies within the area, improving monitoring accuracy. In high-density server scenarios, it can balance monitoring accuracy, resource consumption, and response speed.
[0121] Next, the temperature monitoring system for a server according to an embodiment of the present invention is described with reference to the accompanying drawings.
[0122] Figure 4 This is a schematic diagram of a server temperature monitoring system according to an embodiment of the present invention.
[0123] like Figure 4 As shown, the server temperature monitoring system 10 includes: an acquisition module 100, an adjustment module 200, and a monitoring module 300.
[0124] The acquisition module 100 is configured to acquire a first temperature change rate of at least one temperature monitoring point in the server. The adjustment module 200 is configured to, if the first temperature change rate of any temperature monitoring point is greater than or equal to a first preset change rate, determine a target heat dissipation region where the any temperature monitoring point is located, and adjust a sampling frequency of the temperature monitoring point in the target heat dissipation region to a target sampling frequency. The monitoring module 300 is configured to detect a second temperature change rate of the temperature monitoring point in the target heat dissipation region within a preset time length based on the target sampling frequency, and if the second temperature change rate of a target temperature monitoring point in the target heat dissipation region within the preset time length is greater than a second preset change rate, determine that the target temperature monitoring point is abnormal in temperature, otherwise, obtain a temperature monitoring result of the target heat dissipation region based on a standard deviation of a temperature mean value of the target heat dissipation region within the preset time length.
[0125] Optionally, in some embodiments, before the acquisition module 100 acquires the first temperature change rate of the at least one temperature monitoring point in the server, the acquisition module is further configured to: divide the server into a plurality of heat dissipation regions based on heat dissipation characteristics of hardware regions of the server, wherein the plurality of heat dissipation regions are respectively a backboard heat dissipation region, a power supply heat dissipation region, a central processing unit heat dissipation region, and an interface card heat dissipation region; determine a number of temperature monitoring points of the backboard heat dissipation region according to an area and heat dissipation characteristics of the backboard heat dissipation region, and determine a number of temperature monitoring points of the power supply heat dissipation region according to an area and heat dissipation characteristics of the power supply heat dissipation region, and determine a number of temperature monitoring points of the central processing unit heat dissipation region according to an area and heat dissipation characteristics of the central processing unit heat dissipation region, and determine a number of temperature monitoring points of the interface card heat dissipation region according to an area and heat dissipation characteristics of the interface card heat dissipation region.
[0126] Optionally, in some embodiments, after the monitoring module 300 determines that the target temperature monitoring point is abnormal in temperature, the monitoring module 300 is further configured to: perform temperature abnormality early warning based on the target temperature monitoring point, and send alarm information of the temperature abnormality of the target temperature monitoring point to a preset terminal.
[0127] Optionally, in some embodiments, the monitoring module 300 is further configured to: calculate a standard deviation of a temperature mean value of the temperature monitoring point in the target heat dissipation region within the preset time length; if the standard deviation of the temperature mean value is greater than a preset threshold corresponding to the target heat dissipation region, determine that the target heat dissipation region is abnormal in temperature, and perform temperature abnormality early warning, and send alarm information of the temperature abnormality of the target heat dissipation region to a preset terminal.
[0128] Optionally, in some embodiments, the adjustment module 200 is further configured to: determine a target temperature change interval in which the first temperature change rate of any temperature monitoring point is located, and determine the target sampling frequency according to the target temperature change interval.
[0129] Optionally, in some embodiments, the adjusting module 200 is further configured to: determine whether the first temperature change rate of any temperature monitoring point is greater than or equal to a first preset change rate and less than or equal to a second preset change rate; if the first temperature change rate of any temperature monitoring point is greater than or equal to the first preset change rate and less than or equal to the second preset change rate, determine that the target temperature change interval is a first interval, and set the first sampling frequency as the target sampling frequency of the temperature monitoring point in the target heat dissipation region; if the first temperature change rate of any temperature monitoring point is greater than the second preset change rate, determine that the target temperature change interval is a second interval, and set the second sampling frequency as the target sampling frequency of the temperature monitoring point in the target heat dissipation region.
[0130] Optionally, in some embodiments, the first preset change rate is 0.5℃ / min, the second preset change rate is 5℃ / min, the first interval is greater than or equal to 0.5℃ / min and less than or equal to 5℃ / min, the second interval is greater than 5℃ / min, the first sampling frequency is 1 time per second, and the second sampling frequency is 10 times per second.
[0131] Optionally, in some embodiments, after obtaining the first temperature change rate of at least one temperature monitoring point in the server, the obtaining module 100 is further configured to: in the case where there is no first temperature change rate of any temperature monitoring point greater than or equal to the first preset change rate within a preset time length, control the temperature monitoring point in the heat dissipation region to maintain the current sampling frequency.
[0132] Optionally, in some embodiments, after sending the alarm information of the temperature abnormality of the target temperature monitoring point to the preset terminal, the monitoring module 300 is further configured to: generate a fault diagnosis report and a processing suggestion according to the alarm information of the temperature abnormality of the target temperature monitoring point.
[0133] Optionally, in some embodiments, after detecting the second temperature change rate of the temperature monitoring point in the target heat dissipation region within a preset time length based on the target sampling frequency, the monitoring module 300 is further configured to: determine whether the second temperature change rate of the temperature monitoring point in the target heat dissipation region is less than the first preset change rate; in the case where the second temperature change rate of the temperature monitoring point in the target heat dissipation region is less than the first preset change rate, reduce the target sampling frequency of the temperature monitoring point in the target heat dissipation region to the initial sampling frequency.
[0134] It should be noted that the description of the features in the embodiments of the temperature monitoring system of the server can refer to the related description of the embodiments of the temperature monitoring method of the server, which will not be repeated here.
[0135] Figure 5 The structure schematic diagram of the electronic device provided by the embodiments of the present application is provided. The electronic device can include:
[0136] The memory 501, the processor 502 and the computer program stored in the memory 501 and executable on the processor 502.
[0137] The processor 502 implements the temperature monitoring method of the server provided in the above embodiments when executing the program.
[0138] Further, the electronic device further comprises:
[0139] The communication interface 503 is used for communication between the memory 501 and the processor 502.
[0140] The memory 501 is used for storing the computer program executable on the processor 502.
[0141] The memory 501 can include a high-speed RAM memory, and can also include a non-volatile memory, for example, at least one disk memory.
[0142] If the memory 501, the processor 502 and the communication interface 503 are independently implemented, the communication interface 503, the memory 501 and the processor 502 can be connected to each other through a bus and complete communication between each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 5 In the figure, only one thick line is used to represent, but it does not mean that there is only one bus or one type of bus.
[0143] Optionally, in specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can complete communication between each other through an internal interface.
[0144] The processor 502 can be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0145] The embodiment of the present application also provides a nonvolatile computer readable storage medium, which stores a computer program, and the computer program is arranged to execute the steps in any of the above-mentioned server temperature monitoring method embodiments when running.
[0146] In an exemplary embodiment, the above-mentioned computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0147] The embodiment of the present application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the above-mentioned server temperature monitoring method.
[0148] The skilled person can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0149] The above describes in detail the server temperature monitoring method and the electronic device provided by the present application. The principle and implementation mode of the present application are described by applying specific examples in this paper, and the above-mentioned example is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary skilled person in the art, some improvements and modifications can be made to the present application without departing from the principle of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A temperature monitoring method of a server, characterized by, The method comprises the following steps: obtaining a first temperature change rate of at least one temperature monitoring point in a server; if the first temperature change rate of any of the temperature monitoring points is greater than or equal to a first preset change rate, determining a target heat dissipation region where any of the temperature monitoring points is located, and adjusting the sampling frequency of the temperature monitoring points in the target heat dissipation region to a target sampling frequency; based on the target sampling frequency, detecting a second temperature change rate of the temperature monitoring points in the target heat dissipation region within a preset time length, if the second temperature change rate of a target temperature monitoring point in the target heat dissipation region within the preset time length is greater than a second preset change rate, determining that the temperature of the target temperature monitoring point is abnormal, otherwise, obtaining a temperature monitoring result of the target heat dissipation region based on the standard deviation of the mean temperature of the target heat dissipation region within the preset time length; before obtaining the first temperature change rate of at least one temperature monitoring point in the server, the method comprises: based on the heat dissipation characteristics of the hardware regions of the server, dividing the server into a plurality of heat dissipation regions; and determining the number of temperature monitoring points of the heat dissipation regions according to the area and heat dissipation characteristics of the heat dissipation regions; obtaining the temperature monitoring result of the target heat dissipation region based on the standard deviation of the mean temperature of the target heat dissipation region within the preset time length further comprises: calculating the standard deviation of the mean temperature of the temperature monitoring points in the target heat dissipation region within the preset time length; if the standard deviation of the mean temperature is greater than a preset threshold corresponding to the target heat dissipation region, determining that the temperature of the target heat dissipation region is abnormal, and performing temperature abnormality warning, and sending alarm information of the temperature abnormality of the target heat dissipation region to a preset terminal.
2. The temperature monitoring method of a server according to claim 1, wherein, The plurality of heat dissipation regions are respectively a backplane heat dissipation region, a power supply heat dissipation region, a central processing unit heat dissipation region, and an interface card heat dissipation region, and the number of temperature monitoring points of the heat dissipation regions is determined according to the area and heat dissipation characteristics of the heat dissipation regions, comprising: determining the number of temperature monitoring points of the backplane heat dissipation region according to the area and heat dissipation characteristics of the backplane heat dissipation region, and determining the number of temperature monitoring points of the power supply heat dissipation region according to the area and heat dissipation characteristics of the power supply heat dissipation region, and determining the number of temperature monitoring points of the central processing unit heat dissipation region according to the area and heat dissipation characteristics of the central processing unit heat dissipation region, and determining the number of temperature monitoring points of the interface card heat dissipation region according to the area and heat dissipation characteristics of the interface card heat dissipation region.
3. The temperature monitoring method of a server according to claim 1, wherein, after determining that the temperature of the target temperature monitoring point is abnormal, comprising: based on the target temperature monitoring point, performing temperature abnormality warning, and sending alarm information of the temperature abnormality of the target temperature monitoring point to a preset terminal.
4. The temperature monitoring method of a server according to claim 1, wherein, the method further comprises: determining a target temperature change interval in which the first temperature change rate of any of the temperature monitoring points is located, and determining the target sampling frequency according to the target temperature change interval.
5. The temperature monitoring method of a server according to claim 4, wherein, the method further comprises: determining a target temperature change interval in which the first temperature change rate of any of the temperature monitoring points is located, and determining the target sampling frequency according to the target temperature change interval. determining whether the first temperature change rate of any of the temperature monitoring points is greater than or equal to a first preset change rate and less than or equal to a second preset change rate; if the first temperature change rate of any of the temperature monitoring points is greater than or equal to the first preset change rate and less than or equal to the second preset change rate, determining that the target temperature change interval is a first interval, and taking the first sampling frequency as the target sampling frequency of the temperature monitoring points in the target heat dissipation region; if the first temperature change rate of any of the temperature monitoring points is greater than the second preset change rate, determining that the target temperature change interval is a second interval, and taking the second sampling frequency as the target sampling frequency of the temperature monitoring points in the target heat dissipation region.
6. The temperature monitoring method of a server according to claim 5, wherein, The first preset change rate is 0.5°C / min, the second preset change rate is 5°C / min, the first interval is greater than or equal to 0.5°C / min and less than or equal to 5°C / min, the second interval is greater than 5°C / min, the first sampling frequency is 1 time per second, and the second sampling frequency is 10 times per second.
7. The temperature monitoring method of a server according to claim 1, wherein, After obtaining the first temperature change rate of the monitoring points of each heat dissipation region, the method comprises: if the first temperature change rate of any of the temperature monitoring points is not greater than or equal to the first preset change rate within a preset time length, controlling the temperature monitoring points in the heat dissipation region to maintain the current sampling frequency.
8. The temperature monitoring method of a server according to claim 3, wherein, After sending the alarm information of the target temperature monitoring point temperature abnormality to a preset terminal, the method comprises: generating a fault diagnosis report and a processing suggestion according to the alarm information of the target temperature monitoring point temperature abnormality.
9. The temperature monitoring method of a server according to claim 1, wherein, After detecting the second temperature change rate of the temperature monitoring points in the target heat dissipation region within a preset time length based on the target sampling frequency, the method comprises: determining whether the second temperature change rate of the temperature monitoring points in the target heat dissipation region is less than the first preset change rate; if the second temperature change rate of the temperature monitoring points in the target heat dissipation region is less than the first preset change rate, reducing the target sampling frequency of the temperature monitoring points in the target heat dissipation region to an initial sampling frequency.
10. An electronic device, comprising: The server comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the temperature monitoring method of the server according to any one of claims 1-9.
Citation Information
Patent Citations
Optical module heat dissipation method and device based on 800G transmission, computer equipment and storage medium
CN119012637A